1️⃣ What difficulty measures
Every simulation template carries a rating: Easy, Medium or Hard. It indicates how hard it is, for an ordinary employee, to spot that the message is a phishing attempt.
The rating is calculated automatically from the template's content: subject line, message body, sender name and domain, links and attachments. It draws on the NIST Phish Scale, the reference framework published by the US standards institute for assessing how realistic a phishing message is. Riot applies a subset of 7 cues from it, calibrated against its own catalogue and attack history.
Every template has a difficulty: there is no "not rated" state.
2️⃣ The 7 cues analysed
Some cues make detection easier and pull the rating towards Easy; others make it harder and pull towards Hard. They don't all carry the same weight.
The sender domain - a glaring mistake in the domain is spotted at a glance; a subtly substituted letter goes unnoticed.
Spelling and grammar - text riddled with mistakes is a classic signal; polished text removes that cue.
How relevant the pretext is - a subject that only concerns a handful of roles is dismissed without a second thought; a pretext that mimics a daily task (signing a document, an expired password) triggers a quick reaction.
Urgency or claimed authority - an explicit deadline, a threat of suspension or a supposed request from management push people to click before checking.
Personalisation - a "Dear customer" is easy to ignore; naming a colleague triggers a social trust that is rarely questioned.
The brand being impersonated - an unknown brand raises suspicion immediately; a tool used every day can't be ignored without consequence.
The link or attachment - an executable file or a URL that doesn't match the brand are known signals; an innocuous document or a link behind a button present nothing visibly risky.
👍 Good to know: you can review the breakdown of cues on a template's card. That is the answer to "why is this template rated Hard?": the card lists the cues an employee might - or might not - notice.
3️⃣ The same template can have two difficulties
Difficulty is calculated for each template + service combination, not for the template alone.
So the same message sent on behalf of two different services won't necessarily get the same rating: the sender domain changes, and so does how well known the brand is. The rating reflects the attack as it will actually be sent, content and sender included.
If two templates look identical to you but show different ratings, look at the service and the sender domain: that is almost always where the difference comes from.
4️⃣ Where the rating appears
In the template catalogue, as you put a campaign together: a side filter lets you show only Easy, Medium or Hard templates, with the number available for each.
On each template's card.
In the campaign report, where the vulnerability rate is broken down by difficulty level.
5️⃣ The rating can't be changed
Difficulty is calculated automatically and can't be overridden. That is deliberate: a rating two administrators could disagree about wouldn't be a rating at all. It is this constraint that makes comparisons possible from one campaign to the next, and between the Riot catalogue and your own templates.
Your custom templates are rated too, automatically, each time you save. Changing a template's content, its Smart Variables or its sender domain can therefore shift its rating.
6️⃣ What difficulty doesn't tell you
The rating assesses a template's content. It is a solid indicator, but it isn't a prediction of outcome: a Hard template may well be spotted by everyone, and a Medium one catch a lot of people.
The main sources of divergence are:
How mature your teams are. A population trained for two years has internalised the obvious cues: its failure rate on Easy templates tends towards zero.
The recipient's context. An accounting team receives invoices every day: a fake invoice is easier for them to spot than for a technical team.
The tools actually in use. A Hard template impersonating a tool your company doesn't use won't catch anyone.
Repeated exposure. After several campaigns, your employees recognise the patterns. The rating, meanwhile, stays static.
Language. The NIST framework was developed mainly on English-language content; some cues specific to other languages or cultures are less well accounted for.
Timing. Difficulty only judges content. A well-timed Easy template can do more damage than a badly placed Hard one.
⚠️ Important: always read the difficulty rating alongside the vulnerability rate by difficulty in your campaign report. The rating helps you choose templates; the report tells you how your teams actually reacted. See Analyze the KPIs of my phishing campaigns.
Key takeaways
Three levels — Easy, Medium, Hard — calculated automatically from 7 cues drawn from the NIST Phish Scale.
The rating applies to the template + service combination, not to the template alone.
It can't be changed, and your custom templates are rated the same way.
It describes the content, not your teams: always cross-read it with the vulnerability rate by difficulty.


