Around 22% of UK businesses test their staff with mock phishing exercises. Most of those programmes measure the wrong thing, and a significant body of evidence suggests some of them make the organisation marginally worse off.
That is an uncomfortable opening for an article about how to write a good one. It is also the reason this guide exists: the difference between a phishing simulation that builds resilience and one that quietly corrodes trust is almost entirely in the design, and the design decisions are not intuitive.
Start with the evidence, because it is not what vendors say
The largest and longest study of phishing simulation in a real organisation was run by researchers at ETH Zurich over fifteen months, covering around 14,000 employees in a company of more than 56,000 people. Two findings from it should shape any programme you build.
Finding one: the training most programmes are built around did not work. Embedded training, the page someone lands on immediately after clicking a simulated phish, did not improve resilience. In some conditions employees who received it became *more* susceptible, not less. The researchers suggested a mechanism that should worry anyone running a programme: staff saw that their employer was investing in protection and felt safer as a result.
Finding two: something else worked very well. Giving employees a simple button to report suspicious messages produced accurate, sustained reporting. Reports were correct about 68% of the time for phishing, rising to 79% when spam was counted, and the most active reporters exceeded 80% accuracy. Critically, there was no reporting fatigue over fifteen months. Crowd-sourced reporting let the organisation detect genuine campaigns shortly after they began.
Read together, these point somewhere specific. The value of a phishing programme is the reporting channel, not the gotcha. Design accordingly.
The metric problem
Almost every phishing simulation programme reports click rate, and almost every one is worse for it.
Click rate is attractive because it goes down. The trouble is that it goes down for reasons that have nothing to do with resilience. Staff learn the shape of *your* simulations rather than the shape of attacks. They learn to hesitate over anything unfamiliar, including legitimate email from new suppliers. In organisations where clicking has consequences, they learn to quietly delete rather than report, which actively removes your early warning.
A programme can reach a very low click rate while leaving the organisation less able to detect a real campaign than when it started. That is not a hypothetical failure mode; it is the predictable result of optimising the metric.
What to measure instead
Report rate. What proportion of recipients reported the message? This is the number that maps to a real defensive capability.
Time to first report. How long from delivery to the first person raising it? This is the single most operationally useful figure you will collect, because it determines how quickly you could respond to a genuine campaign.
Critical action rate. Clicking a link is a weak signal. Entering credentials, approving a multi-factor prompt or replying with information is where real harm begins. Track those separately and treat them as a different category of event.
Repeat rate over time. Individuals who repeatedly take critical actions are a coaching conversation, not a statistic.
A reasonable dashboard has report rate and time to first report at the top, critical actions second, and click rate present but demoted, because it is diagnostic rather than a target.
Writing the simulation itself
Make it plausible, not clever
The goal is a message that resembles what your people actually receive. Signals that used to distinguish phishing, poor spelling and awkward phrasing, have been removed by generative AI at essentially no cost. A simulation containing them is testing people against an attacker who no longer exists.
Base the pretext on a real process in your organisation: an expenses system, a document sharing notification, an HR request, a supplier portal. Realism comes from context, not from sophistication.
Calibrate difficulty deliberately
Run a mix. Easy simulations confirm the basics are holding and give people a win. Hard ones, using real internal context, tell you what a targeted attack would achieve.
The point of a hard simulation is emphatically not to catch people out. It is to establish the ceiling: if a well-researched message succeeds against most of the finance team, that is a finding about your process, not about your finance team.
Do not use themes that cause real harm
This is where programmes damage themselves badly and, occasionally, publicly. A simulation promising a bonus that does not exist, announcing redundancies that are not happening, or referencing a genuine bereavement or crisis will achieve a high click rate and cost you far more than it measures.
The line is straightforward: never simulate something that would cause genuine distress or genuine financial expectation if believed. No fake bonuses, no fake redundancy notices, no fake emergencies involving family. If the reveal would make a reasonable person feel humiliated or deceived about something that matters to them personally, pick another pretext.
The damage is not only ethical. Once staff associate the security team with deception that hurt them, reporting rates fall, and the reporting channel is the part that actually works.
Include the sender you are hardest to defend against
Most programmes simulate external senders. Attackers increasingly do not bother. Messages from a compromised supplier account, or inside a collaboration tool that feels internal and vetted, are both more effective and less tested. If your tooling allows it, include at least one.
Running the programme
Announce that the programme exists
Not the individual simulations, the programme. People should know that their employer runs phishing simulations, why, and what happens if they click. Concealing the programme's existence buys a marginally more authentic first result and costs you trust permanently.
Never make clicking punitive
Programmes that tie simulation results to appraisals, bonuses or public naming produce exactly one reliable outcome: staff stop reporting. They cannot stop clicking, because clicking is a function of attention and workload as much as knowledge, but they can certainly stop telling you.
Treat a click as a system finding. If 40% of a department clicked, the interesting question is what about that department's workload, tooling or process made the message land.
Reward reporting, including when it is wrong
A false positive is a person doing exactly what you asked. If reporting a legitimate email produces mild embarrassment, reporting stops. Thank people for reports regardless of outcome, and make the response fast enough that reporting feels worthwhile.
Frequency: often enough to be routine, rarely enough to be tolerable
Monthly is a reasonable default for most organisations, with higher frequency for high-risk roles such as finance and executive assistants. Quarterly is too infrequent to build habit. Weekly becomes noise and breeds resentment.
What a phishing simulation cannot tell you
This is worth stating plainly, because phishing simulation is frequently sold as a complete answer to human risk and it is not.
It tells you very little about how your organisation behaves after a successful attack. It cannot tell you whether someone will escalate at 6pm on a Friday, whether your finance team would refuse an urgent phone call from a cloned executive voice, whether your help desk would reset credentials for a convincing caller, or whether leadership can make a containment decision under time pressure.
Those are the failures that turn an incident into a crisis, and they are exercised in a room, not in an inbox. The UK retail attacks of 2025 began with phone calls to help desks, not with emails. A phishing programme would not have tested that at all.
The honest framing is that phishing simulation covers one high-volume attack path well and says nothing about the decisions that follow. Both need exercising, and most organisations do only the first.
A programme worth running
Announce the programme. Send plausible monthly simulations grounded in real internal processes. Put a report button in front of every employee and make it one click. Measure report rate and time to first report, track critical actions separately, and demote click rate to a diagnostic. Never punish a click and always thank a report. Keep the pretexts away from anything that would cause real distress.
Then accept that you have covered the inbox, and exercise the decisions separately.