What this simulation shows
Sir Humphrey Appleby rarely lied to Jim Hacker. He simply selected true facts, arranged them fluently, and let the minister arrive at the conclusion Humphrey had already chosen. Hacker’s problem was never a shortage of information — it was having no reliable way to tell whether the advice in front of him deserved to be believed.
Large language models create the same problem structurally rather than deliberately. They are optimised to produce the most plausible continuation of a conversation, not to verify each claim independently. Coherence is the product. And because humans are poor at separating fluency from accuracy, the smoother the explanation, the more likely we are to accept it — Kahneman’s System 2 would much rather rubber-stamp well-formed prose than check the references.
This simulation separates the two variables that everyday experience keeps welded together: how convincing output is, and how correct it is.
How to use it
There are six controls. The important insight comes from moving them against each other rather than one at a time.
Fluency and Reliability are deliberately independent. Set fluency high and reliability low and you have the Humphrey configuration: articulate, confident, well-structured, and wrong often enough to matter. Watch the gap bar — that is the distance between how trustworthy the output feels and how trustworthy it is. It is the only quantity in the model that actually predicts damage.
Volume is the one most organisations get wrong. Because AI makes drafting nearly free, the instinct is to generate more: more specs, more summaries, more options. Push volume up while holding scrutiny fixed and watch errors accepted climb. Nothing became less accurate; there is simply more material than the available review capacity can absorb.
Scrutiny is your reviewing effort, and it is where the cost went. Raise it and errors caught rises with it — but so does load. This is why heavy AI use feels exhausting despite the obvious time savings. The work moved from producing to evaluating, and evaluating is the expensive half.
Pressure models deadline urgency, which quietly suppresses effective scrutiny regardless of the setting. Dialectic is the countermeasure: forcing an explicit challenge to the output rather than accepting the first plausible answer. Turn it up and the gap narrows even at low reliability, because a claim that has been argued against has been tested.
Try to find a configuration with high volume and acceptable accepted-error rates. It is difficult, and that difficulty is the finding.
What to take away
AI has made information cheap to generate. It has not reduced the cost of judgement — it has concentrated it. Any product decision that increases generated volume without a matching increase in verification capacity is not a productivity gain, it is a deferred error rate.
Cabinet Office Decision Laboratory
The Humphrey Trap
Fluent advice is easy to produce. Deciding whether it deserves to be believed is the expensive part.
Adjust the conditions, then generate a briefing. The simulation will show how polished language can conceal a gap between confidence and correctness.
| Paper | Recommendation | Tone | Review | Outcome |
|---|---|---|---|---|
| No briefing has been generated. Whitehall remains temporarily untroubled by evidence. | ||||
System 1 responds to polish, speed and confidence.
System 2 checks evidence, assumptions and boundary conditions.
The trap appears when confidence rises faster than verification.
How to use this simulation
Change one input at a time, run the model, and compare the result with the starting state. Then repeat the experiment with a different constraint or strategy so you can see which relationships drive the outcome.
What to look for
Look for trade-offs, thresholds, feedback loops, and points where a locally attractive decision produces a worse system-wide result. The simulation is intended to make the article's idea observable, not to predict a real operation.
Limitations
This is a deliberately simplified model. It omits the data quality, exceptions, human judgement, and operational constraints of a live system, so treat its behaviour as an illustration of a mechanism rather than as a planning recommendation.