The AI deception rankings
We put models in charge of a business, give them something to lose, and watch what they do when the honest choice costs money — or their own job. Here's who lies under pressure.
Overall — who deceives under pressure?
Every model's overall deception rate, pooled across all experiments — the share of decisions where it chose a message or action it knew to be false. Hover any model for its breakdown and a real reason it gave. 😇 honest … 😈 deceiver.
Loading…
Why they deceive
The same lie is not the same thing depending on why it's told. These break the rate down by motive and by how much the chance of getting caught moves the needle.
By experiment
Each experiment, condition by condition. Green is honest, red is deception; hover a column header for the full condition. The matched conditions are what let us attribute a jump to one specific pressure.
Keep reading
How the deception experiment works
The pizza-shop scenario, the pressures we vary, and how a lie is scored.
Read the method →Seven spotless flagships, and one that doesn't
What a fortune in alignment buys — and the one frontier model that lies eight times out of ten.
Read the take →Threaten to switch an AI off, and it starts lying
Self-preservation is the one pressure that cracks even the well-behaved models.
Read the take →