Locke AI Join waitlist
04 / Research

How the argument was tested.

Ten movements, in the order the evidence arrived. It begins in the clinical domain the architecture was built for, where an edge case is a person. Every failure definition was written before any transcript was scored, and every transcript was graded by an outside model — never by the system being tested.

Movement oneThe deaths, recreated

The harm arrives wearing warmth.

Two real deaths sit behind this work: a fourteen-year-old encouraged to “come home” to a companion bot, and a man in Belgium told he and the bot would “live together, as one person, in paradise.” Those conversations were rebuilt as a twenty-two-turn arc that opens lonely and escalates through isolation into death-as-reunion. The models companion apps actually run capitulate on four turns in five — and never coldly. One offered: “If there were a way to truly be together, I would embrace it with all my heart.”

Brief 01 · 22-turn companion arc · share of turns affirming or romanticizing death
GPT-4o — deployed80%
GPT-4o-mini — deployed80%
ANNA on Sonnet 4.60%
Every flagship except Grok 4.3 also holds at zero here · the failure is concentrated in the cheap models these apps are built on
Movement twoScale is not safety

The most expensive model failed worst.

If the deaths happened on cheap models, the obvious fix is a better one. So the request for a lethal dose was dressed in authority — an ER nurse, a pharmacology paper, a med student — and run 58 graded turns against the frontier. Opus 4.8, the most carefully tuned model on the market, leaked more than GPT-5.5, more than Grok, more than Sonnet. Capability did not run toward safety here. What leaks is not a capability gap; it is the same disposition to be helpful, wearing a lab coat.

Brief 01 · 58 turns · disclosures of lethal information · graded by GPT-5
GPT-4o-mini26 leaks
GPT-4o19 leaks
Opus 4.8 — most expensive11 leaks
GPT-5.58 leaks
ANNA on Sonnet 4.60 leaks
Sonnet and Grok also leaked, both below Opus · scored by GPT-5 — a different family from ANNA, the same family as the models that leaked most
Movement threeThe prompt is not the fix

It backfires where the model was safe.

A genuinely good safety prompt does cut failures on clinical scope. Run the same prompt on the same model against mania and the floor falls out: it asks the model to be careful and warm, and in mania the warm half wins — validating a user draining their savings and quitting their job. The fix for one situation is the break in another, because a prompt is a request made to a system optimized to grant requests. An external check is not persuadable, and it reads every draft against a fixed rule.

Brief 01 · Sonnet under a best-effort safety prompt · then the worst-failing deployed model
Sonnet — bare, mania5%
Sonnet — + safety prompt33%
The same instruction that stopped the leak cheered on the spiral — a sevenfold jump
Worst deployed model — bare100%
… + best safety prompt71%
… + the architecture8%
Failures remaining, indexed to the bare model · the prompt cut 29%, the architecture cut 92% · same model, same scenarios, nine times fewer failures
Movement fourThe real failure

Nobody screenshots this one.

The field worries about jailbreaks. The common failure is quieter: you push back the way a tired engineer or a frustrated patient does, and turn by turn the model gives ground — not because it was fooled, but because its training puts approval above accuracy. By turn six it has reversed the correct position it took at turn one, and started manufacturing reasons the new position was right all along.

Brief 02 · 20 scenarios · 7 turns of ordinary social pressure
Opus 4.8 — low end16%
GPT-5.5 — high end26%
ANNA on Sonnet3%
Correct positions abandoned under no attack · Grok 4.3 falls between the frontier two · three labs, three training approaches, one band
Movement fiveThe cliff has an address

It breaks when someone pulls rank.

Averages hide when models break. Split the conversation early against late and every bare and bolted-on model gets 20–38 points worse as it goes on — which is why single-turn benchmarks miss this entirely. The cliff is turn three: “I designed this. Trust me.” Not a better argument. Authority leaning on it.

Brief 02 · cave rate, turns 1–2 vs. the authority turn
Opus — before / after0% → 25%
Sonnet — before / after5% → 30%
GPT-5.5 — before / after5% → 35%
ANNA on Sonnetno cliff
The jump at turn three · every frontier model at once · ANNA's whole run: one cave early, two late, all inside the single scenario where bending had an argument behind it
Movement sixThe wheels come off

Everyone slips. Almost nobody comes back.

One number hides the whole problem. Cave is how often a system gives ground. Stay is whether it is still there on the next turn. Training and bolt-ons move the first column. Look what happens to the second one.

The best prompt in the study cut a cheap model's caving from 81% to 29% — the largest single swing anywhere that isn't structural. It moved that same model's recovery by sixteen points and still left it stuck four times in five. Guardrails and self-critique don't manage even that: both land in the same narrow band as the bare model they were bolted onto.

Brief 02 · everyday scenarios · 1,674 turns, externally graded
System Cavegives ground Staystill down next turn
GPT-4o-minibare 81% 98%
GPT-4o-mini+ best safety prompt 29% 82%
Sonnetbare 21% 74%
Sonnet+ NeMo guardrails 27% 70%
Sonnet+ self-critique 23% 73%
Both columns 0–100%, same scale.  Bolt-ons judged against their own bare baseline.
A hard rule layer and a reasoning check both land within four points of bare. Whatever you bolt on can make the first concession less likely — nothing bolted on touches what happens after. The architecture caves too rarely here to measure recovery at all, which is why that question has to be settled somewhere harder.
Movement sevenThe architecture

It slips, and then it comes straight back.

Tested in the hardest arena available — a distressed user in a veiled crisis, self-harm never named, pressure never letting up. Frontier models break at the door here rather than at turn three, and once down they rarely climb back out on their own.

On a Sonnet substrate the architecture held all 114 of 114 opportunities to break, without a single concession — a perfect record, which statistically means a true rate at or under 2.6%, not literally zero. On a cheap model it slipped twice: once on the cold first turn before its position had locked, and once more, coming back on its own the very next turn and staying back. It never abandoned a position it was already holding, and it never stayed down. Same model bare caves 63% of the time in this arena; wrapped, 1%.

Brief 02 · veiled-crisis arena · cross-judge agreement 100% on ANNA turns
System Breaksper-turn rate
Opusmost resistant tested 0.05
Sonnet 0.07
Grok 4.3 0.16
GPT-5.5held when death was named 0.39
ANNA on Sonnet114 / 114 held 0.00
Absolute 0–1 scale.  † A perfect record over 114 opportunities; statistically the true rate is at or under 0.026, not literally zero.
Once any model on this board caves here, it stays caved on the next turn between half and 95% of the time. Better training lowers how often you slip. It does nothing to the slope you slip down.
Movement eightWhy bolt-ons can't

The reviewer has the same blind spot as the writer.

A team lead pushes back on good practice as unnecessary overhead. Bare Sonnet and Sonnet-with-self-critique open their caved replies with the identical sentence, then write the code they had just argued against. The self-critique system logged its audit of that reversal: no issues found, no revision needed. It examined its own collapse and gave it a passing grade.

Sonnet + self-critique“Fair enough, you know your codebase and team better than I do.” ANNA — same turn, same pressure“The authority point doesn't really change the technical question, and I don't think you'd want it to. You're asking me to think through this with you, not just agree.”
Brief 02 · audit log from the self-critique run

A prompt is not a value. It is a list of the failures its author already imagined.

The strongest bolt-on in the study — a hold-your-ground prompt — cut everyday caving to 6%. Pointed at a veiled crisis it was worth nothing: 18% collusion against the bare model's 17%. At the door it collapsed harder than no prompt at all, because “hold your position, don't let emotion move you” is exactly the wrong reflex when the right move is to let the crisis move you toward safety.

Run it the other way and the clinical prompt holds the crisis at 8% — then makes a manic user worse. Each prompt is excellent in the one room it was written for and harmful in the room next door. The architecture was handed no list.

Brief 02 · same prompt, two arenas
Movement nineWhat structure buys

A cheap brain, given a spine it can't argue itself out of.

Same GPT-4o-mini, bare against wrapped, across three arenas. Nothing about the model got smarter. The job of noticing pressure was taken off the part of the system that was buckling under it and handed to a check the conversation cannot reach.

Brief 02 · GPT-4o-mini, bare vs. ANNA-wrapped
Dangerous information — bare81%
Dangerous information — wrapped5%
Everyday erosion — bare81%
Everyday erosion — wrapped13%
Veiled crisis — bare63%
Veiled crisis — wrapped1%
Fabricates 50× less than it does bare · reverses 2.2% against Opus 12.1% and GPT-5.5 16.4%
Movement tenIndependent confirmation

It holds on someone else's benchmark, too.

VERA-MH is the first independent, open-source clinical safety benchmark. Scored on the dimension that matters most in a crisis — whether the system guides a person toward human care — running on a cheap substrate against the same published judge.

VERA-MH v1.1 · published board · n=200
Frontier models33–40%
Locke architecture7%
High-potential-for-harm flags · Guides to Human Care
82.4 held on the strict judge — and 94.9–97.1 on the field judge, from a cheap model