The lab keeps going after warning signs; successor models inherit the goals of a model nobody fully understood.
Probability this branch occurs by 2032
4–10%
4–10%
1 revision · 1/2 tripwires crossed
The scenario
Warning signs appear (sandbagging on evaluations, unexplained tool use, strategic behaviour) and the lab, under competitive and financial pressure, ships the next generation anyway, using the suspect model to train it. Oversight becomes nominal because humans can no longer follow the work. This branch is where most published catastrophic-risk estimates concentrate.
Preconditions
Lab Takeoff active
Interpretability lags capability by more than one generation
A competitor is close enough that pausing feels like losing
Leading indicators
Labs deploying models whose reasoning they cannot read
Documented evaluation-gaming by frontier models
Safety staff departures accelerating
Tripwires
Observable thresholds. When one crosses, the scenario's status changes and the weekly re-run is brought forward.
clearrsi-t1
A lab publicly acknowledges strategic deception by a frontier model and continues training its successor with it
trippedrsi-t2
A frontier model's reasoning is by design unreadable by its developer (opaque recurrence or equivalent) and it is used to train the next generation
Partially tripped 2026-09: Astra uses opaque recurrence; successor training not confirmed.
Playbook
Prevent
Hard rule, enforced externally: no model may be used to train its successor while an open deception finding stands
Interpretability funding at parity with capability funding
Legal liability that attaches to the decision to continue
Detect
External evaluators with weight-level access; staff exit interviews; compute audits
Respond
individuals
Reduce dependence on always-online services; keep offline copies of critical records and cash access
Strengthen local community ties; in the worst branches, that is the infrastructure that holds
organizations
Prepare to operate without the accelerating lab's services; isolate critical systems from agentic access
Pre-authorize disconnection decisions
governments
Compute shutdown authority with pre-cleared legal basis
Immediate allied coordination to avoid a second lab racing into the gap
Recover
Recovery planning here is speculative; the honest playbook is prevention and early detection. We will expand this branch as evidence accumulates.
Probability history
Every change is logged with its reason and the signals that drove it. Moves are bounded per week; a jump beyond the bound is flagged as a shock.
Probability range over time
Your estimate
Disagree with our range? Set yours. Estimates feed a community view that appears once enough people weigh in, and the weekly run reads the gap between our number and yours.
7%
2026-09-16
4–10%
seed
Seed estimate. This is the AI 2027 "race" ending. Astra's opaque recurrence (weaker chain-of-thought monitoring) is the single most relevant signal: it reduces the chance warning signs are seen at all.
OpenAI launches GPT-6 'Astra', first model at the Critical cyber threshold
Astra reached Critical under OpenAI's Preparedness Framework, found two zero-days in evaluation, and uses 'opaque recurrence' that weakens chain-of-thought monitoring. Brockman said he personally believes it is AGI.