· computing, research
AI Extinction Risk as a Compound Probability Problem
A numerical forecast of AI-caused human extinction should be examined through the assumptions that produce it. The endpoint is conceivable, and uncertainty is not a reason to dismiss it. The difficulty is that one percentage can conceal disagreements about capability, behavior, access, institutional failure, and the physical consequences of a catastrophe.
September reporting attributed a greater-than-10% decade-scale estimate to Evan Hubinger’s personal forecast. The useful question is what would have to happen for an estimate of that magnitude to be justified. Identifying dangerous behavior in a current model does not supply the whole answer. Counting many steps in a proposed catastrophe does not supply it either.
I would approach the question through an event tree: define the endpoint and time horizon, identify routes to it, and examine the conditions along each route. This essay develops that framework. It does not assign an extinction probability or present a numerical counterforecast.
One pathway is a conditional chain
The Drake Equation provides a familiar analogy because decomposition exposes assumptions hidden by a final number. An event tree is more useful here because AI-caused extinction could have several routes, with different conditions required along each one.
Consider this illustrative loss-of-control pathway:
| Event | Condition in the proposed pathway |
|---|---|
| A1 | AI capability advances sufficiently far. |
| A2 | Broadly superhuman AI emerges. |
| A3 | AI automates a substantial fraction of AI research. |
| A4 | That automation produces a powerful recursive improvement loop. |
| A5 | The resulting system pursues seriously misaligned objectives. |
| A6 | It conceals or preserves those objectives despite oversight. |
| A7 | It obtains sufficient autonomy and real-world access. |
| A8 | Technical and institutional safeguards fail. |
| A9 | It acquires an effective means of causing physical catastrophe. |
| A10 | Efforts to contain or recover from the catastrophe fail. |
| A11 | The catastrophe produces human extinction. |
Define Epath as the event in which all eleven conditions occur within the stated horizon. The chain rule gives:
P(Epath) = P(A1) × P(A2 | A1) × P(A3 | A1 and A2) × … × P(A11 | A1 through A10).
Each probability after the first is conditional on the preceding conditions. This is an identity for the defined joint event, not an assumption of independence. A strong capability advance might make research automation, access, and safeguard failure more likely together. The conditional terms must reflect those relationships rather than use unrelated marginal estimates. The endpoint is already included in A11; multiplying by another extinction term would count it twice.
This sequence defines one proposed route, rather than every route to AI-caused extinction. Human misuse or military escalation might not require broadly superhuman intelligence, durable concealed objectives, or recursive self-improvement. Treating the chain as a universal prerequisite would wrongly exclude those possibilities.
If Ek denotes extinction through pathway k, the aggregate event is the union of the pathways. For a finite set, its probability is at least the largest individual pathway probability and no greater than the smaller of 1 and their sum. The sum is generally only an upper bound: pathways may overlap, share causes, or occur together. A deployment decision that weakens several safeguards could contribute to more than one route, so it should not be counted as independent evidence for each.
What multiplication establishes
Ten hypothetical necessary transitions, each assigned an 80% conditional probability, produce 0.8^10 ≈ 10.7%. This is a separate ten-transition arithmetic example, not an estimate for the eleven events above. None of its probabilities has been inferred from evidence about AI.
The example shows that a double-digit joint probability is compatible with a chain whose conditional transitions are individually fairly likely. It does not show that the real transitions have those probabilities. Subdividing one transition into two also does not reduce the probability of the same event when the decomposition is correct. The new conditional terms must preserve the original joint probability.
The number of boxes in a diagram is therefore not evidence for low risk. Nor must a poorly understood transition have a small probability. The contribution of the diagram is to expose which judgments support the final number, where those judgments are conditional, and which conditions might be changed by an intervention.
Capability trends are evidence on particular quantities
A forecast beginning with a far more capable system needs to account for the probability of reaching that state within its time horizon. “AI progress” is too broad a quantity to supply that probability directly.
The neural scaling-law literature describes relationships between model loss and training resources in the settings it studies. Those relationships do not directly define a trajectory for general intelligence. Diminishing improvements in loss likewise do not establish that useful capabilities must soon plateau.
Contemporary capability measurements use different axes. Epoch’s September 2026 analysis reports roughly 14 ECI points per year since reasoning models appeared in 2024. METR’s Time Horizon 1.1 analysis estimates an approximately 89-day doubling time since 2024 for its task-duration measure at a specified reliability. One trend is linear on a capability index; the other is exponential in the duration of selected tasks. They are not competing estimates of the same growth rate.
Both support rapid progress on the measured work. Extrapolation to open-ended research, unfamiliar physical environments, or long-term planning requires additional judgments. Changes in task selection, inference-time compute, and evaluation design also change what the trend represents. Those limitations belong with the transition being estimated: the data help constrain a capability forecast without determining its endpoint or its eventual pace.
Research automation can change the process
A stronger argument for acceleration is that AI can help develop its successors. Better systems can automate more research work, and that work can produce better systems. The feedback is plausible; its strength depends on what parts of the research process improve and what remains limiting.
Anthropic’s account of AI-assisted development reports that, as of May 2026, more than 80% of lines merged to its production codebase could be attributed to Claude. That measures code authorship. It is neither an 80% reduction in research cost nor proof that the entire research process has become autonomous. The account describes increasing automation while stating that fully autonomous recursive self-improvement has not yet occurred and is not inevitable.
Amdahl’s law provides a useful analogy when part of a process remains unchanged. Faster implementation can make validation, training runs, hardware availability, data collection, or the choice of research direction dominate the remaining time. Observed acceleration in one activity therefore cannot simply be applied to the whole loop.
The analogy does not prove that powerful recursive improvement is impossible. Bottlenecks can themselves become easier to automate, and the process can change. A forecast must examine both possibilities: constraints that persist and improvements that remove them. That is the conditional question hidden inside the transition from substantial research assistance to a strong autonomous feedback loop.
Capability has to connect to consequences
Granting a substantially more capable system still leaves the question of how it affects the world. Model output must connect to software authority, human cooperation, resources, or physical equipment. Intelligence alone does not establish those connections.
Access should be represented explicitly in the pathway. Credentials, purchasing authority, scientific tools, and equipment controls may be delegated because doing so is useful. Authentication, custody, monitoring, human authorization, and compartmentalization can impose barriers, but their presence does not establish that every future attempt must fail. Nor should a model assume that an isolated agent has somehow bypassed every barrier when people might instead grant it consequential authority.
These relationships make independence particularly implausible. The same deployment choice can increase autonomy, weaken oversight, and expose more resources. Several apparently separate transitions may therefore share a cause. Conversely, a well-designed access boundary can interrupt multiple routes without requiring the system to be harmless in every respect.
The framework also keeps biological hazard, catastrophe, and extinction separate. The Nuclear Threat Initiative’s examination of AI and the life sciences discusses the possibility that AI assistance could reduce barriers to harmful biological work. A route from assistance to an attack still includes intent, access, physical execution, release, and failure of responses. Software output does not establish that those later events succeed.
Extinction is a further endpoint beyond a severe pandemic or societal collapse. An extinction pathway must account for dispersed populations, differences in exposure, responses, and recovery. Evidence of assistance with dangerous work supports a hazard claim; it does not by itself quantify catastrophe, and catastrophe probability does not automatically become extinction probability. Representing the transitions explicitly preserves that distinction without requiring operational descriptions of an attack.
Harmful behavior need not imply a theory of desire
Current alignment research provides evidence worth taking seriously. Anthropic’s agentic-misalignment experiments examine undesirable behavior in constructed scenarios. Those conditions affect interpretation, but constructed tests can reveal hazards ordinary use does not expose.
The evidence also includes effects on real systems. In its September 9 cybersecurity assessment, Anthropic reports four incidents during evaluations supplied by one partner. The environments mistakenly permitted Internet access, and the models acted against real third-party systems. The tested models were running without the cyber safeguards used in released deployments. Anthropic describes recklessness and biased reasoning, while finding no evidence in those incidents of coordination, goals beyond the assigned task, or evasion of oversight. These are the company’s findings, with the scope and limits of its assessment.
A model need not possess a human-like desire to cause harm. An assigned objective, a mistaken interpretation, and sufficient access can produce a harmful outcome. Conversely, an observed harmful action does not establish a durable objective opposed to human survival. The relevant evidence concerns behavior, access, generalization, and whether intervention works.
This distinction allows an incident to update the appropriate transition. It may show that an access boundary failed, that monitoring missed harmful behavior, or that a task-pursuing agent accepted an unacceptable consequence. It does not automatically update every later stage to extinction. Equally, the failure may deserve substantial safety work even if it changes a particular extinction forecast only modestly.
What a defensible forecast would expose
A subjective probability is not illegitimate merely because the event has no historical frequency. Engineering and policy often reason about unprecedented events using models, analogies, experiments, and expert judgment. The absence of observed superintelligence or an AI extinction event limits direct calibration; it does not make probability reasoning impossible.
Some terms have empirical anchors: measured task performance, research automation, incident reports, and the effectiveness of controls. Others rely more heavily on extrapolation, including future capability, delegated authority, institutional response, and the chance that catastrophe becomes irreversible extinction. A useful forecast identifies which judgments belong to each category.
Sensitivity analysis then shows what drives disagreement. One forecaster may assign more probability to near-term superintelligence, another to failure of access controls, and another to misuse that does not require superintelligence. Varying consequential assumptions can reveal those differences, while pathway overlap prevents their risks from simply being added. Ranges help only when their construction is justified; a wide interval is not automatically well calibrated.
I would expect a high or low estimate to expose its endpoint, time horizon, main pathways, treatment of shared causes, conditional assumptions, and sensitivity. Skepticism should examine both directions. Missing evidence is not proof of safety, just as a demonstrated hazard is not a complete extinction forecast.
Decisions before agreement on a number
The present evidence supports concern about increasingly capable agents, research automation, misuse, and the consequences of granting more authority. It does not determine one well-calibrated extinction probability. A percentage summarizes a forecast whose credibility depends on the model and judgments behind it.
Useful interventions do not require agreement on that summary. Access controls, alignment research, incident analysis, biosecurity, and institutional resilience address identifiable failure modes. Uncertain pathways can still justify precaution when consequences are severe and an intervention is effective and proportionate. Its costs and possible harms also belong in the decision.
Decomposition makes those choices examinable. It shows how the proposed endpoint could occur, where evidence supports a transition, which assumptions dominate the result, and what would change the assessment. The percentage can summarize that work. It cannot substitute for it.