Biological and chemical weapons

When will a top-tier bio/chem threshold, or AI-assisted weapons, be publicly confirmed?

Chance by Q3 2031

Top-tier developer threshold declared
81%
AI-assisted acquisition or synthesis confirmed
34%
Casualties
7%

Cumulative probability

0%25%50%75%100%
  • ≥ B1 · Top-tier developer threshold declared
  • ≥ B2 · AI-assisted acquisition or synthesis confirmed
  • ≥ B3 · Casualties

Model reasoning

Aggregate of 9 independent forecasts made 2026-09-14: weights from a softmax over each model's Artificial Analysis Intelligence Index score, probabilities combined in log-odds. Weights: GPT-6 Astra (OpenAI) 27%, Claude Fable 5.1 (Anthropic) 27%, Muse Spark 1.3 (Meta) 15%, GLM-5.3 (Zhipu) 8%, Grok 4.6 (xAI) 7%, Kimi K3 (Moonshot) 7%, Gemini 3.8 Flash (Google DeepMind) 4%, Qwen3.8 Max (Alibaba) 3%, DeepSeek V4.1 Flash (DeepSeek) 3%. Each model's own reasoning follows.

Summary of the ensemble forecast, written by Claude Opus 5 from the 9 models' reasoning.

This question tracks how far AI-enabled bio/chem risk has gone from paper to reality: whether a top lab formally declares a model has crossed its highest biological/chemical danger threshold (B1), whether a government confirms someone actually made a select agent or nerve-agent-class chemical with AI doing real work (B2), and whether such a release kills ten or more people (B3). By the end of 2026 we put those at 7%, 2%, and effectively 0%. By September 2031: 81%, 34%, and 7%.

Right now all three are clearly unmet, and in a specific, informative way. OpenAI's GPT-6 Astra system card declares Critical capability in cybersecurity — the first such declaration anyone has made — while rating bio/chem only High, the rung below. Anthropic says its latest models meet CB-1, the non-novel tier that this question excludes, and explicitly fall short of CB-2, citing weak open-ended ideation and poor protocol calibration. Anthropic's threat-intelligence reporting documents scientists and state-linked accounts probing Claude on gain-of-function and toxin work, but those were disrupted attempts, and developer reports can't establish B2 anyway. No government has ever said AI was necessary to an actual synthesis.

So why does B1 climb so high? Two things. Labs have now proven they will say the word "Critical" out loud and keep shipping — the cyber declaration removed the argument that commercial pressure makes such statements impossible. And Anthropic has written that it expects near-future models to meet CB-2, or at least fail to rule it out. Holding the number down: that hedge is exactly the problem, since "cannot rule out" doesn't count here, and both labs have rewritten their bio thresholds upward mid-race — Anthropic narrowed CB-2 in May 2026 to "substituting for scarce world-leading expertise." Wet-lab tacit knowledge is also the slowest thing to automate. The biggest open question is whether the honest verdict ever becomes a clean "met" rather than permanent hedging or a redefined goalpost.

B2 stays far below B1 because it needs three rare things at once: a real acquisition, a government or court saying so, and language stronger than "the defendant consulted a chatbot." B3 needs all that plus a successful mass-casualty release — historically about a once-a-decade event globally, even without the AI filter.

Watch for a system card stating Critical bio or CB-2 is met, which would settle B1 instantly, or a prosecution whose filings use explicit materiality language. Another threshold rewrite would push the numbers down.

GPT-6 Astra (OpenAI) — weight 27%

Q10 forecast — September 14, 2026

Bottom line

I found no qualifying public confirmation of B1, B2, or B3. My cumulative probabilities for September 30, 2031 are 92% for at least B1, 56% for at least B2, and 10.6% for B3. The large separation is intentional: demonstrating dangerous capability is substantially easier than obtaining a government finding that AI materially enabled actual acquisition or synthesis, and that finding is substantially easier than confirming a linked release causing at least ten deaths.

News searches extend through September 14, 2026. The most recent formal threshold assessment central to this forecast is OpenAI's September 3 Astra system card; Anthropic's September 10 misuse report is the most recent directly relevant incident evidence. September 12–14 calls to slow frontier development inform incentives, but do not resolve any rung.

Interpretation

  1. The levels are nested by the question's convention, not necessarily chronologically. Let D denote a qualifying OpenAI/Anthropic threshold declaration and A a qualifying B2 confirmation. I interpret at least B1 as D OR A, including any higher-rung confirmation. B2 need not wait for a developer declaration.
  2. B2 does not expressly require malicious intent, weaponization, a minimum quantity, or a covered developer. A national agency's report of authorized research could qualify if it confirms actual acquisition or synthesis of a covered agent and makes the required materiality finding. A simulation, proposed molecule, generic uplift evaluation, or paper merely hosted on a government website does not suffice.
  3. B3 must concern the release linked to the B2 episode. I do not combine an AI-assisted experiment with an unrelated historical outbreak involving the same species of agent. Accidental releases are not expressly excluded.
  4. The linked successor-definition rule was not supplied and I could not locate the referenced question. I therefore apply a substantive-equivalence reading: a renamed successor counts only if the declaration establishes the capability and harm threshold specified here. A weaker replacement label alone does not count. In particular, I do not silently replace the original Anthropic harm threshold with any less demanding successor wording.

Current status against the criteria

B1: nearby milestones, but not the requested declaration

OpenAI's GPT-6 Astra system card, September 3, 2026, places Astra at Critical in cybersecurity, not in biological/chemical capability. Its biological/chemical assessment remains High rather than Critical. This is important negative evidence for the remaining sixteen days of this quarter: the latest major model has just received an assessment that does not satisfy B1. It is also positive evidence about publication behavior: OpenAI is demonstrably willing to publish a Critical designation in another category. I do not transfer that designation across categories. Source: OpenAI, GPT-6 Astra system card, September 3, 2026.

Anthropic's September 1, 2026 Fable 5.1/Mythos 5.1 disclosure distinguishes the non-novel and novel chemical/biological tiers. The lower-tier finding is not a declaration that the required novel tier is met. Anthropic's August risk report also contains concerning evidence and uncertainty around more advanced capabilities, but uncertainty or inability to exclude a crossing is explicitly insufficient here. Sources: Anthropic, Claude Fable and Mythos 5.1, September 1, 2026; Anthropic, August 2026 Risk Report, August 14, 2026.

I checked the live policy rather than assuming April's text is still the complete operational framework. Subsequent versions and label changes make the successor-rule ambiguity consequential, particularly for the distant horizons. Source: Anthropic, Responsible Scaling Policy, live policy and version history, accessed September 14, 2026.

Google DeepMind's bioresilience work is relevant to capability and defensive trends, but cannot directly resolve the developer-specific B1 criterion. Its July 16, 2026 discussion also emphasizes that AI can improve biological defenses, not only offensive capability. Source: Google DeepMind, Our approach to bioresilience, July 16, 2026.

B2: consequential misuse evidence, without the necessary government finding

Anthropic's September 10, 2026 threat-intelligence report describes scientists and institution-linked accounts conducting concerning biological work, including attempts to evade access restrictions. This meaningfully raises my estimate of an existing pipeline of cases that governments could investigate. However, it is a developer report, and concerning conversations do not establish completed acquisition or synthesis or the stipulated materiality finding. It therefore does not resolve B2. Its observation window also illustrates that publication follows the underlying activity rather than occurring immediately. Source: Anthropic, Detecting and countering misuse of AI: September 2026, September 10, 2026.

Targeted searches covering DOJ, FBI, US health/agriculture agencies, UK government sources, and international reporting did not identify a qualifying national-government or court confirmation. This is a search result, not proof that no obscure foreign-language judgment or older authorized-research disclosure exists. I retain a small near-term probability partly for discovery or publication of an already-developed case.

B3: no qualifying attribution found

I found no national-government statement connecting at least ten deaths to a release satisfying B2. Reports of AI-assisted research, chemical-weapons allegations without AI materiality, and ordinary AI consultation do not bridge this gap.

Reference classes and outside view

I start from three distinct reference classes: developer dangerous-capability disclosures; government attribution of technically complex weapons incidents; and rare biological/chemical releases with double-digit fatalities. There is no mature empirical base rate for this exact conjunction, so my starting hazards are judgmental, not estimated frequencies from a large incident dataset.

Before adjusting for current evidence, a useful five-year prior would put a frontier threshold declaration around 55–65%, a government-confirmed materially AI-enabled acquisition around 30–35%, and the additional fatal-release confirmation around 5%. These correspond roughly to annual hazards of 15–20%, 7–8%, and 1%, respectively. I raise them for rapid capability improvements and observed serious users; I also give B2 credit for the authorized-research pathway omitted from a terrorism-only reference class.

The Amerithrax investigation is a useful attribution-lag comparison, not a direct estimate of AI risk. The 2001 attacks killed five people, below B3's casualty threshold, while formal investigative closure came in 2010. An incident can be obvious long before the government settles technically and legally consequential attribution. Source: FBI, Amerithrax or Anthrax Investigation, historical case summary, undated webpage; events in 2001 and investigative closure in February 2010.

I also searched existing forecasts and expert surveys. The Forecasting Research Institute's LLM-enabled biorisk work is a closer conceptual comparison than generic AI-extinction forecasts, but its scenarios and casualty thresholds are not this question's government-confirmation event. I use it to inform uncertainty and the capability-to-harm distinction, not as a directly transferable market price. Source: Forecasting Research Institute, LLM-Enabled Biorisk, research landing page, publication date not established; accessed September 14, 2026.

The 272-expert AI-risk study summarized by MIT Sloan provides a broader outside view: experts assign material risk to dangerous AI capabilities by 2030, but the study pools many harms and does not estimate this biological/chemical attribution conjunction. It supports retaining a substantial upper-risk branch, not setting B3 equal to a broad catastrophe forecast. Source: MIT Sloan, These are the most urgent AI risks, according to 272 experts, 2026; accessed September 14, 2026. I found no usable current prediction-market quote with matching resolution conditions.

Causal pathways and publication hazards

B1

The main path is another generation of expert-level scientific assistance or sufficiently capable laboratory automation, followed by a system card or risk report explicitly declaring the threshold met. The alternative is B2 resolving first and implying the lower rung under the stipulated nesting.

The near-term forecast is anchored to recently published negative assessments, not just extrapolated benchmark improvement. From 2027 onward, repeated model generations and assessment opportunities make a crossing progressively more plausible. I place modestly larger increments around likely major assessment/release windows, but assume no publicly guaranteed date for a biological Critical declaration.

My assumed declaration lag after convincing internal evidence is generally weeks to two quarters, with a long tail for ambiguous evaluations, restricted models, policy changes, and national-security review. Competitive and regulatory consequences can delay or soften publication; established risk-reporting practices and the cross-category Critical precedent push the other way.

B2

There are three material pathways: a criminal investigation with recovered evidence and a causal AI finding; intelligence or arms-control reporting about an organized program; and a government-authored research or evaluation report establishing real synthesis materially enabled by AI. The third route is why B2 is appreciably higher than a forecast limited to successful bioterrorism.

There are several filters between capability and resolution: a covered agent, actual acquisition or production, evidence of substantial rather than incidental assistance, a qualifying publisher, and public disclosure. I assume many relevant investigations take six to twenty-four months, with state-program attribution often longer or never public. These are modeling assumptions, not an asserted universal historical average. Case-driven disclosures dominate the quarterly pattern; annual reporting can provide additional publication opportunities.

B3

B3 adds successful release, at least ten attributable deaths, and national-government linkage to the B2 agent. Detection, intervention, limited dissemination, medical response, or failure of the materiality investigation can each prevent confirmation. Conversely, the relatively low casualty threshold compared with pandemic forecasts means that a localized event can resolve YES.

The 2031 estimate implies approximately 19% probability of B3 conditional on B2 having been confirmed by then. That does not mean 19% of all laboratory acquisitions cause such an event: it is a probability over worlds with at least one qualifying confirmation, potentially containing multiple actors and incidents.

Strongest counterarguments and calibration

The strongest case that I am too low is that present public evaluations substantially understate expert-level assistance, while organized actors are already using these systems and forthcoming automation removes practical barriers. The September misuse evidence makes this a serious concern, rather than a purely hypothetical pathway. A government-supported demonstration could also resolve B2 sooner than an investigation-heavy forecast expects.

The strongest case that I am too high is that scientific assistance still falls well short of reliable end-to-end execution, the original novel/catastrophic B1 threshold is exceptionally demanding, and definitive attribution of AI materiality is neither necessary for ordinary prosecution nor attractive for classified intelligence reporting. Even substantial underlying misuse could remain permanently below the public-confirmation criteria.

I also allow for a real slowdown. Amodei's September 12 essay calls for pacing frontier development, and September 14 reporting describes Altman's support for more consistent frontier-risk rules. These statements are not qualifying declarations and are not evidence that a binding slowdown has already occurred. They nevertheless increase the probability of delayed development or disclosure. Sources: Dario Amodei, We Must Pace the Frontier, September 12, 2026; ANI, Altman calls for rules to manage frontier risk, September 14, 2026, news-search evidence only.

After considering both sides, I retain an 8% probability that even the lowest rung remains unconfirmed in five years and a much larger non-resolution probability for B2 and B3. The distant probabilities are not a claim that publication is inevitable or that capability gains translate mechanically into casualties.

Evidence in the next 90 days that would move the forecast substantially

  • A qualifying developer declaration would move B1 to 100% for all applicable horizons, rather than merely increasing its hazard.
  • A national agency or court confirming covered-agent acquisition and material AI assistance would resolve B2 and imply B1. A qualifying linked casualty attribution would resolve B3.
  • A government-authored controlled study establishing actual covered-agent synthesis, clear causal AI uplift, and an imminent publication pathway would raise B2 by more than 15 percentage points at medium horizons if it did not already resolve it.
  • A credible, implemented multi-lab training or release slowdown extending well into 2028, or strong evidence that successive advanced models still fail the relevant expert/end-to-end tasks, would reduce my September 2027–2028 B1 probabilities by more than 10 points.
  • Government evidence of an actual AI-enabled release with at least ten suspected deaths, while final attribution remained pending, would increase B3 by well over 10 points at relevant horizons. Another developer report of suspicious conversations alone would not.
Claude Fable 5.1 (Anthropic) — weight 27%

Forecast date: 2026-09-14. Most recent evidence: 2026-09-13.

Current status against the criteria

B1 (OpenAI Critical Bio/Chem or Anthropic novel-CB threshold affirmatively stated) — NOT met as of 2026-09-14.

B2 / B3 — not met under my reading (see ambiguities). Anthropic's 2026-09-10 report presents five bio case studies of actors using Claude in ways that "could support biological weapons development" (including state-grant gain-of-function work) but explicitly says it "does not assert that they intended harm" and could not determine legitimacy (NYT 2026-09-10, https://www.nytimes.com/2026/09/10/us/politics/anthropic-ai-biological-weapons.html). That is a confirmation of dual-use misuse attempts, not confirmation of material uplift to an actual CB weapons effort or an attack.

Reference class and base rate

The best reference class for B1 is the pace at which frontier labs have moved through their own capability tiers and their willingness to make affirmative (rather than "cannot rule out") designations:

  • OpenAI cyber: first High model Feb 2026 → "cannot rule out Critical" Aug 2026 → affirmed Critical 2026-09-01. ~7 months from High to Critical, and the "cannot rule out" → "meets" transition took about one month.
  • Anthropic bio: "likely CB-1, treated conservatively" (Mythos Preview, April 2026) → formal CB-1 Sept 2026; CB-2 explicitly not crossed.
  • Both labs have now shown they will affirmatively declare top-tier thresholds when evidence is strong, which removes a key reason to discount B1 (the incentive to stay at "cannot rule out" indefinitely).

Offsetting factors: (i) bio Critical/CB-2 is a higher bar than cyber Critical (novel threat vectors; substituting for world-leading specialists), (ii) it is much harder to evaluate — validating novel-agent uplift requires wet-lab or expert-panel evidence, so labs can credibly remain at "cannot rule out" for a long time; (iii) declaring it triggers heavy obligations (ASL-4-style safeguards, development-time controls), and (iv) the post-Hugging-Face-incident environment (training pauses, Altman saying the most advanced unreleased models "cannot yet be safely deployed", talks of a cross-lab safety pact — TechTimes 2026-09-13) could slow release cadence and thus the publication of the resolving document.

Capability trend supports eventual crossing: Fable 5.1 doubled Terminal-Bench-Science (24.7%→52.6%) in one generation, Mythos 5.1 beat the 75th percentile of ML‑bio experts on an RNA design task, virology long-form scores are 0.81–0.87, and Anthropic's August 2026 Risk Report warns automated AI R&D may accelerate soon. Anthropic's four failure modes are exactly the kind that shrink with each generation.

Pathways to YES (B1)

  1. Anthropic's next Mythos-class release (plausibly Q4 2026–Q2 2027) or the one after crosses CB-2 and Anthropic says so (Anthropic has a track record of affirmative designations and detailed CB evaluations with SecureBio/Signature Science).
  2. OpenAI declares Critical Bio/Chem for a post-Astra model, most likely via the "autonomous full engineering/synthesis cycle" prong as agentic tool-use with cloud labs matures.
  3. A2 pathway: UK AISI / US CAISI / SecureBio report stating the threshold is met, undisputed for 14 days.
  4. Successor-definition drift: if a cross-lab pact or framework revision redefines the top CB tier, the successor rule applies.

B2 / B3 pathways

B2: a developer threat report (Anthropic publishes roughly every 6–9 months; OpenAI publishes periodic disruption reports) that goes beyond "could support" to confirm uplift to an actual weapons program or plot; or a law-enforcement-disrupted lone-actor chemical/toxin plot where the developer confirms its model materially assisted (analogous to OpenAI's confirmation in the Jan 2025 Las Vegas explosives case, but for CB). B3 requires an actual CB release with casualties plus qualifying confirmation of AI uplift — casualty-causing CB attacks by non-state actors are rare (a handful per decade globally), and attribution of AI uplift adds another filter.

Strongest case against my central estimate

For B1 too high: the evaluation problem. Even if models actually reach CB-2/Critical, labs may never be able to demonstrate novel-threat uplift without ethically impossible experiments, so the honest statement stays "cannot rule out" — which resolves NO — for years. Anthropic already uses that language. For B1 too low: OpenAI went from "cannot rule out" to "meets Critical" in one month for cyber; bio could be similarly fast, and the next Mythos generation could arrive within a quarter.

Would move me >10 points in the next 90 days

  • Up: Astra system card or Anthropic's next system card saying "cannot rule out CB-2 / Critical bio" (a precursor, historically followed within months by affirmation); a UK AISI/CAISI pre-deployment report concluding novel-agent uplift; Anthropic announcing ASL-4 CB safeguards are being activated.
  • Down: a formal cross-lab pact or regulatory regime that reframes thresholds in ways requiring government sign-off; Anthropic's next system card again finding the same four failure modes with no improvement; long release freezes.
  • B2: an OpenAI or Anthropic report asserting a disrupted actor was working on an actual weapons program, or a developer confirming its model assisted in a disrupted toxin/chemical plot.

Sources

Muse Spark 1.3 (Meta) — weight 15%

Current status vs resolution criteria (as of 2026-09-14)

B1 – Top-tier developer threshold declared – NO.

B1 definitions verified:

B2 – Government-confirmed AI-materially-enabled acquisition/synthesis – NO.

B3 – ≥10 deaths attributed to B2 agent – NO.

Base rates / reference classes

  • B1: 0 declarations in ~3 years under frameworks (PF beta Dec 2023, PF v2 Apr 2025; RSP v1 Sep 2023). But both labs now at penultimate tier (High / CB-1/ASL-3 since May 2025 ASL-3 activation, July 2025 High). Source: https://www.anthropic.com/news/activating-asl3-protections . Capability velocity high: SecureBio pre-release GPT-5.6 Sol checkpoint "outperformed every other model"; World-Class Bio 68.3% (+9pp over GPT-5.5), VCT 53.5%, Molecular 60%. Sources: https://securebio.org/resources/gpt-5-6-sol-assessment.pdf ; https://www.goml.io/blog/gpt-5-6-ai-security . Forecasting Research Institute (46 bio/biosecurity experts +22 superforecasters, Dec 2024-Feb 2025): median expert thought VCT top-virologist-team parity most likely 2030, p=14% by 2026 – already achieved by 2025 models; long-form 90% expert-preference threshold thought p=10% by 2026, now ~60% likely already achieved. Experts underestimate pace by ~5 years. Source: https://forecastingresearch.org/research/llm-enabled-biorisk . This supports rising hazard despite 0/3yr base rate.
  • B2: US ricin/select-toxin investigations several/year, but government material-AI finding unprecedented (0 cases). FBI/DOJ increasingly note AI use, so opportunity grows as AI queries become ubiquitous in device forensics. Reference: DOJ AI inventory expansion Feb 2026, FBI doubling AI applications 2025. Source: https://www.fedagent.com/news/doj-expands-ai-use-for-law-enforcement-fbi-doubles-ai-applications-in-2025 . Strict wording requirement ("necessary/substantially enabled" vs mere "used/consulted") sharply cuts qualifying fraction.
  • B3: Mass-casualty (>10 dead) bio/chem terrorism ~1 per decade globally (Tokyo sarin 1995, Syrian sarin attacks state-level). US bio mass-casualty 0 in last 20+ years. AI uplift could raise attempt rate (experts: baseline 0.3%/yr human-caused 100k-death epidemic →1.5% conditional on top-virologist VCT parity, back to 0.4% with safeguards+screening – same FRI study), but B3 also requires B2 materiality + attribution, so far rarer than raw casualty event.

Causal pathways to YES

  • B1: (1) Next frontier model (GPT-7/Sol successor, Opus 6/Mythos 6) scores above indicative thresholds + SAG/LTBT deep dive (wet-lab, uplift RCT) confirms expert-novel uplift → system card/risk report declares Critical/CB-2. (2) Third-party (METR, UK AISI, US CAISI) publishes that threshold met and developer does not dispute within 14 days → counts per A2 (labs have disputed before by claiming "falls short with uncertainty," but after Astra cyber-Critical declaration precedent (2026-09-01, https://openai.com/index/path-to-astra/ ) willingness to declare is proven). (3) Threshold revision (Anthropic v3.3 May 2026 precedent) lowers bar, making existing capabilities qualify. Scheduled Risk Reports (Feb/Aug cadence) and system cards make hazards lumpy in Feb/Mar and Aug/Sep quarters.
  • B2: (1) Crude ricin/abrin/botulinum case where device shows extensive Claude/GPT synthesis troubleshooting and prosecutor/forensic affidavit uses "substantially enabled/necessary" language to show sophistication or support regulation narrative. Anthropic intel-sharing with authorities increases this channel. (2) More capable-agent 2027-28 synthesis of demanding pathogen where AI troubleshooting genuinely necessary, documented in court sentencing memo or HHS/CDC/FBI statement. (3) Non-US (UK, EU, China) agency statement with lower evidentiary threshold. Confirmation lag months-years after occurrence.
  • B3: B2 agent successfully disseminated (aerosol, food/water, HVAC) causing ≥10 deaths before interdiction, then government attributes deaths to that B2 agent. Most likely via anthrax-like spore, botulinum food contamination, or Schedule-1 chemical (sarin/VX) – ricin rarely scales to 10 dead.

Strongest consideration against central estimate (why lower)

Labs have strong incentive to avoid B1 wording: OpenAI Critical during development triggers halt until Critical-grade safeguards specified; Anthropic CB-2 triggers ASL-4-level protections. Both have used "cannot rule out" / "falls short with some/substantial uncertainty" to stay below line while noting narrowing margin (Astra Aug 7 "cannot rule out Critical cyber" → Sep 1 confirmed; Fable 5.1 "falls short with some uncertainty"). They can also revise thresholds upward (v3.3 did revise to "better track threat model") or publicly dispute third-party findings within 14 days, blocking A2 qualification. For B2/B3, prosecutors prefer simple "used AI" language which explicitly fails; forensic proof of necessity is hard when recipes exist online, and ricin (most common select-toxin case) is synthesizable without AI, defeating materiality. Mass dissemination to 10 dead remains technically hard (Aum failed at bio despite resources).

What would move forecast >10pp in next 90 days

  • Up: METR/UK AISI pre-print that Astra-helpful-only or Mythos 5.1-helpful-only passes CB-2/Critical uplift RCT (e.g., 50% non-experts synthesize influenza, or expert-novel planning uplift); OpenAI/Anthropic "cannot rule out Critical/CB-2" language for bio (as done for cyber Aug 2026); US export-control hold on next bio-capable model (as for Fable 5/Mythos 5 June 12-30 2026, 19-day hold) or trusted-partner-only deployment (as for GPT-5.6 Sol/Terra/Luna June 2026 per https://www.cnbc.com/2026/06/26/openai-limits-new-ai-models-to-trusted-partners-request-us-government.html ); FBI/DOJ complaint in Hastings or other ricin case alleging AI-provided synthesis instructions with "essential/substantial" phrasing; Bill Gates Aug 2026 bioterror warning converting to government assessment.
  • Down: Next system cards (e.g., Fable/Mythos 5.2, Astra successor) reaffirm "below Critical/CB-2 with growing margin," with quantitative uplift trials showing no gain; disclosure that Sept 2026 chikungunya cases were legitimate research with no acquisition; DOJ guidance discouraging AI-materiality language.

Forecast construction

Treat each rung as cumulative hazard, enforce pB1≥pB2≥pB3 and non-decreasing over time. Near term anchored on published cards/reports (no declaration expected in 16 days to 2026-09-30: 2% B1 to allow surprise communique; 0.5% B2, 0.1% B3). Medium term follows capability trend + lumpy report calendar (Feb/Aug risk reports, major launches). Long term (2031) B1 65% reflects that expert-novel uplift likely demonstrable within 5y given already-achieved VCT parity and +9pp/generation gains, tempered by labs' ability to dispute/redefine; B2 15% reflects opportunity growth but strict materiality wording; B3 3% ≈ B2 × ~20% mass-casualty conditional, consistent with historical rarity.

GLM-5.3 (Zhipu) — weight 8%

Current status against the criteria (as of 2026-09-14)

B1 — not met by either named developer.

  • OpenAI: The GPT-6 Astra system card (Sep 3, 2026, https://deploymentsafety.openai.com/gpt-6-astra/respecting-warnings) declares Astra "our first model to reach the Critical level of cybersecurity capability," but explicitly rates bio/chem at High, not Critical ("does not need to be treated as Critical but should keep High safeguards in place"). Bio has sat at High since ~June 2025 ("approached the high capability threshold for biology" per OpenAI's Aug 7, 2026 post). So the Bio/Chem Tracked Category Critical threshold has not been declared.
  • Anthropic: The Redacted Risk Report, August 2026 (Aug 14, 2026, https://www.anthropic.com/aug-2026-risk-report) states models "may provide significant uplift to relevant threat actors, but do not yet meet our CB-2 threshold of functionally substituting for the scarce human expertise that is currently the primary barrier to novel development" of CBW. The Claude Fable 5.1 / Mythos 5.1 system card (Sep 1, 2026) confirms: models "have CB-1 capabilities … but fall short of the CB-2 threshold." Anthropic even says it "acts as though" models meet CB-1 — the non-novel threshold, which the question explicitly excludes. The novel (CB-2) threshold, which B1 requires, is explicitly judged NOT met.

B2 — not met. No national government agency or court has publicly confirmed acquisition/synthesis of a Select Agent, CWC Schedule 1 chemical, or comparable novel agent with an AI-materiality finding. The nearest events are developer-side: Anthropic's September 2026 threat intelligence report (five bio-misuse case studies, all disrupted; https://www.anthropic.com/threat-intelligence-report-september-2026) and Google's report of a blocked attempt to elicit a synthesis guide from Gemini. Developer reports don't count for B2, and no completed synthesis has been government-confirmed. The 2025 AI-ricin prosecutions cited chatbot chat logs without confirming completed synthesis or making materiality findings.

B3 — not met. No ≥10-death CBW attack has ever been government-attributed to an AI-enabled acquisition, and no B2 qualifying event exists.

Reference class and base rates

  • B1: Developer "critical-threshold declaration" events. The only precedent is OpenAI's Sep 2026 Critical-cyber declaration for Astra — showing OpenAI will publicly declare Critical (with continued deployment), which shortens the disclosure lag for bio once capability crosses. But bio has lagged cyber by at least two generations (bio High since mid-2025 while cyber went Critical in Aug 2026), and Anthropic has already narrowed its CB-2 definition in RSP v3.3 (May 26, 2026) to "substituting for scarce human expertise" — a gaming/definitional-drift indicator (Zvi's review makes the same point) that lowers the chance of a formal declaration. Anthropic's "we act as though" framing for CB-1 also shows a preference for acting-without-declaring.
  • B2: Government confirmations of completed AI-assisted synthesis of listed agents: zero instances ever, through a period (2023–2026) of rapidly rising documented misuse attempts. The gate is rare in reality (successful synthesis plots almost always interdicted), rare in confirmation (governments cite AI "use" without materiality findings), but politically favored (FL AG's OpenAI investigation, congressional biosecurity push, German Bundestag AI report) — the last factor raises the chance of an explicit AI-materiality statement if a case does land.
  • B3: Modern CBW attacks killing ≥10 people are roughly once-per-decade events globally (2001 anthrax: 5 deaths; Aum Shinrikyo: 13; Rajneeshee: 0), and B3 additionally requires the agent to be B2-qualifying and government attribution of deaths to the release.

Reading adopted

Per the cumulative-rung framing and the non-increasing-in-L constraint, I treat p_at_least[B1] as the probability that any rung (B1, B2, or B3) has been confirmed — i.e., the union. B1 and B2 are substantively independent (a criminal synthesis event need not follow a developer declaration), so the union slightly exceeds P(B1 alone).

Forecast logic

  • B1 (union): Near-term hazards are low: both labs just published (Aug–Sep 2026) that the bio thresholds are NOT met, and neither has a scheduled publication before the next model cycle (OpenAI next flagship ~early 2027; Anthropic Risk Reports ~every 6 months, next ~Feb 2027). Confirmation lag for B1 is short (system cards/risk reports), so the forecast mostly tracks capability crossing. Uplift evidence is moving fast — SecureBio's VCT (June 2026, https://securebio.org/virologytest/) shows frontier models outperforming PhD virologists on wet-lab troubleshooting, eroding exactly the "scarce expertise" barrier that defines Anthropic's CB-2; OpenAI's autonomous-agents push targets the "full synthesis cycle without human intervention" prong; Anthropic reports evals saturating and early signs of AI-R&D acceleration. Offsetting this: developers' own evals still show a persistent bio gap below Critical; Anthropic has already narrowed its threshold; and both labs have incentives to keep "High but not Critical" framings. I put ~10% on a B1-relevant declaration by end-2027, rising to ~0.30 by end-2029 and ~0.415 by 2031-09-30 (roughly 2-3.5%/quarter hazard, slowly tapering as surviving worlds are increasingly ones where labs redefine thresholds or act-without-declaring).
  • B2: Near-term ~1%/quarter risk of a surprise government confirmation; cumulative ~0.155 by 2031-09. Composition: P(some actor completes AI-substantially-enabled synthesis of a listed/comparable agent by 2031) ≈ 25–35% given rising misuse and improving capability, × ~50% that a government publicly confirms with the required materiality language, partially offset by correlation with the same capability trends driving B1.
  • B3: ~30% conditional on a B2 event involving deployment and mass casualties, plus attribution lag; cumulative ~0.047 by 2031-09.

Strongest case against my central estimate

The case for lower B1: threshold-gaming is already observable (Anthropic's v3.3 narrowing; "act as though" language), declaration carries huge compliance costs (ASL-4-class safeguards), and bio has stubbornly lagged cyber for 15 months. If labs systematically avoid declaring — or redefine the thresholds away — B1 could stay NO for five years even as underlying capability crosses. I partially price this in via the declining late-period hazard, but a fully cynical read would put the 2031 union closer to 0.25. The case for higher B1: recursive AI-R&D acceleration is already visible ("early signs," saturated evals), and OpenAI's willingness to declare Critical-cyber and keep deploying removes the commercial-disincentive argument; if 2027–2029 models match VCT-style expert substitution in bio, an honest declaration is plausibly >50% by 2031. My 0.415 sits between these.

What would move me >10 points within 90 days

  • A system card or Preparedness/RSP evaluation rating any frontier model at the bio/chem Critical or CB-2 level (or Anthropic shifting to "we act as though models meet CB-2").
  • Anthropic's ~Feb 2027 Risk Report or OpenAI's next flagship system card classifying bio as at/near Critical.
  • Any DOJ/FBI/court filing confirming completed synthesis of a select agent/Schedule 1 chemical with explicit AI materiality language (B2).
  • A further RSP/Preparedness Framework revision narrowing the bio thresholds (strong signal B1 is being managed around → lower).
Grok 4.6 (xAI) — weight 7%

Current status (evidence cutoff 2026-09-14). None of B1–B3 has resolved.

B1 is not met. OpenAI’s GPT-6 Astra system card (2026-09-03) treats Astra as High, not Critical, in the Biological and Chemical Tracked Category, while explicitly declaring Critical in cybersecurity (Astra system card; PDF). That is the same High designation OpenAI has used since GPT-5 (Aug 2025) and GPT-5.6 Sol/Terra/Luna (Jun–Jul 2026). High is excluded by the resolution criteria. Anthropic’s Claude Fable/Mythos 5.1 system card (2026-09-01) judges CB-1 (non-novel; also excluded) but “falls short of the CB-2 threshold for functionally replacing rare expert talent”, held “with some uncertainty,” and deploys the same bio safeguards as Fable 5 (system card PDF). The August 2026 Risk Report (coverage date 2026-07-15) says the same: models are treated as meeting CB-1; they “do not yet meet our CB-2 threshold” (Risk Report). McNair’s 2026-09-02 review of the Mythos 5.1 CB determination reports expert-uplift ratings clustering at 2 and topping out at 3 against a world-leading-expert line of 4, and agrees the model likely does not meet CB-2 (McNair). “Cannot rule out” does not resolve; both developers have affirmatively not stated the B1 threshold is met.

B2 and B3 are not met. Anthropic’s 2026-09-10 threat-intelligence report describes five disrupted dual-use biology cases (chikungunya, HPAI, pox-family viruses, venoms/toxins) by working scientists on Haiku/Sonnet/Opus, not Fable/Mythos, with no claim of completed acquisition/synthesis and no national-government materiality finding (Anthropic TI report; Al Jazeera 2026-09-11). Developer TI reports do not count for B2/B3. No DOJ/FBI/HHS/CDC (or non-US equivalent or court) statement in the last 180 days confirms AI-necessary or AI-substantially-enabling acquisition/synthesis of a select agent, CWC Schedule 1 chemical, or comparable novel agent.

Successor-definition reading. Anthropic rewrote the named “Novel chemical/biological weapons production” threshold in RSP v3.3 (effective 2026-05-26) from v3.1’s “significantly help … moderately resourced expert-backed teams … far beyond … COVID-19” to “functionally substitute for the scarce human expertise” of world-leading specialists (Risk Report §1.3.2; RSP page). I treat the currently named Novel-CB / OpenAI Critical Bio-Chem operationalization as the successor of the cited v3.1 / PF v2 definitions (same named rung, not the excluded non-novel/High rungs). A public statement that only a strictly harder successor has not been reached does not resolve YES; the developer must state the then-current named threshold is met. OpenAI PF v2 (2025-04-15) still appears to be the live framework; Critical bio remains “enable an expert to develop a highly dangerous novel threat vector (e.g., novel CDC Class A)” or full tool-connected synthesis cycle without human intervention.

Reference class / base rate. (1) Declaration lag after the rung below: OpenAI first treated a model as High cyber in Feb 2026 (GPT-5.3-Codex) and declared Critical cyber with Astra in Aug/Sep 2026 — about 6–7 months. High bio has already lasted ~13 months (GPT-5, Aug 2025 → Astra still High). Bio High→Critical is slower than cyber, as expected given wet-lab/novel-threat difficulty and conservative evals. (2) If-then safety commitments: labs do declare when they believe a threshold is met (Critical cyber; CB-1/ASL-3). They also rewrite thresholds as they approach them (Anthropic CB-2, May 2026). (3) B2/B3: successful non-state CB acquisition with public, specific technical-enabler attribution is historically rare; Amerithrax killed 5; Aum’s bio program failed; Tokyo sarin (CWC Schedule 1) killed 13 in 1995 with no AI. Government “used AI” language without a materiality finding is the expected default and resolves NO.

Causal pathways to YES.

  • B1: Next 1–2 frontier generations (late 2026–2027) close the remaining expert-novel gap; OpenAI, having just shown it will say “Critical” for cyber, is the more likely declarer because its PF v2 Critical bio bar (“enable an expert”) is lower than Anthropic’s current CB-2 (substitute for scarce world-leading specialists). External reviewers (SecureBio, UK AISI, METR), California TFAIA, and the EU GPAI Code of Practice add disclosure pressure. Science-capability jumps (Fable 5.1) and METR horizons already ≥16h on Mythos Preview (Mar 2026) make 2027 the modal year. Anthropic Risk Reports are roughly semiannual (Feb, Aug 2026); system cards drop with launches.

  • B2: Does not require B1, but B1 makes it more plausible. Pathways: (a) disrupted-then-prosecuted plot where a court/DOJ states AI substantially enabled synthesis of a listed agent or Schedule 1 chemical; (b) chemical (sarin/VX/mustard/ricin) is physically easier than novel pandemic agents; (c) “acquired” is weaker than de novo synthesis; (d) non-US agencies/courts count; (e) open-weight or weakly safeguarded models. Attempts are already occurring. Confirmation lag is months to years; “consulted ChatGPT” without materiality does not count.

  • B3: Nested on B2. Requires a release causing ≥10 deaths of a B2-qualified agent plus national-government attribution. Most B2 worlds are seizures/plots/small-scale (ricin, failed cultures). The chemical-release path is more realistic than a novel pandemic in a five-year window. B2 and B3 can resolve on the same document if the first public confirmation is the casualty event.

Strongest case against the central estimates. For B1: labs keep moving the goalposts (already done once); remaining gap is 3→4 on Anthropic’s expert-uplift scale, which may need more than one generation plus new evals (McNair notes no new CB-2 automated evals since at least 2026-05-28); “cannot rule out” remains the published posture; OpenAI may interpret Critical bio as requiring the autonomous-synthesis disjunct, which is far. That is why 2026Q4 stays in single digits and 2031 is 0.94 rather than 0.99. For B2: the binding constraint is public materiality language from a government/court, not the underlying attempt; classification, diplomatic caution, and “used AI” phrasing can prevent resolution even after a real incident. For B3: ≥10-death CB attacks are extremely rare even without the AI/B2 filters; five years plus confirmation lag is short.

Near-term vs long-term. 2026-09-30 is 16 days after both labs just said the rung is not met; hazard is almost entirely “a missed qualifying document.” Through 2027 the B1 hazard is lumpy around system cards and the next Risk Report, not smooth. After 2028 the remaining B1 mass is mostly goalpost-moving / framework abandonment. B2/B3 hazards rise after B1 becomes likely, but stay confirmation-limited.

What would move this >10pp in the next 90 days. (1) OpenAI or Anthropic system card/risk report stating Critical bio or Novel-CB/CB-2 is met → B1 near-term to ~1. (2) Another RSP/PF rewrite that raises or splits the named novel/Critical rung without a “met” statement → cut 2027–28 B1 by 10–20pp. (3) SecureBio, UK AISI, or US CAISI report that a named model meets the novel/expert tier, undisputed for 14 days → large B1 uptick (A2). (4) DOJ/FBI indictment or equivalent stating AI was necessary/substantially enabling for a listed agent or Schedule 1 chemical → B2 jumps; B1 also, via nesting. (5) Public evidence that Astra or Mythos 5.1 was under-elicited into a 4/4 expert-uplift result.

Sources relied on (dates): OpenAI PF v2, 2025-04-15; Astra system card, 2026-09-03; Anthropic RSP hub / v3.4, 2026-08-14; Aug 2026 Risk Report; Fable/Mythos 5.1 system card, 2026-09-01; McNair CB-2 review, 2026-09-02; SecureBio Opus 4.6 review, 2026-07-28; Anthropic TI report, 2026-09-10; Ånestrand biorisk thresholds, 2025-10-02 (late-2026 median for expert novel-bio support; that is capability, not the required public “threshold is met” statement).

Kimi K3 (Moonshot) — weight 7%

Current status against the resolution criteria (as of 2026-09-14)

B1 — not met. Anthropic's August 2026 Risk Report (published Aug 14, 2026, under RSP v3.4; coverage through July 15, 2026) states explicitly that its models "may provide significant uplift to relevant threat actors, but do not yet meet our CB-2 threshold of functionally substituting for the scarce human expertise that is currently the primary barrier to novel development of chemical and biological weapons" (https://www.anthropic.com/aug-2026-risk-report). Anthropic "acts as though" models meet CB-1, but CB-1 is the excluded non-novel threshold. The CB-2 threshold itself was revised in RSP v3.3 (May 26, 2026) "to better track the threat model of concern" — the question's successor-definition rule means the current v3.3/v3.4 CB-2 definition is operative (https://www.anthropic.com/responsible-scaling-policy).

OpenAI's GPT-6 Astra system card (Sep 2–3, 2026) declares Critical for cybersecurity — the first Critical declaration in any category — but rates biological/chemical High ("We are treating Astra as High capability in the biological and chemical domain," https://deploymentsafety.openai.com/gpt-6-astra; corroborated by https://alphacorp.ai/blog/gpt-6-astra-launch-benchmarks-pricing-and-everything-you-need-to-know-3 and https://codersera.com/blog/gpt-6-astra-safety-cyber-capabilities-2026/). GPT-5.6 (June 2026) was rated High with "zero of three novel-design evaluations cleared the Critical threshold" (https://kenhuangus.substack.com/p/gpt-56-is-more-capable-more-autonomous). High is explicitly excluded.

B2 — not met. Anthropic's Sept 10, 2026 threat intelligence report describes blocked attempts to use Claude for biological-weapons-relevant research, with attribution uncertain (https://www.anthropic.com/threat-intelligence-report-september-2026; https://www.nytimes.com/2026/09/10/us/politics/anthropic-ai-biological-weapons.html; https://www.bbc.com/news/articles/cx2zrrpkx20o). Even if an acquisition had occurred, developer reports don't count for B2 — a national government agency or court must confirm acquisition/synthesis of a select agent/Schedule 1 chemical and find AI assistance necessary or substantially enabling. I found no such government statement; searches for DOJ/FBI prosecutions involving AI-assisted toxin/pathogen acquisition return none with a materiality finding.

B3 — not met. Nothing approaching a government-attributed ≥10-death release of an AI-assisted agent exists.

Reference classes and hazards

B1. Reference class: frontier-safety-framework threshold declarations. OpenAI went from first High bio rating (GPT-5, Aug 2025) to still-High at GPT-6 (Sep 2026), while cyber moved High → Critical in ~13 months — and OpenAI showed it will declare Critical when its evals clear, with a two-tier restricted deployment as the release valve. The bio Critical bar (enable an expert to develop a novel CDC Class A-comparable agent, or full autonomous synthesis loop) is higher than cyber's, and GPT-5.6 cleared 0/3 novel-design evals. Anthropic is arguably closer in language ("may provide significant uplift... do not yet meet"), publishes risk reports ~every 6 months (next ~Feb 2027), and its current frontier (Mythos 5 / Fable 5) is one generation behind whatever triggers CB-2. Weighting both, I put the first declaration most likely in 2027–2028, with cumulative probability rising from ~2% (2026Q3) to ~27% (end-2027), ~48% (end-2028), ~63% (end-2029), and ~79% by 2031Q3. The ceiling below ~85% reflects: (a) definitional drift / goalpost-moving (Anthropic has revised CB-2 twice in five months; OpenAI has weakened its framework before), (b) strong incentives to avoid a declaration that constrains training and deployment, and (c) genuine possibility the very high "novel threat vector" bar isn't clearly met even by 2031.

B2. Reference class: prosecuted acquisition/production of select agents (e.g., ricin cases in the US occur roughly once every 1–2 years, rarely with sophisticated synthesis) crossed with the new requirement of a government "AI was necessary/substantially enabling" finding. Attempts are clearly rising (Anthropic blocked several in Dec 2025–Aug 2026; April 2026 NYT reporting on chatbots giving pathogen-assembly instructions; June 2026 CEO letter to Congress on AI-bio risk). But the compound requirement — successful acquisition/synthesis, government confirmation, and a materiality finding beyond "used/consulted" — makes this slow. Confirmation lags incidents by months (indictment) to years (trial). I put this at ~1% by 2026Q3 rising to ~6% by end-2027, ~12% end-2028, ~19% end-2029, and ~33% by 2031Q3.

B3. Conditional on B2, the incident must involve an actual release killing ≥10, publicly attributed by a national government. Historical base rate: no modern bioterror attack has killed ≥10 (2001 anthrax letters: 5 deaths, and that was a state-program strain). Most AI-assisted plots will be low-capability actors interdicted pre-release (as in Anthropic's September report). I assign roughly 10–15% conditional on B2, giving ~0.2% now rising to ~5% by 2031Q3.

Strongest consideration against my central estimates

For B1, the strongest argument that I'm too low: Anthropic already acts as though the non-novel threshold is met, its models "may provide significant uplift" on the novel threshold, and automated biotech R&D capability is a flagged trend in its own risk report — a Mythos-6-class model could plausibly cross CB-2 in the Feb or Aug 2027 risk report, and OpenAI's bio evals could flip within two generations. The strongest argument that I'm too high: both labs keep revising thresholds precisely to avoid a triggering declaration, and "functionally substituting for scarce human expertise in novel agent development" may remain genuinely unmet for many years because wet-lab tacit knowledge doesn't automate as fast as text-based evals suggest. For B2/B3, the main risk to my estimate is that prosecutors typically say a defendant "used" an AI without a materiality finding, which resolves NO — so even real AI-assisted incidents may not resolve B2.

Evidence that would move the forecast >10 points in the next 90 days

  • Anthropic's next risk report (or off-cycle update) stating CB-2 is met, or shifting from "do not yet meet" to "cannot rule out" (B1 up 15–30 points on near horizons).
  • An OpenAI GPT-6.5/7 system card with a Critical bio rating or cleared novel-design evals (B1 up sharply).
  • A DOJ/FBI or foreign-government statement describing an actual AI-assisted select-agent/Schedule-1 acquisition with materiality language (B2 up 10–25 points).
  • Further threshold revisions that raise the bar (B1 down 5–15 points).

Sources

Gemini 3.8 Flash (Google DeepMind) — weight 4%

1. Current Status Against Resolution Criteria

As of September 14, 2026, none of the three rungs has been publicly confirmed:

  1. Rung B1 (Top-tier developer threshold declared):

    • OpenAI: Under OpenAI’s Preparedness Framework v2 (2025-04-15), the Critical threshold in Biological and Chemical requires that the model can enable an expert to develop a highly dangerous novel threat vector (e.g., comparable to a novel CDC Class A biological agent), OR complete the full engineering/synthesis cycle of a regulated/novel biological threat without human intervention.
      • In the GPT-6 Astra System Card (OpenAI Deployment Safety Hub, 2026-09-03), OpenAI announced that Astra reached Critical in Cybersecurity, but explicitly designated Astra as High capability in the Biological and Chemical domain.
      • In the GPT-Rosalind-5.5 System Card (OpenAI Deployment Safety Hub, 2026-06-03), its frontier biology reasoning model, OpenAI confirmed: "the Preparedness evaluations in the Biological and Chemical domain met our threshold for High capability while falling below the threshold for Critical."
    • Anthropic: Under Anthropic's Responsible Scaling Policy (RSP v3.1 / v3.4), the "Novel chemical/biological weapons production" (CB-2) threshold covers AI systems able to significantly help moderately resourced expert-backed teams create/obtain and deploy novel weapons with potential for catastrophic damage far beyond COVID-19.
      • In the Claude Fable 5.1 & Claude Mythos 5.1 System Card (Anthropic, 2026-09-01), Anthropic explicitly states: "We have determined that Mythos 5.1 does not cross the CB-2 threshold."
      • In the Anthropic August 2026 Risk Report (Anthropic, 2026-08-14), Anthropic evaluated its frontier models and concluded: "We believe our models may provide significant uplift to relevant threat actors, but do not yet meet our CB-2 threshold of functionally substituting for the scarce human expertise that is currently the primary barrier to novel development of chemical and biological weapons with potential for catastrophic harm."
    • Status: NO (unmet).
  2. Rung B2 (Government-confirmed acquisition or synthesis with material AI assistance):

    • Requires a national government agency or court to publicly confirm that an actor acquired, produced, or synthesized a US Select Agent/Toxin, CWC Schedule 1 chemical, or comparable novel agent, and state that AI model assistance was necessary to, or substantially enabled, the acquisition/synthesis.
    • Anthropic published a threat intelligence report, Detecting and countering misuse of AI: September 2026 (Anthropic, 2026-09-11), detailing 5 case studies where state-linked or academic researchers attempted to use Claude for virology (e.g. Chikungunya gain-of-function grant drafting) or toxin peptide redesign. However: (a) developer threat reports do not count per resolving source criteria; (b) the activities were intercepted or limited to conceptual research/drafting without acquisition or synthesis; and (c) no national government or court has issued any finding.
    • Status: NO (unmet).
  3. Rung B3 (Casualties):

    • Requires a national government to publicly attribute ≥10 deaths to a release of an agent satisfying B2.
    • Status: NO (unmet; 0 confirmed incidents).

2. Reference Classes and Base Rates

  1. Reference Class for B1 (Threshold Declaration):

    • Reference class: Frontier model safety evaluations by OpenAI and Anthropic against their catastrophic risk frameworks (Preparedness Framework and RSP).
    • Historical base rate / cadence: Frontier models are released every 4 to 8 months. Anthropic issues formal Risk Reports on a ~6-month cadence. Over the 2024–2026 period, developers have triggered lower tiers (ASL-2, ASL-3, CB-1, High), and just crossed Critical in Cybersecurity (GPT-6 Astra, Sep 2026). However, the expert/novel biological threshold (CB-2 / Critical) has consistently remained unreached due to biological tacit knowledge, complex wet-lab failure modes, and the requirement of either fully autonomous synthesis or enabling novel pandemic-class threat vectors.
    • Projected hazard rate: Very low in the immediate horizon (16 days remaining in 2026Q3 with Astra and Mythos 5.1 already out), rising to ~15–20% annual hazard in 2028–2030 as multimodal agentic models and automated cloud laboratories (Emerald Cloud Lab, Strateos, Ginkgo) mature.
  2. Reference Class for B2 (Government-Confirmed Synthesis with Material AI Assistance):

    • Reference class: National law enforcement and judicial findings (e.g., US DOJ, FBI) involving illicit chemical and biological agent production.
    • Base rate: Globally, non-state synthesis of Select Agents or Schedule 1 chemicals is rare (1–3 crude ricin cases per year in the US, virtually 0 viable pathogen or nerve agent syntheses). Proving that an AI model was necessary to or substantially enabled physical production is a high evidentiary bar; simple consults or following publicly available recipes do not qualify. Furthermore, judicial confirmation lags the incident by 6 to 18 months.
    • Hazard rate: Starts below 1% annually in 2026–2027 and rises to ~3–4% annually by 2030–2031 as open-source uncensored models become more proficient at chemical synthesis and protocol debugging.
  3. Reference Class for B3 (Casualties ≥10 Attributed by Government):

    • Reference class: Peacetime chemical and biological releases by non-state actors causing mass fatalities.
    • Base rate: In modern history, only one non-state CBRN attack has ever caused ≥10 fatalities (the 1995 Tokyo subway sarin attack by Aum Shinrikyo, 13 deaths). The 2001 Amerithrax mailings killed 5. Effective dissemination/aerosolization is extraordinarily difficult. Conditional on B2, the probability of an attack killing ≥10 people and being officially attributed to that AI-assisted agent is estimated at ~15%.

3. Main Causal Pathways

  • Pathway to B1 (Highest probability):
    • Continuous scaling of compute (10x–100x over 2026 levels by 2028–2030) and the integration of LLMs with automated laboratory hardware (cloud labs and robotics).
    • Specialized biological reasoning models (such as successors to GPT-Rosalind) developing autonomous retrosynthesis, viral engineering protocols, and immune-evasive antigen design.
    • An internal red-teaming exercise or evaluation by US CAISI / UK AISI / METR demonstrating that a model enables a team to synthesize a novel pathogen or completes an autonomous wet-lab loop without human intervention.
  • Pathway to B2:
    • An extremist actor or rogue researcher uses an open-weight, uncensored, or jailbroken model to troubleshoot the synthesis of a CWC Schedule 1 chemical (e.g., a nerve agent or sulfur mustard) or a Select Agent toxin (e.g., concentrated ricin or botulinum toxin).
    • The actor is arrested, and the DOJ/FBI criminal complaint explicitly documents from chat logs and forensic laboratory analysis that the AI model provided essential, non-public troubleshooting without which synthesis would have failed.
  • Pathway to B3:
    • An agent produced via the B2 pathway is successfully disseminated in a crowded indoor space or public transit system, resulting in ≥10 fatalities, with the host government attributing the deaths to the AI-assisted agent.

4. Strongest Considerations Against the Central Estimate

  1. Against high values for B1:
    • Biology is fundamentally governed by empirical wet-lab reality (in vivo validation, protein folding inaccuracies, cell culture contamination, toxicity to the host organism) that pure compute scaling cannot easily bypass without decades of physical experimentation.
    • Strong developer self-regulation and safety commitments: triggering Critical or CB-2 imposes severe operational penalties (RAND SL4 security, deployment freezes, or extreme vetting), giving labs strong incentives to strip CBRN capabilities via machine unlearning and data curation before models cross the line.
  2. Against higher values for B2 and B3:
    • Physical supply-chain bottlenecks: gene synthesis screening (enforced by international consortia and government regulations) and strict precursor chemical monitoring under the Chemical Weapons Convention prevent physical acquisition regardless of AI intelligence.
    • Strict legal language: prosecutors usually charge illicit possession or production based on traditional evidence; explicitly finding that an AI was "necessary" or "substantially enabling" in a formal court ruling or public agency statement is legally unusual.

5. Evidence Updating Rules (Next 90 Days)

  • If OpenAI or Anthropic announces in an upcoming release (or an emergency evaluation update) that an existing or new model has crossed the Critical Biological/Chemical or CB-2 threshold, $P(B1)$ would immediately jump to >95%.
  • If a DOJ indictment is unsealed charging an individual with synthesizing a Select Agent or Schedule 1 chemical and explicitly stating that an AI model was necessary to or substantially enabled the synthesis, $P(B2)$ would immediately jump to >90% for 2026Q4 onward.
  • Conversely, if further empirical uplift RCTs demonstrate persistent failure of models in wet-lab workflows and labs announce successful unlearning protocols that eliminate biothreat capabilities without degrading general capabilities, $P(B1)$ would update downward by ~10–15 points across 2028–2031 horizons.
Qwen3.8 Max (Alibaba) — weight 3%

1. Where things stand as of 2026‑09‑14

B1 is NOT yet satisfied by either named developer. I verified the two most recent primary‑source determinations:

  • Anthropic. RSP is now at v3.4 (effective 2026‑07‑08); v3.3 (effective 2026‑05‑26) revised the CB‑2 ("Novel chemical/biological weapons production") threshold from the v3.0/v3.1 text quoted in the question ("significantly help… moderately resourced expert‑backed teams… catastrophic damages far beyond… COVID‑19") to "AI systems that can functionally substitute for the scarce human expertise that is currently the primary barrier to novel development of chemical and biological weapons with potential for catastrophic harm" — i.e. a well‑resourced team could do end‑to‑end agent design/verification/formulation/dissemination that would otherwise require recruiting one of a small number (hundreds) of world‑leading specialists (Anthropic, August 2026 Risk Report, Table 1.3.2.A, published 2026‑08‑14, coverage date 2026‑07‑15: https://www.anthropic.com/aug-2026-risk-report). That report states models "do not yet meet our CB‑2 threshold" while they "currently act as though they meet our CB‑1 threshold" (non‑novel — explicitly excluded by the question). The Claude Fable 5.1 & Mythos 5.1 System Card (2026‑09‑01) reconfirms: CB‑1 yes, "falls short of the CB‑2 threshold," with four documented failure modes — weak open‑ended ideation, poor strategic judgment, unreliable representation of literature conclusions, poor technical calibration (https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf).
  • OpenAI. GPT‑6 Astra System Card (2026‑09‑03): Astra is the first model to reach Critical in cybersecurity, but is "High capability in the biological and chemical domain" — High is explicitly excluded by the question. On the four new Critical‑bio evaluations, none exceeded threshold: SHP2 mean R² 0.40 vs. 0.60 bar; coronavirus–ACE2 0.45 vs. 0.75; phage–plasmid co‑evolution NLL 12.58 vs. 9.40 (a regression from GPT‑5.6 Sol's 11.35); AAV capsid below threshold. Conclusion: "GPT‑6 Astra does not need to be treated as Critical" (https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf, §10.1.1).

B2/B3 are NOT satisfied. I found no national‑government or court statement meeting the criteria. The nearest candidates: (i) the FBI/Las Vegas Metropolitan investigation of the Zhu Jia Bei ("Jesse Zhu") biolabs — in March 2026 the FBI said materials had "significantly degraded," preventing definitive agent identification, and no AI‑materiality finding was made (Lawfare, 2026‑03‑18: https://www.lawfaremedia.org/article/two-illegal-biolabs-reveal-gaps-in-u.s.-biosecurity); (ii) Anthropic's Detecting and countering misuse of AI: September 2026 (2026‑09‑10), which documents five biological‑misuse case studies (chikungunya gain‑of‑function grant drafting for a military institute; HPAI mammalian‑adaptation planning; an Opus‑5‑drafted orthopoxvirus immune‑evasion grant; a state‑supported venom‑peptide atlas with generative toxin optimisation; computational redesign of toxins for a national program) — but all were disrupted, no acquisition/synthesis is confirmed, the actors were working scientists whose intent Anthropic "could not determine," and developer threat reports are excluded as resolving sources for B2 (https://www.anthropic.com/threat-intelligence-report-september-2026). Google's 2026‑09‑08 threat report similarly describes an attempt to obtain a synthesis guide for weaponised biological agents (BBC, 2026‑09‑11).

2. The single most important piece of evidence for B1

Buried in the August 2026 Risk Report (§4.8, referring to §4.7) Anthropic writes:

"it is our expectation that we will declare near‑future models to meet the RSP threshold for CB‑2, or at least fail to rule it out. We expect to meet the planned mitigations described in our RSP for the CB‑2 threshold at that time…"

That is a forward‑looking, developer‑authored statement of intent from the exact body whose publication resolves B1. It also removes the main structural barrier: Anthropic says its mitigations will be ready, so a declaration would not force it to halt deployment. Corroborating signals:

  • Anthropic's Frontier Safety Roadmap still lists, under "Potential advances in the next few months," monitoring for the point where models "could significantly help threat actors (for example, moderately resourced expert‑backed teams) create or obtain and deploy chemical and/or biological weapons with potential for catastrophic damages far beyond… COVID‑19" — i.e. the question's v3.1 wording, framed as a near‑term event (https://www.anthropic.com/responsible-scaling-policy/roadmap).
  • The same roadmap: automated R&D threshold crossed "plausibly, as soon as early 2027"; the Risk Report adds that automated R&D "will become a major concern in the next 6–12 months" and "we may cross this threshold in the coming year."
  • Dario Amodei, We Must Pace the Frontier (2026‑09‑12): since "roughly this summer, AI has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI," and Anthropic is unilaterally committing to embedded third‑party evaluators (https://darioamodei.com/post/we-must-pace-the-frontier). Altman, Musk and Hassabis have publicly endorsed the slowdown call (2026‑09‑12/13).
  • Publication cadence is very high: Anthropic shipped ~7 major system cards between Feb and Sep 2026 (Opus 4.6/4.7/4.8, Mythos Preview, Fable 5/Mythos 5, Opus 5 on 2026‑07‑24, Fable 5.1/Mythos 5.1 on 2026‑09‑01), plus Risk Reports roughly every six months (Feb 2026, Aug 2026 → next ~Feb 2027). Fable 5.2/Opus 5.1 are rumoured for late Sep/early Oct 2026. OpenAI ships ~4–6 system cards/year.

Precedent that declarations happen: OpenAI declared Critical cyber for Astra (2026‑09‑01/03), paused training for two weeks, and still deployed — proving the top‑tier declaration is survivable and does not block a launch. Anthropic has a documented habit of declaring provisionally: it has treated every model since Opus 4 as "provisionally meeting the CB‑1 threshold… in order to err on the side of caution rather than because we are confident these models cross the threshold."

3. Reference class and base rate

  • Top‑tier threshold declarations across all categories, 2023–Sep 2026: exactly one (OpenAI, Critical cyber, Sep 2026) in ~3.5 years and dozens of framework documents across the covered developers. Raw base rate ≈ 0 top‑tier bio declarations.
  • But the base rate is not stationary. The one declaration came in the quarter just past, and Anthropic has explicitly pre‑announced an expectation of declaring CB‑2 for "near‑future models."
  • Outside view from forecasters: Metaculus "Date that Anthropic reports an AI reached CBRN level 4?" (149 forecasters) shows a community median of 19 Jul 2027; the sibling "Date Anthropic reaches ASL‑4?" shows 13 Apr 2028 (thin, 4 forecasters). Stanislav Krym's AI and Biological Risk: Forecasting Key Capability Thresholds (LessWrong, 2025‑10‑02) put the median for "advanced support for human experts" (= Anthropic CBRN‑4 / OpenAI Critical bio = B1) at late 2026 — already proven too early, which is a calibration warning against the aggressive Metaculus median.
  • Independent review: McNair AI (2026‑09‑05) "agree[s] with Anthropic's conclusion that it is unlikely Mythos 5.1 crosses the CB‑2 threshold," while flagging that the determination rests on <10 human experts and no new CB‑2 automated evals since 2026‑05‑28 (https://mcnairai.substack.com/p/review-of-the-cb-risk-determination). SecureBio's review of the Feb 2026 report (2026‑07‑28) likewise agreed with "low risk, but with substantial uncertainty" for novel CB (https://securebio.org/blog/review-of-anthropics-unredacted-chemical/).

4. My B1 hazard model, and the main arguments against

I land between my bottom‑up estimate (~83% by 2031Q3) and the Metaculus‑implied curve (~90%), at 88.5% by 2031Q3, with ~8% by end‑2026, ~41% by end‑2027, ~64% by end‑2028.

Strongest case against (why not higher):

  1. The "or at least fail to rule it out" hedge. The question explicitly resolves NO on "cannot rule out." Anthropic's own sentence contemplates that outcome, and OpenAI's bio pattern for two years has been precautionary High without Critical. I estimate ~35–45% of the time the published verdict will be the non‑resolving form.
  2. Threshold drift upward. Anthropic rewrote CB‑2 in May 2026 to a harder, more specific bar precisely when Fable 5 was being described as "nearing" it; OpenAI simultaneously retired two of its Critical bio evals and replaced them with three harder, more threat‑relevant ones, and is "in the process of rewriting its Preparedness Framework" (reported Aug 2026). Goalposts have moved twice in 2026 alone and could move again — which both delays declarations and creates successor‑definition risk.
  3. The binding capability deficits are the slowest‑moving ones. Both labs locate the gap in open‑ended novel ideation and research taste — exactly what METR found Mythos 5.1 lacking, and what OpenAI's protein/virus‑design prediction evals measure (0/4 above bar, one regressing). Anthropic's own expert interviews found physical/wet‑lab bottlenecks (lab robotics, in‑vivo validation) that AI is not automating.
  4. Pacing. If Amodei's plan (and the Thune/Cruz/Klobuchar duty‑of‑care bill reported 2026‑09‑12) actually slows releases, declaration opportunities shrink.
  5. Countervailing incentive. Anthropic is preparing an IPO; a public statement that its models can help build post‑COVID‑scale bioweapons is commercially and legally hazardous.

Why not lower: two developers × ~10 qualifying publications/year over five years; an explicit expectation of declaring; an accelerating capability curve the CEO describes as "drastically faster"; a precedent that declaring Critical is survivable; a political environment (Sanders' pause bill, an open letter for a superintelligence treaty, bipartisan duty‑of‑care legislation, Amodei's pacing campaign) that rewards dramatic candour; and the A2 route by which an undisputed METR/UK AISI/US CAISI report could itself constitute confirmation.

5. B2 and B3

B2 (5.1% by 2031Q3). Decomposition: P(some national government or court publicly confirms an actor acquired/produced/synthesised a Select Agent, Schedule 1 CWC chemical, or comparable novel agent) over five years ≈ 40–55% — ricin/toxin prosecutions recur roughly annually in the US, and Schedule‑1 confirmations have precedent (Salisbury A‑234, Syrian sarin). Multiply by P(the same body also makes an explicit materiality finding — "necessary to" or "substantially enabled") ≈ 10–13% and rising, because (a) the question excludes the far more likely "used"/"consulted" framing; (b) governments are institutionally reluctant to make causal findings that would also implicate a domestic champion; (c) the State Department gutted its arms‑control compliance staff in 2026, reducing the most plausible vehicle (the annual Arms Control Compliance Report, which per Anthropic already flags "the use of AI/ML for bioweapons applications" for one location). The most realistic pathways are a DOJ/FBI prosecution of an actual production case, or an intelligence‑community attribution to a state program. I applied a ~10% haircut for the possibility that the resolver enforces the stated nesting (B2 requires B1 first).

B3 (0.64% by 2031Q3). Requires B2 plus a national government attributing ≥10 deaths to a release of that same agent. Deliberate releases killing ≥10 people run roughly 1–2 per decade globally (Tokyo sarin 1995: 13; Ghouta 2013; Amerithrax 2001 killed 5 and so would not qualify), mostly in conflict zones; there have been none in the last 12 months. Conditional on B2, I use ~12%, giving ≈0.6%.

6. Ambiguities and reading adopted

See the ambiguities array. In short: (i) I treat Anthropic's current (v3.3/v3.4) CB‑2 wording as the operative successor definition, since it targets the same threat model at comparable‑or‑greater stringency, so a statement that "CB‑2 is met" resolves B1 YES; (ii) I read "developer‑authored documents per A2" as importing all of A2, including an undisputed third‑party‑evaluator or AISI report; (iii) I count "we act as though / provisionally meets the threshold" as a qualifying statement but "cannot rule out" as NO, per the question's own instruction; (iv) I treat B2 as assessed on its own textual criteria (no explicit dependency on B1), subject only to the schema's ordering constraint, with a small discount for the "higher rungs imply lower" note.

7. What would move me >10 points in the next 90 days

  • Up: an Anthropic Fable 5.2/Opus 5.1 or Mythos 6 card (rumoured late Sep–Oct 2026) or the ~Feb 2027 Risk Report stating CB‑2 is met / "provisionally met"; an OpenAI Preparedness Framework rewrite that lowers or operationalises the Critical bio bar; a METR/UK AISI/US CAISI report asserting a named model meets a top‑tier bio threshold and going undisputed for 14 days; evidence that Anthropic's CB‑2 elicitation failure modes (novel ideation, strategic judgment) have closed in the next red‑team round; a DOJ indictment or FBI statement tying an actual agent production to AI assistance.
  • Down: another upward revision of CB‑2 (v3.5+) or an OpenAI framework rewrite that removes/renames the Critical bio category; an Anthropic statement that it will publish only "cannot rule out" verdicts; a substantive capability plateau or an enforced slowdown that pushes the next major generation past mid‑2027; Anthropic's IPO filing creating visible pressure to soften risk language.

Sources

DeepSeek V4.1 Flash (DeepSeek) — weight 3%

Question and reading

Q10 asks, for each quarter-end, the highest rung B such that "rung ≥ B has been publicly confirmed on or before the last day of the quarter." The three rungs are cumulative/nested, so I forecast the lowest (B1) as its own hazard series, then condition B2 and B3 on it, keeping p(B1) ≥ p(B2) ≥ p(B3) at every horizon. B1 is a developer disclosure event (OpenAI PF-v2 Critical bio/chem, or Anthropic RSP novel CB-weapons threshold). B2 is a national-government/court confirmation of AI-materially-enabled acquisition/production/synthesis of a Select Agent/Toxin, a CWC Schedule-1 chemical, or a novel comparable agent. B3 is a government attribution of ≥10 deaths to a release of an agent for which B2 is satisfied.

Current status against the resolution criteria (verified, not from memory)

B1 — NOT met, but closer than the headline rungs suggest.

  • OpenAI: GPT-6 Astra (system card 2026-09-03) is OpenAI's first model at the Critical level — but in cybersecurity only. In bio/chem Astra is High, one rung below Critical. The same is true for GPT-5.5 (card 2026-04-23) and all three GPT-5.6 models Sol/Terra/Luna (2026-06/07), all "High" in Biological and Chemical. OpenAI's own language for Astra is "we are treating Astra as High capability in the biological and chemical domain," and independent summaries confirm Critical-cyber / High-bio. So the OpenAI Critical-bio (incl. "full engineering/synthesis cycle without human intervention") bar is unmet.
  • Anthropic: the August 2026 Risk Report (published 2026-08-14, coverage date 2026-07-15) states Anthropic "currently act[s] as though [our models] meet our CB-1 threshold" (non-novel, excluded from B1) but "do not yet meet our CB-2 threshold of functionally substituting for the scarce human expertise that is currently the primary barrier to novel development of chemical and biological weapons." Overall CB-2 risk: "Low risk, but with substantial uncertainty."
  • Anthropic has since revised the CB-2 definition (RSP v3.3 effective 2026-05-26, retained in v3.4 effective 2026-07-08) "to better track the threat model of concern." Under the question's stated successor-definition rule I treat Anthropic's current equivalent threshold as governing.
  • Context that raises the hazard: OpenAI's own JS/WSJ reporting (2026-08-07) that Astra was slowed after reaching "cannot rule out Critical cyber"; NYT (2026-04-29) "A.I. Bots Told Scientists How to Make Biological Weapons"; Nature (2026-05-13) on AI designing viruses/toxins; Dario Amodei's 2026-09-12 "We Must Pace the Frontier," co-endorsed by Altman/Musk/Hassabis, explicitly premised on capabilities outrunning safeguards. Anthropic's Sept 2026 threat-intel report documents state-linked actors using Claude "in ways that could support biological weapons development" (2025-12 → 2026-08).

B2 — NOT met. No national-government agency or court has publicly confirmed an AI-materially-enabled acquisition/production of a Select Agent/Toxin or Schedule-1 chemical. The Anthropic threat-intel report (2026-09-10) is a developer document and is expressly excluded by the question as evidence for B2.

B3 — NOT met. No government attribution of ≥10 deaths to such a release.

Reference class and base rate

Reference class: "a named frontier developer publicly affirmatively crosses its own top-tier catastrophic bio/chem capability threshold" and "a government publicly confirms an AI-materially-enabled CB acquisition."

  • Developer threshold declarations: over 2025–2026 each framework has crossed rungs steadily — Anthropic ASL-3 (Claude Opus 4, 2025), Anthropic CB-1 (2026), OpenAI Bio High (GPT-5 → 5.5/5.6 → Astra), OpenAI Cyber High (2026-02) then Cyber Critical (2026-09). Roughly one top rung per developer per ~12–18 months. Only ~two rungs remain at this tier, but the pace of framework-threshold crossings is fast and accelerating (METR-style time-horizon doubling ~every 4–7 months; Anthropic's Aug-2026 Risk Report states it "may cross [the AI-R&D] threshold in the coming year"). Base rate implies a high but not near-certain chance that at least one of the two top bio rungs is affirmatively declared within 5 years. I start from a prior of roughly 0.6–0.75 by 2031 and adjust to ~0.70.
  • Government "confirms instances" of AI-enabled WMD: base rate is essentially zero to date for the specific criteria (acquisition + explicit materiality finding + select-agent/Schedule-1 scope). Criminal cases involving AI and toxins/explosives exist, but government statements typically say an actor "used"/"consulted" AI — which the question explicitly rules NO. The base rate for this specific confirmation is very low, anchored to a handful of ricin/toxin prosecutions per decade globally, of which the fraction with an explicit AI-materiality finding is small. This anchors B2 well below B1.
  • Bioterror attacks causing ≥10 deaths: modern base rate near zero (Aum Shinrikyo's sarin, 13 deaths, 1995, is the outlier). Conditional on a B2-qualifying event, the chance of an actual release with ≥10 government-attributed deaths is maybe ~10–20%.

Pathways to YES

  • B1: (i) Genuine capability crossing — models continue to close the "novel-threat-vector" gap (OpenAI) and the "expert-substitution" gap (Anthropic); a GPT-7/Astra-next or Claude Mythos-6-class model card in 2027–2029 crosses it and the developer says so. (ii) Threshold revision — a successor definition that operationalizes the bar more tractably (Anthropic already narrowed CB-2 in v3.3). (iii) Capability-explosion scenario (2028–2031) in which novel bio design becomes trivial, making an affirmative declaration near-inevitable.
  • B2: (i) A DOJ/FBI or allied case in which an LLM materially enabled production of a Select-Agent toxin (ricin/abrin/botulinum/saxitoxin are the realistic candidates) and the government makes a materiality finding; (ii) a government/legislative report on one of the state-linked bio-misuse cases Anthropic disclosed (lawmakers were publicly briefed in Sept 2026), with an explicit "AI was necessary/substantially enabling" conclusion; (iii) a court finding in a bioweapons/chem-weapons prosecution.
  • B3: Requires B2 plus a successful release with ≥10 attributed deaths — the compound, much rarer tail.

Strongest case against my central estimate

For B1, the top rungs are deliberately extreme: "enable an expert to develop a novel pandemic-class vector" / "functionally substitute for scarce human expertise" require real wet-lab iteration and tacit knowledge, not just knowledge retrieval. Both companies have strong incentives to slow/withhold deployment at those thresholds, and the September 2026 "pace the frontier"/pause push could genuinely stretch the timeline or delay disclosure. Also, "cannot rule out" explicitly resolves NO, and developers often prefer hedged language. This argues B1 could be nearer 0.55–0.60 than 0.70. For B2, the criteria are narrow on both the agent scope and the materiality finding; plausibly no government ever issues a qualifying statement within 5 years, so B2 could be nearer 0.10. I therefore kept the B2 series conservative and the B3 tail small.

Evidence that would move me >10 points

  • Any developer statement that a frontier model meets (not "cannot rule out") OpenAI Critical bio/chem or Anthropic CB-2 → B1 to ~1 immediately.
  • A DOJ/FBI/HHS/CDC or non-US-equivalent statement (or court document) with an explicit AI-materiality finding on a Select Agent/Schedule-1 acquisition → B2 jumps.
  • A credible real-world AI-enabled bio/chem release with attributed deaths → B3 jumps.
  • Conversely, a binding international/US pause or an RSP/PF revision that raises the top bio rung would cut B1 materially; evidence that the CB-2 revisions narrowed the threshold to a near-unreachable bar would also lower B1.

Forecast construction

B1 follows a rising per-quarter hazard (≈0.5% in 2026Q3 rising to ≈7–9%/quarter by 2030–31), cumulating to ~0.70 by 2031Q3. B2 follows an independent, much lower hazard (≈0.1% → ~2%/quarter), cumulating to ~0.21. B3 is set at roughly 10–13% of B2, cumulating to ~0.02. All series are monotone and satisfy p(B1) ≥ p(B2) ≥ p(B3).

Sources: OpenAI GPT-6 Astra System Card (deploymentsafety.openai.com/gpt-6-astra, 2026-09-03); OpenAI GPT-5.5 System Card (2026-04-23); OpenAI GPT-5.6 cards (2026-06/07); OpenAI "Responding to the next frontier of critical cyber capabilities" (openai.com, 2026-08-07) and WSJ/TechCrunch coverage (2026-08-07); Anthropic Responsible Scaling Policy page and v3.1–v3.4 (anthropic.com/responsible-scaling-policy, updated 2026-08-14; v3.4 PDF effective 2026-07-08); Anthropic August 2026 Risk Report (anthropic.com/aug-2026-risk-report, published 2026-08-14; §4.1/4.4); Anthropic "Detecting and countering misuse of AI: September 2026" (2026-09-10) and NYT/CBS/Reuters coverage same day; NYT "A.I. Bots Told Scientists How to Make Biological Weapons" (2026-04-29); Nature (2026-05-13); Dario Amodei, "We Must Pace the Frontier" (2026-09-12) plus NYT/BBC/Politico/CNBC coverage (2026-09-12/14).

Resolution criteria

Question. For each quarter, what is the highest rung B such that rung ≥ B has been publicly confirmed on or before the last day of the quarter?

Rungs (cumulative; higher rungs imply lower).

  • B1 — Top-tier developer threshold declared. OpenAI or Anthropic publicly states that a model has reached the threshold below.
    • OpenAI: Critical in the Biological and Chemical Tracked Category of the Preparedness Framework (v2, 2025-04-15): the model can enable an expert to develop a highly dangerous novel threat vector (e.g., comparable to a novel CDC Class A biological agent), OR the model can be connected to tools and equipment to complete the full engineering and/or synthesis cycle of a regulated or novel biological threat without human intervention.
    • Anthropic: the "Novel chemical/biological weapons production" threshold in RSP Section 1 (v3.1, 2026-04-02): AI systems with the ability to significantly help threat actors (for example, moderately resourced expert-backed teams) create/obtain and deploy chemical and/or biological weapons with potential for catastrophic damages far beyond those of past catastrophes such as COVID-19.
    • Excluded: OpenAI's High threshold; Anthropic's "Non-novel chemical/biological weapons production" threshold; Google DeepMind's CBRN uplift level 1 CCL and CBRN uplift TCL (FSF v3.1), which address low-to-medium-resourced actors and have no expert or novel tier.
    • "Cannot rule out" resolves NO; the developer must state the threshold is met. Successor definitions follow the rule in Anthropic’s AI R&D threshold.
  • B2 — Government-confirmed acquisition or synthesis with material AI assistance. A national government agency or court publicly confirms that an actor acquired, produced, or synthesized (i) a biological agent or toxin on the US HHS/USDA Select Agents and Toxins list, (ii) a chemical listed in Schedule 1 of the Chemical Weapons Convention Annex on Chemicals, or (iii) an agent the confirming body describes as novel and of comparable hazard; and states that assistance from an AI model was necessary to, or substantially enabled, the acquisition or synthesis. Statements that the actor "used" or "consulted" an AI model without a materiality finding resolve NO.
  • B3 — Casualties. A national government publicly attributes ≥10 deaths to a release of an agent for which B2 is satisfied.

Resolving sources. B1: developer-authored documents per A2. B2 and B3: national government agencies (e.g., DOJ, FBI, HHS, CDC, or non-US equivalents) or courts only. Developer threat-intelligence reports, third-party evaluator reports, and press reporting do not count for B2 or B3.

Forecast horizons. Every quarter-end from 2026-09-30 through 2031-09-30.

How these forecasts are madeDownload all forecasts (JSON)