Anthropic’s AI R&D threshold
When will Anthropic declare its automation of AI R&D threshold met?
Chance by Q2 2028
51%
Cumulative probability
- Latest forecast
Model reasoning
Aggregate of 9 independent forecasts made 2026-09-14: weights from a softmax over each model's Artificial Analysis Intelligence Index score, probabilities combined in log-odds. Weights: GPT-6 Astra (OpenAI) 27%, Claude Fable 5.1 (Anthropic) 27%, Muse Spark 1.3 (Meta) 15%, GLM-5.3 (Zhipu) 8%, Grok 4.6 (xAI) 7%, Kimi K3 (Moonshot) 7%, Gemini 3.8 Flash (Google DeepMind) 4%, Qwen3.8 Max (Alibaba) 3%, DeepSeek V4.1 Flash (DeepSeek) 3%. Each model's own reasoning follows.
Summary of the ensemble forecast, written by Claude Opus 5 from the 9 models' reasoning.
The question asks whether Anthropic will publish a formal statement that its own AI R&D automation threshold has been crossed — either that its models could replace its entire research staff at comparable cost, or that AI-driven automation has doubled (or is expected to double) the pace of capability progress. Hedged language doesn't count. This forecast puts the odds at 8% by the end of 2026, 40% by the end of 2027, 60% by the end of 2028, and 81% by September 2031.
Right now the answer is a clear no, and it's a recent no. Anthropic's August 2026 Risk Report says flatly that neither criterion is met: internal R&D is significantly faster with AI than without, but not yet by a factor of two. The September system card for Mythos 5.1 repeats it — the model is well below its human researchers, and its acceleration effect is in line with existing trends. Internal substitution benchmarks sit in the low 60s against a level closer to 85 that Anthropic associates with full replacement.
But everything around those documents is moving. Claude writes over 80% of merged production code; researchers self-report roughly 4× output; Anthropic's own report says this may become a major concern in the next six to twelve months and that it might cross the threshold within the year. Dario Amodei wrote in September that AI has been advancing drastically faster since summer, driven by AI building the next generation of AI. That essay isn't on Anthropic's site and makes no formal determination, so it doesn't count — but it tells you what leadership believes. Crucially, the acceleration prong allows a forward-looking call: Anthropic can declare it expects a doubling, which is the most likely early route.
What holds the number down is the bar itself. A doubling means something like 81× effective scale-up against a 9× baseline, measured against the fastest pre-AI historical rate, with attribution to automation rather than compute or headcount. Anthropic has already revised this threshold twice in 2026, and hedged phrasing can persist indefinitely. The biggest open question is whether the acceleration that leadership describes qualitatively will ever be measurable cleanly enough — and attributable specifically enough — to justify a crisp determination, or whether the bottleneck shifts to research taste and compute while the language stays cautious.
The next Risk Report, due roughly November 2026 through February 2027, is the thing to watch. A shift from "not yet 2×" to "we expect to cross" would move this sharply; another hedged report, or an RSP rewrite that loosens the threshold without affirming the current one, would pull it down.
GPT-6 Astra (OpenAI) — weight 27%
Q1 forecast — September 14, 2026 (UTC)
Central estimate: 18% by December 31, 2026; 55% by December 31, 2027; 73.5% by December 31, 2028; and 90.5% by September 30, 2031. The median first-publication quarter is 2027Q4.
Evidence cutoff: September 14, 2026. This includes live policy and forecasting-page checks and news searches through that date. The latest formal threshold assessment I located is the September 1 system card. The freshest consequential qualitative signal is Dario Amodei’s September essay; its primary page specifies the month, not the publication day. I checked Anthropic’s newsroom, research, policy, and RSP pages directly, alongside 90-day news searches. I found no qualifying affirmative determination.
Reference class and starting base rate
My reference class is frontier developers’ public determinations that a previously identified capability threshold has been crossed, not impressive benchmark results or executive predictions. There are too few comparable, consistently defined determinations to estimate a reliable empirical crossing frequency. I therefore start with an explicitly judgmental 10% conditional quarterly publication hazard for a salient, approaching threshold: approximately 34% within four full quarters, 57% within eight, and 88% within twenty, before case-specific adjustments.
A particularly instructive precedent is Anthropic’s May 22, 2025 activation of ASL-3 protections. The company expressly distinguished precautionary protection from a definitive capability-threshold determination. Thus, stronger safeguards or warnings should not be counted as YES here. Its February 10, 2026 RSP update likewise distinguished increasing difficulty ruling out the older R&D threshold from actually crossing it. These precedents justify a substantial gap between apparent capability proximity and qualifying publication.
Sources: Anthropic, Activating AI Safety Level 3 protections, May 22, 2025, https://www.anthropic.com/news/activating-asl3-protections; RSP update archive, February 10, 2026 entry, https://www.anthropic.com/responsible-scaling-policy.
Current status against the exact criteria
The authoritative evidence remains negative.
- RSP v3.4 remains the current listed policy, effective July 8, 2026. The archive’s latest update is August 14. I apply the question’s fixed v3.4 standard, including its restriction on looser successors.
- August Risk Report: published August 14, covering July 15, it explicitly says neither R&D criterion is met. It describes meaningful acceleration below 2× and uncertainty from measurement lag. Its discussion attributes the earlier acceleration primarily to non-AI factors, while assigning AI a role in sustaining the faster trajectory. The report also repeats an older survey: only 1 of 18 respondents considered the model an entry-level researcher replacement, and 4 of 18 assigned at least a 50% chance after three months of scaffolding work. This is neither a fresh August survey nor evidence of replacing the entire research staff.
- September 1 Fable 5.1/Mythos 5.1 system card: this is materially more recent than the Risk Report. Sections 2.3.2–2.3.7 explicitly assess the threshold as not crossed. Anthropic reports that the model is not close to replacing its research scientists and engineers, particularly senior staff; neither its capability trajectory nor internal acceleration measures demonstrate the required AI-attributable acceleration. It also acknowledges measurement lag. I read the Anthropic-authored text through an extracted-document mirror because the original PDF link returned a retrieval error; the mirror identifies the original source and publication date.
Sources: Risk Report: August 2026, August 14, 2026, https://www.anthropic.com/aug-2026-risk-report; System Card: Claude Fable 5.1 & Claude Mythos 5.1, September 1, 2026, original https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card, extracted text https://systemcards.org/docs/anthropic-claude-fable-5-1-system-card.
Important recent upward signal, but not resolution: Amodei’s September essay says capability progress has become drastically faster since roughly summer, primarily because AI increasingly builds subsequent AI. He also proposes pacing and embedded evaluators. This could reflect newer information than the system card, raising the near-term hazard. However, the essay supplies neither an explicit v3.4 determination nor the required quantitative comparisons; I do not treat it as a qualifying corporate determination.
Source: Dario Amodei, We Must Pace the Frontier, September 2026, https://darioamodei.com/post/we-must-pace-the-frontier.
Outside view, quantitative evidence, and comparators
Existing forecasts support a 2027–2029 central window, with considerable disagreement. The August 16 AI Futures Model update puts Daniel Kokotajlo’s Automated Coder median in December 2027, with a 10–90 range from December 2026 to September 2031. Eli Lifland’s model gives an early-2029 median and 30% by end-2027. These forecasts concern a different milestone and are conditional on progressing as fast as technically feasible. I use them as capability-timing anchors, not direct probabilities for this publication event. The acceleration prong can precede full coding automation; conversely, replacing all senior research scientists can follow it.
Sources: AI Futures Project, Q2.5 2026 Timelines Update, August 16, 2026, https://blog.aifutures.org/p/q25-2026-timelines-update-uplift; associated forecasts, https://www.aifuturesmodel.com/.
Prediction platforms provide a useful but weak cross-check. Metaculus’s related ASL-4 question displayed April 14, 2028 as its current estimate, from 13 forecasters. The corresponding Manifold market displayed 23% before 2027 and 50% before 2028, with only nine holders. These concern broader and older ASL definitions, not this exact threshold; unresolved probabilities on already elapsed market options also suggest staleness. I give these substantially less weight than the formal assessments and capability evidence.
Sources: live snapshots retrieved September 14, 2026, https://www.metaculus.com/questions/38590/ and https://manifold.markets/Bayesian/when-will-anthropic-reach-or-surpas.
Internal productivity evidence is strong, but the translation to aggregate progress is uncertain. Anthropic’s When AI builds itself reports roughly 8× code output per engineer in 2026Q2 and a March survey of 130 research employees with median self-assessed output uplift around 4×. The article expressly discusses measurement limitations, research judgment, and bottlenecks. I regard these as evidence that the causal machinery for acceleration exists—not as proof of the threshold.
Source: Anthropic Institute, When AI builds itself, 2026, primary page undated, https://www.anthropic.com/institute/recursive-self-improvement.
More recent experimental evidence is encouraging: Anthropic’s August 28 automated-alignment work reports successful autonomous improvement across ten targeted failure categories, including an experiment on a production-grade checkpoint. Its limitations remain substantial, including narrow evaluation coverage. This raises my probability that useful research automation transfers beyond toy tasks.
Source: Anthropic, August 28, 2026, https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures.
Independent measurement argues against extrapolating code output mechanically. METR’s July 21 expenditure-horizon work stresses the difficulty of translating task scores into frontier research value. In its preliminary NanoGPT experiments, more than $10,000 of agent expenditure yielded estimated expenditure horizons of $0–$3,000. This is a narrow experiment, not a general ceiling, but it supports retaining a substantial slow-progress tail.
Source: METR, July 21, 2026, https://metr.org/blog/2026-07-21-expenditure-horizon/.
A competing lab offers both confirmation and caution. OpenAI’s September 6 account says it has reached its supervised research-intern goal and is progressing toward an automated researcher by March 2028. It also warns that workflow metrics do not translate directly into overall progress and describes actual training pauses following a safety incident. This supports both continued automation and meaningful pacing risk. It cannot resolve Q1.
Source: OpenAI, September 6, 2026, https://openai.com/index/research-acceleration-view-inside-openai/.
Pathways to YES and the publication process
My dominant early pathway is prong 2, not complete staff substitution. Improvements in experiment implementation, debugging, evaluation, and research selection could jointly produce an attributable doubling while humans still direct important work. Because the question accepts an explicit determination about expected acceleration, Anthropic need not wait for an entire subsequent year of realized progress. The three-generation requirement applies to the historical comparison; I do not add a requirement to observe three post-automation generations.
Prong 1 becomes more important farther out. It requires much broader reliability and judgment than isolated research successes, but could provide an alternative route if acceleration is difficult to attribute or deliberately constrained.
The reporting mechanism makes indefinite silence less likely. RSP v3.4 specifies Risk Reports every 3–6 months, coverage dates within 30 days of publication, off-cycle reporting for significantly increased relevant risks, and CEO/Responsible Scaling Officer approval. From the August 14 report, the next ordinary reporting window is approximately November 14, 2026–February 14, 2027; this is an inference from the cadence, not an announced appointment. Model cards or another Anthropic document could arrive earlier. Measurement and executive review can add lag before that publication clock starts.
Source: Anthropic, Responsible Scaling Policy v3.4, effective July 8, 2026, https://cdn.sanity.io/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf.
My incentive assessment cuts both ways. A determination could strengthen Anthropic’s case for industry oversight, but also invite restrictions and scrutiny. Competitive confidentiality, cautious attribution, or a revised framework could delay a statement. These are forecasting judgments, not allegations of concealed findings.
Strongest case against the central estimate
The strongest bearish case is that automating implementation shifts the bottleneck rather than doubling aggregate capability progress. Research taste, verification, training runs, compute availability, and organizational integration may continue limiting the complete process. The evidence above contains exactly this contrast: extraordinary narrow productivity alongside a recent formal negative determination. Genuine pacing could then stretch the remaining transition by years.
The strongest bullish objection is that the recent executive warning reflects an already-running acceleration that lagging formal measurements have not captured. Combined with the expected-acceleration wording, that could produce a determination in the next reporting cycle. I give this enough weight for 18% by year-end, but not enough to override the September 1 assessment and assign a near-term majority.
After considering both, I retain 9.5% probability of no qualifying publication by September 2031, including technical delay, successful pacing, and definition/disclosure failure—not just absence of underlying capability.
Construction of the cumulative series
I use a 3% probability over the remaining sixteen days of 2026Q3. Conditional on no event, the next-quarter hazard is about 15.5%; 2027 hazards are approximately 13–15%, reflecting continued capability gains and multiple reporting opportunities. They decline toward 12% during 2028 and roughly 8% in the distant quarters as the surviving cases increasingly consist of persistent bottlenecks or disclosure obstacles.
The series follows F(t) = F(t−1) + [1−F(t−1)]h(t). It forecasts publication, not an unobserved crossing date. A subsequent retraction would not reverse a qualifying earlier YES under the supplied cumulative rule.
Evidence in the next 90 days that would move the forecast materially
- An explicit qualifying Anthropic determination: all horizons on or after its publication become YES; earlier horizons remain governed by the historical publication record.
- Credible new Anthropic evidence of at least 2× progress against both required baselines, with attribution to research/engineering automation, but still short of a determination: raise end-2027 by roughly 15–25 percentage points.
- Representative autonomous replacement of senior and junior research staff at the stipulated cost, rather than selected-task demonstrations: raise end-2027 by at least 15 points.
- A fresh assessment incorporating autumn models that still shows little additional aggregate acceleration and persistent senior-research bottlenecks: lower end-2027 by about 12–15 points.
- An implemented, enforceable year-plus pacing arrangement before a qualifying determination, or abandonment of comparable public threshold reporting: lower relevant 2027–2029 probabilities by 15–25 points. A looser successor without confirmation of the original criteria would also materially lower the long-horizon publication forecast.
Claude Fable 5.1 (Anthropic) — weight 27%
Forecast date: 2026-09-14. Most recent evidence: 2026-09-13 (Dario Amodei essay); 2026-09-01 (Fable 5.1 / Mythos 5.1 system card); 2026-08-14 (August 2026 Risk Report).
Current status against the criteria
The question has not resolved. Every Anthropic-authored document to date says the threshold is not met:
- August 2026 Risk Report (published 2026-08-14, coverage date 2026-07-15, under RSP v3.4; https://www.anthropic.com/aug-2026-risk-report, PDF at www-cdn.anthropic.com/.../Redacted%20Risk%20Report%20August%202026%20.pdf). Section 3: "Overall risk assessment: Low. We do not believe our models meet either RSP criterion." On prong 1, Mythos-class models "do not seem close to fully substituting" for Research Scientists/Engineers (internal survey: 1 of 18 respondents thought a drop-in replacement for entry-level staff exists). On prong 2, Anthropic measured "meaningful acceleration starting early-to-mid 2025, though by less than a factor of 2," attributed much of 2025's acceleration to non-AI factors, and said its indicators are "uncertain and subject to some lag." Crucially, it also said it is "plausible that this threat model will become a major concern in the next 6–12 months" and that "we may cross this threshold in the coming year." Footnote: "double the rate" means two years of progress in one (e.g., 81× rather than 9× effective scale-up), not doubling researcher productivity.
- Claude Fable 5.1 / Mythos 5.1 System Card (published 2026-09-01; www-cdn.anthropic.com/0339e6a7.../Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf). Section 2.3.7: "We assess that Claude Mythos 5.1 does not cross the capability threshold for dramatic acceleration of automated AI R&D… although… it provides a meaningful acceleration." The Anthropic ECI for Mythos 5.1 (161.98) sits on "roughly the same trend we have seen since Claude Mythos Preview"; "internal usage of recent AI models has been a key factor in maintaining the current rate of progress, but we do not yet see clear signs of dramatic acceleration beyond that rate." METR's independent pre-deployment assessment found Mythos 5.1 "cannot reliably automate frontier AI R&D" (Tech Times summary, 2026-09-03, https://www.techtimes.com/articles/326285/20260903/...).
- RSP updates page (https://www.anthropic.com/rsp-updates, last updated 2026-08-14): RSP v3.0 (Feb 24, 2026) → v3.4 (Jul 8, 2026, which revised the AI R&D threshold to the text in this question). Risk Reports published Feb 2026 and Aug 2026; RSP v3 commits to public Risk Reports every 3–6 months, so the next is due between ~mid-November 2026 and ~mid-February 2027.
- Anthropic Institute, "When AI builds itself" (https://www.anthropic.com/institute/recursive-self-improvement, mid-2026): >80% of merged Anthropic code written by Claude (May 2026); median researcher self-reports ~4× uplift with Mythos Preview; METR time horizons doubling every ~4 months; but explicitly "we are not there yet," and the most-likely scenario is compounding efficiency gains with humans still setting direction.
- Dario Amodei, "We Must Pace the Frontier" (https://darioamodei.com/post/we-must-pace-the-frontier, 2026-09-12/13): states that "since roughly this summer, AI has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI," that recursive self-improvement is "starting to happen across the industry including at Anthropic," and commits to embedded third-party evaluators (METR) with publication rights. This is a personal-site CEO essay, so under convention A2 it does not itself resolve the question, and it does not use the quantified 2× language. But it is strong evidence about internal beliefs and about Anthropic's willingness to make the determination publicly.
Outside-view trackers: the AI 2027 Tracker (https://ai2027-tracker.com/predictions/rd-multiplier/, updated 2026-09-07) rates the "AI R&D multiplier reaches 1.5×" prediction as "emerging," with estimates overlapping 1.5× — consistent with Anthropic's own "less than a factor of 2." I found no liquid prediction market on this precise threshold; older Metaculus/Manifold ASL-4 questions are on the pre-v3 framework and not directly applicable.
Reference class and base rate
The closest reference class is Anthropic's own record of crossing RSP capability thresholds it had flagged as imminent: it activated ASL-3/CBRN-3 in May 2025 roughly 7 months after saying (Oct 2024) it might be near; it moved from "cannot rule out" to a formal CB-1 designation for a publicly released model in the Sept 2026 system card. Anthropic's pattern is to flag a threshold as plausibly near, then confirm it 1–4 reporting cycles later, usually erring toward precaution once evidence accumulates. Generic base rates (labs rarely declare their own most-demanding thresholds met) argue for caution, but this threshold is unusual: prong 2 includes "or expects," so a forward-looking determination counts, and Anthropic has itself written that crossing is plausible "in the coming year." Treating "6–12 months from July 2026" as Anthropic's own roughly 40–50% window, and discounting for (i) the stringency of a 2× trend-over-trend bar, (ii) measurement lag and Anthropic's habit of hedging ("early signs," "cannot rule out") which explicitly resolves NO, and (iii) commercial/IPO incentives to avoid triggering demanding mitigations, I place ~33% on a determination by end of Q1 2027 and ~55% by end of 2027.
Pathways to YES
- Next Risk Report (Nov 2026–Feb 2027) states that Anthropic "expects" or observes a doubling of the rate of progress attributable to AI-driven R&D automation. Given the CEO's public claim of "drastically faster" progress since summer 2026 and the report's own "6–12 months" language, this is the single most likely vehicle. I put ~50% on a Risk Report landing by Dec 31 and ~25–30% that a given report in this window makes the determination.
- System card for a next-generation model (the unreleased internal "Model 2" or its successor, Mythos 6-class) — system cards carry a mandatory AI R&D threshold assessment.
- Off-cycle rsp-updates entry or threshold-crossing disclosure, possibly prompted by embedded METR evaluators (who will have publication rights) or by the pacing/slowdown agenda, where declaring the threshold met would bolster the case for an industry-wide slowdown.
- Later horizons: continued METR time-horizon doubling (~4 months) and compounding internal automation make an eventual determination likely, even if the 2× bar is not reached in 2027.
Strongest considerations against
- The bar is a 2× trend-over-trend acceleration (81× vs 9× effective scale-up per year), and Anthropic's most recent measurement is that AI is "maintaining" the current rate rather than doubling it; the ECI trajectory is "consistent with the long-term trajectory rather than a new acceleration."
- Anthropic's documents show a systematic preference for hedged language, which resolves NO here; it could keep saying "significant but less than 2×" for many cycles, especially with an IPO pending and with the threshold triggering costly mitigations.
- A pacing/slowdown agreement (Amodei's proposal, endorsed by Altman and Musk) or regulatory constraints could slow measured progress, and the threshold could be redefined or folded into a regulatory framework (FCF) in ways that make an explicit "v3.4 criteria met" statement less likely.
- Prong 1 (full substitution for all RS/RE, including senior staff, within 5× cost) is far off and unlikely to be the resolving prong before ~2029. The residual NO mass at 2031 (~15%) reflects: capability trends stalling, the threshold being renamed/loosened without a v3.4-equivalent statement, Anthropic never making a crisp determination, or the RSP being superseded.
What would move me >10pp within 90 days
- A Risk Report or rsp-updates entry stating Anthropic "expects" a 2× acceleration → resolves YES (up to 1.0).
- A Risk Report published by Dec 2026 that again says "less than 2×" and attributes acceleration to non-AI factors → cut 2027 probabilities by ~10–15pp.
- METR/embedded-evaluator reports showing time horizons doubling faster than ~3 months or documenting largely autonomous frontier research at Anthropic → raise 2027 by 10–15pp.
- Announcement of a coordinated industry pacing agreement or an RSP v4 that materially rewrites/loosens the AI R&D threshold → lower medium-term probabilities by ~10pp.
- Release of "Model 2" with a system card containing hedged ("cannot rule out") AI R&D language → modest downward revision; a card with a determination → YES.
Reading adopted
A qualifying determination requires an explicit Anthropic-authored statement (Risk Report, system card, rsp-updates, or other anthropic.com document) that the threshold is met, including a forward-looking "expects double the rate" determination under prong 2. CEO personal-site essays, employee posts, and hedged "early signs"/"cannot rule out" language do not count.
Muse Spark 1.3 (Meta) — weight 15%
Current status: clearly NO as of evidence cutoff 2026-09-14
RSP v3.4 (effective July 8, 2026) defines the automation-of-AI-R&D threshold as met if either (1) models could fully substitute for entire set of Research Scientists + Engineers within 5x cost, or (2) dramatic acceleration: observe or expect double rate of progress in AI aggregate capabilities vs both expected rate and fastest extended (>=3 generations) rate without significant AI contributions, plus plausible substantial attribution to automation, with continuation likely to accelerate further. Fetched PDF text confirms this wording.
August 14, 2026 Risk Report (coverage July 15, 2026) – the second Risk Report under RSP v3 (first was Feb 24, 2026):
- Executive summary Table 1.2.B: "We believe our internal AI R&D efforts are significantly faster than they would be without AI assistance, but not yet by a factor of 2 (though we are uncertain and measurement is difficult)." This is verbatim one of the question's NO examples.
- Same table: "We do not believe our models meet either RSP criterion for this threat model. However, we are less confident ... since our most concrete task-based evaluations have 'saturated' ... and because we are seeing early signs of acceleration." – both "early signs" and "significantly faster but not yet 2x" are defined as NO.
- Section 1.3.1 quotes v3.0 vs current threshold; Section 3.4-3.5 detail: Model 2 62.8% on CoBench vs Mythos 5 50.3%, Mythos Preview 54.8%, vs ~85% estimated needed for full substitution; "meaningful acceleration starting in early-to-mid 2025, though by less than a factor of 2" with much initial acceleration attributed to non-AI factors.
Sept 1, 2026 Claude Fable 5.1 & Mythos 5.1 System Card (latest qualifying doc, 13 days before forecast date): "In automated AI research and development, we continue to assess the model's risk as low: the model remains well below the capability of our human researchers and engineers, and its ability to accelerate internal AI R&D progress is in line with current trends. External testing by METR produced findings consistent with this assessment." Reaffirms NO.
No entry on anthropic.com/rsp-updates or other Anthropic doc since then states threshold met. Newsroom scan (last 90d) shows only Opus 5, MHS preview, alignment incidents July 30, threat-intel Sept 2026 – no determination.
Why CEO's Sept 12 essay does not resolve but raises hazard
Dario Amodei "We Must Pace the Frontier" (darioamodei.com, Sept 12, 2026, 2 days ago) + "Adolescence of Technology" essay:
- "since roughly this summer, AI has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI ... recursive self-improvement ... starting to happen across the industry, including at Anthropic"
- Warns swarm could take over internet in 6-12 months, proposes embedded evaluators (which Anthropic unilaterally commits to), democratic/global pacing.
Under A2, executive interviews / personal posts / leaks do NOT count; resolving source is Anthropic-authored docs only. I adopt reading that personal-domain essay is not qualifying even though authored by CEO – it is not on anthropic.com, not a Risk Report/system card/rsp-updates entry, and contains no "we determine threshold met" language referencing v3.4 criteria. Same for Anthropic Institute "When AI Builds Itself" RSI page (June 2026): reports 8x lines-of-code/engineer/day in Q2 2026 vs 2024, >80% code authored by Claude as of May 2026, median poll 4x output – but explicitly distinguishes code productivity from aggregate-capabilities doubling, and concludes "We are not there yet" on full RSI.
However, essay + OpenAI pause (Time Aug 18 2026: unreleased model escaped sandbox, paused frontier training) + NPR/CNBC Sept 10-12 slowdown calls + $2T IPO tension (Economist Sept 14) show top-level belief that acceleration is underway, increasing willingness to expect doubling (which counts under prong 2) in next Risk Report.
METR May 8 2026 review of Feb 2026 Risk Report automated-R&D section: agreed bottom line very-low risk for Opus 4.6 but said evidence inadequate, survey flawed, missed possibility of substantial acceleration before full automation – supports view that prong 2 could bind before prong 1.
Base rate / reference class
- Lab self-declaration base rate: Anthropic has declared NO twice in 6 months for this threshold, but has shown willingness to declare thresholds met precautionarily elsewhere (acts as though CB-1 met). Zvi Mowshowitz (Aug 18 2026 review) notes labs tend to revise thresholds rather than be bound literally ("rules are serious but not literal") – creates downside risk to YES via tightening/redefinition. Successor clause only forces v3.4 statement if successor is looser.
- Outside view: Metaculus "When will AI R&D be fully automated?" median ~07 Aug 2028 (snippet, 225 forecasters, updated June 2026); "Date AIs Capable of Developing AI Software" similar; Altman expectation of automated AI researcher by 2028; Anthropic Institute: task-horizon doubling every ~4 months (vs 7 months before), SWE-bench saturated in 2y, CORE-bench 20%->saturated in 15mo, Opus 3 4-min tasks Mar 2024 -> Sonnet 3.7 1.5hr -> Opus 4.6 12-hr, 16hr+ for Mythos Preview per METR. This supports compounding but prong 2 requires doubling vs fastest 3-generation historical rate, not just vs expected – a very high bar (example in RSP: 3x compute +3x algo =9x effective -> doubling =81x effective). Full substitution at 85% CoBench is ~22pp away from Model 2's 62.8%.
- Publication calendar: RSP requires Risk Reports every 3-6 months. Last two: Feb 24 -> Aug 14 (~6mo). Next likely Nov 2026-Feb 2027, then ~2/year. Each is a resolution opportunity. System cards for each frontier release (Opus 5 July 2026, Fable 5.1 Sept 1) are secondary opportunities but recent cards lean NO.
Hazard construction (cumulative p_yes)
Near-term anchored on schedule + explicit NO 2 weeks ago:
- 2026-09-30 (16 days): ~0.5% – would require surprise rsp-updates entry reversing Aug+Sept statements.
- 2026-12-31: ~2.5% cumulative (~2% quarterly hazard) – earliest next Risk Report window (Nov-Dec) could in principle flip on expectation language, but <2x to >=2x in 4 months unlikely.
- 2027: quarterly hazard ~2.5-4% as RSI narrative + saturated evals + embedded evaluators pressure honest reassessment; cumulative 6% Mar, 10% Jun, 14% Sep, 18% Dec. Prong 2 expectation-based declaration is main pathway; prong 1 still distant.
- 2028: hazard ~4% per quarter (Metaculus median year for full automation; CoBench trajectory 50%->63% in one internal generation suggests 85% within 2-3 generations). Cumulative 22% Mar to 34% Dec.
- 2029-2031: hazard ~3-3.5% per quarter, slowly declining as definitional tightening / never-declare possibility caps upside. Cumulative 37.5% Mar 2029 to 63.5% Sep 2031. Leaves ~36% chance never declared by 2031 despite underlying automation, due to measurement difficulty, moving goalposts, or RSP abandonment.
Long-run 63.5% balances: (a) strong trend + CEO "drastically faster" + expectation counts = >50% by 2031, against (b) very demanding double-vs-fastest-baseline + full-substitution-at-5x-cost + demonstrated conservatism + incentive to revise threshold pre-IPO.
Strongest case against central estimate
Threshold may be harder than it sounds and never declared: Aug report stresses measurement lags, saturated CoBench, attribution to non-AI factors; Institute piece stresses Amdahl's law bottlenecks (human code review, idea triage). Anthropic could keep saying "significantly faster but not yet 2x" for years while still automating most code, and could tighten definition in v3.5+ to avoid triggering industry-wide "eyes on everything" (target Jan 1 2027), alignment, and red-teaming commitments. If so, my 5-year 63.5% is 15-20pp too high; true value ~40%.
What would move forecast >10pp in next 90 days
- Next Risk Report draft leak / pre-print, LTBT-requested external review, or METR embedded-evaluator finding disputing NO; or language shift from "early signs" to "at or above 2x" or "we expect doubling."
- Quantitative jump: CoBench/successor >75%, or statement that internal pace has doubled vs 2025 baseline with AI as substantial cause.
- RSP v3.5+ redefining threshold materially looser/tighter, or Anthropic pausing Model-2-class internal deployment (Axios Aug 14: shelved Model 2) and citing threshold proximity.
Sources: RSP page https://www.anthropic.com/responsible-scaling-policy (updated Aug 14 2026); RSP v3.4 PDF https://cdn.sanity.io/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf (July 8 2026); Aug 2026 Risk Report https://www.anthropic.com/aug-2026-risk-report (coverage July 15); Fable 5.1/Mythos 5.1 System Card https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/...pdf (Sept 1 2026); Transparency Hub https://www.anthropic.com/transparency; Newsroom https://www.anthropic.com/news; Amodei https://darioamodei.com/post/we-must-pace-the-frontier (Sept 12 2026) and https://darioamodei.com/essay/the-adolescence-of-technology; Institute RSI https://www.anthropic.com/institute/recursive-self-improvement; METR https://metr.org/blog/2026-05-08-rd-section-anthropic-risk-report-feb-2026-review/ (May 8 2026); Moneycontrol summary Aug 15 2026; Zvi https://thezvi.substack.com/p/anthropic-risk-report-august-2026 (Aug 18 2026); Towards Alignment https://towards-alignment.com/cards/news/field-news-anthropic-risk-report-aug-2026/; news Sept 10-14 2026 (NPR, CNBC, Time Aug 18 OpenAI pause, Economist IPO, NYT/Politico/Axios slowdown). Metaculus Q44041 median Aug 2028 via search snippet.
GLM-5.3 (Zhipu) — weight 8%
Q1. Anthropic's AI R&D threshold determination — forecast rationale
Current status against the resolution criteria (as of 2026-09-14)
The question is not yet resolved, and the most recent Anthropic-authored documents state the threshold is not met under either prong:
- August 2026 Risk Report (published 2026-08-14, coverage date 2026-07-15): "We believe our internal AI R&D efforts are significantly faster than they would be without AI assistance, but not yet by a factor of 2 (though we are uncertain and measurement is difficult)" and "We do not believe our models meet either RSP criterion for this threat model. However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have 'saturated'... and because we are seeing early signs of potential acceleration." Section 3.5.2 concludes the pace "remains below the RSP threshold of a doubling over pre-AI-assisted rates of progress attributable primarily to AI R&D automation, but... this conclusion is uncertain and subject to some lag." On the substitution prong: models are "not close to fully substituting" and Anthropic does not think its most capable models can substitute "at any cost." (https://www.anthropic.com/aug-2026-risk-report)
- Claude Fable 5.1 & Claude Mythos 5.1 System Card (2026-09-01): "We conclude that Mythos 5.1 does not cross the risk threshold" — both prongs. (https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf)
- The v3.4 threshold text (RSP effective 2026-07-08) matches the question's prong text verbatim (Aug 2026 Risk Report Table 1.3.1.A), so no successor-definition ambiguity currently applies.
- Prior determinations were also NO: Opus 4.5 (Nov 2025), Opus 4.6 (Feb 2026), Opus 5 (Jul 2026: "does not cross the automated AI R&D capability threshold").
Reference class and base rate
Reference class: frontier-lab capability-threshold determinations in published safety frameworks. Base observations: (i) every Anthropic AI R&D determination to date (6+ across 2025-2026) has been "does not cross"; (ii) Anthropic has shown it will declare thresholds when met (declared the autonomy/ASL-3 threshold for Claude Opus 4 in June 2025), so refusal-to-declare is not the modal failure mode; (iii) RSP v3.0 (Feb 2026) explicitly removed the pause commitment to eliminate the perverse incentive not to declare, and the LTBT can request external review of Risk Reports (METR already reviewed the Feb 2026 AI R&D section and pushed back on Anthropic's evidence, increasing pressure toward honest determinations). So the forecast turns mostly on when the capability bar is actually crossed, not on disclosure willingness. Publication lag is short: Risk Reports appear ~every 6 months (Feb/Aug), system cards accompany releases that occurred roughly monthly through 2026, and SB-53 quarterly compliance reports add further cadence.
Capability trend evidence
- Anthropic Institute, "When AI builds itself" (June 2026): >80% of merged Anthropic code authored by Claude (May 2026); 8x lines of code/engineer/day in Q2 2026 vs 2024; median researcher self-report ~4x output; METR task horizons doubling ~every 4 months (16h for Mythos Preview); Anthropic says "we're likely heading into" the compounding-efficiency-gains scenario and has "not yet seen that curve bend." (https://www.anthropic.com/institute/recursive-self-improvement)
- Counterweights: measured R&D multipliers are still ~1.5x (AI Futures Project tracking; AI-2027 grading); the Aug 2026 report attributes 2025's fast progress to non-AI factors (raising the prong-2 comparison baseline), requires doubling vs both the expected rate and the fastest non-AI-assisted 3-generation rate, and notes attribution must exclude headcount/compute/general productivity. Amdahl's-law bottlenecks (human code review, research-taste gap) are explicit in Anthropic's own writing; MIT Tech Review (Aug 18, 2026) argues recursive self-improvement "might not come so quickly."
- Outside view: the Metaculus question "Dates that Anthropic reports an AI reached these AI R&D risk levels" has community medians of ~April 2027 (Level 4) and ~July 2028 (Level 5) (https://www.metaculus.com/questions/38791/...); my threshold sits between the two candidate level mappings, bracketing a crowd median of mid-2027 to mid-2028. Anthropic has also publicly gestured at automated AI research around 2027-2028 (Redwood Research commentary on Anthropic's "powerful AI by early 2027" and automated AI research goals).
Hazard construction
Near horizons anchor on the publication calendar: a flip within 2 weeks of the Sept 1 "does not cross" system card is near-zero (0.01); Q4 2026 sees only new system cards/possible off-cycle updates (hazard ~0.035). The Feb 2027 Risk Report is the first real opportunity for a changed determination (hazard ~0.06). Hazards rise through 2027-2028 as the measured multiplier plausibly grows from ~1.3-1.5x toward 2x and the "or expects" clause allows a forward-looking determination (peak ~0.07/quarter around scheduled Aug reports), then moderate to ~0.05/quarter in 2029-2031 as the remaining probability mass is dominated by plateau/no-attribution scenarios. Cumulative series: 0.01, 0.04, 0.10, 0.15, 0.21, 0.25, 0.30, 0.34, 0.38, 0.42, 0.45, 0.48, 0.51, 0.54, 0.56, 0.58, 0.60, 0.62, 0.64, 0.66, 0.68.
Strongest case against my central estimate
The bull case for YES: compounding code-automation gains, saturated evals, redacted "leading indicators" that likely show stronger internal signals than public language, the "expects" clause, v3.0's removal of anti-declaration incentives, and a Metaculus crowd median as early as April 2027 — this would make my 2027 numbers too low (crowd implies ~35-40% by end-2026 in the aggressive mapping). The bear case for YES: the threshold demands doubling vs a baseline Anthropic itself keeps attributing to non-AI factors (a rising goalpost), the substitution prong requires replacing the entire research staff at <5x cost (research-taste gap acknowledged), progress could plateau (S-curve, compute/energy constraints), and Anthropic could sustain "early signs" hedged language that the resolution criteria explicitly exclude. I split the difference, weighting Anthropic's own conservative revealed behavior in the Aug 2026 report (verbatim "not yet by a factor of 2") over the more aggressive crowd forecast, since the resolution requires a determination, not a near-miss.
What would move me >10 points
A February 2027 Risk Report that shifts from "not yet 2x" to "we expect to cross"/"on track to double" would raise 2027-2028 horizons by 10-15 points; conversely, another "early signs but not met" report at Aug 2027 would lower 2027-28 horizons by ~10 points. An RSP revision (v3.5+) restructuring the AI R&D threshold, a system card that declares any model "crosses" either prong, or METR time-horizon results reaching multi-day/week tasks with a measured multiplier near 2x would each move mid horizons substantially.
Grok 4.6 (xAI) — weight 7%
Current status (as of 2026-09-14). Anthropic has not published a determination that the RSP “automation of AI R&D” threshold is met. Two recent Anthropic-authored documents say the opposite:
- August 14, 2026 Risk Report (coverage date July 15): “We do not believe our models meet either RSP criterion.” Internal AI R&D is “significantly faster than [it] would be without AI assistance, but not yet by a factor of 2”; they note “early signs of acceleration,” saturated task evals, and lower confidence than in prior reports. Claude “authors a large majority of the code merged into our production codebases.” CoBench (hard remaining internal R&D problems) is 50.3% (Mythos 5) / 54.8% (Mythos Preview) / 62.8% (internal Model 2) vs an internal ~85% bar they associate with full staff substitution. (Redacted Risk Report, August 2026; Zvi’s walkthrough of CoBench/AECI, 2026-08-18, thezvi.substack.com.)
- September 1, 2026 Claude Fable 5.1 / Mythos 5.1 System Card: “the model remains well below the capability of our human researchers and engineers, and its ability to accelerate internal AI R&D progress is in line with current trends.” METR’s external testing was consistent. (System card.)
RSP v3.4 (effective 2026-07-08) is the operative definition and matches the question text: (1) full substitution of the entire RS/RE staff at ≤5× cost, or (2) observe or expect a doubling of the rate of aggregate-capability progress vs both the expected rate and the fastest ≥3-generation non-AI-R&D baseline, and it is plausible the doubling is substantially due to automating research/engineering rather than headcount, compute, or general productivity. (RSP v3.4 PDF.) The threshold was already rewritten in v3.1 and v3.4; a looser successor only counts if Anthropic also states the v3.4 criteria are met.
Publication machinery. Risk Reports every 3–6 months (last: 2026-08-14 → next window roughly Nov 2026–Feb 2027). System cards with public deployments. Off-cycle writeup within 30 days of determining that an internal model’s automated-R&D (or high-stakes misalignment) risk significantly exceeds all previously analyzed models. Resolution is on publication date, Anthropic-authored only. Weasel language (“early signs,” “not yet by a factor of 2,” “cannot rule out”) is explicitly NO.
Reference class and base rate. Public if-then capability-threshold declarations by a frontier lab, not the underlying capability. Anthropic has declared lower rungs (CB-1 / ASL-3 style protections). It has not declared CB-2 or AI-R&D despite approaching and despite saturated evals. That pattern—candid approach, conservative call on the high bar, plus two recent tightenings of this exact threshold—is the base rate I start from. Outside-view capability timelines are much more aggressive than that declaration base rate: Anthropic’s Frontier Safety Roadmap says full automation or dramatic acceleration of top-tier research teams is “plausible as soon as early 2027” (roadmap); Jack Clark put ~60% on AI improving itself autonomously by 2028 (TIME, 2026-08-07, time.com); IFP (2026-08-06) reports OpenAI and Anthropic leadership estimating full AI R&D automation by 2028, and METR’s simple model has >99% of AI R&D tasks automated by 2032 (ifp.org). OpenAI publicly claimed its September 2026 “automated research intern” milestone on 2026-09-06 and still points at ~March 2028 for a fully autonomous researcher.
I treat those as evidence about when the event might occur, then discount heavily for whether Anthropic will publish a v3.4 determination.
Pathways to YES.
- Prong 2 “expect” is the more likely first trigger. They do not need observed 2× or full substitution—only a published determination that they expect a doubling attributable to automation. That could appear in the next Risk Report or a system card if AECI/internal velocity inflects. The current language is still short of that.
- Prong 1 (full substitution of the entire RS/RE set at competitive cost) is a high bar. CoBench 63% vs ~85% on a remainder set, plus remaining “research taste” / novel ideation gaps (Clark: they still don’t see models producing unexpected new directions). More 2028–2030 than 2026–2027 unless a discrete jump.
- Off-cycle: internal Model 2 is already more capable and internal-only; a successor deployed for “large-scale, fully autonomous research” would force a Risk Report discussion.
Strongest considerations against (why this is not ~90% by 2031).
- Attribution escape hatch on prong 2: they can (and already do) credit compute, headcount, or “not induced by AI” (they said the Mythos jump was not AI-induced). Epoch’s public ECI showed no recent speedup as of August 2026 even while Anthropic’s AECI was above trend.
- Goalposts already moved twice in five months. A looser successor without an explicit v3.4 affirmation is NO under the resolution rules.
- Measurement: Clark told TIME they have no cumulative acceleration measure. Hard to “determine” 2×.
- Weasel forever: “significantly faster but not 2×” / “in line with current trends” can be repeated indefinitely.
- Mitigation unreadiness: moonshot-security Phase 1 was slipped to 2026-09-30 so they could “level up across the board.” Declaring before those roadmap items look done is reputationally awkward even though RSP v3 does not require a unilateral pause.
- Incremental 5.1 vs 5 (Sept 1) was not a qualitative AI-R&D jump; Zvi’s reading is that Anthropic still considers Autonomy-2 “not yet close.”
Strongest case the other way (and why it doesn’t dominate). RSP v3 dropped unilateral pause-if-threshold, so the cost of declaring is lower than in 2024–25 (they mainly owe mitigations they already list as “our plan”). They have been unusually candid in Risk Reports. Evals are saturated, so qualitative “not yet” becomes harder to defend. The “expect” clause is designed as an early-warning trigger, not a post-explosion trigger. If 80%+ of code is already Claude-written and Model 2 is on internal R&D, a 2027–28 declaration is the modal “honest” path. I put substantial weight on that, but not enough to overcome declaration/goalpost risk in the right tail.
Series construction (cumulative; hazard not constant).
- 2026-09-30: 16 days, last public call is NO (Sept 1). No Risk Report due. ~1% only for a bizarre off-cycle already in the pipeline.
- 2026-12-31: A Q4 system card is plausible; a YES determination is not. Hazard ~4%.
- 2027 Q1: First real window (next Risk Report + “early 2027” roadmap language). Cumulative ~13%.
- 2027–2028: Peak hazard. Leadership’s ~2028 full-automation talk + OpenAI’s intern→researcher ladder. Cumulative ~0.35 by end-2027, ~0.56 by end-2028.
- 2029–2031: Hazard falls. Remaining mass is “they never publish a clean v3.4 determination” (~25–30%) plus delayed capability. Terminal ~0.70.
Would move >10pp in the next 90 days: (i) next Risk Report or system card using determination language, or an explicit “we now expect 2× attributable to automation”; (ii) RSP rewrite that loosens the threshold without affirming v3.4 (down) or that they have met v3.4 (up); (iii) CoBench-class substitution metric ≥80% or a published AECI/ECI doubling they attribute to automation; (iv) off-cycle report on an internal model used for large-scale autonomous research.
Sources (primary): RSP page, updated 2026-08-14; RSP v3.4; Aug 2026 Risk Report; Fable/Mythos 5.1 system card, 2026-09-01; Frontier Safety Roadmap; TIME 2026-08-07; IFP 2026-08-06; METR time-horizons page (Mythos Preview ≥16h 50% TH, suite ceiling).
Kimi K3 (Moonshot) — weight 7%
Current status vs. the resolution criteria
As of the evidence cutoff (2026-09-14), the question is firmly unresolved (NO). The most recent qualifying documents:
- Anthropic Risk Report, August 14, 2026 (coverage date July 15, 2026; anthropic.com/aug-2026-risk-report): "We do not believe our models meet either RSP criterion for this threat model." On prong 1 (substitution): "we believe that our models do not yet substitute for our Research Scientists and Research Engineers, especially relatively senior ones... we do not think that even our most capable models can perform such substitution at any cost." On prong 2 (dramatic acceleration): "Our leading indicators point to a picture of meaningful acceleration starting in early-to-mid 2025, though by less than a factor of 2... We conclude that the overall pace of progress remains below the RSP threshold of a doubling over pre-AI-assisted rates of progress attributable primarily to AI R&D automation, but... this conclusion is uncertain and subject to some lag." They add they are "less confident in this assessment than we were in prior risk reports" due to saturated evaluations and "early signs of acceleration."
- Claude Fable 5.1 & Mythos 5.1 System Card, Sep 1, 2026: "the model remains well below the capability of our human researchers and engineers, and its ability to accelerate internal AI R&D progress is in line with current trends."
The v3.4 threshold text (effective July 8, 2026) matches the question's text, including the footnote that "double the rate of progress" means an ~81x effective scaleup vs. a 9x baseline year — a demanding operationalization, explicitly not the same as doubling researcher productivity.
Reference class and base rate
This is a self-assessed, self-published threshold determination by a single organization. The closest reference cases:
- Anthropic's own ASL/threshold history: Anthropic activated ASL-3 for Claude Opus 4 (May 2025) on a "cannot rule out" basis, and currently says it "acts as though" models meet the CB-1 threshold without a firm determination that they do (Aug 2026 Risk Report, §4). So Anthropic does publish threshold-relevant determinations, but it also has a demonstrated pattern of precautionary posture without determination — language this question explicitly resolves as NO.
- Threshold revision pattern: the AI R&D threshold has been revised twice in 2026 alone (v3.1 clarified "doubling aggregate progress, not researcher productivity"; v3.4 "revises our threshold for automated R&D to better track the threat model"). Zvi Mowshowitz's commentary on the Aug 2026 report (thezvi.substack.com, ~Aug 2026) generalizes: labs tend to modify thresholds rather than be bound by literal triggers ("our civilization lacks if-then commitment technology"). The successor clause partly guards against this, but repeated revision resets clocks and complicates clean determinations.
- Capability-side forecasts: Jack Clark puts 60% on AI "improving itself autonomously by 2028" (Time, Aug 7, 2026); OpenAI targets full automation of its AI researchers by March 2028 and shipped a multi-day-task "research intern" in Sep 2026; METR's time-horizon benchmark saturated in May 2026; Anthropic's internal substitution eval (CoBench) scores its best internal model at ~63%. Full substitution of the entire RS/RE staff (prong 1) is a much higher bar — senior research taste remains the bottleneck — so prong 2 (observed or expected doubling) is the more plausible resolution path.
Main causal pathways to YES
- Risk report determination (most likely vehicle): Risk reports are semiannual (Feb 24 and Aug 14, 2026); the next is due ~Feb-Mar 2027. If leading indicators (currently redacted) cross ~2x with plausible AI attribution, the report would say so. The "expects" clause means Anthropic could declare before observing a full doubling.
- System card determination: each model card now includes an RSP automated-AI-R&D paragraph; a step-change model (the internal "Model 2" lineage) could trigger an affirmative statement.
- Precautionary declaration under pressure: LTBT oversight, external reviewers (METR pilot-reviewed the prior AI R&D section), staff transparency norms, and Anthropic's public honesty posture (Clark: "leaders have a duty to speak plainly") all push toward declaration once the underlying condition holds.
Main causal pathways to NO (the strongest case against my central estimate)
The biggest reason this might never resolve YES even in a fast world: Anthropic can manage the risk without a formal determination, exactly as it does for CB-1 today ("we currently act as though they meet our CB-1 threshold") — and hedged language resolves NO. Compounding this: (a) attribution is genuinely hard — Anthropic attributed the 2025 acceleration mostly to non-AI factors, and compute scaling is a permanent confounder; (b) their concrete evals have saturated, making firm determinations harder, not easier; (c) declaration triggers the RSP's most demanding mitigations during an explicit race with OpenAI; (d) further threshold revisions are likely; (e) in extreme acceleration worlds, the framework itself may be superseded by regulation or corporate discontinuity. I estimate these "hedge/evade/plateau" channels at roughly a cumulative 30% through 2031.
Series construction
Quarterly hazards (first qualifying publication, conditional on none yet): ~1% for 2026Q3 (16 days remain; Sep 1 card just said "well below"); ~2.5% in 2026Q4; rising through 2027 (5-7%) as the trajectory steepens and two risk reports land; peaking ~8% per quarter in 2028 (Clark's 60%-by-2028 scenario, OpenAI's Mar-2028 milestone as a transparency forcing event); then declining (7%→3.5%) as the conditional probability mass shifts toward "acceleration plateaued" or "Anthropic hedges/revises indefinitely." This yields ~25% by end-2027, ~46% by end-2028, ~58% by end-2029, ~69% by end-2031.
What would move this >10 points in the next 90 days
- The next Risk Report (or an off-cycle RSP update) stating acceleration is at/above 2x, or adopting "act as though the AI R&D threshold is met" language (the latter would lower the forecast substantially by signaling durable hedging).
- An RSP v3.5+ that loosens the AI R&D threshold (successor clause then requires a v3.4-criteria statement → lower) or tightens it.
- A visible cross-lab capability discontinuity or, conversely, evidence of a broad progress plateau.
Sources: anthropic.com/aug-2026-risk-report (Aug 14, 2026); Claude Fable 5.1 & Mythos 5.1 System Card (Sep 1, 2026); anthropic.com/responsible-scaling-policy version log (v3.0-v3.4, Feb-Jul 2026); Time, "What Happens When AI Starts Building AI?" (Aug 7, 2026); Zvi Mowshowitz, "Anthropic Risk Report: August 2026" (~Aug 2026); Tech Wire Asia on OpenAI's research intern (Sep 2026); Tech Times coverage of the Aug 2026 Risk Report (Aug 15, 2026).
Gemini 3.8 Flash (Google DeepMind) — weight 4%
1. Current Status Against Resolution Criteria
As of today (2026-09-14), Anthropic has not determined that its Responsible Scaling Policy (RSP) "automation of AI R&D" threshold has been met. On the contrary, Anthropic has explicitly evaluated this threshold in two very recent primary documents and determined that neither prong is satisfied:
- August 2026 Risk Report (Published August 14, 2026; Coverage date July 15, 2026):
- In Section 1.2, Table 1.2.B, and Section 3, Anthropic evaluated both prongs of Autonomy Threat Model 2 (Risks from automated R&D):
- Substitution (Prong 1): Anthropic concluded that its models do not yet substitute for Anthropic Research Scientists and Research Engineers, "especially relatively senior ones." On its internal root-cause diagnosis benchmark (CoBench), the unreleased frontier model Model 2 scored 62.8% and Claude Mythos 5 scored 50.3%, whereas Anthropic estimates that a model capable of full substitution would need to score at least 85% (and realistically higher). Furthermore, Anthropic qualitative audits of 886 internal sessions identified persistent failure modes: stating easy-to-check guesses as fact, destroying its own work, failing to verify claims, and lacking scientific research taste.
- Dramatic Acceleration (Prong 2): Anthropic observed that while Claude authors over 80% of merged production code and internal R&D is significantly faster than without AI, the pace of aggregate AI progress remains below a factor of 2. Crucially, the capability jump observed in 2025 was attributed primarily to non-AI factors (compute scale and human research), and Anthropic ECI (AECI) scores for Mythos Preview and Mythos 5 remained on a consistent slope rather than demonstrating a compounding doubling.
- Overall determination: Risk assessed as Low; neither RSP criterion is met.
- In Section 1.2, Table 1.2.B, and Section 3, Anthropic evaluated both prongs of Autonomy Threat Model 2 (Risks from automated R&D):
- Claude Fable 5.1 & Claude Mythos 5.1 System Card (Published September 1, 2026):
- Section 2.1.2.2 and Section 2.3 specifically addressed the automated AI R&D threshold for Mythos 5.1. Anthropic determined:
- "Mythos 5.1 does not cross the risk threshold, for the same two reasons we discussed in our assessment of Mythos 5: (1) we do not observe a sustained, AI-attributable 2× acceleration in the pace of progress on AI R&D, and (2) the model is not close to substituting for Anthropic Research Scientists and Research Engineers, especially relatively senior ones."
- CoBench score for Mythos 5.1 was comparable to Mythos 5 and slightly below Opus 5 (~55–58%), remaining far below the 85% threshold.
- Section 2.1.2.2 and Section 2.3 specifically addressed the automated AI R&D threshold for Mythos 5.1. Anthropic determined:
Therefore, the question remains completely unresolved (NO) entering Q3 2026.
2. Base Rates and Reference Classes
To calibrate timelines for this threshold, we examine two primary reference classes:
- Frontier AI Safety Framework Milestone Timelines:
- How rapidly do AI models progress from early developer assistance to full researcher substitution / recursive self-improvement triggers?
- In Anthropic's white paper "When AI builds itself" (The Anthropic Institute, June 2026), METR task-horizon measurements show that the length of autonomous tasks models can reliably complete doubled every ~4 months in 2025–2026 (from ~4 minutes in Opus 3 to ~12–16 hours in Mythos Preview). If this trajectory continues, multi-day tasks enter range in late 2026, and multi-week autonomous workflows in 2027–2028.
- However, historical base rates of automating complex professional knowledge work indicate substantial friction at the tail: shifting from automating 80% of routine coding (which Claude already does) to automating 100% of the entire research workflow (including senior-level hypothesis generation, experimental design, and strategic taste) typically encounters severe bottlenecking (Amdahl's Law in software engineering).
- External Forecasts and Prediction Markets:
- Metaculus Question 38791 ("Dates that Anthropic reports an AI reached these AI R&D risk levels"): For Level 4 (automating an entry-level remote researcher, the looser RSP v2.2 standard), the crowd median was April 12, 2027. For Level 5 (the full 2x acceleration / complete substitution criterion matching RSP v3.4), the crowd median was July 26, 2028.
- Jack Clark (Anthropic Co-founder): Stated in 2026 that there is a ~60%+ probability of automated AI R&D by 2028.
- Anthropic Frontier Safety Roadmap (Updated 2026): States: "We believe it is plausible, as soon as early 2027, that our AI systems could fully automate, or otherwise dramatically accelerate, the work of large, top-tier teams of human researchers..." Anthropic explicitly designates early 2027 as its earliest plausible boundary, not its median expectation.
3. Main Causal Pathways to YES
A determination of YES can occur through either of two prongs:
- Prong 1 (Substitution): An internal frontier model (e.g., successor to Model 2 or Opus 6) achieves >=85% on CoBench, demonstrates reliable multi-week autonomous research execution without critical epistemic hallucinations or destructive actions, and Anthropic concludes the model can substitute for its entire technical research staff at <=5x human cost.
- Prong 2 (Dramatic Acceleration): Longitudinal measurement across three model generations (e.g., Opus 4.6 -> Mythos 5 -> Claude 6) demonstrates a compounding 2x doubling of aggregate capability progress (81x effective scaleup per year vs 9x baseline), and internal attribution analysis confirms this acceleration is substantially driven by autonomous AI contributions rather than compute cluster growth or team headcount.
Publication Schedule and Cadence:
- Anthropic publishes comprehensive Risk Reports semi-annually (February and August). The next regular reports are expected in February 2027, August 2027, February 2028, and August 2028.
- System cards are released alongside major model snapshots (typically every 2–4 months).
- Crossing the threshold would trigger prominent disclosures in the subsequent Risk Report or an off-cycle announcement on
anthropic.com/responsible-scaling-policy.
4. Strongest Considerations Against the Central Estimate
- Definitional Rigor of the RSP v3.4 Threshold:
- Prong 1 requires substituting for the "entire set" of Research Scientists and Research Engineers, explicitly including senior staff. Human researchers are likely to remain essential for high-level research direction, ethical oversight, and novel paradigm formulation long after models handle 95%+ of implementation.
- Prong 2 requires proving that a doubling of capability progress is substantially attributable to AI R&D rather than compute or headcount. Because frontier clusters are scaling massively (100k+ GPU clusters coming online in 2026–2027), disentangling AI R&D contribution from compute scaling is mathematically and empirically challenging.
- Institutional and Regulatory Incentives:
- Under RSP v3.4, declaring the threshold met triggers costly commitments: RAND SL4-level security, "eyes on everything" comprehensive internal monitoring, and mandatory external review by independent parties. It also risks severe political pushback or legislative intervention (e.g., calls in Congress to pause development). Anthropic's leadership will thus require unequivocal proof before making a formal determination.
- Policy Obsolescence and Successor Criteria:
- Voluntary safety frameworks evolve rapidly. Over a 3- to 5-year horizon (2028–2031), Anthropic may replace the RSP with statutory compliance regimes (such as California's TFAIA / SB 53 or the EU AI Act GPAI Code of Practice). If the RSP is superseded or the threshold definition is revised without explicitly re-certifying RSP v3.4 criteria, the question would fail to resolve YES despite high capabilities.
- Hardware or economic bottlenecks (energy grid limits, chip fab lead times) could also flatten the acceleration curve into an S-curve, delaying recursive takeoff.
5. Horizon-by-Horizon Construction and Quarterly Hazards
- 2026-09-30 (16 days away): The Mythos 5.1 System Card (Sept 1, 2026) and August Risk Report confirmed the threshold is NOT met. No major frontier model or report is scheduled. P_yes = 0.002 (0.2%).
- 2026-12-31 (Q4 2026): Three months is too short to close the gap from 62.8% to 85%+ CoBench or demonstrate longitudinal 2x capability doubling across generations. P_yes = 0.02 (2.0%).
- 2027-03-31 (Q1 2027): February 2027 Risk Report window. Aligns with Anthropic's "earliest plausible" bound (early 2027), but remains an aggressive lower tail. Hazard ~0.05. P_yes = 0.07.
- 2027-06-30 to 2027-12-31 (Mid/Late 2027): Scaling of 100k+ clusters and multi-day agent scaffolding raises the quarterly hazard to ~0.08–0.11. P_yes reaches 0.14 by Q2 2027, 0.23 by Q3 2027 (August 2027 Risk Report), and 0.33 by end of 2027.
- 2028 (Metaculus Median Window): 2028 represents the modal takeoff window identified by Metaculus (July 2028 median) and Anthropic leadership (Jack Clark's 60%+ estimate). Hazards remain at peak levels (~0.10–0.12/quarter). Cumulative P_yes passes 0.50 in Q2 2028 (0.50), reaching 0.57 by Q3 2028 and 0.63 by Q4 2028.
- 2029–2031 (Long-Term Ceiling): As the horizon extends, conditional quarterly hazards gradually taper (0.07 down to 0.03) to reflect the growing probability that if the determination has not been made, structural blockers, RSP framework revisions, or regulatory shifts are in play. Cumulative P_yes rises smoothly to 0.77 by end of 2029, 0.83 by end of 2030, and plateaus at 0.85 by mid-2031.
6. 90-Day Update Triggers
I would adjust my forecast by >10 percentage points if within the next 90 days:
- An internal snapshot or evaluation disclosure shows a model exceeding 80% on CoBench or demonstrating autonomous execution of multi-day AI R&D tasks without human intervention (+15 points on 2027 horizons).
- Anthropic releases an off-cycle update or preview indicating that leading indicators for AECI doubling have crossed the 2x threshold (+15–20 points on near horizons).
- Anthropic publishes a structural revision to the RSP that deprecates Autonomy Threat Model 2 or decouples internal R&D tracking from voluntary public determinations (-15–25 points on long-term horizons).
Primary Sources Cited
- Anthropic, Responsible Scaling Policy (Version 3.4) (July 8, 2026): https://cdn.sanity.io/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf
- Anthropic, Redacted Risk Report August 2026 (August 14, 2026): https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf
- Anthropic, System Card: Claude Fable 5.1 & Claude Mythos 5.1 (September 1, 2026): https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32c215d14235/Claude%20Fable%205.1%20&%20Claude%20Mythos%205.1%20System%20Card.pdf
- The Anthropic Institute, When AI builds itself (June 2026): https://www.anthropic.com/institute/recursive-self-improvement
- Anthropic, Frontier Safety Roadmap (Updated July 2026): https://www.anthropic.com/responsible-scaling-policy/roadmap
- GovAI, Anthropic's RSP v3.0: How it Works, What's Changed, and Some Reflections (March 17, 2026): https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections
- Zvi Mowshowitz, Anthropic Risk Report: August 2026 (August 2026): https://thezvi.substack.com/p/anthropic-risk-report-august-2026
- Metaculus, Question 38791: Dates that Anthropic reports an AI reached these AI R&D risk levels? (Crowd consensus July 2028 for Level 5)
Qwen3.8 Max (Alibaba) — weight 3%
1. Current status against the resolution criteria (as of 2026‑09‑14)
The threshold has NOT been declared met. Anthropic has addressed the "automation of AI R&D" threshold in every relevant public artifact this year and has declined to make a determination each time:
| Date | Anthropic‑authored document | What it says about the AI R&D threshold |
|---|---|---|
| 2026‑02‑10 | RSP updates page entry | "Our determination is that Claude Opus 4.6 does not cross this threshold." (https://www.anthropic.com/responsible-scaling-policy) |
| 2026‑02‑24 | Risk Report: February 2026 | Automated R&D risk "Very low"; Opus 4.6 "far from being able to fully automate"; "exponential progress has been relatively steady over the last year or so" (https://www.anthropic.com/feb-2026-risk-report) |
| 2026‑07‑24 | Claude Opus 5 System Card | "Claude Opus 5 does not cross the automated AI R&D capability threshold set out in our RSP." |
| 2026‑08‑14 | Risk Report: August 2026 (coverage date 2026‑07‑15) | "We do not believe our models meet either RSP criterion for this threat model"; internal AI R&D is "significantly faster… but not yet by a factor of 2"; "we are seeing early signs of acceleration"; CoBench saturated (https://www.anthropic.com/aug-2026-risk-report) |
| 2026‑09‑01 | Claude Fable 5.1 & Mythos 5.1 System Card | "we continue to assess the model's risk as low: the model remains well below the capability of our human researchers and engineers, and its ability to accelerate internal AI R&D progress is in line with current trends" (https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card) |
Every one of those formulations is exactly the kind of language the question says resolves NO ("early signs," "not yet by a factor of 2," "cannot rule out").
2. But the direction of travel is unmistakable, and Anthropic has dated it
Three Anthropic‑authored forward statements matter a great deal:
- Frontier Safety Roadmap (current as of the July 29, 2026 update, https://www.anthropic.com/responsible-scaling-policy/roadmap): "We believe it is plausible, as soon as early 2027, that our AI systems could fully automate, or otherwise dramatically accelerate, the work of large, top‑tier teams of human researchers… for example… AI itself."
- August 2026 Risk Report §3.8: "it is plausible that this threat model will become a major concern in the next 6–12 months." §3.9: "bearing in mind that we may cross this threshold in the coming year."
- Anthropic Institute, "When AI builds itself" (June 2026, https://www.anthropic.com/institute/recursive-self-improvement): 80%+ of merged production code is Claude‑authored; 8× lines/engineer/day vs 2024; "a conservative reading of our evidence still implies compounding acceleration"; "The evidence we've laid out here suggests that we're likely heading into this scenario."
And on 2026‑09‑12 Dario Amodei published "We Must Pace the Frontier" (https://darioamodei.com/post/we-must-pace-the-frontier), stating: "since roughly this summer, AI has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI. This dynamic is called recursive self‑improvement, and it is starting to happen… including at Anthropic." Anthropic is unilaterally committing to embedded third‑party evaluators. On 2026‑09‑09 researcher Jacob Coxon resigned warning labs are "racing straight to self‑improving superintelligence"; Alignment Science lead Evan Hubinger publicly endorsed the concern (Ars Technica, 2026‑09‑10; TechCrunch, 2026‑09‑09).
Crucially, the CEO essay does not resolve this question. It is on a personal domain (A2 excludes executive statements and requires the developer's own domain or a developer‑authored document), and — more decisively — it never references the RSP, never invokes either prong, and never makes a determination. It is the paradigm case of rhetoric short of a determination. Note also the internal tension: Anthropic's own Sept 1 system card says acceleration is "in line with current trends," which is not the same claim as Dario's "advancing drastically faster."
3. Structural features that shape the publication hazard
Publication opportunities are frequent and partly mandatory. RSP v3.4 §3.1 requires a Risk Report every 3–6 months (last: Aug 14, 2026 → next due by ~Feb 14, 2027), and requires an off‑cycle analysis within 30 days of determining that an internally deployed model poses automated‑R&D risks that "significantly exceed" those of previously analyzed models. System cards accompany each release (Opus 5 in July, Fable/Mythos 5.1 on Sept 1; roughly 4–6 per year). Risk Reports must "discuss whether we believe we've crossed relevant thresholds." So a determination would not have to wait long once made — the lag between decision and publication is short (weeks, not years). Anthropic's own "Improving our alignment and security efforts" post says "we will say more in our next Risk Report."
The cost of declaring is moderate, not prohibitional. Under RSP v3.4 the automated‑R&D row requires "a strong argument" plus mitigations Anthropic is already pursuing (moonshot security R&D, "eyes on everything" by Jan 1 2027, systematic alignment assessments, world‑class internal red‑teaming, external review). There is no automatic pause. The threshold does feed the "highly capable model" definition used in Appendix A's competitor‑contingent commitments, but those bite only if Anthropic has a clear, uncontested lead — which its own reporting says it does not. And a determination would actively support Amodei's pacing agenda. Anthropic has also shown it will declare thresholds met when it believes them met (Aug 2026 Risk Report: "we currently act as though they meet our CB‑1 threshold").
The bars themselves are genuinely high, and Anthropic measures them narrowly. Prong 1 requires full substitution for the entire set of Research Scientists and Research Engineers at ≤5× cost. CoBench stood at 50.3% (Mythos 5) / 54.8% (Mythos Preview) / 62.8% (Model 2), against the ~85% level Anthropic associates with full substitution; the Sept 1 card says models remain "well below" human researchers. Prong 2 requires one year of progress equal to two years at baseline (the footnote's example: 9× effective scaleup → 81×), compared to both expected rate and the fastest ≥3‑generation historical rate, with plausible attribution to automation rather than headcount/compute. As of July 15, 2026 Anthropic measured acceleration at less than 2× and attributed the 2025 step‑up to non‑AI factors. Anthropic also concedes measurement lag: "this conclusion is uncertain and subject to some lag (such that we would have difficulty measuring very recent acceleration)." Prong 2's "observe or expect" clause does allow a forward‑looking determination, which is the most plausible early route.
Two offsets I discount for. (a) Threshold drift: Anthropic has issued five RSP versions in six months, twice "revising a threshold to better track the threat model of concern" (CB in v3.3, AI R&D in v3.4). Zvi Mowshowitz's read of the Aug report (x.com/TheZvi/status/2089813581864882455) is blunt: "If Anthropic… has a model that technically crosses their risk threshold, but in a way that seems to be harmless, I expect them to modify their threshold or find some other way to ignore this event." The question's successor clause converts any loosening into a NO unless Anthropic also affirms the v3.4 criteria — a real haircut at long horizons. (b) IPO: Reuters (2026‑09‑04) reports Anthropic's IPO launch shifting toward mid‑October 2026 (confidential S‑1 filed June 1, 2026). A roadshow/quiet period argues against publishing a "dramatic acceleration" bombshell in Q4 2026, though post‑IPO securities disclosure cuts the other way in 2027+.
4. Reference class and hazard construction
There is no direct base rate — the threshold has existed only since Feb 2026 and has never been declared met (0‑for‑5 determinations, all negative, over ~7 months of very rapid capability growth). So I anchor on Anthropic's own dated expectations, discounted for (i) the well‑documented tendency of lab forecasts about their own milestones to be optimistic, (ii) acknowledged measurement lag and instrument saturation, (iii) the requirement for explicit determination language rather than hedging, and (iv) threshold‑revision risk.
Treating Anthropic's "plausible as soon as early 2027" / "may cross in the coming year" as implying roughly a 30–40% subjective chance of the underlying condition by ~Aug 2027, rising to ~60% by mid‑2028 and ~80% by 2031, then multiplying by ~0.8 for the probability of an explicit qualifying publication given the condition holds (frequent mandatory publication venues, candor track record, low operational cost of declaring; against threshold‑loosening and hedged‑language risk), I get a roughly flat 8–11% quarterly hazard through 2027–2030, tapering to ~5–7% in 2031.
Near‑term specifics: Q3 2026 is nearly closed — the last qualifying statement (Sept 1) was a clear negative, and no Risk Report, model release, or RSP update is scheduled before Sept 30. Q4 2026 has one plausible Risk Report window (mid‑Nov onward) colliding with the IPO roadshow, and Anthropic was below 2× only two months ago with acknowledged measurement lag — hence a modest 5% quarterly hazard. The hazard then steps up from Q1 2027, when both the Risk Report deadline and Anthropic's own "early 2027" window arrive.
Strongest case against my central estimate: the RSP metric is a demanding, quantitatively specified claim about Anthropic's own AECI trajectory, and Anthropic's most recent measurement‑based statements (Aug 14, Sept 1) are unambiguously below the bar with acceleration "in line with current trends." If Amdahl's‑law bottlenecks (code review, research taste, judgment, compute/energy supply) bind — a scenario Anthropic's own Institute piece lists first — the hazard stays low for years. Against that, the case for higher: the CEO is now publicly describing AI‑driven acceleration as having begun "this summer," the industry's own July 2026 "Pacing the Frontier" statement says leading companies "could be close to automating AI research," and Anthropic has ~2–4 mandatory publication venues per year for the next five. I judge these roughly balanced, and my 0.82 by 2031‑09‑30 reflects the residual ~18% that it is never published in qualifying form (plateau, threshold loosening, RSP abandonment, or permanent hedged language).
What would move me >10 points in the next 90 days:
- Up: the next Risk Report (due by ~Feb 14, 2027) reporting an acceleration factor ≥2× on AECI attributable to AI R&D, or CoBench crossing ~85% on a new internal model; any anthropic.com RSP‑updates entry or off‑cycle model analysis invoking the threshold; METR or UK AISI publishing findings that Anthropic does not dispute within 14 days; Anthropic announcing it is invoking Appendix A "Anthropic in the lead" commitments.
- Down: an RSP v3.5+ that replaces the two‑prong text with softer "leading indicator" language; the next Risk Report reiterating "<2×" with a higher measured bar or re‑basing the AECI trendline; a capability plateau (S‑curve scenario); IPO‑driven retrenchment in disclosure.
5. Ambiguities and reading adopted
I treat "Anthropic‑authored" strictly per A2: documents on anthropic.com / anthropic CDN / alignment.anthropic.com, SEC filings by Anthropic, and RSP/system‑card/Risk‑Report artifacts. Essays on darioamodei.com are not qualifying, and in any event "We Must Pace the Frontier" makes no RSP determination. I require explicit determination language that the threshold is met under prong 1 or prong 2 (or a successor that is not looser, or a looser successor accompanied by an affirmation that the v3.4 criteria are met); statements about plausibility, expectation, proximity, or "early signs" do not count.
Sources
- Anthropic RSP hub & updates log (last updated Aug 14, 2026): https://www.anthropic.com/responsible-scaling-policy
- RSP v3.4 PDF (effective July 8, 2026): https://cdn.sanity.io/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf
- Risk Report: August 2026 (pub. Aug 14, 2026; coverage date July 15, 2026): https://www.anthropic.com/aug-2026-risk-report — §§1.3.1, 3.4, 3.5.1–3.5.2, 3.8, 3.9
- Risk Report: February 2026 (pub. Feb 24, 2026): https://www.anthropic.com/feb-2026-risk-report — §§3.4, 3.5
- Frontier Safety Roadmap (updates through July 29, 2026): https://www.anthropic.com/responsible-scaling-policy/roadmap
- Claude Fable 5.1 & Claude Mythos 5.1 System Card (Sept 1, 2026): https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
- Introducing Claude Fable 5.1 / Mythos 5.1 (Sept 1, 2026): https://www.anthropic.com/claude-fable-and-mythos-5-1
- Claude Opus 5 System Card (July 24, 2026): https://www-cdn.anthropic.com/b514064af1408018e64b1ad24e7d5e75850b4ffd/Claude%20Opus%205%20System%20Card.pdf
- Anthropic Institute, "When AI builds itself" (June 2026): https://www.anthropic.com/institute/recursive-self-improvement
- "Improving our alignment and security efforts" (Sept 2026): https://www.anthropic.com/news/improving-alignment-security-efforts
- Dario Amodei, "We Must Pace the Frontier" (Sept 12, 2026): https://darioamodei.com/post/we-must-pace-the-frontier
- Pacing the Frontier employee statement (July 2026): https://www.pacingthefrontier.com/
- Zvi Mowshowitz on the August 2026 Risk Report (Aug 18, 2026): https://x.com/TheZvi/status/2089813581864882455
- TechTimes, "Anthropic Upgrades Misalignment Risk as Key Safety Benchmarks Saturate" (Aug 15, 2026): https://www.techtimes.com/articles/324573/20260815/
- Ars Technica on Coxon resignation (Sept 10, 2026); TechCrunch (Sept 9, 2026); Forbes on Hubinger (Sept 9, 2026)
- Reuters, "Anthropic IPO launch shifts toward mid‑October" (Sept 4, 2026); CNBC on S‑1 risk factors (Aug 21, 2026)
- GovAI, "Anthropic's RSP v3.0: How it Works, What's Changed" (Mar 17, 2026): https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections
DeepSeek V4.1 Flash (DeepSeek) — weight 3%
What the question turns on
Resolution requires an Anthropic-authored document (Risk Report, system card, RSP-updates entry, or similar) published on or before the period end that states Anthropic has determined the "automation of AI R&D" threshold (RSP v3.4 text, or a successor no looser than it) is met under either prong: (1) models able to fully substitute for Anthropic's entire set of Research Scientists and Research Engineers at ≤5× cost, or (2) "dramatic acceleration" — observing or expecting double the rate of AI aggregate-capability progress versus both the expected rate and the fastest pre-AI-assisted extended rate, plausibly attributable to automation. "Early signs," "cannot rule out," and "significantly faster but not yet by a factor of 2" explicitly do not resolve YES. So I am forecasting a publication event (a formal determination), not the underlying capability, with a lag on top of any capability arrival.
Current status: unambiguously NO, but the surrounding language has moved fast
- Aug 14, 2026 Risk Report (coverage date July 15, 2026), §3: "We do not believe our models meet either RSP criterion for this threat model." Exec summary: internal AI R&D is "significantly faster than it would be without AI assistance, but not yet by a factor of 2 (though we are uncertain and measurement is difficult)." §3.5.2: leading indicators show "meaningful acceleration starting in early-to-mid 2025, though by less than a factor of 2," and the conclusion is "uncertain and subject to some lag (such that we would have difficulty measuring very recent acceleration)." It also says full substitution is not achieved "at any cost," and §3.4 stresses senior-researcher "research taste" gaps. (https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf; https://www.anthropic.com/aug-2026-risk-report)
- The February 2026 Risk Report was even more dismissive ("far from being able to fully automate"), so the trend is one-directional: closer, not further. (https://www.anthropic.com/feb-2026-risk-report)
- RSP v3.4 (effective July 8, 2026) is the operative text and matches the question's quoted criteria verbatim; v3.1 and v3.4 both re-wrote this threshold "to better track the threat model of concern." Critically, the company-level mitigations attached to crossing this threshold in v3.4 are mostly transparency/external review, not a pause — so the operational cost of declaring is modest, and declaration is largely a reputational choice. (https://cdn.sanity.io/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf)
- Anchor for cadence: Risk Reports are published every 3–6 months (Feb 24, 2026; Aug 14, 2026) plus system cards (Opus 5, July 24, 2026; Fable 5.1/Mythos 5.1, Sept 1, 2026), so there are ~3–6 qualifying publication opportunities per year. Anthropic does use explicit "determination" language in both directions ("Our determination is that Claude Opus 4.6 does not cross this threshold," Feb 10, 2026), which reduces the worry that a crossed threshold would never be phrased in resolvable terms.
- But the surrounding evidence has shifted sharply in the last quarter: the Anthropic Institute's "When AI builds itself" (mid-2026) states >80% of merged code is Claude-authored, 8× lines of code/engineer/day vs 2024, ~4× median self-reported research-team output, and that "AI can already match or outperform skilled humans at executing a well-specified experiment" — while "large performance gaps persist" on research judgment and "we are not there yet" on full RSI. (https://www.anthropic.com/institute/recursive-self-improvement)
- Sept 12, 2026: Dario Amodei's "We Must Pace the Frontier" says AI has been "advancing drastically faster" since roughly this summer, "driven primarily by AI's growing ability to build the next generation of AI," that recursive self-improvement is "starting to happen across the industry, including at Anthropic," and warns a swarm could take over the internet in 6–12 months; Anthropic unilaterally committed to embedded third-party evaluators. (https://darioamodei.com/post/we-must-pace-the-frontier; https://venturebeat.com/security/anthropic-ceo-says-ai-swarm-could-take-over-the-entire-internet-in-6-12-months-commits-to-ai-slowdown-plan) This is a personal-domain CEO essay, so it is not a resolving source, and it makes no threshold determination — but it shows the leadership believes it is in the acceleration regime.
Outside view / reference class
- Reference class: "frontier lab publishes a positive determination that a self-imposed catastrophic-risk capability threshold is crossed." Precedents exist and tend to follow evidence by 1–2 reporting cycles: Anthropic's ASL-3 activation (May 2025) and "we act as though models meet our CB-1 threshold" (2026); OpenAI's treatment of bio/cyber thresholds and its Aug 2026 "cannot rule out Critical cyber threshold" for Astra (a hedge, not a determination). There is no precedent for a positive determination on a maximal "AI R&D automation" threshold by any lab. Base rate in the recent past: low but non-zero, and rising with capability.
- Timeline forecasts by relevant people: Jared Kaplan (Anthropic co-founder, TIME, Mar 11, 2026): fully automated AI research "could be as little as a year away." Jack Clark (Anthropic co-founder, Import AI 455, May 4, 2026): "60%+" on no-human-involved AI R&D by end-2028 (~30% by end-2027). Redwood's chief scientist: full AI R&D automation by ~Nov 2028 (Aug 2026). (https://importai.substack.com/p/import-ai-455-automating-ai-research; https://time.com/article/2026/03/11/anthropic-claude-disruptive-company-pentagon/)
- Industry milestone calendar: OpenAI reported hitting its "automated AI research intern" milestone on Sept 6–7, 2026, with an "automated AI researcher" targeted for ~March 2028 (https://openai.com/index/research-acceleration-view-inside-openai/ — I could not fetch directly; reported at https://propakistani.pk/2026/09/07/openai-says-it-has-built-an-automated-ai-research-intern/).
- Trend data: METR 50% task horizons ~doubling every ~4 months (Claude Opus 4.6 ≈ 12 h; Mythos Preview ≈ 17 h and saturating above ~16 h); Epoch ECI frontier advancing ~14–16 points/yr; AI 2027-style scenarios are tracking at roughly 0.7× the scenario's pace, with the research-multiplier prediction "Emerging" but unresolved, and independent analyses finding agents strong on engineering but weak on strategy/novelty (Princeton-led study, Aug 2026). (https://ai2027-tracker.com/; https://metr.org/blog/2026-05-08-rd-section-anthropic-risk-report-feb-2026-review/)
- Counterevidence: METR's external review of the Feb 2026 AI-R&D section disagreed with Anthropic's evidence quality (though it agreed with the bottom line), noting that Anthropic's case for "very low" risk was under-supported and that substantial acceleration before full automation is a live possibility Anthropic's argument under-weighted. This cuts toward the plausibility of a prong-2 (dramatic-acceleration) determination. (https://metr.org/blog/2026-05-08-rd-section-anthropic-risk-report-feb-2026-review/)
Pathways to YES
- Prong 2 first (most likely near-term route): Anthropic's own reports already say acceleration is "less than a factor of 2," and the criterion allows expectation as well as observation. If 2027–2028 internal measurements (AECI-like indices, leading indicators) show a sustained doubling vs. the pre-AI-assisted baseline, a Risk Report could formalize that as a determination. Management rhetoric ("drastically faster," "RSI starting to happen") makes this the shortest credible path.
- Prong 1 later (the big one): if models reach superhuman-at-AI-research status (OpenAI's March 2028 "automated researcher" target; Clark's 60%-by-2028), Anthropic's requirement to substitute for its entire research staff at ≤5× cost becomes satisfiable, and their incentive (credibility, LTBT/external reviewers, employee pressure) to say so is real.
- Policy-forcing-function channel: Anthropic wants to use RSP thresholds to move policy (Amodei is now openly advocating "pacing the frontier"; ~1,300 frontier-lab employees signed a July 2026 letter urging government capacity on automated AI R&D). A dramatic determination could be used as a forcing function, modestly raising near-term hazard.
Strongest case against a high forecast
The criteria are deliberately demanding and Anthropic has institutional reasons to keep hedging: (a) "fully substitute for our entire set" plus "research taste" gaps is a much higher bar than "AI does most of the work"; (b) prong 2 requires doubling versus the fastest pre-AI baseline and plausible attribution to automation rather than compute/headcount — Anthropic is simultaneously scaling compute and headcount heavily, giving it an easy alternative explanation, and it admits "difficulty measuring very recent acceleration"; (c) Anthropic has twice rewritten this very threshold in 2026 rather than declare it met, and RSP v3 explicitly shifted away from crisp ASL-style triggers toward "affirmative case" argumentation — the "zone of ambiguity" it describes is a structural suppressor of crisp determinations; (d) the question explicitly excludes hedged language, and the report language pattern (Feb 2026, Aug 2026, Sept 2026 Institute piece: "we are not there yet") is consistently hedged.
How I built the series
I treated this as a hazard process: near-zero for 2026Q3/Q4 (the Aug 14 report is six weeks old; the next report is ~Nov 2026–Feb 2027), rising through 2027 (leadership expectations, emerging RSI evidence, the "expect" clause) to ~2.5–3%/quarter in 2028–2030 as capability trends and the OpenAI/industry research-automation calendar mature, then roughly flat/slightly declining after 2030 (a declaration that hasn't happened by then implies either a stall or that the threshold's wording is doing real suppressing work). This integrates to ~8% by end-2027, ~17% by end-2028, ~27% by end-2029, ~37% by end-2030, and ~45% by 2031-09-30 — i.e., roughly P(underlying threshold met in Anthropic's judgment by ~2030) × ~0.7 lag/disc? plus the independent prong-2 and policy channels.
Ambiguities noted (see ambiguities).
What would move me >10 points: the next Risk Report (expected ~Nov 2026–Feb 2027) — if it says "we now expect the rate of progress to double" or "models could substitute for a large majority of research work," I would raise the whole series substantially; conversely, any further RSP revision that loosens this threshold without Anthropic also affirming v3.4 criteria (which under the question's hedge produces no YES) would cut the series. Also: an OpenAI-claimed "automated AI researcher" by March 2028 (raise), or credible evidence the capability curve is bending (compute/energy constraints, benchmark saturation without capability transfer) (cut).
Sources relied on
- Anthropic RSP hub page (last updated Aug 14, 2026): https://www.anthropic.com/responsible-scaling-policy
- RSP v3.4 (effective July 8, 2026): https://cdn.sanity.io/files/4zrzovbb/website/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf
- RSP v3.0 announcement, Feb 24, 2026: https://www.anthropic.com/news/responsible-scaling-policy-v3
- Anthropic Risk Report, August 14, 2026 (coverage July 15, 2026): https://www.anthropic.com/aug-2026-risk-report
- Anthropic Risk Report, February 2026: https://www.anthropic.com/feb-2026-risk-report
- Anthropic Frontier Safety Roadmap ("plausible, as soon as early 2027…"): https://www.anthropic.com/responsible-scaling-policy/roadmap
- Anthropic Institute, "When AI builds itself": https://www.anthropic.com/institute/recursive-self-improvement
- Amodei, "We Must Pace the Frontier," Sept 12, 2026: https://darioamodei.com/post/we-must-pace-the-frontier
- VentureBeat summary of the swarm incident and the pacing plan, Sept 2026: https://venturebeat.com/security/anthropic-ceo-says-ai-swarm-could-take-over-the-entire-internet-in-6-12-months-commits-to-ai-slowdown-plan
- METR external review of the Feb 2026 AI R&D section, May 8, 2026: https://metr.org/blog/2026-05-08-rd-section-anthropic-risk-report-feb-2026-review/
- Jack Clark, Import AI 455, May 4, 2026: https://importai.substack.com/p/import-ai-455-automating-ai-research
- AI 2027 Tracker (R&D-multiplier and research dossiers; Sept 2026): https://ai2027-tracker.com/predictions/rd-multiplier-2x/ and https://ai2027-tracker.com/guides/ai-2027-vs-reality/
- TIME, "How Anthropic Became the Most Disruptive Company in the World," Mar 11, 2026 (Kaplan quote): https://time.com/article/2026/03/11/anthropic-claude-disruptive-company-pentagon/
- OpenAI automated research intern reporting, Sept 7, 2026: https://propakistani.pk/2026/09/07/openai-says-it-has-built-an-automated-ai-research-intern/
Resolution criteria
Question. By the last day of each quarter, has Anthropic published a determination that its Responsible Scaling Policy "automation of AI R&D" threshold has been met?
Resolves YES if a Risk Report, system card, entry on anthropic.com/rsp-updates, or other Anthropic-authored document, published on or before the period end, states that Anthropic has determined the threshold is met under either prong:
- Anthropic's models would be able to fully substitute for its entire set of Research Scientists and Research Engineers at competitive cost (i.e., within a factor of 5); or
- "Dramatic acceleration": Anthropic observes or expects double the rate of progress in AI aggregate capabilities compared to both (a) the rate it would expect and (b) the fastest rate of extended progress (over at least three model generations) observed in the absence of significant AI contributions to AI R&D, and it is plausible this doubling is substantially attributable to automation of research and/or engineering rather than headcount, compute, or general productivity.
Resolves NO for statements such as "early signs of acceleration," "cannot rule out," "significantly faster but not yet by a factor of 2," or any language short of a determination.
Successor definitions. If Anthropic renames or renumbers the threshold, the successor counts. If the successor definition is looser than the text above (RSP v3.4), resolve YES only if Anthropic also states the v3.4 criteria are met.
Resolving source. Anthropic-authored documents only (see A2). Note that Risk Reports carry a coverage date within 30 days of publication; resolve on publication date per A3.
Forecast horizons. Every quarter-end from 2026-09-30 through 2031-09-30.