{
  "question_id": "Q1",
  "slug": "anthropic-rd-threshold",
  "title": "Anthropic’s AI R&D threshold",
  "type": "binary",
  "resolution": {
    "levels": [],
    "values": [],
    "annotations": []
  },
  "forecasts": [
    {
      "question_id": "Q1",
      "forecast_date": "2026-09-14",
      "evidence_cutoff": "2026-09-14",
      "reading_adopted": null,
      "ambiguities": [
        "Whether posts on darioamodei.com count as 'Anthropic-authored documents' (adopted: no, per A2 exclusion of executive statements; the Sept 2026 essay also lacks the quantified 2x determination).",
        "Whether a Risk Report statement that Anthropic 'expects' a 2x acceleration, without yet observing it, counts as a determination (adopted: yes, per the 'observes or expects' wording of prong 2).",
        "If RSP v4 or a regulatory Frontier Compliance Framework replaces the v3.4 threshold with looser language, resolution requires an additional explicit statement that the v3.4 criteria are met, which Anthropic may not provide.",
        "Whether only publicly available Anthropic-authored documents count, or whether internal/leaked documents would qualify; I adopt the public-publication reading, consistent with A2 and the named resolving sources.",
        "Whether an RSP revision that changes (rather than renames) the threshold while remaining equally strict counts as a 'successor'; I treat any non-looser successor as counting, with the looser-successor guard applied as written.",
        "Prong 2's baseline ('fastest rate of extended progress in the absence of significant AI contributions') is set by Anthropic's own attribution judgments, which are forecast-resolvable only via their published reports.",
        "How explicit an 'expect' statement must be to count as a determination rather than a roadmap-style plausibility claim.",
        "Whether applying the associated mitigations 'as if' the threshold were met, without saying it is met, would count (adopted: no).",
        "Whether a future stricter successor that they declare met automatically counts (adopted: yes, per successor-definition clause).",
        "Whether language like 'we act as though the threshold is met' (Anthropic's current CB-1 posture) would count as a determination; I treat it as NO since the question resolves NO for 'any language short of a determination'.",
        "Whether Anthropic could satisfy prong 2 via a purely forward-looking 'expects' statement without observed doubling; the criteria allow it ('observes or expects'), and I treat a clear, firm expectation statement as YES, but Anthropic's demonstrated style ('early signs', 'uncertain') makes near-term use of this route less likely.",
        "Whether an essay by Anthropic's CEO on darioamodei.com counts as an 'Anthropic-authored document' under A2. Adopted: no (A2 excludes executive statements and requires the developer's own domain or a developer-authored document). Moot as of the evidence cutoff, since 'We Must Pace the Frontier' makes no RSP determination either way.",
        "Whether Anthropic's IPO prospectus/S-1 would qualify. Adopted: yes, as a developer-authored document, though it is unlikely to contain a threshold determination.",
        "Whether a statement that the threshold is met for a *successor* threshold whose looseness is arguable (rather than clearly looser) counts. Adopted: any renumbered/renamed threshold with substantively equal-or-stricter criteria counts; only clearly looser criteria require the additional v3.4 affirmation.",
        "Whether an affirmative determination made about an *internal-only* model (e.g. 'Model 2') counts. Adopted: yes — the criteria reference 'Anthropic's models' generally and do not require public deployment.",
        "The question specifies that a renamed/renumbered successor 'counts', but spells out the consequence only for a LOOSER successor (v3.4 criteria must also be stated met). I adopt the reading that a determination under a successor that is the same or tighter counts as YES; a looser successor requires Anthropic to also state the v3.4 criteria are met.",
        "Whether a CEO essay on a personal domain (darioamodei.com, Sept 12, 2026) is an 'Anthropic-authored document'. I read A2 as excluding executive statements and personal-domain writings, and in any case that essay makes no threshold determination.",
        "Because the question restricts resolving sources to Anthropic-authored documents, third-party reports (e.g. METR reviews, UK AISI findings) do not resolve even where A2 would allow undisputed third-party reports.",
        "Whether a functionally equivalent statement (e.g. 'our models can now fully substitute for our entire research staff at competitive cost') counts even without the words 'the threshold is met'. I read it as counting.",
        "The question cites 'anthropic.com/rsp-updates'; the operative page is now anthropic.com/responsible-scaling-policy, which carries the same RSP update log."
      ],
      "key_drivers": [
        "The September 1, 2026 Anthropic system card explicitly assesses the AI R&D threshold as not crossed, anchoring the imminent horizon low.",
        "The dramatic-acceleration prong, including an explicit determination about expected acceleration, is the most likely early route to YES.",
        "Large internal productivity gains and recent automated-research results support increasing capability, but do not themselves establish doubled aggregate capability progress.",
        "Amodei's September acceleration warning raises the probability of newer developments not yet captured in formal lagging assessments.",
        "Risk Reports every 3–6 months, off-cycle disclosures, and model cards create repeated publication opportunities; measurement, approval, and confidentiality can delay a qualifying statement.",
        "Senior research judgment, compute and integration bottlenecks, deliberate pacing, and successor-definition or disclosure changes preserve a substantial late tail.",
        "August 2026 Risk Report explicitly says Anthropic 'may cross this threshold in the coming year' and the threat model is plausibly a 'major concern in the next 6-12 months', while currently measuring acceleration 'by less than a factor of 2'.",
        "Sept 1, 2026 Fable/Mythos 5.1 system card reaffirms NO: no clear signs of dramatic acceleration; ECI on the same trend since Mythos Preview; METR finds Mythos 5.1 cannot reliably automate frontier AI R&D.",
        "Dario Amodei's Sept 12-13, 2026 essay claims AI progress is 'drastically faster' since summer 2026 due to AI building AI, including at Anthropic, and commits to embedded METR evaluators with publication rights - signaling willingness to declare acceleration.",
        "Prong 2's 'observes or expects' wording lowers the bar: a forward-looking determination in a Risk Report counts.",
        "Publication cadence: Risk Reports every 3-6 months (next due ~Nov 2026-Feb 2027); system cards for each frontier release (unreleased 'Model 2' is more capable than Mythos 5).",
        "Countervailing: 2x trend-over-trend bar is stringent; measurement lag; Anthropic's habit of hedged language (which resolves NO); IPO/commercial incentives; possible pacing/slowdown agreement or threshold redefinition.",
        "Aug 2026 Risk Report explicitly NO on both prongs with <2x acceleration and CoBench 62.8% vs 85% needed",
        "Sept 1 2026 Mythos/Fable 5.1 system card reaffirms NO (well below human researchers, in line with trends)",
        "CEO Sept 12 2026 pacing essay says RSI starting since summer, raising near-term hazard but not a qualifying determination",
        "Outside view Metaculus full AI-R&D automation median ~Aug 2028; publication lags and high bar (double vs fastest 3-generation baseline) push determination later",
        "Publication calendar: Risk Reports every 3-6 months are main resolution vehicle; incentives to revise threshold or delay vs incentive to declare to justify pacing",
        "Aug 2026 Risk Report explicitly states neither prong is met ('significantly faster... but not yet by a factor of 2'; 'early signs of potential acceleration'; saturated evals; reduced confidence) — question currently unresolved.",
        "Strong internal acceleration trend (Claude authors >80% of merged code, 8x code/engineer/day, ~4x self-reported output, METR task-horizon doubling every ~4 months, ~1.5x measured R&D multiplier) makes capability-side crossing plausible in 2027-2029, most likely via the acceleration prong.",
        "Publication cadence: Risk Reports ~every 6 months (Feb/Aug), system cards on frequent model releases, and SB-53 quarterly compliance reports — short lag (<1 quarter) once a determination is made.",
        "RSP v3.0 removed the pause commitment and thus the perverse incentive to avoid declaring; LTBT oversight and METR-style external review pressure toward honest determinations once criteria are plausibly met.",
        "The threshold is demanding: doubling vs BOTH the expected rate and the fastest non-AI-assisted 3-generation baseline, with attribution to automation (not compute/headcount) — and Anthropic attributed 2025's fast progress to non-AI factors, raising the comparison baseline; substitution prong described as 'not close... not at any cost' in Aug 2026.",
        "Plateau/S-curve scenario (MIT Tech Review skepticism, Amdahl's-law bottlenecks like human code review) is the main reason the far horizon stays well below 1.",
        "Prong 2 'expect' can fire before full researcher substitution; that is the main 2027–28 path.",
        "Declaration friction: attribution-to-automation clause, two recent threshold rewrites, and saturated-but-still-negative evals.",
        "Next Risk Report window (Nov 2026–Feb 2027) and cadence of system cards.",
        "Gap from CoBench ~63% / 'well below' staff to full substitution of the entire RS/RE set.",
        "Leadership/roadmap timelines (early 2027 plausible; ~2028 full automation) vs Epoch ECI showing no speedup as of Aug 2026.",
        "Steep internal automation trajectory: Claude writes 80%+ of Anthropic's code, 8x code volume per person, METR benchmark saturated May 2026, internal substitution eval (CoBench) at ~63% for the internal frontier model, and Jack Clark's stated 60% on autonomous AI self-improvement by 2028 (Time, Aug 7 2026).",
        "Current official position is firmly NO but softening: Aug 2026 Risk Report says neither prong is met, acceleration is 'less than a factor of 2' with 'early signs of acceleration', saturated evals, and reduced confidence vs prior reports.",
        "Publication vehicles and cadence: semiannual Risk Reports (Feb 2026, Aug 2026; next ~Feb 2027) plus per-model system cards; coverage-date-within-30-days rule keeps publication lag short once a determination is made.",
        "Counter-incentives to declaration: crossing triggers the RSP's most demanding mitigations amid an explicit race (OpenAI's March 2028 full-automation target); the threshold has already been revised twice in 2026; Anthropic's CB-1 precedent shows a durable 'act as though met' hedge that would resolve NO.",
        "Attribution difficulty: prong 2 requires the doubling be plausibly attributable to AI R&D automation rather than compute/headcount; Anthropic attributed 2025 acceleration mostly to non-AI factors and notes measurement lag, which can honestly delay a determination for years.",
        "Anthropic's latest Risk Report (August 14, 2026) and Claude Fable 5.1/Mythos 5.1 System Card (September 1, 2026) explicitly determined that neither prong of the automated AI R&D threshold is met. Model 2 achieves 62.8% on CoBench against an estimated threshold of >=85% for researcher substitution, and aggregate capability progress (Anthropic ECI) remains below the 2x AI-attributable acceleration threshold.",
        "Anthropic's Frontier Safety Roadmap explicitly states that the earliest plausible timeline for automated AI R&D is early 2027, with external consensus (Metaculus Level 5 median July 2028, Jack Clark estimate of 60%+ by 2028) placing the median around mid-to-late 2028.",
        "The resolution criteria strictly require a published determination in an Anthropic-authored document under either full substitution of all Research Scientists and Research Engineers at competitive cost or observed/expected 2x progress acceleration attributable to AI R&D. Institutional incentives to avoid premature declaration, difficulty in attributing capability gains to AI rather than compute/headcount, and potential future revisions or obsolescence of the RSP cap the long-term cumulative probability at ~85%.",
        "Anthropic's most recent qualifying statements are explicit negatives: the Aug 14, 2026 Risk Report ('We do not believe our models meet either RSP criterion'; acceleration 'not yet by a factor of 2'; 'early signs of acceleration') and the Sept 1, 2026 Fable 5.1/Mythos 5.1 System Card ('risk as low… well below the capability of our human researchers… in line with current trends').",
        "Anthropic's own dated forward guidance: the Frontier Safety Roadmap says full automation or dramatic acceleration is 'plausible, as soon as early 2027'; the Aug 2026 Risk Report says the threat model may become 'a major concern in the next 6-12 months' and that 'we may cross this threshold in the coming year.'",
        "Publication venues are frequent and partly mandatory: a Risk Report every 3-6 months (next due by ~Feb 14, 2027), a system card per model release, and an off-cycle analysis within 30 days of determining an internal model's automated-R&D risks significantly exceed prior ones. Decision-to-publication lag is therefore weeks, not years.",
        "The operational bars are demanding and narrowly measured: prong 1 needs full substitution of the entire RS/RE staff at <=5x cost (CoBench 62.8% for Model 2 vs ~85% associated with substitution); prong 2 needs one year of aggregate-capability progress equal to two years at baseline, versus both expected and fastest historical rates, attributed to automation. Anthropic concedes measurement lag and that CoBench has saturated.",
        "Declaring is not operationally prohibitive: RSP v3.4 requires a 'strong argument' plus mitigations already underway, with no automatic pause; and a determination would reinforce Amodei's Sept 12, 2026 'pacing the frontier' campaign. Anthropic has shown willingness to declare thresholds met (CB-1).",
        "Offsetting risks: five RSP revisions in six months, twice narrowing a threshold, create a real chance a future declaration falls under a looser successor definition (which would resolve NO); Amdahl's-law bottlenecks (code review, research taste, compute/energy) could delay the underlying condition; and the mid-October 2026 IPO roadshow discourages an alarming Q4 2026 disclosure.",
        "External corroboration that the field thinks this is near: the July 2026 'Pacing the Frontier' statement signed by Anthropic/OpenAI/DeepMind/Meta leadership says leading companies 'could be close to automating AI research'; Anthropic's Institute reports 80%+ of merged code Claude-authored and 8x lines/engineer/day vs 2024; Amodei states recursive self-improvement is 'starting to happen… including at Anthropic' since 'roughly this summer.'",
        "Current status is an explicit NO: the Aug 14, 2026 Risk Report (coverage July 15, 2026) states 'We do not believe our models meet either RSP criterion for this threat model' and that internal acceleration is 'not yet by a factor of 2'.",
        "Very fast capability trend: METR 50% task horizons doubling roughly every 4 months (Claude Opus 4.6 ~12h, Mythos Preview ~17h); Anthropic Institute reports >80% of merged code Claude-authored, 8x lines of code per engineer per day vs 2024, ~4x median self-reported research-team output.",
        "Leadership expectations running ahead of formal determinations: Kaplan (Mar 2026) 'as little as a year away'; Jack Clark (May 2026) 60%+ on no-human AI R&D by end-2028; Amodei (Sept 12, 2026) 'advancing drastically faster' since summer, RSI 'starting to happen'.",
        "Industry milestone calendar: OpenAI claimed the automated AI 'research intern' milestone Sept 2026, targeting an automated AI researcher by ~March 2028; Redwood's chief scientist publicly forecasts full AI R&D automation by ~Nov 2028.",
        "Demanding criteria: prong 1 needs full substitution for the ENTIRE research staff at <=5x cost (Anthropic says models cannot substitute 'at any cost' and cites persistent research-taste gaps); prong 2 needs doubling vs the fastest pre-AI-assisted baseline PLUS plausible attribution to automation rather than compute/headcount.",
        "Institutional hedging: the threshold was rewritten twice in 2026 (v3.1, v3.4) rather than declared met; RSP v3 shifted away from crisp ASL triggers toward qualitative 'affirmative case' argumentation; Anthropic says evaluation science is not dispositive and it has 'difficulty measuring very recent acceleration'.",
        "Modest declaration cost and many opportunities: v3.4 company-level mitigations for crossing this threshold are mainly transparency plus external review (no pause), and Risk Reports (3-6 month cadence) plus system cards give ~3-6 qualifying publication windows per year.",
        "Counterevidence on the capability itself: a Princeton-led study (Aug 2026) and METR's independent review find current agents strong on engineering but weak on strategy/novelty, and Anthropic's own Institute piece says 'we are not there yet' on full RSI."
      ],
      "would_update_on": [
        "An Anthropic-authored explicit determination satisfying either v3.4 prong.",
        "New quantified aggregate capability acceleration against both required baselines, with substantial attribution to research or engineering automation.",
        "Representative entire-staff substitution evidence, including senior researchers and the within-factor-five cost condition.",
        "A fresh autumn assessment showing persistent absence of aggregate acceleration despite stronger models and heavier internal AI use.",
        "Implemented year-plus pacing restrictions or material changes to public reporting and successor-threshold definitions.",
        "Next Risk Report (Nov 2026-Feb 2027) content: an 'expects 2x' or 'observed 2x' statement resolves YES; a repeated 'less than 2x, partly non-AI factors' finding would lower 2027 probabilities by 10-15pp.",
        "METR or embedded-evaluator reports showing time-horizon doubling faster than ~3 months or largely autonomous frontier research at Anthropic (raise 2027 by 10-15pp).",
        "Release of 'Model 2' or a Mythos 6-class model and the AI R&D language in its system card.",
        "An RSP v4 or regulatory framework that renames/loosens the AI R&D threshold, or a coordinated industry pacing agreement slowing measured progress (lower medium-term by ~10pp).",
        "Anthropic IPO timing and any signs the company is deferring threshold determinations for commercial reasons.",
        "Next Risk Report shifting from 'early signs'/'significantly faster but not yet 2x' to 'at or approaching 2x' or raising confidence",
        "CoBench or successor eval jumping from 62.8% toward 75-85% or Anthropic stating it expects doubling attributable to automation",
        "New RSP version loosening/tightening automation threshold or LTBT-requested external review disputing Anthropic NO assessment",
        "February 2027 Risk Report language: a shift from 'not yet by a factor of 2' to 'expect to cross' or 'approaching the threshold' would raise near-horizon probabilities by 10+ points; continued 'early signs' hedging would lower them.",
        "Any new RSP revision (v3.5+) that renames, restructures, or loosens the AI R&D threshold, or changes determination procedures.",
        "System cards for next releases (e.g., Model 2, Opus 5.x, Mythos 5.2): a first-ever 'crosses' determination or 'close to crossing' framing.",
        "METR time-horizon and R&D-multiplier measurements: weeks-long task horizons or measured multipliers approaching 2x would substantially raise mid-horizon probabilities.",
        "Evidence of attribution disputes in future risk reports or pushback from LTBT/external reviewers (e.g., METR arguing acceleration is being under-credited) would raise the probability of an earlier forced determination.",
        "Signs of capability plateau (AECI trendline flattening, benchmark saturation without frontier gains, compute constraints) would lower far-horizon probabilities significantly.",
        "Next Anthropic Risk Report or frontier system card using determination language, or an explicit 'we expect 2× attributable to automation.'",
        "RSP rewrite that loosens the threshold without affirming v3.4 (down) or that states v3.4 is met (up).",
        "Published substitution metric (CoBench-class) ≥80% or AECI/ECI doubling they attribute to automation rather than compute.",
        "Off-cycle report on an internal model used for large-scale autonomous research.",
        "Sustained stall in METR-like horizons or a public statement that they cannot measure acceleration (down on prong 2).",
        "The next Risk Report (~Feb-Mar 2027) stating acceleration is at or above a factor of 2, or announcing leading-indicator evidence consistent with doubling attributable to AI.",
        "Any RSP v3.5+ change that loosens the AI R&D threshold (lowers p_yes via the successor clause) or tightens operationalization/reporting commitments (raises it).",
        "Anthropic adopting CB-1-style 'we act as though the AI R&D threshold is met' language without a determination — would signal durable hedging and lower all horizons.",
        "External shocks: a visible, rapid capability jump across labs in late 2026/2027 (raises near-term hazard), or a slowdown/pause regime or governance disruption at Anthropic (lowers it).",
        "METR or AISI third-party measurements showing multi-week autonomous AI R&D task completion, which Anthropic would have trouble not acknowledging in subsequent reports.",
        "A new internal model checkpoint demonstrating >=80% on CoBench or an official Anthropic publication documenting an acceleration factor approaching 2x on AECI.",
        "A statement in Anthropic's February 2027 Risk Report declaring the automated AI R&D threshold formally crossed or imminent.",
        "A formal restructuring or replacement of Anthropic's RSP that removes the AI R&D threshold in favor of statutory compliance frameworks without specifying v3.4 equivalence.",
        "Publication of the next Anthropic Risk Report (due by ~Feb 14, 2027) reporting an AECI acceleration factor at or above 2x attributable to AI R&D automation, or CoBench crossing ~85% for an internal model — would raise near-term horizons by 15-25 points.",
        "Any new entry on anthropic.com/responsible-scaling-policy (rsp-updates) or an off-cycle model risk analysis invoking the automated-R&D threshold — would raise all horizons sharply.",
        "An RSP v3.5+ that replaces the two-prong v3.4 text with softer 'leading indicator' or checkpoint language — would cut 2028+ horizons by 10-20 points, since a declaration under a looser successor does not resolve YES.",
        "METR, UK AISI, or another government AI safety institute publishing findings that Anthropic's models meet the threshold, undisputed by Anthropic within 14 days — would raise near-term horizons (though note the resolving source must still be Anthropic-authored, so this mainly signals an imminent Anthropic determination).",
        "Anthropic invoking its Appendix A 'Anthropic in the lead' competitor-contingent commitments, or announcing it is delaying development pending a 'strong argument' that catastrophic risk is contained.",
        "Evidence of a capability plateau or S-curve bend in Anthropic's own reporting (e.g. AECI trendline flattening, Mythos/Model successors landing on rather than above trend) — would lower all horizons.",
        "Anthropic's IPO pricing and the resulting disclosure regime: post-IPO securities litigation risk could either accelerate candor (higher) or encourage hedged framing (lower).",
        "The next Risk Report (expected ~Nov 2026-Feb 2027): language like 'we now expect the rate of progress to double' or 'models could substitute for a large majority of our research work' would raise the whole series by >10 points; a repeat of 'early signs only' would trim it.",
        "Any further RSP revision of the AI R&D threshold: a loosening without a simultaneous statement that v3.4 criteria are met would resolve nothing under the question's hedge and would lower my series; a tightening with a declared determination would raise it.",
        "OpenAI (or Google DeepMind) publicly declaring an 'automated AI researcher'/superhuman-coder milestone, especially around the March 2028 target.",
        "Evidence that frontier models produce genuinely novel research contributions judged by independent experts (moves up) versus evidence of a capability S-curve, benchmark saturation without transfer, or compute/energy constraints binding (moves down).",
        "Explicit Anthropic statements about its own expectations for the determination, e.g. that it does not expect the threshold to be met for several years (moves down), or an off-cycle RSP update asserting accelerating progress attributable to automated R&D (moves up).",
        "Whether the embedded-evaluator commitment (Amodei, Sept 12, 2026) leads external reviewers to publish findings that Anthropic declines to dispute, changing what Anthropic feels obliged to state in its own documents."
      ],
      "forecasts": [
        {
          "period_end": "2026-09-30",
          "p_yes": 0.0155
        },
        {
          "period_end": "2026-12-31",
          "p_yes": 0.082
        },
        {
          "period_end": "2027-03-31",
          "p_yes": 0.1755
        },
        {
          "period_end": "2027-06-30",
          "p_yes": 0.2553
        },
        {
          "period_end": "2027-09-30",
          "p_yes": 0.3304
        },
        {
          "period_end": "2027-12-31",
          "p_yes": 0.3963
        },
        {
          "period_end": "2028-03-31",
          "p_yes": 0.4578
        },
        {
          "period_end": "2028-06-30",
          "p_yes": 0.5097
        },
        {
          "period_end": "2028-09-30",
          "p_yes": 0.5543
        },
        {
          "period_end": "2028-12-31",
          "p_yes": 0.595
        },
        {
          "period_end": "2029-03-31",
          "p_yes": 0.6278
        },
        {
          "period_end": "2029-06-30",
          "p_yes": 0.6571
        },
        {
          "period_end": "2029-09-30",
          "p_yes": 0.6832
        },
        {
          "period_end": "2029-12-31",
          "p_yes": 0.7035
        },
        {
          "period_end": "2030-03-31",
          "p_yes": 0.7218
        },
        {
          "period_end": "2030-06-30",
          "p_yes": 0.7386
        },
        {
          "period_end": "2030-09-30",
          "p_yes": 0.7544
        },
        {
          "period_end": "2030-12-31",
          "p_yes": 0.7691
        },
        {
          "period_end": "2031-03-31",
          "p_yes": 0.7828
        },
        {
          "period_end": "2031-06-30",
          "p_yes": 0.7958
        },
        {
          "period_end": "2031-09-30",
          "p_yes": 0.8079
        }
      ]
    }
  ]
}