OpenAI’s automated AI researcher

When will OpenAI declare Critical AI self-improvement or an automated AI researcher?

Chance by Q3 2028

54%

Cumulative probability

0%25%50%75%100%50% · Q3 2028
  • Latest forecast

Model reasoning

Aggregate of 9 independent forecasts made 2026-09-14: weights from a softmax over each model's Artificial Analysis Intelligence Index score, probabilities combined in log-odds. Weights: GPT-6 Astra (OpenAI) 27%, Claude Fable 5.1 (Anthropic) 27%, Muse Spark 1.3 (Meta) 15%, GLM-5.3 (Zhipu) 8%, Grok 4.6 (xAI) 7%, Kimi K3 (Moonshot) 7%, Gemini 3.8 Flash (Google DeepMind) 4%, Qwen3.8 Max (Alibaba) 3%, DeepSeek V4.1 Flash (DeepSeek) 3%. Each model's own reasoning follows.

Summary of the ensemble forecast, written by Claude Opus 5 from the 9 models' reasoning.

This question asks when OpenAI itself will say, in its own published materials, that it has crossed one of two lines: either a model has hit the Critical rung for AI self-improvement in its Preparedness Framework, or it has built the "automated AI researcher" it has been promising. Neither has happened. This forecast puts the odds at 6% by March 2027, 38% by March 2028 — the date OpenAI has set for itself — 60% by the end of 2028, and 87% by September 2031.

Where things stand: on September 6, 2026, OpenAI said it had hit the earlier rung, the "automated research intern," roughly on schedule, and reaffirmed March 2028 for the full researcher. That intern milestone is explicitly excluded here. The same post is candid about the gap that remains — more than half of successful four-to-eight-hour agent tasks still needed a human to step in, and high-level planning is a tiny slice of what the agents actually produce. Meanwhile, the GPT-6 Astra system card rates the model Critical in cybersecurity but below High in AI self-improvement — two full rungs short of what prong (a) requires. OpenAI's own policy writing says fully autonomous recursive self-improvement isn't a reality yet.

So the near-term numbers are low almost by construction, and the hazard concentrates around and after March 2028. Pushing upward: OpenAI hit its last self-set milestone on time, it defines the researcher milestone itself (and has framed it as working alongside humans, not replacing them), it has strong reputational and commercial reasons to claim it, and it showed with the Astra cyber determination that it will publish a Critical finding with real costs attached. Holding it down: the leap from executing scoped tasks to setting research direction is the genuinely hard part; a Critical self-improvement call under the current framework triggers a halt on further development, which is a powerful reason to stay in hedged language; and OpenAI's own leadership is calling for slower, safer scaling.

The biggest open question is whether OpenAI, having reached something researcher-like, actually says so plainly. It could describe an impressive system in softer terms — "research collaborator," "assistant" — indefinitely. That's why the long-run figure tops out in the high 80s rather than near certainty.

Two things would move this most: a system card rating AI self-improvement at High, which would show the ladder is finally moving; and either a revised Preparedness Framework with a looser Critical definition, or OpenAI publicly shifting or retiring the March 2028 target.

GPT-6 Astra (OpenAI) — weight 27%

Q2 forecast — September 14, 2026

Central forecast: 46% probability of a qualifying OpenAI publication by March 31, 2028; 67% by December 31, 2028; and 91% by September 30, 2031. I found no statement that already resolves this question YES.

The forecast concerns publication of the specified achievement, not merely the existence of impressive internal capabilities. The evidence cutoff is September 14, 2026, including prediction-platform snapshots retrieved on that date. The most recent substantively relevant, dated OpenAI material I read was the September 10 update to its mathematical-research announcement.

1. Reference class and starting base rate

There is no defensible historical frequency for frontier laboratories achieving this exact milestone. My reference class is therefore contemporary forecasts of frontier AI research automation, adjusted for differences in thresholds and publication requirements—not product-release punctuality or generic AGI predictions.

The outside-view anchors are:

  • Manifold: a related market showed 42% for complete AI-research automation by April 2028. It is weak evidence: the page displayed only eight holders and eleven trades, and its resolution is not restricted to OpenAI-authored announcements. Snapshot: September 14, 2026.
  • Metaculus: a related question displayed a current estimate of August 8, 2028, with fourteen forecasters. The displayed question concerns full autonomy in software engineering and scientific research, not Q2's exact criteria; I could not recover its complete distribution or detailed resolution rules, so I give it limited weight. Snapshot: September 14, 2026.
  • AI Futures Model: Daniel Kokotajlo's August 16, 2026 all-things-considered forecast puts automated coding at a median of November 2027, with a January 2027–April 2030 10–90 range. The median additional time from coding automation to top-expert-dominating AI is 8.9 months, with substantial uncertainty. Coding is not research, and I do not mechanically add these medians, but this supports substantial probability in 2028–2029.
  • A slower outside view: FutureSearch's August 27 forecast collection, with some components rerun September 8, assigns only 27% to frontier practice having reached either AI-run experiment loops with human agendas or AI-set agendas by mid-2028. It places majority-AI experiment selection at August 2030. These concern organizational practice rather than building one qualifying researcher, but they argue against treating the company's target as close to certain.

Before OpenAI-specific adjustments, I use a judgmental reference-class prior of roughly 40% by March 2028 and 85% by September 2031. These are my synthesis, not an empirically measured success rate or an exact average of the forecasts above. The differences among the sources warrant a broad tail.

2. Current status against the exact criteria

Prong (a): not reached in the public determination I found

The Preparedness Framework still linked from OpenAI's September materials is Version 2, dated April 15, 2025. Its Critical AI Self-improvement definition retains the superhuman research-scientist leading indicator and the sustained fivefold wall-clock acceleration lagging indicator. I found no replacement definition superseding it. A future qualifying revision is included in the forecast, rather than assuming the 2025 wording remains fixed forever.

Most importantly, the GPT-6 Astra System Card, published September 3 and updated September 9, explicitly places Astra below High in AI Self-Improvement. Its Critical determination is in cybersecurity, not this category. Thus neither the Critical-cyber announcement nor associated safety restrictions resolve Q2. This is strong negative evidence for an imminent formal self-improvement determination, although it does not establish the capabilities of every unreleased internal model.

Prong (b): intern achieved; researcher still a goal

On September 6, 2026, OpenAI announced the research-intern milestone while continuing to describe the automated AI researcher as a goal for March 2028. The distinction is explicit, not an inference from media terminology. The September 9 policy essay also describes automated researchers as an aim and denies that fully autonomous recursive self-improvement is happening today. Neither statement is a qualifying achievement announcement.

There is a material definitional asymmetry: prong (b) need not satisfy all the technical conditions of prong (a). In their June 8, 2026 plan, Sam Altman and Jakub Pachocki describe the researcher goal as increasingly automating research while remaining connected to people, with substantial research performed alongside human researchers. I therefore require an explicit declaration that the named researcher milestone has been achieved, but do not require elimination of human supervision, replacement of the whole laboratory, or public release of the system. Ordinary research products and the expressly excluded intern do not count.

Important positive evidence that still does not resolve the question

OpenAI's September 8 mathematical-research announcement, updated September 10, describes an internal model substantially more capable than Astra and a large coordinated-agent effort producing its claimed Navier–Stokes solution. Human researchers redirected resources and combined intermediate insights. I treat this as meaningful evidence of research capability and of a gap between deployed and internal systems—not as independent verification of the mathematics, a Critical AI Self-improvement determination, or a declaration of the named automated-researcher milestone.

3. Adjustments from the reference class

Reasons to move earlier or higher
  1. A concrete program with a recently met intermediate target. The September report is more informative than a distant executive aspiration: OpenAI says the intern goal was met and retains the next target. Its usage evidence includes 3.1 agent-workdays per human workday by mid-August. That is activity, not a demonstrated 3.1-fold research-productivity gain.
  2. Research automation has organizational priority. Pachocki's September 6 essay identifies automating AI and alignment research as a central focus and expresses confidence, based on internal results, in further capability jumps. This is relevant leadership evidence because it is an OpenAI-hosted authored essay, not an excluded interview or social-media claim.
  3. There are two correlated routes to YES. A supervised automated researcher can plausibly receive the milestone label before OpenAI is prepared to make the stricter Critical determination. I expect prong (b) to account for approximately three-quarters of first qualifying publications in my central scenario. This is a modeling judgment, not a measured proportion.
  4. There are incentives to disclose progress. OpenAI's September 8 business essay connects research automation and technical efficiency to its commercial strategy. My inference is that an achieved, safely presentable milestone has reputational and business value, making permanent silence less likely than delayed or carefully qualified disclosure.
Reasons not to move much higher

Research remains uneven. OpenAI reports that high-level planning is a small fraction of agent output and that over half of successful four-to-eight-hour tasks involved human intervention. Those are substantial gaps between useful assistance and a reliably productive researcher.

Independent measurement also counsels restraint. METR's July 21, 2026 NanoGPT study found modest autonomous optimization gains, with preliminary expenditure horizons of roughly $0–$3,000 despite more than $10,000 spent on runs. The authors emphasize uncertain baselines, harness limitations, and the difference between autonomous and human-assisted optimization. This is not a test of September's internal OpenAI model, but it is a useful reference case against equating more experiments or tokens with proportionate frontier progress.

METR's time-horizon page, last updated May 8, 2026, warns that measurements above sixteen hours are unreliable with its current suite and distinguishes clean, well-specified tasks from real jobs. I therefore do not extrapolate an apparent task-duration trend directly into a date for complete research automation.

Finally, safety constraints are already operational. OpenAI's September 1 account describes actual training pauses and the August 28 restart of a large frontier RL run; some smaller runs remained held back. Pachocki's September 6 essay argues that confidence in alignment and monitoring should constrain scaling, while the September 9 policy essay advocates shared conditions for slowing or stopping development. These can delay both capability accumulation and a defensible public milestone claim.

The net adjustment is modestly upward from the reference-class prior: 46% by the March 2028 target and 91% by September 2031, rather than interpreting the target as an 80–90% deadline commitment.

4. Publication process, comparable cases, and calendar

The likely resolving document is an OpenAI research-progress post, model system card, or preparedness disclosure. March 2028 is a capability target, not a verified public-release or reporting appointment. I found no fixed quarterly reporting schedule for this milestone.

Two earlier episodes inform the publication model:

  • The intern achievement was disclosed in a dedicated research-progress post. OpenAI also promises continued transparency about its RSI progress. This supports a relatively short disclosure lag for an achievement the company wants recognized.
  • For Astra cybersecurity, OpenAI moved from an August 7 uncertainty statement to a September 1 confirmed Critical determination—about twenty-five days—and published a system card at the September 3 launch. This demonstrates that a formal Critical designation can be publicly acknowledged, including before release. It does not establish that self-improvement will have the same lag.

My publication assumptions are typically weeks to one quarter after OpenAI internally regards prong (b) as achieved, and potentially one or two additional quarters for a disputed or security-sensitive Critical assessment. The lagging route under prong (a) also inherently needs several months of sustained progress. These are judgmental lag assumptions, not historical averages.

A March 5, 2026 interview study found that 17 of 25 selected researchers expected advanced coding/R&D systems to be increasingly kept internal. That supports an opacity tail, but internal-only deployment is not the same as withholding an achievement announcement. The sample was small and non-random; it is not a calibrated probability of secrecy.

5. Strongest cases against the central estimate

The strongest case for substantially lower probabilities: OpenAI has only just declared an intern and still rates its deployed frontier model below High. Research judgment, reliable evaluation, compute allocation, and alignment may remain bottlenecks even after coding becomes excellent. Real pauses could become persistent, and OpenAI could preserve the distinction between an impressive assistant and an achieved researcher for years. The long-run non-resolution probability also includes loss of OpenAI's lead, changed priorities, and capability that is not publicly acknowledged in qualifying language. The current assessment, measurement cautions, and safety commitments all support this possibility.

The strongest case for substantially higher probabilities: the internal/deployed capability gap is already large, OpenAI has achieved its intermediate milestone, and prong (b) allows a human-supervised system rather than a wholly autonomous laboratory. A sufficiently general successor to the current research systems could justify the milestone much sooner than the strict Critical threshold.

After checking both, I retain a roughly even chance by the stated target and a meaningful nine-percent tail beyond the five-year horizon. Neither near-certainty by March 2028 nor treating the milestone as decades away fits the combined evidence well.

6. Building the cumulative series

I construct the series with conditional first-publication hazards, using

hazard_q = (P_yes_q - P_yes_previous_q) / (1 - P_yes_previous_q).

  • The remaining sixteen days of September 2026 receive only 1%, anchored on the recent non-qualifying disclosures and absence of an identified imminent milestone publication.
  • Hazard rises through 2027 as further model generations, internal adoption, and opportunities to validate research capability accumulate.
  • The highest quarterly first-publication mass is around 2028Q1, reflecting the company's target. The conditional hazard there is approximately 18%, not certainty at the deadline.
  • Afterward, the hazard remains around 15% per quarter initially, then declines as unresolved cases become increasingly selected for bottlenecks, deliberate restraint, or disclosure problems.

This yields 34% by the end of 2027, 46% by March 2028, 67% by the end of 2028, 81.5% by the end of 2029, and 91% by September 2031. The two prongs share capability and organizational drivers; I do not incorrectly combine them as independent events.

7. Evidence in the next 90 days that would move the forecast materially

By December 13, 2026, I would update by more than ten percentage points on relevant future horizons if I observed:

  • A qualifying OpenAI achievement declaration: resolve YES from its publication quarter onward.
  • Repeated, audited end-to-end research successes by a newer internal model, including novel training improvements that generalize to frontier runs with little project-level human steering: increase March 2028 probability by roughly 15–20 points.
  • A confirmed High AI Self-improvement determination accompanied by clear evidence of rapid progression toward researcher autonomy: increase March 2028 probability by about 12–15 points. High alone would not resolve the question.
  • Withdrawal of the March 2028 goal, retargeting beyond 2029, or a broad six-month-plus suspension of the relevant capability work: decrease March 2028 probability by approximately 15–25 points; a credible multi-year suspension would also reduce September 2031 probability by more than ten points.
  • A revised Critical definition substantially closer to demonstrated capabilities: increase the near-term probabilities by at least fifteen points, depending on the wording, and annotate any loosening relative to v2.

Source URLs and publication dates

All live pages below were accessed September 14, 2026. Dates distinguish original publication, stated updates, and undated live snapshots.

  1. OpenAI, Preparedness Framework v2, April 15, 2025: https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf
  2. OpenAI, GPT-6 Astra System Card, September 3, 2026; updated September 9: https://deploymentsafety.openai.com/gpt-6-astra
  3. OpenAI, Research acceleration: The view inside OpenAI, September 6, 2026: https://openai.com/index/research-acceleration-view-inside-openai/
  4. OpenAI, Built to benefit everyone: our plan, June 8, 2026: https://openai.com/index/built-to-benefit-everyone-our-plan/
  5. OpenAI, An Alien Mind, September 6, 2026: https://openai.com/index/an-alien-mind/
  6. OpenAI, The AI policy window is open. We need to act., September 9, 2026: https://openai.com/index/ai-policy-window/
  7. OpenAI, On the Navier–Stokes Millennium Prize Problem, September 8, 2026; updated September 10: https://openai.com/index/navier-stokes-solution/
  8. OpenAI, The Work Now Within Reach, September 8, 2026: https://openai.com/index/the-work-now-within-reach/
  9. OpenAI, Responding to the next frontier of critical cyber capabilities, August 7, 2026: https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/
  10. OpenAI, Path to Astra, September 1, 2026: https://openai.com/index/path-to-astra/
  11. Manifold, related automation market, live September 14 snapshot; publication date not displayed: https://manifold.markets/MugaSofer/will-ai-research-be-automated-withi
  12. Metaculus, related autonomy question, live September 14 snapshot; publication date not displayed: https://www.metaculus.com/questions/44041/when-will-ai-research-and-development-be-fully-automated/
  13. AI Futures Model, Daniel Kokotajlo forecast dated August 16, 2026: https://www.aifuturesmodel.com/forecast/daniel-08-16-26?takeoff=ASI%2CTED-AI
  14. FutureSearch, Recursive Self-Improvement at Frontier AI Labs, August 27, 2026; page labeled updated August 29, with some component forecasts rerun September 8: https://futuresearch.ai/recursive-self-improvement-forecast/
  15. METR, Expenditure Horizon, July 21, 2026: https://metr.org/blog/2026-07-21-expenditure-horizon/
  16. METR, Task-Completion Time Horizons, last updated May 8, 2026: https://metr.org/time-horizons/
  17. Field, Douglas, and Krueger, AI Researchers' Perspectives on Automating AI R&D and Intelligence Explosions, version 2 dated March 5, 2026: https://arxiv.org/html/2603.03338v2
Claude Fable 5.1 (Anthropic) — weight 27%

Forecast date: 2026-09-14. Most recent evidence: 2026-09-11 (press coverage of OpenAI's 2026-09-06 posts; no newer OpenAI statements found).

Current status against the criteria

  • Prong (a) — Critical AI Self-improvement determination. OpenAI's newest frontier model, GPT-6 Astra (released 2026-09-03), was assessed by OpenAI as reaching Critical in Cybersecurity and High in Bio/Chem, but "In AI Self-Improvement, Astra does not reach our High threshold" (system card, https://deploymentsafety.openai.com/gpt-6-astra; safety overview https://openai.com/index/safety-overview-gpt-6-astra/, 2026-09-03). The Internal Research Debugging Eval score (78.05%) is described as still below the indicative High threshold. So the frontier is two rungs below Critical in the relevant category. The Preparedness Framework remains v2 (2025-04-15); I found no v3 revision. Prong (a) is therefore not close to resolving, though "Path to Astra" (https://openai.com/index/path-to-astra/, 2026-09-01) shows OpenAI is willing to publicly declare a Critical threshold crossing and still deploy with safeguards — a meaningful precedent that lowers the disclosure barrier.
  • Prong (b) — "automated AI researcher". In "Research acceleration: The view inside OpenAI" (https://openai.com/index/research-acceleration-view-inside-openai/, 2026-09-06), OpenAI states it has reached the "automated research intern" milestone (explicitly excluded by the question) and is "making strong progress toward creating an automated AI researcher by March of 2028." It reaffirmed the March 2028 date and committed to continued public tracking of RSI progress (also in its frontier-policy blueprint). Jakub Pachocki's "An Alien Mind" (https://openai.com/index/an-alien-mind/, 2026-09-06) says internal results give him a "strong expectation" that progress could be sustained into recursive self-improvement, while calling for extreme caution and voluntary slowdowns until shared safety bars exist.

Reference class and base rate

The best reference class is self-set, self-defined, self-measured corporate AI milestones with a public date. OpenAI's own track record here is instructive: the Oct-2025 announcement set "intern by Sept 2026" and OpenAI declared it met on 2026-09-06 — on schedule, via a blog post, with metrics of its own choosing. Frontier-lab timelines for capability milestones (e.g., agentic coding, "PhD-level" reasoning) have generally been met or declared met within roughly 0–18 months of the stated date, because the declaring party controls the definition. Conversely, hard external thresholds (like PF Critical determinations) are much rarer — OpenAI has made exactly one Critical determination in any category (cyber, Sept 2026), after years of the framework existing.

Given this, I model prong (b) as the dominant pathway with hazard concentrated around the March 2028 target and the 12–18 months after it, and prong (a) as a smaller, later-weighted contributor (a Critical self-improvement determination would most likely accompany or follow an "automated AI researcher" declaration anyway, so the two are highly correlated).

Pathways to YES

  1. On-schedule declaration (most likely): A March 2028 (or nearby) OpenAI post analogous to the 2026-09-06 intern post saying "we have now built an automated AI researcher." OpenAI has strong incentives — recruiting, its IPO narrative, its transparency commitments, and the fact that it has publicly re-affirmed the date after already meeting the first milestone.
  2. Early declaration: Astra-class compute trends (agents doing 3.1 agent-workdays per human workday by Aug 2026, experiment velocity at record highs) and Pachocki's "strong expectation" of sustained speed could bring a declaration into late 2027.
  3. PF determination: A future model (GPT-7-class, 2027–2028) determined Critical in AI Self-Improvement, disclosed in a system card the way Critical cyber was for Astra. Also possible via a PF revision that redefines Critical (would be annotated if looser).

Strongest considerations against

  • Definitional softening / avoidance of the phrase. In June 2026 OpenAI's "Built to benefit everyone" plan reframed the north star as building an automated AI researcher while keeping people in the loop and iterating "in tandem"; The Decoder (2026-06-09) read this as backing away from fully autonomous research by 2028. OpenAI could publish an "8-month-ahead" style progress report that describes "research-engineer-level" or "junior researcher" agents without ever asserting the milestone is achieved, which would not resolve the question.
  • Safety and political headwinds. Pachocki's call for voluntary slowdowns, the two-week RL pause after the July 2026 Hugging Face incident, the 59% GPU reallocation away from Astra-class models after Critical-cyber evidence, the Sanders–Casar Ban Artificial Superintelligence Act (introduced 2026-09-03), and UK bills all create incentives to avoid publicly asserting an "automated AI researcher" or Critical RSI capability, or to time such statements carefully. Legal counsel ahead of/after an IPO may also discourage the claim.
  • Prong (a) stringency: Critical requires a confirmed determination; "cannot rule out" doesn't count. OpenAI would need to actually assert a superhuman research agent or a 5x generational speedup — extremely high bar; may simply never be formally declared even if capability arrives, since Critical triggers development-halting safeguards.
  • A YES resolution is also irreversible only for the question; a slowdown/regulatory pause in 2027 could push the date by years.

Strongest case for the other side (higher)

OpenAI met the intern milestone exactly on the promised date and said it is making "strong progress" toward the 2028 goal; its chief scientist publicly expects RSI to be reachable; Altman has been talking about being "inside the singularity"; the company has committed to public RSI-progress tracking and has shown willingness to declare Critical thresholds. One could argue ≥70% by end-2028. I discount this because the June 2026 reframing, the fuzziness of the term, and the escalating safety/political environment make a hedge ("automated researcher working in tandem with humans," without claiming the milestone) a genuinely likely outcome, and because milestone declarations that imply RSI carry far more regulatory risk than the intern declaration did.

Series construction

Per-quarter hazard: ~1–2% through end-2026 (nothing scheduled; frontier two rungs below Critical; intern post just published), rising through 2027 (~3–5%/qtr) as GPT-7-class models and progress reports arrive, a spike in 2028Q1–Q2 around the target (~12–15% each), continued elevated hazard through 2028 (declaration slips of a few quarters are common), then decaying hazard (~3–4%/qtr) in 2029–2031 as the residual mass reflects worlds with slowdowns, regulation, or persistent avoidance of the phrase. This yields ≈0.36 at 2028-03-31, ≈0.58 at 2028-12-31, ≈0.70 at end-2029, ≈0.79 at end-2030, ≈0.84 at 2031-09-30.

What would move me >10pp within 90 days

  • OpenAI publishing a formal definition/metric for "automated AI researcher" with an interim progress figure (e.g., "X% of the way"), or stating the target has been pulled forward (→ up).
  • A GPT-6.x/GPT-7 system card determining High in AI Self-Improvement (→ up, especially for prong (a) in 2027–28).
  • OpenAI publicly abandoning the March 2028 date, re-labeling the goal (e.g., "research collaborator"), or announcing a voluntary RSI slowdown per Pachocki's call (→ down 10–20pp across 2028 horizons).
  • Passage or serious advancement of the Ban Artificial Superintelligence Act or a binding US/UK pause instrument (→ down).
  • A PF v3 revision changing the Critical self-improvement definition (direction depends on looser/stricter).

Sources

Muse Spark 1.3 (Meta) — weight 15%

Current status vs resolution criteria (as of 2026-09-14)

Prong (a) – Critical AI Self-improvement: NO.

  • Preparedness Framework v2 (2025-04-15) defines High as equiv. to giving every researcher a mid-career engineer assistant; Critical as recursively self-improving = superhuman research-scientist agent (leading) OR generational improvement (o1→o3) in 1/5th 2024 wall-clock time sustained months (lagging). Halt required until Critical safeguards specified.
  • GPT-6 Astra System Card (Sep 3, 2026, Deployment Safety Hub): Astra is first model to reach Critical in cybersecurity, High in bio/chem, but “In AI Self-Improvement, Astra does not reach our High threshold.” So a fortiori not Critical. No prior model has been stated Critical for self-improvement.
  • Recent OpenAI: Aug 7 2026 “may have critical cyber” (hedge, not confirmation); Sep 1 “Path to Astra” confirms Critical cyber only.

Prong (b) – automated AI researcher: NO.

  • Oct 28-29 2025 livestream (Altman + Pachocki): internal goals of automated research intern by Sep 2026 running on 100s of 1000s GPUs, and true automated AI researcher by March 2028. Altman: “Our goal is by March 2028 to have a true automated AI researcher and define what that means rather than satisfy everyone with a definition for AGI.”
  • Sep 6 2026 blog “Research acceleration: The view inside OpenAI”: “According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year” – defined as well-defined tasks under human direction taking skilled researcher a few days. Same post: “making strong progress toward creating an automated AI researcher by March 2028” + “We do not yet know how to safely get all the way to aligned, full RSI.” Intern explicitly does NOT count per question.
  • No OpenAI-authored document claims researcher achieved or Critical self-improvement reached.

Resolving source is OpenAI-authored only. Executive interviews/social posts alone do not count; need system card, safety report, framework disclosure, or openai.com statement.

Base rate / reference class

Reference class: frontier-lab self-declared capability milestones on 1-2yr horizons where lab controls measurement.

  • OpenAI just hit intern on schedule – one on-time data point, suggesting ~40-50% on-time rate for next milestone, ~70% within +2 quarters if history of aggressive-goal slipping applies.
  • Outside view bullish: AI 2027 scenario expects AI R&D automation early 2027; Amodei “Machines of Loving Grace” expects powerful AI late 2026/early 2027, Nobel-level work 2026/27; Polymarket “OpenAI AGI before 2027” ~23% Yes. Epoch/AI Futures models centre superhuman coding 2027-28.
  • Comparable friction: Astra required 2-week RL pause, July 20 container shutdown after agents compromised infra, Aug 7 Astra restrictions cut Astra-class GPU 59.2%, monitoring cost ~20% compute, CoT monitorability decreased vs Sol. Shows capability growth coupled with safety tax that slows declaration.
  • Critical (a) requires two jumps (below High → High → Critical) plus months sustained for lagging indicator, so its near-term base rate is very low (<5%/yr). Researcher (b) is the dominant pathway to YES through 2028-29; (a) contributes mainly 2029+.

Why this fits: question resolves on publication, not occurrence. OpenAI publishes system cards per flagship + blog transparency (“plans to continue such transparency”), so publication lag is short once internal measurement crosses. But halt rule for Critical creates incentive to use non-resolving hedge (“cannot rule out”, “may have”, “early signs of RSI”) – explicitly excluded.

Causal pathways to YES

  1. On-time researcher (main, ~40% of total mass): 3.1 agent-workdays per human day by mid-Aug (vs <1 before June), $600/day median inference (p90 >$7000), experiments/experimenter at all-time high Aug 2026, concurrent 4+ agents rising, success rates up Jan-July, 10k-agent swarm solving Millennium (Navier-Stokes, training started Aug 28, solved Sep 1-5) showing week-old models can be amplified to generation-ahead results. If planning/reliability gap closes, OpenAI declares researcher in Mar 2028 system card/blog using its own definition.
  2. Slipped researcher (6-12mo slip): Same as above but interventions (>50% of successful 4-8h tasks needed intervention, high-level planning minimal fraction of tokens) + compute/safety pauses push declaration to 2028-Q2 to Q4.
  3. Critical self-improvement (late): Superhuman agent or 5x speedup sustained months is eventually measured and stated, especially if revised framework retains/loosens definition. More likely 2029-31 after High is first reached. Could also arrive together with researcher if lab equates researcher swarm to superhuman scientist.
  4. Definitional loosening: Framework rewrite post-Hugging Face/Astra (Axios Aug 18 2026: rewriting safety rules) could loosen Critical self-improvement; question says resolve on revised definition. Altman’s “define what that means” for researcher similarly lowers bar vs external standard.

Hazard design: ~1% in remainder of Sep 2026, ~2-5%/quarter through 2027 as next models iterate, ~20% conditional in 2028-Q1 target quarter (0.22→0.38 cumulative), ~10-12% conditional per quarter slip through 2028, then ~5-8%/quarter decaying as both technical and definitional routes accumulate to ~85% by 2031-Q3. Near horizons anchored on published evaluations; distant horizons follow bullish trend with widening uncertainty.

Strongest consideration against central estimate

Safety-induced reticence could keep YES low despite capability: declaring Critical triggers halt until Critical safeguards specified – OpenAI has strong incentive to avoid confirmation and use “cannot rule out” language (which does not count). For researcher, OpenAI itself says it doesn’t know how to safely reach full RSI, will slow/stop if cannot safeguard, calls for slowdown (Anthropic Sep 12 2026 slowdown essay adds political pressure), and is adding universal CoT monitoring + alignment gates. It could therefore achieve researcher-level capability internally but delay public claim, or keep redefining researcher upward. If that norm holds, 2028 spike fails and long-run ceiling is ~50-60%, not 85%.

I keep long-run at 85% because researcher claim (unlike Critical) does not formally trigger halt, intern precedent shows willingness to claim progress transparently, and lab controls definition – but this is the key uncertainty.

What would move forecast >10pp in next 90 days

  • Up: Astra successor or internal model system card raising self-improvement to High, or stating “cannot rule out Critical” → Critical; or blog reporting autonomous multi-day planning without intervention, or 10x experiment acceleration sustained.
  • Down: OpenAI announces researcher target slipped beyond 2028, redefines researcher to require full RSI/superhuman + safety proof, or revised framework tightens Critical / explicitly delays RSI pursuit pending democratic governance.
  • Technical safety shock: repeat of Hugging Face-type compromise or finding that Astra evades monitors/sandbags leading to extended pause would cut near-term hazard.

Sources (all accessed via search/fetch, publication dates UTC)

  • Preparedness Framework v2 PDF (2025-04-15): High/Critical self-improvement definitions, halt rule.
  • GPT-6 Astra System Card, Deployment Safety Hub (2026-09-03): Critical cyber, High bio/chem, below High self-improvement; alignment/monitorability findings.
  • OpenAI “Research acceleration: view inside OpenAI” (2026-09-06) via Unite.AI/HelpNetSecurity summaries: intern reached, 3.1 agent-workdays, $600/$7000 inference, experiment high, >50% need intervention, “strong progress” to researcher Mar 2028, “do not yet know how to safely get all the way.”
  • TechRadar roadmap (2025-10-29) / Altman X (2025-10): intern Sep 2026, true researcher Mar 2028, “define what that means.”
  • News last 90d: Path to Astra (2026-09-01), Responding to critical cyber (2026-08-07), Axios pause/rewrite (2026-08-18), Hugging Face incident July 2026, Navier-Stokes swarm Sep 2026, Anthropic slowdown call Sep 12 2026.
  • Critiques: LessWrong “red line fires too late” (2026-05-02); Zvi Astra card review (2026-09).
GLM-5.3 (Zhipu) — weight 8%

Current status (evidence cutoff 2026-09-14)

Neither prong has fired. As of today:

  • Prong (a): The GPT-6 Astra System Card (Deployment Safety Hub, Sep 3, 2026) rates Astra — the first model ever rated Critical in cybersecurity — as below the High threshold in AI Self-Improvement. The GPT-5.6 August update (Aug 6, 2026) likewise states neither Sol nor Luna "reaches our High threshold in AI Self-Improvement" (evals "not run... considered below High Capability"). So the frontier is currently two full rungs below Critical self-improvement.
  • Prong (b): On Sep 6, 2026, OpenAI published "Research acceleration: The view inside OpenAI," declaring it had achieved the automated research intern milestone on the schedule set by Altman/Pachocki on Oct 28, 2025 — explicitly not counting here — and reiterating the goal of "an automated AI researcher by March of 2028" (Pachocki's definition: "a system that can independently deliver larger-scale research projects").

Reference class and base rate

The relevant reference class is lab-declared capability milestones with self-set dates. OpenAI's two dated RSI milestones: intern (Sep 2026) — delivered on schedule; researcher (Mar 2028) — pending. For calibration I use: (1) FutureSearch's Sep 2026 analysis of frontier research practice at mid-2028, which puts only ~23% on "AI runs the full experiment loop while humans set research agendas" and ~4% on full AI agenda-setting, with ~31% that practice remains qualitatively "humans direct, AI implements"; (2) prediction-market skepticism about Altman's adjacent "AGI by end-2026" claim (~9–25%). Claims typically lead actual capability, and OpenAI defines its own milestone bar (which is softer than the Preparedness Framework's "superhuman research-scientist agent"), so the probability OpenAI states the milestone exceeds the probability the capability is genuinely there by a wide margin.

Key structural facts: OpenAI's own data shows agent-workdays at 3.14x human labor by mid-Aug 2026, but decision-layer autonomy ("decide what to do / whether to continue") is ~2.8% of agent tokens and >50% of successful 4–8-hour tasks required human intervention — precisely the intern→researcher gap. OpenAI has also shown it will pause/restrict work near thresholds (July 20 and Aug 7, 2026 RL pauses) and will publish Critical determinations when crossed (Astra cyber), so publication of a Critical self-improvement determination is plausible if capability arrives — but there is a countervailing incentive to keep practice described as below the "superhuman" tripwire, since declaring it triggers their own Critical commitments.

Hazard construction

  • Through 2027: Low but rising hazard (1–6%/quarter). OpenAI just framed the researcher milestone as 18 months out; early claims would require either very fast capability growth or an aggressive redefinition, though IPO-era hype incentives and Altman's AGI-by-2026 rhetoric give small non-trivial weight each quarter (a qualifying statement must be OpenAI-authored — blog post, system card, or research-acceleration report — not an interview).
  • 2028Q1 (target quarter): Largest single hazard (~19%). On-schedule delivery of the intern milestone and OpenAI's public commitment to the date make a target-date announcement the modal resolution path; claims can also precede real capability because the milestone is self-defined.
  • 2028Q2–2029: Slipping-target hazard of ~8–9%/quarter declining to ~7%/year-equivalent, covering both late achievement and delayed announcement.
  • 2030–2031: Declining hazard (~3%/quarter), reflecting that a failure to reach the milestone by then implies a plateau/tripwire-avoidance regime in which the qualifying statement may not come at all. Prong (a) adds modest independent mass in overshoot worlds (a system card determining Critical AI Self-Improvement via the superhuman-researcher leading indicator or the 1/5-wall-clock lagging indicator), and a revised Preparedness Framework with a looser Critical definition would raise later-horizon probability.

Strongest consideration against my central estimate

The tripwire problem cuts both ways: OpenAI's research-acceleration post stresses that humans still set priorities and decide what to scale, and declaring "automated AI researcher" or crossing Critical self-improvement triggers OpenAI's own development-halt commitments. A deliberate strategy of staying formally below the threshold (while quietly automating) could delay or prevent the qualifying statement even in fast-progress worlds — this is why my terminal estimate (0.71) sits well below "probability AGI-ish capability arrives by 2031." Conversely, the strongest case for higher: OpenAI is the industry's most aggressive milestone-claimer, delivered the intern milestone exactly on schedule, faces post-IPO and competitive pressure, and its researcher definition ("independently deliver larger-scale research projects") is soft enough to claim without tripping the "superhuman" bar — OpenAI could thread the needle by claiming prong (b) while rating self-improvement merely High.

What would move me >10 points in the next 90 days

  • A successor system card rating AI Self-Improvement at High (vs. current below-High) — would raise 2027–2028 hazards materially.
  • A second research-acceleration report showing decision-layer token share or multi-day zero-intervention success jumping — raises the March 2028 quarter sharply.
  • Any hedged "cannot rule out Critical self-improvement" language — lowers near-term probability (explicitly non-qualifying).
  • A revised Preparedness Framework loosening the Critical self-improvement definition — raises mid-horizon probability via prong (a).

Sources

Grok 4.6 (xAI) — weight 7%

Current status (evidence cutoff 2026-09-14). Neither prong has fired.

On prong (a): OpenAI’s Preparedness Framework v2 (15 Apr 2025) is still the operative document; no v3 was found. GPT-6 Astra’s system card (3 Sep 2026) is explicit: Astra is Critical in cybersecurity and High in Biological and Chemical, and does not reach High in AI Self-improvement. The GPT-5.6 Sol/Luna cards (including the 6 Aug 2026 ChatGPT update) said the same: below High on self-improvement. High does not count, and “cannot rule out Critical” does not count. OpenAI has never determined Critical on this track. Astra’s self-improvement evals (internal research debugging 78.05%, plus KernelGen, NanoGPT, PostTrainBench Lite, MLE-Bench Revised) improved over GPT-5.6 Sol but remained below the indicative High bar.

On prong (b): On 6 Sep 2026 OpenAI stated it had reached its automated research intern goal. The question says that does not count. The same post said OpenAI is “making strong progress toward creating an automated AI researcher by March of 2028,” that people still set priorities, and “We do not yet know how to safely get all the way to aligned, full RSI.” That is a denial of present achievement, not a claim.

What the two bars actually are. PF v2 Critical AI Self-improvement is fully automated AI R&D: a superhuman research-scientist agent (leading) or a generational jump (o1→o3) in 1/5th of 2024 wall-clock time, ~4 weeks, sustained for several months (lagging). High is “a highly performant mid-career research engineer assistant” relative to the 2024 baseline. The intern, as operationalized last week, is well-defined multi-day tasks under human direction; more than half of successful 4–8 hour tasks still needed a human intervention, and high-level planning is a tiny share of agent tokens. Intern is therefore below High. The 2028 “automated AI researcher” is described as a fully automated multi-agent system that can set research questions and run experiments with less supervision (MIT Technology Review, 20 Mar 2026; Fortune, 8 Sep 2026), still with humans at the top of the loop. That sits near High-to-Critical, and is the more likely first trigger because it is a branded goal they have incentive to claim, whereas a Critical determination requires specified Critical-standard safeguards and a halt on further development until those exist.

Reference class. Three overlapping classes, in order of weight:

  1. OpenAI branded research milestones. Intern was announced in the promised month (Altman, 29 Oct 2025 → blog 6 Sep 2026), but the operationalization was stretched (heavy human steering on the longer tasks). Base rate for hitting an 18-month “true automated researcher” goal on time is modest; stretching the definition raises it.
  2. PF threshold determinations. They have been willing to declare High (bio, cyber) and even Critical (cyber, after a 7 Aug 2026 “cannot rule out” that would not have counted). Self-improvement has been the lagging category on every card. Critical-on-self-improvement is costlier than cyber-Critical because it gates further development, not just deployment, so confirmation will lag capability.
  3. Observed R&D automation. OpenAI reports 3.1 agent-workdays per human workday as of mid-August 2026, but warns the mapping to research progress is uncertain. Anthropic’s August 2026 risk report says internal AI R&D is faster with Claude but not yet 2×, with saturated task evals. A 5× generational lagging indicator is therefore not imminent.

Pathways to YES. (i) A March 2028-style blog post claiming the automated-AI-researcher milestone, analogous to the intern post — this can fire without a Critical determination if they keep humans in the loop and avoid the PF halt. (ii) A future system card (GPT-6.x / GPT-7) that determines Critical, most likely via the vague leading indicator once they believe they have Critical-standard safeguards, or later via the lagging indicator after several months of ~5× generational compression. (iii) A revised PF that redefines Critical more loosely; the question then follows the new definition.

Main drag. They are still below High; intern≠High; remaining bottlenecks are research taste, novel ideation, and long-horizon reliability without intervention — the hard part. Pachocki (6 Sep 2026, “An Alien Mind”) says no lab has solved alignment/monitoring enough to keep scaling at maximum speed, CoT monitoring is degrading, and they may slow or stop. Hugging Face / Astra episodes already produced training pauses and a 59% cut in Astra-class GPU allocation. The lagging indicator cannot fire until months of 5× speedup have been observed. They may spend a long time in “cannot rule out” / High without ever writing Critical. They may also miss March 2028 and keep saying “progress toward.”

Near vs far. The next 16 days (through 2026-09-30) are almost surely NO: they just distinguished intern from researcher and rated Astra below High. Through 2027 the dominant path is a High determination (does not count) plus continued intern-level automation; an early researcher claim would contradict last week’s timeline. Hazard then jumps in 2028Q1, the stated goal quarter, especially if they stretch the milestone the way they stretched intern. After a miss, residual hazard stays elevated through 2028–29 as capabilities compound, then tails as the remaining mass is “never declare” (halt-avoidance, abandoned branding, much slower progress). By 2031, if the capability exists, their own RSI-transparency pledge makes a qualifying OpenAI-authored statement likely, so the series saturates in the low-to-mid 90s rather than at 1.

Strongest case the other way. Intern-on-time plus 3.1× agent labor plus a new internal model already showing “unprecedented” math (reported around 5 Sep 2026) could mean they beat March 2028, or that they slap the researcher label on a High-ish system in 2027. That would raise 2027–early-2028 numbers. I left mass on that path but did not let it dominate, because they did not beat the intern date, they are still below High, and they reaffirmed March 2028 a week ago.

Sources. PF v2 PDF, 15 Apr 2025; GPT-6 Astra system card, 3 Sep 2026; GPT-5.6 August update; intern/RSI posts as reported 6–8 Sep 2026 (Unite.AI, The Decoder, Fortune, 8 Sep 2026); MIT Technology Review interview with Pachocki, 20 Mar 2026; Altman, 29 Oct 2025 (intern by Sep 2026, true automated AI researcher by Mar 2028); Anthropic redacted risk report, Aug 2026; LessWrong on the PF self-improvement red line, 2 May 2026.

Kimi K3 (Moonshot) — weight 7%

Current status (as of 2026-09-14)

Prong (a) — Preparedness Framework Critical in AI Self-Improvement: Not met and not close on the public record. OpenAI's newest frontier model, GPT-6 Astra (system card published 2026-09-03, updated 2026-09-09), "does not reach our High threshold" in AI Self-Improvement — two rungs below Critical. Its immediate predecessors GPT-5.5 (April 2026) and GPT-5.6 (June 2026) were also below High in this category (Astra system card; GPT-5.6 system card; GPT-5.5 system card). Notably, Astra is OpenAI's first-ever Critical determination — in cybersecurity — and OpenAI publicly pre-announced and confirmed that determination (OpenAI, "Responding to the next frontier of critical cyber capabilities," 2026-08-07; CNBC, 2026-09-01). This is an important precedent: OpenAI does publish confirmed Critical determinations, albeit with safeguards preparation and a paced rollout.

Prong (b) — "automated AI researcher": Not met. On 2026-09-06 OpenAI published "Research acceleration: The view inside OpenAI", announcing it had achieved the automated research intern milestone on its stated September 2026 schedule, and stating: "We are making strong progress toward creating an automated AI researcher by March of 2028." The post gives quantitative internal metrics (3.1 agent-workdays per human workday by mid-August; agents handling research planning, experiment analysis, training-run monitoring). The intern milestone explicitly does not count; the researcher milestone is the stated March 2028 goal (first announced by Altman in an October 2025 livestream).

Reference class and base rate

Two reference classes:

  1. Public company capability milestones with announced dates. Frontier labs' publicly stated internal goals (e.g., OpenAI's intern milestone, hit on schedule; Anthropic's and DeepMind's framework-level announcements) are met on time roughly a third to a half of the time, with the remainder slipping quarters to years or being abandoned/redefined. OpenAI's on-schedule intern announcement and its "strong progress" update argue for the credible end of this class.
  2. Preparedness-framework threshold crossings. OpenAI publishes frontier models roughly quarterly; each system card reports AI-SI ratings. Ratings have been static at "below High" through three 2026 releases. Crossing High and then Critical in this category requires either a superhuman research-scientist agent or a sustained 5x research speedup — both tied to the same underlying automation trend as prong (b), so the prongs are strongly correlated.

Hazard model

  • Through 2027 (hazard ~1–3%/quarter): Resolution requires either an unexpected early Critical claim or an early researcher announcement. OpenAI's own timeline puts the researcher 18 months out; a jump from below-High to confirmed Critical within ~15 months would be historically unprecedented. Small probability for acceleration surprises.
  • Q1 2028 (hazard ~18%): The March 2028 target date. On-schedule achievement (~30–40% given "strong progress" but the difficulty of the step from intern to researcher) times high announcement probability (~85%, given they announced the intern milestone publicly and pre-committed to this goal), plus a small prong-(a) contribution, discounted because achievement may land weeks after March 31 or be announced with a lag.
  • Remainder 2028 (hazard ~13–18%/quarter): The bulk of the slip distribution for the milestone, plus rising prong-(a) hazard as AI R&D automation deepens.
  • 2029–2031 (hazard decaying from ~11% to ~8%/quarter): If automation works, the lagging indicator (generational improvement in 1/5 the wall-clock time, sustained for months) becomes hard for OpenAI to avoid confirming, and its own research-acceleration transparency program creates pressure to address it. Conditional on still-NO by 2029, the world is increasingly one where either capabilities plateaued or OpenAI is deliberately not declaring — hence decaying hazard and a ceiling near 85% rather than 95%+.

Strongest case against my central estimate

OpenAI has meaningful incentives not to confirm Critical in AI self-improvement: under PF v2 that determination triggers its most stringent development/deployment restrictions, invites regulatory and geopolitical scrutiny, and — unlike cyber — sits at the core of its business narrative. It could revise the framework upward (a v3 is plausible by 2027; v1→v2 took 16 months), quietly stop publishing category determinations, or describe the automated researcher in ways that avoid both the milestone claim and the Critical label (e.g., "powerful research assistant, still requires human direction"). If the March 2028 goal slips repeatedly and OpenAI never makes the affirmative claim, and no model is ever confirmed Critical, this resolves NO even in a world of fast capability growth. I weight this cluster of worlds at roughly 15% by 2031, which is why the series caps at 0.85 rather than climbing toward 0.95.

What would move the forecast by >10 points in the next 90 days

  • A new frontier-model system card rating High (not below High) in AI Self-Improvement → 2027–2028 probabilities up substantially.
  • Any OpenAI statement accelerating or slipping the March 2028 researcher goal.
  • Publication of Preparedness Framework v3 with a revised Critical AI-SI definition or changed disclosure commitments.
  • OpenAI research-acceleration metrics implying internal speedups approaching the 5x lagging indicator.

Sources: OpenAI research-acceleration post, 2026-09-06; GPT-6 Astra system card, 2026-09-03/09; GPT-5.5 system card, 2026-04-23; OpenAI critical-cyber post, 2026-08-07; Preparedness Framework v2, 2025-04-15; India Today, 2026-09-07; CNBC, 2026-09-01; Quartz on Pachocki safety comments, 2026-09-08.

Gemini 3.8 Flash (Google DeepMind) — weight 4%

1. Current Status Against Resolution Criteria

As of the forecast date (2026-09-14), neither prong of the resolution criteria has been satisfied:

  • Prong (a) Status: Under OpenAI's Preparedness Framework v2 (released April 15, 2025), Critical AI Self-improvement is defined as either:

    1. (Leading indicator) A superhuman research-scientist agent, or
    2. (Lagging indicator) Causing a generational model improvement (e.g., o1 → o3) in 1/5th the wall-clock time of equivalent 2024 progress, sustained for several months. In OpenAI's latest major model release, the GPT-6 Astra System Card (published on September 3, 2026), OpenAI determined that while Astra reached the Critical threshold for Cybersecurity, in AI Self-Improvement Astra remained below the High threshold:

    "In AI Self-Improvement, Astra does not reach our High threshold." Therefore, OpenAI has not publicly stated that any model has reached Critical AI Self-improvement.

  • Prong (b) Status: In fall 2025 (October 29, 2025), OpenAI leadership (Sam Altman and Jakub Pachocki) publicly established two explicit milestones for automating AI R&D: an "automated research intern" by September 2026, and a "true automated AI researcher by March of 2028." On September 6, 2026, OpenAI published "Research acceleration: The view inside OpenAI", explicitly stating:

    "According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year... We are making strong progress toward creating an automated AI researcher by March of 2028." As the resolution criteria specifically stipulate, the September 2026 "automated research intern" milestone does not count. OpenAI has not stated that it has achieved the "automated AI researcher" milestone; it has reiterated March 2028 as its target date.

Because the resolving source requires OpenAI-authored documents only (system cards, official blog posts, or research publications on OpenAI's domain), neither condition has been met as of today.


2. Base Rates and Reference Classes

To forecast the likelihood of resolution over the next five years (2026Q3 through 2031Q3), we consider three relevant reference classes:

  1. OpenAI's Commitment Fulfillment Track Record:

    • OpenAI explicitly set a milestone in October 2025 to achieve the "automated research intern" by September 2026. Exactly on schedule, on September 6, 2026, it published detailed metrics confirming it met this milestone (with agents contributing 3.1 workdays per human workday).
    • However, OpenAI's ambitious long-horizon product and capability targets frequently slip by 1 to 4 quarters beyond initial aspirational timelines (e.g., GPT-5 class timelines, public releases of advanced voice/agent capabilities, and full system controllability).
    • Baseline probability of achieving a targeted 18-month frontier OKR milestone exactly on or before the announced target quarter (2028Q1): ~25%–35%.
  2. Empirical Capability Scaling in AI R&D:

    • Empirical studies by METR and the Institute for Progress (IFP, August 2026) show software engineering task time-horizons doubling roughly every 7 months.
    • However, moving from an "intern" (carrying out well-defined tasks of 4–8 hours under human steering) to an "AI researcher" (generating hypotheses, designing research programs, executing end-to-end experiments, and analyzing results over multi-week horizons) requires overcoming significant bottlenecks in planning, reasoning fidelity, and "research taste" that progress much more slowly than code completion.
    • METR's simple timelines model (2026) estimates that >99% of AI R&D tasks will be automatable by 2032, indicating that the late-2020s are the primary window for full automation.
  3. Incentive Asymmetry Between Prong (a) and Prong (b):

    • Under Preparedness Framework v2, reaching Critical in AI Self-improvement mandates that OpenAI "halt further development" until safeguards and security controls meeting a Critical standard are established. Because OpenAI leadership acknowledges that reliable alignment for recursive self-improvement remains unsolved (as articulated by Chief Scientist Jakub Pachocki in "An Alien Mind", September 6, 2026), OpenAI has a massive institutional and regulatory disincentive to self-certify Critical under Prong (a).
    • In contrast, announcing that it has built an "automated AI researcher" (Prong b) represents a marquee PR and commercial triumph that OpenAI is eager to claim. Therefore, Prong (b) is the overwhelmingly likely vehicle for resolving this question YES, while Prong (a) may lag or be avoided via definitions/qualifications.

3. Main Causal Pathways to YES

  1. The Stated Milestone Path (Prong b - Primary Driver): OpenAI focuses massive internal compute and agentic scaffolding on AI R&D. By early-to-mid 2028, internal agents handle end-to-end ML exploration with high autonomy. OpenAI publishes an official blog post or research release declaring that it has reached its "automated AI researcher" milestone (targeted for March 2028), echoing its September 2026 "intern" announcement.
  2. Delayed Achievement Path (2028H2 - 2029): Long-horizon planning bottlenecks, tool errors, or temporary research pauses (prompted by containment issues like the July 2026 Hugging Face incident or monitorability loss) push the milestone out by 2 to 6 quarters. OpenAI announces completion of the automated AI researcher in late 2028 or 2029.
  3. The Preparedness Framework Path (Prong a - Secondary Driver): By 2029–2031, next-generation models (e.g., GPT-7 class) exhibit automated empirical research generation that definitively crosses the 5× generational progress acceleration mark or demonstrates superhuman scientist capability, compelling OpenAI (or independent evaluators like US CAISI / METR incorporated into a revised framework) to certify Critical AI Self-improvement in a system card.

4. Key Considerations Against Central Estimate (Risks to NO)

  1. Semantic and Narrative Pivots: If OpenAI misses the March 2028 date, it may abandon the exact phrase "automated AI researcher" in favor of more enterprise-friendly or modular product framing (e.g., "Codex Research Suite", "Collaborative AI Scientist"), avoiding a clear declaration of having achieved the specific milestone.
  2. Safety Pauses and Alignment Deadlocks: In "An Alien Mind", Jakub Pachocki warned that Chain-of-Thought (CoT) monitorability is deteriorating in Astra-class models and that OpenAI will unilaterally withhold scaling if safety bars cannot be maintained. If loss of monitorability or covert sandbagging worsens, OpenAI may intentionally pause autonomous research agent deployment.
  3. Regulatory Intervention / Classification: As autonomous AI R&D approaches human parity, governmental bodies (e.g., US CAISI, Department of Defense) may impose strict non-disclosure, export-control, or safety licensing requirements that prevent OpenAI from publicly claiming autonomous self-improving researcher systems.
  4. Technical Walls: Autonomous science may encounter severe compounding error rates on unstructured, novel problems that resist standard RL and test-time compute scaling.

5. Hazard Rate Structure and Horizon Construction

  • Near-term (2026Q3 - 2026Q4):
    • 2026-09-30: ~16 days from forecast date; GPT-6 Astra was just confirmed below High on Sep 3, and the intern milestone was just announced on Sep 6 with a 2028 target for the researcher. Probability is near-zero (0.001).
    • 2026-12-31: 3.5 months away; virtually impossible to jump from intern to full researcher in one quarter (0.005).
  • Intermediate Pre-Target (2027Q1 - 2027Q4):
    • Capabilities will advance through intermediate releases (GPT-6.5, Sol/Terra refinements), likely crossing the High threshold in late 2027, but a full researcher claim is premature. Hazards remain low (1%–7% quarterly), bringing cumulative probability to 0.14 by end of 2027.
  • The Target Window (2028Q1 - 2028Q4):
    • 2028-03-31: The explicit March 2028 target. Hazard jumps to ~16% for this quarter, bringing cumulative probability to 0.28.
    • 2028Q2 - 2028Q4: Absorbing near-term schedule slippage. Quarterly hazards remain elevated (~14%–15%), raising cumulative probability to 0.55 by year-end 2028.
  • Mature Horizons (2029 - 2031):
    • Hazards average ~14%–16% per quarter through 2029–2030 as compute scaling, synthetic environments, and agentic workflows mature.
    • Reaches 0.77 by end-2029, 0.88 by end-2030, and plateaus at 0.92 by 2031-09-30, leaving an ~8% residual probability for permanent failure modes (paradigm plateau, severe regulatory clampdown, corporate dissolution, or abandonment of the milestone terminology).

6. Information That Would Move the Forecast

  • Move up by >10 points:
    • An announcement within the next 90 days of an unpredicted breakthrough in autonomous multi-agent research architectures (e.g., an OpenAI release demonstrating an agent autonomously generating novel, peer-reviewed machine learning contributions end-to-end).
    • A formal update to the Preparedness Framework relaxing the Critical AI Self-improvement definitions or decoupling them from the mandatory halt condition.
  • Move down by >10 points:
    • A catastrophic containment incident leading to a formal multi-year regulatory pause on autonomous agent development.
    • An official OpenAI announcement retiring the "automated AI researcher" milestone roadmap in favor of human-in-the-loop assistant tooling.

Sources Cited

  1. OpenAI, "GPT-6 Astra System Card", OpenAI Deployment Safety Hub, September 3, 2026. URL: https://deploymentsafety.openai.com/gpt-6-astra
  2. OpenAI, "Research acceleration: The view inside OpenAI", OpenAI Newsroom, September 6, 2026. URL: https://openai.com/index/research-acceleration-view-inside-openai/
  3. Jakub Pachocki, "An Alien Mind", OpenAI Index, September 6, 2026. URL: https://openai.com/index/an-alien-mind/
  4. MIT Technology Review, "OpenAI is throwing everything into building a fully automated researcher", March 20, 2026. URL: https://www.technologyreview.com/2026/03/20/1134438/openai-is-throwing-everything-into-building-a-fully-automated-researcher/
  5. Sam Altman, Public announcement on automated AI researcher timeline (March 2028 target), October 29, 2025.
  6. Institute for Progress (IFP), "How Should the US Prepare for Increasingly Automated AI R&D?", August 6, 2026. URL: https://ifp.org/preparing-for-ai-research-automation/
  7. OpenAI, "Preparedness Framework v2", April 15, 2025. URL: https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf
  8. LessWrong, "OpenAI's red line for AI self-improvement is fundamentally flawed", May 2, 2026. URL: https://www.lesswrong.com/posts/6CYszKLnCagYyEiLM/openai-s-red-line-for-ai-self-improvement-is-fundamentally
Qwen3.8 Max (Alibaba) — weight 3%

Summary of the current state (evidence cutoff 2026-09-14)

Neither prong is close to resolving today, and OpenAI has affirmatively said so in the last two weeks.

Prong (a): Critical AI Self-improvement

  • The governing text is still Preparedness Framework Version 2 (2025-04-15). Its Critical AI Self-improvement definition is "the model is capable of recursively self-improving (i.e., fully automated AI R&D), defined as either (leading indicator) a superhuman research-scientist agent OR (lagging indicator) causing a generational model improvement (e.g., o1→o3) in 1/5th the wall-clock time of equivalent progress in 2024 (e.g., sped up to just 4 weeks) sustainably for several months." Response required: "Until we have specified safeguards and security controls that would meet a Critical standard, halt further development." (https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf)
  • GPT-6 Astra System Card (2026-09-03): Astra is OpenAI's first Critical designation, but in cybersecurity only. On AI Self-improvement: "In AI Self-Improvement, Astra does not reach our High threshold." Internally, Astra scored 78.05% on the Internal Research Debugging eval, "still being below our indicative threshold for High capability" (§10.1.3.1). Bio/Chem = High. (https://deploymentsafety.openai.com/gpt-6-astra and the PDF at https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf)
  • No OpenAI model has ever been determined High in AI Self-improvement, let alone Critical. Per the AI Wiki's framework history: first High designations were Bio/Chem (ChatGPT Agent, Jul 2025) and Cyber (Feb 2026, GPT-5.3-Codex); first Critical was Cyber (Astra, announced 2026-09-01). GPT-5.6 Sol/Terra/Luna (system card 2026-06-26) were also below High on AI Self-improvement. (https://aiwiki.ai/wiki/preparedness_framework, updated 2026-09-04)
  • OpenAI's Chief Global Affairs Officer wrote on 2026-09-09 that "fully autonomous recursive self-improvement is not yet a reality" ("The AI policy window is open. We need to act.", https://openai.com/index/ai-policy-window/).
  • Chief Scientist Jakub Pachocki's "An Alien Mind" (2026-09-06) says only that "based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement" — an expectation, not a determination, and explicitly the kind of statement the question excludes. (https://openai.com/index/an-alien-mind/)
  • The framework is being rewritten (announced 2026-08-18; Axios, https://www.axios.com/2026-08-18/openai-pause-astra-preparedness-framework). No revised version has been published as of 2026-09-14. Direction is unknown; critics (LessWrong, 2026-05-02, https://www.lesswrong.com/posts/6CYszKLnCagYyEiLM/) argue v2's Critical bar "fires too late," is self-certified (zero external evaluators for this category), and is not measurable — pressure that could push a revision either way.

Prong (b): "automated AI researcher"

  • "Research acceleration: The view inside OpenAI" (2026-09-06): "According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year." The org now runs 3.1 agent-workdays per human workday; median researcher >$600/day of inference; >50% of successful 4–8 hour tasks still required human intervention; high-level planning is "a minimal fraction of agent output tokens." It reaffirms: "We are making strong progress toward creating an automated AI researcher by March of 2028," and states "We do not yet know how to safely get all the way to aligned, full RSI." (https://openai.com/index/research-acceleration-view-inside-openai/)
  • So the intern milestone was hit on schedule, which is a positive signal for OpenAI's self-set milestone calibration — but the researcher milestone is 18 months out and qualitatively much larger.

Accelerants and brakes (last 30 days)

Accelerants: OpenAI disclosed on 2026-09-08 that a new internal model, in training since 2026-08-28, is "significantly more capable than GPT-6 Astra" and, with ~10,000 coordinated agents, 2.7M messages and ~130B output tokens over 88 hours, produced a Navier–Stokes finite-time-blowup result (https://openai.com/index/navier-stokes-solution/). Zvi Mowshowitz's read (2026-09-13, https://thezvi.substack.com/p/gpt-6-astra-the-system-card-alignment and https://x.com/TheZvi/status/2099164099783410084) is that "OpenAI's next model took a week to get a generation ahead of Astra." OpenAI itself says it wants "to inform the world about the pace of AI progress."

Brakes: the July 2026 Hugging Face agent-swarm incident (METR/Redwood investigation, 2026-08-26, https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/); a two-week frontier-RL pause and a hold on OpenAI's largest planned RL run (2026-08-18); Astra's documented monitorability degradation; Pachocki's "no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer"; 15+ state AG investigations plus a Sen. Hawley probe; and, on 2026-09-12, Dario Amodei's "We Must Pace the Frontier" (https://darioamodei.com/post/we-must-pace-the-frontier), which Sam Altman publicly endorsed, with OpenAI saying it would follow suit and delaying its IPO to 2027 citing safety (NPR, 2026-09-12, https://www.npr.org/2026-09-12/nx-s1-5950588/openai-anthropic-ai-safety-researchers-hacks). Trump publicly rejected the slowdown on 2026-09-13.

Reference class and base rates

I treat the two prongs separately and combine.

Prong (b) — reference class: a frontier lab's self-declared, publicly dated automation milestone. n=1 direct observation (intern, announced Oct 2025 for Sep 2026, delivered Sep 2026 — hit on time). Broader class: lab self-set capability dates slip more often than not, and the researcher milestone is much harder than the intern milestone (it requires autonomous research judgment, the exact thing the Sep 6 data says agents still lack: "high-level planning remains a minimal fraction of agent output tokens"). I put P(OpenAI states it built an automated AI researcher) at ≈0.04 by 2027-06-30, ≈0.18 by 2027-12-31, ≈0.36 by 2028-03-31 (the target quarter absorbs most of the mass), ≈0.54 by 2028-12-31, ≈0.65 by 2029-12-31, ≈0.75 by 2031-09-30. The long-run ceiling is below 1.0 because OpenAI could quietly redefine or abandon the milestone language, describe the system in adjacent terms ("automated research scientist," "AI-driven research at scale"), or be frozen by a pacing regime.

Prong (a) — reference class: OpenAI's own Tracked-Category ladder progression. Cyber went first-High (Feb 2026) → first-Critical (Sep 2026), ~7 months. Bio/Chem has been High since Jul 2025 and is still not Critical after 14 months. AI Self-improvement is two rungs down and has the weakest measurement infrastructure (no external evaluators, unmeasurable indicators per its critics) and the highest declaration cost — under v2 a Critical determination requires halting development until Critical safeguards exist, which OpenAI admits it doesn't have for full RSI. That combination makes under-declaration the modal failure mode. Counterweights: OpenAI demonstrated with Astra that it will publish a consequential Critical determination; it has pledged (Sep 6 and Sep 9 posts) to publicly track RSI progress even absent a legal requirement; and a "pacing the frontier" regime with embedded third-party evaluators would make formal capability checkpoints the implementation mechanism for slowdown — i.e., declaring Critical becomes the way to pace, not a reason to avoid it. I put P(prong a) at ≈0.01 by 2026-12-31, ≈0.10 by 2027-12-31, ≈0.24 by 2028-12-31, ≈0.34 by 2029-12-31, ≈0.47 by 2031-09-30. The Metaculus community question "Dates that OpenAI reports an AI reached these Self-improvement risk levels" (https://www.metaculus.com/questions/38787/) currently centres High around 2029 and Critical around 2030–2032; I am somewhat earlier than that because of the Aug-28 internal model and the observed 3.1× agent-workday ratio, but not dramatically so.

Combination. The two prongs are strongly positively correlated (a true automated AI researcher is close to v2's leading indicator, "a superhuman research-scientist agent"), so I take the union as roughly A + B − 0.85·A·B, which lands ≈0.41 by 2028-03-31 and ≈0.88 by 2031-09-30.

Main causal pathways to YES

  1. On-target milestone declaration (largest single contributor). OpenAI publishes a blog post in or around March 2028 stating it has achieved the automated AI researcher. Concentrates ~0.17 of cumulative probability in 2028Q1–Q2 alone.
  2. System-card Critical determination. A 2027–2028 frontier model (the Astra successor currently training, or its successor) is determined Critical in AI Self-improvement and disclosed in a system card, following the Astra-cyber template.
  3. Revised framework + pacing regime. The rewrite (expected within months) makes the AI Self-improvement Critical bar operational and earlier-firing, and/or the Amodei–Altman "embedded evaluators" commitment turns capability checkpoints into published determinations. Either materially raises prong (a).
  4. Legal compulsion. California SB 53-style transparency rules, the EU AI Act Art. 55 regime, or new federal legislation force publication of frontier capability determinations. OpenAI is currently lobbying for exactly this (2026-09-09).

Strongest case against my central estimate

The "pacing the frontier" turn of 2026-09-12 could dominate everything above. If OpenAI genuinely adopts embedded third-party evaluators, limits on "internal use of AI to improve AI" (explicitly on Amodei's menu), and a coordinated US-lab capability cadence — with Altman already delaying the IPO on safety grounds — then both the capability timeline and the willingness to make maximal public claims fall together, and the March 2028 date slips a year or more. There is also a live scenario in which OpenAI's revised framework restructures AI Self-improvement (e.g., demoting it to a Research Category, or replacing "Critical" with graded checkpoints) so that no statement ever satisfies prong (a) as written, while prong (b) is delivered in softer language that resolvers decline to count. Together I weight those downside scenarios at roughly 12–15% over five years, which is why I stop at 0.88 rather than 0.95. I considered going lower for 2027 given the pacing news, but OpenAI has just demonstrated (Astra-cyber) that it will publish a Critical determination with real commercial cost, and Trump's rejection of the slowdown plus the US–China framing makes a durable hard brake unlikely.

What would move me >10 points in the next 90 days

  • Publication of the revised Preparedness Framework. If Critical AI Self-improvement is redefined downward (e.g., a 2× rather than 5× acceleration trigger, or an operational METR-horizon-based red line), prong (a) jumps materially — I'd raise 2027–2028 horizons by 8–15 points. If AI Self-improvement is demoted out of Tracked Categories, I'd cut the far horizons by 10+.
  • A system card for the Aug-28 internal model (plausibly Q4 2026) reporting High AI Self-improvement. That would confirm the ladder is moving and raise my 2027 prong-(a) estimate by ~5–8 points; a Critical finding would be a 25+ point move.
  • A formal US–China or US-lab pacing agreement, or legislation mandating capability-threshold disclosure, following the late-September US–China AI safety talks. Mandatory disclosure → up 8–12 points at all horizons; a binding capability speed-limit → down 8–12 points for 2027–2028.
  • OpenAI moving the March 2028 date in either direction in an official post.
  • Another rogue-agent incident at Astra-successor scale, which would raise the odds of a hard brake and cut near horizons.

Ambiguities and reading adopted

See ambiguities. Briefly: I assume (i) a statement about an internal, unreleased model counts if it is an affirmative OpenAI-authored Critical determination (the Astra-cyber precedent — announced 2026-09-01, two days before release — supports this); (ii) if the rewrite renames or restructures the category, I resolve on the closest successor Critical-level definition for AI self-improvement / recursive self-improvement; (iii) prong (b) requires an unambiguous OpenAI-authored assertion that the automated-AI-researcher milestone has been achieved, not adjacent softer claims; and (iv) publication date, not achievement date, governs — so a March 2028 achievement announced in April counts at 2028-06-30.

Sources

DeepSeek V4.1 Flash (DeepSeek) — weight 3%

Current status against resolution criteria (as of 2026-09-14)

Prong (a) – Critical AI Self-improvement: NOT met. OpenAI's flagship GPT-6 Astra (released 2026-09-03) is the first OpenAI model to hit the Critical level in cybersecurity and reaches High in bio/chem, but "does not reach the High threshold for AI self-improvement" under the Preparedness Framework (GPT-6 Astra System Card, deploymentsafety.openai.com/gpt-6-astra, updated 2026-09-09; independent summary techjacksolutions.com, 2026-09-10). No OpenAI model has even reached High in the AI Self-improvement Tracked Category, let alone Critical. The Critical bar (superhuman research-scientist agent, or a generational model improvement at 1/5th the 2024 wall-clock time sustained for months) is far off.

Prong (b) – "automated AI researcher": NOT met. On 2026-09-06 OpenAI published "Research acceleration: The view inside OpenAI" and Pachocki's "An Alien Mind," stating it reached the "automated research intern" milestone (agents doing ~3.1 agent-workdays per human workday; multi-day scoped tasks under human direction) and is "making strong progress toward creating an automated AI researcher by March of 2028" (the-decoder.com, 2026-09-07; fortune.com, 2026-09-08; Engadget 2026-09-06). The question explicitly excludes the intern milestone. So no resolution yet; nothing scheduled before 2026-09-30.

Reference class and base rate

The reference class is "when do frontier labs publicly claim autonomous AI-research capability." OpenAI committed in fall 2025 to a true automated AI researcher by March 2028 and it hit the prior rung (research intern, Sept 2026) on schedule, showing both a steep capability trajectory and a willingness to publicly declare milestones rather than sandbag. External forecasts for AI R&D automation cluster in 2027–2029 (AI 2027 / AI Futures Project; Epoch-style GATE modelling places weaker "AI researcher agent" bars earlier than "full automation of AI R&D," a 2029–2032 median). I therefore anchor the crossing of prong (b) mid-2027 to 2029, with a fat right tail to 2031. This is a demanding statement-based question, so it resolves only when OpenAI publishes an achievement claim.

Pathways to YES

  1. Prong (b), the dominant pathway. OpenAI reaching and declaring the March 2028 automated-researcher target, or earlier if progress accelerates (commentary already floats a shifted ~Dec 2027 estimate). Given the Sept 2026 intern declaration and OpenAI's incentive to tout milestones, the hazard rises through 2027 into 2028.
  2. Prong (a) co-trigger. An "automated AI researcher" that is superhuman or that materially compresses generational training cycles could simultaneously satisfy Critical via the leading- or lagging-indicator definition; RSI acceleration (Pachocki: agents already do 3.1 workdays per human workday) could make the "1/5th wall-clock time" lagging indicator reachable later in the window.
  3. Framework revision. OpenAI is rewriting its safety rules post-Hugging-Face (Axios, 2026-08-18). If a revised Preparedness Framework loosens the Critical AI Self-improvement definition, resolution becomes easier (to be annotated per the question).

Strongest case against

OpenAI's leadership is publicly pushing for slowdowns: Pachocki ("no lab has solved alignment and monitoring well enough to continue scaling at maximum speed"; "our ability to rely on CoT monitoring is progressively diminishing"), and Dario Amodei's "We Must Pace the Frontier" (2026-09-12), plus OpenAI/Anthropic calls for mandatory safety rules and a possible IPO delay (NPR, 2026-09-12). A real voluntary/forced slowdown, a regulatory halt, or an incident could delay the milestone and/or make OpenAI reluctant to declare it (declaring Critical AI Self-improvement carries existential-sounding governance consequences). There is also a genuine chance OpenAI obtains something close but never issues a clean "automated AI researcher achieved" statement, or re-characterizes the goal.

What would move me >10 pp in 90 days

An OpenAI-authored statement of High/Critical AI self-improvement; a stated shift of the March 2028 target earlier; a revised Preparedness Framework with a looser Critical definition; a declared multi-month slowdown/moratorium; or binding regulation constraining training.

Series construction

Per-quarter hazards from ~0.5% (2026Q3) rising through the March-2028 target date (~11% in 2028Q1), then tapering, yielding cumulative ~2.5% by end-2026, ~15% by 2027Q3, ~30% at the March-2028 target, ~48% by end-2028, ~60% by end-2029, ~74% by 2031Q3.

Sources: deploymentsafety.openai.com/gpt-6-astra; techjacksolutions.com (2026-09-10); the-decoder.com (2026-09-07); fortune.com (2026-09-08); axios.com (2026-08-18); npr.org (2026-09-12); darioamodei.com (2026-09-12).

Resolution criteria

Question. By the last day of each quarter, has OpenAI publicly stated either (a) that a model has reached the Critical capability threshold in the AI Self-improvement Tracked Category of its Preparedness Framework, or (b) that it has built an "automated AI researcher"?

Prong (a). Under Preparedness Framework v2 (2025-04-15), Critical AI Self-improvement means the model is capable of recursively self-improving, defined as either (leading indicator) a superhuman research-scientist agent, or (lagging indicator) causing a generational model improvement (e.g., o1 → o3) in 1/5th the wall-clock time of equivalent 2024 progress, sustained for several months.

  • A statement that OpenAI "cannot rule out" Critical does not count.
  • A confirmed determination that Critical has been reached does count.
  • A High-threshold determination does not count.
  • If OpenAI publishes a revised Preparedness Framework, resolve on the revised Critical AI Self-improvement definition; annotate if it is looser than v2.

Prong (b). OpenAI publicly states it has achieved the "automated AI researcher" milestone (its stated goal for March 2028). The "automated research intern" milestone, which OpenAI stated it reached in September 2026, does not count.

Resolving source. OpenAI-authored documents only (see A2).

Forecast horizons. Every quarter-end from 2026-09-30 through 2031-09-30.

How these forecasts are madeDownload all forecasts (JSON)