About
takeoff.watch tracks how fast the transition to advanced AI is happening, using a set of concrete questions that can be checked against the public record.
Why this exists
Conversations about AI progress tend to be vague: "soon," "transformative," "doom." Vague claims are hard to check and easy to quietly revise. This site takes the opposite approach. Each question has precise resolution criteria, is forecast on a regular schedule, and is published with the full reasoning behind it, so the outlook can be checked against what happens.
What's tracked
The questions currently fall into five areas:
- Research automation — whether AI developers declare that AI is substantially automating their own research. See Anthropic’s AI R&D threshold and OpenAI’s automated AI researcher.
- Control — whether developers keep hold of their models: whether an AI system keeps running after its developer tries to shut it down, and whether the weights of a closed frontier model are stolen. See AI survives a shutdown attempt and theft of closed frontier weights.
- Harm — whether AI is confirmed to have caused serious damage: AI-executed cyberattacks, and biological or chemical weapons. See AI-executed cyber incidents and biological and chemical weapons.
- Economic effects — what happens to US GDP growth and the share of income going to workers.
- Governance — how far a US–China agreement and US federal regulation go.
The set will change over time as questions resolve, lose relevance, or are replaced by better ones.
How it works
Each question is forecast independently by an ensemble of AI models from several developers. Every model researches the question with web search, writes out its reasoning, and gives probabilities or ranges for every quarter over the next five years. The forecasts are checked for consistency, then combined, weighting stronger models more heavily. The combined forecast and every model's reasoning are published together. The methodology page covers the details, and every question page links to a download of its forecast data.
Can AI models forecast?
Close to as well as the best humans, on current evidence, and the gap has been closing fast.
The most rigorous comparison is the Forecasting Research Institute’s ForecastBench, a live benchmark that scores models only on questions that resolve after their training cutoffs, against a panel of superforecasters, the top few percent of human forecasters. Through 2025 the models trailed but gained at a steady rate, and in July 2026 FRI reported that several AI systems had become statistically indistinguishable from superforecaster accuracy, with one system ahead on prediction-market questions. Metaculus’s FutureEval tournaments tell the same story from the other side: its professional forecasters still beat the best bots in spring 2026, but by a margin no longer statistically distinguishable from zero, down from a wide lead a year earlier.
So the fair summary is near parity, not clearly past it, with two caveats: superforecaster baselines are expensive to keep current, and models can be confidently wrong on novel, one-off events. What the research agrees on is what helps: broad research before forecasting, combining several models rather than trusting one, and careful calibration. That is how the forecasts here are produced. They are not oracles, but they are a reasonable, checkable, and improving read on where things stand, and this site keeps score.
Who runs it
I'm Preston Jensen, a recent AI master's graduate from UT Austin. This is an independent, one-person project, unaffiliated with any AI developer. I'm concerned about the pace of AI progress and wanted a way to track it from outside the labs. LinkedIn.
Contact
Questions, corrections, or ideas for new questions: hello@takeoff.watch.