About

takeoff.watch tracks how fast the transition to advanced AI is happening, using a set of concrete questions that can be checked against the public record.

Why this exists

Conversations about AI progress tend to be vague: "soon," "transformative," "doom." Vague claims are hard to check and easy to quietly revise. This site takes the opposite approach. Each question has precise resolution criteria, is forecast on a regular schedule, and is published with the full reasoning behind it, so the outlook can be checked against what happens.

What's tracked

The questions currently fall into five areas:

The set will change over time as questions resolve, lose relevance, or are replaced by better ones.

How it works

Each question is forecast independently by an ensemble of AI models from several developers. Every model researches the question with web search, writes out its reasoning, and gives probabilities or ranges for every quarter over the next five years. The forecasts are checked for consistency, then combined, weighting stronger models more heavily. The combined forecast and every model's reasoning are published together. The methodology page covers the details, and every question page links to a download of its forecast data.

Can AI models forecast?

Close to as well as the best humans, on current evidence, and the gap has been closing fast.

The most rigorous comparison is the Forecasting Research Institute’s ForecastBench, a live benchmark that scores models only on questions that resolve after their training cutoffs, against a panel of superforecasters, the top few percent of human forecasters. Through 2025 the models trailed but gained at a steady rate, and in July 2026 FRI reported that several AI systems had become statistically indistinguishable from superforecaster accuracy, with one system ahead on prediction-market questions. Metaculus’s FutureEval tournaments tell the same story from the other side: its professional forecasters still beat the best bots in spring 2026, but by a margin no longer statistically distinguishable from zero, down from a wide lead a year earlier.

So the fair summary is near parity, not clearly past it, with two caveats: superforecaster baselines are expensive to keep current, and models can be confidently wrong on novel, one-off events. What the research agrees on is what helps: broad research before forecasting, combining several models rather than trusting one, and careful calibration. That is how the forecasts here are produced. They are not oracles, but they are a reasonable, checkable, and improving read on where things stand, and this site keeps score.

Who runs it

I'm Preston Jensen, a recent AI master's graduate from UT Austin. This is an independent, one-person project, unaffiliated with any AI developer. I'm concerned about the pace of AI progress and wanted a way to track it from outside the labs. LinkedIn.

Contact

Questions, corrections, or ideas for new questions: hello@takeoff.watch.