How close are we to artificial general intelligence?When will AI learn almost any thinking task?
or, simply: When will AI learn almost any thinking task?or, precisely: How close are we to artificial general intelligence?
Frontier models saturate most static benchmarks, yet reliable long-horizon autonomy, calibrated reasoning and a shared definition of general intelligence all remain open - so any single 'percent to AGI' is a guess dressed as a number.AI can already help with many kinds of thinking, but it still makes confident mistakes and struggles to carry a big job through on its own.
MirrorCode tests whole-program reconstruction - MirrorCode asked agents to recreate complete programs from their behavior without source code. Its strongest model passed 56% across 25 targets and nearly reproduced the 16,000-line gotree toolkit, but used unusually precise executable feedback and large budgets; the claimed weeks of human work were estimates, not completed trials. Next up - Forecast: "weakly general" AI (expected 2028).
State of playWhere agi stands right now
The current stage, the honest metric, and the single threshold that gates the next stage. Each threshold is a falsifiable claim with a named next test.How far up the ladder we've climbed, the honest verdict, and the one thing blocking the next step.
AI can already help with many kinds of thinking, but it still makes confident mistakes and struggles to carry a big job through on its own.Frontier models saturate most static benchmarks, yet reliable long-horizon autonomy, calibrated reasoning and a shared definition of general intelligence all remain open - so any single 'percent to AGI' is a guess dressed as a number.
No single honest number for this one - read where it stands by the milestones and the next test.No single honest scalar - progress is read by milestones and the named next test, not a headline figure.
Learn many different kinds of workBroad cognitive competence Next test: ARC-AGI-2 and held-out expert evals resistant to training-data leakage
The thresholds that gate the next stageWhat has to happen next
Each threshold is a falsifiable claim with a named next test; the gap chart shows how far today's metric sits from the goal.Each row is one thing that has to be proven — and how far today's number is from the target.
The record behind the verdict
Major events set large; context events set small but never hidden. Everything below the TODAY rule is a schedule, not a result.
The symbolic dream
The symbolic dream moved the field from turing asks "can machines think?" to the perceptron learns from data. The results narrowed the next question without closing it.
Winters and the wilderness
Winters and the wilderness moved the field from the first ai winter sets in to the second ai winter. The results narrowed the next question without closing it.
Machines that learn
Machines that learn moved the field from deep blue beats kasparov to alexnet ignites the deep-learning era. The results narrowed the next question without closing it.
Foundation models and agents
Foundation models and agents moved the field from the transformer architecture to forecast: "weakly general" ai. The results narrowed the next question without closing it.
Events outside the declared eras
Events outside the declared eras moved the field from alphago defeats lee sedol to projected: ai handles week-long tasks. The results narrowed the next question without closing it.
Why the meters read the way they do
The learning curves and comparisons that justify each threshold's percentage. Every series is measured, with the source event linked in the timeline above.
Read the evidence more closely
Definitions, system boundaries and experimental caveats behind the headline record.
01Autonomy gains depend on the time window
TH1.1's full 2019-2025 fitted doubling time remains about 196 days, but its post-2023 estimate shortened from 165.3 to 130.8 days and its post-2024 estimate from 108.9 to 88.6 days. These rates are sensitive to task selection and should not be extrapolated as physical laws.
02Long-task baselines are mostly estimated
Only 5 of TH1.1's 31 tasks estimated at eight hours or longer have measured human baselines; the rest use expert estimates. The suite covers mainly software engineering, machine learning and cybersecurity, not general intellectual work.
03What a time horizon measures
A time horizon denotes the human duration of task difficulty at which fitted agent success is 50%, not how long the agent literally operates. Agents are often several times faster than the associated human time.
04MirrorCode gives unusually rich feedback
MirrorCode gives agents execute-only access to the reference program, documentation, shell tools and very large inference budgets-up to one billion tokens, about $550 per task in preliminary experiments. This unusually precise, testable specification helps explain its much longer apparent horizons.
05The gotree baseline is uncertain
The gotree human-duration range-2-17 weeks-comes from four informal estimates, not completed human trials. A shorter 2,000-line baseline was still only 42% complete after 20 hours.
06Interactive adaptation remains weak
ARC-AGI-3 tests continual interactive adaptation rather than static answer production. In later replay analysis, GPT-5.5 scored 0.43% and Opus 4.7 scored 0.18%; common failures included learning a local action effect without forming a correct world model and failing to transfer a successful level strategy.
07Why the 2025 IMO result is stronger
The 2025 IMO result is technically stronger than the 2024 silver result: it worked end-to-end in natural language within 4.5 hours, whereas the 2024 systems needed human translation into formal languages and as much as three days. The IMO validated the submitted proofs, not Google's model or experimental process.
08Productivity estimates remain inconclusive
METR's attempted late-2025 productivity follow-up produced raw estimates of 18% speedup for returning developers and 4% for new recruits, but both confidence intervals included zero and METR judged selection and time-reporting biases too severe for a reliable magnitude.
09Benchmark errors limit headline scores
Static benchmark quality is itself a constraint: Stanford's 2026 review reports invalid-question estimates ranging from 2% on MMLU Math to 42% on GSM8K, reinforcing that saturated scores should not directly set an AGI percentage.
Every test we can still hold it toThe tests — and who leads each one
Each row is an independent benchmark: its verified frontier, any named human mark, and how much ruler remains. Solid dots are steward-verified; dashed circles are lab claims. No row is averaged into another.Each row is one test. A solid dot is independently checked; a dashed circle is a company claim. We never add the rows into one big number.
7 rulers tracked · 5 still standing · 2 outgrown or withdrawn
READ THIS BOARD — a dot near the right edge means one ruler is nearly used up, not that AGI is nearly done. The rows that remain stubborn are often the more useful evidence.
If the remaining tests pass
Downstream capabilities, drawn dashed because they depend on results not yet in.
Who is building it-and what the money saysCapital, institutions and the global race
The teams doing the work, where they are based, and whether the money points to real delivery or only a plan.Company finance, public programmes, institutional leadership and market evidence-kept separate from valuations, forecasts and announced capacity.
The frontier is concentrated in a few heavily financed US labs, with Google and Meta funding research from corporate balance sheets and Mistral anchoring Europe's independent model effort. Capital and compute have scaled much faster than evidence of robust general intelligence: no lab has demonstrated the page's cross-domain, long-horizon threshold.
Who is building itCompanies, laboratories and programmes
OpenAI
USAFrontier multimodal and reasoning models, agent products and large-scale training infrastructure.
Anthropic
USAFrontier Claude models with interpretability, alignment and model-safety research.
xAI
USAGrok model family and vertically integrated large-scale training infrastructure.
| Player | Country | What they are doing | Funding | Named investors | Source |
|---|---|---|---|---|---|
| OpenAIcompany | USA | Frontier multimodal and reasoning models, agent products and large-scale training infrastructure. Committed capital is not cumulative funding, revenue or money already spent on compute. | $122B committed capital · Mar 2026 · $852B post-money valuation | Amazon · NVIDIA · SoftBank · Microsoft | Source · openai.com |
| Anthropiccompany | USA | Frontier Claude models with interpretability, alignment and model-safety research. The reported $5B August 2025 revenue figure is an annualized run-rate, not audited full-year revenue. | Series F · $13B · Sep 2025 · $183B post-money valuation | ICONIQ · Fidelity · Lightspeed · BlackRock affiliates · Qatar Investment Authority | Source · anthropic.com |
| xAIcompany | USA | Grok model family and vertically integrated large-scale training infrastructure. No cumulative total is encoded because the release does not reconcile equity, debt and infrastructure financing. | Series E · $20B · Jan 2026 | Valor Equity Partners · Fidelity · Qatar Investment Authority · NVIDIA · Cisco | Source · x.ai |
| Mistral AIcompany | France | European frontier and open-weight models, enterprise deployment and sovereign AI tooling. ASML invested €1.3B; valuation is stated separately from capital raised. | Series C · €1.7B · Sep 2025 · €11.7B post-money valuation | ASML · Bpifrance · NVIDIA | Source · presse.bpifrance.fr |
| Google DeepMindlab | UK / USA | Gemini frontier models, reinforcement learning, AI-for-science and robotics research. Funded inside Alphabet; corporate capex is not a standalone lab round. | Not disclosed | Not disclosed | Source · deepmind.google |
| Meta AIlab | USA | Open-weight Llama models and research in multimodal, world-model and embodied AI. Funded inside Meta; company-wide infrastructure spending is not lab funding. | Not disclosed | Not disclosed | Source · ai.meta.com |
Where every number comes from
4 sources — every figure on this page traces to one.