AGI: Timelines & Definitions
Why the People Building It Can't Agree When It Arrives
TL;DR:
AGI has at least three competing working definitions (OpenAI's economic-output framing, DeepMind's five-level performance matrix, Amodei's capability-list “powerful AI”) and no agreed arrival date. Researchers track progress with concrete metrics instead — METR's task-length horizon (doubling roughly every 7 months) and Epoch AI's compute growth (~4–5x per year) — while the 2023 AI Impacts survey of 2,778 researchers put 50% odds of human-level AI at 2047, thirteen years sooner than the same survey found just one year earlier. This course teaches you to read that spread instead of anchoring on any single date.
Who this course is for
This course is for anyone who wants to evaluate AGI claims — in a headline, an investor pitch, a policy debate, or a colleague's hot take — without either dismissing or accepting them on reputation alone. It's especially useful for founders and investors making multi-year bets, policymakers and analysts tracking frontier AI governance, and anyone planning a career or curriculum around an uncertain timeline.
No technical AI background is required. This is a course about how to read evidence and disagreement, not about how transformer models work internally.
What you'll learn
Competing Definitions
OpenAI's economic-output framing, DeepMind's five-level matrix, and Amodei's capability-list "powerful AI" — and why the choice changes the answer.
Real Progress Metrics
METR's task-length time horizon and Epoch AI's compute-growth trend — concrete, tracked numbers instead of vibes.
Why Experts Disagree
The AI Impacts researcher survey, public lab-leader forecasts, and the Existential Risk Persuasion Tournament's finding of persistent, un-bridged disagreement.
A Fact-Check Framework
Four questions to ask about any AGI headline before you update your beliefs or your plans.
Planning Under Uncertainty
Why a 20-year-wide probability range is more useful for real decisions than a single confident date.
What to Actually Track
Which specific metrics to bookmark and revisit instead of re-reading the same debate every few months.
Academy review: July 17, 2026
Model names and capabilities change frequently. We use provider families (GPT, Claude, Gemini) in Academy copy — verify current tiers on provider sites or our LLM Model Family Guide.
For live model picks, cross-check with our The Coming Wave course.
Living document — every number here will be superseded
Module 1 — What “AGI” actually means (and why that's contested)
There is no single, universally accepted definition of “artificial general intelligence.” Three influential framings currently compete, and each implies a different way of deciding whether AGI has “arrived.”
OpenAI — Economic output
OpenAI's Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” A single, output-based bar — but one that's hard to measure until it happens.
DeepMind — Levels & breadth
DeepMind's “Levels of AGI” paper proposes five performance levels — Emerging, Competent, Expert, Virtuoso, Superhuman — crossed with how broad the capability is. AGI is a spectrum, not a finish line.
Amodei — “Powerful AI”
Anthropic's Dario Amodei avoids the term “AGI” in “Machines of Loving Grace” and instead lists concrete capabilities — proving unsolved theorems, writing difficult codebases from scratch — comparing it to “a country of geniuses in a datacenter.”
DeepMind's framework is worth understanding in more detail because it's the most explicit about grading, not just declaring. Its five performance levels are defined by percentile against skilled human adults: Emerging (equal to or slightly better than an unskilled human), Competent (outperforms 50% of skilled adults), Expert (90th percentile), Virtuoso (99th percentile), and Superhuman (outperforms 100% of humans). By this scheme, the paper's authors classify current large language models like ChatGPT as “Emerging AGI” — broad in the tasks they attempt, but not yet reliably better than an unskilled human across most of them, let alone at the “Competent” level. No system has reached a higher level as general-purpose AI as of the paper's analysis.
The practical takeaway: “has AGI arrived?” is not a single yes/no question. The more useful question is “AGI by which definition, at which performance level, in which domain?” — and most headlines that claim otherwise are quietly picking one definition without saying so.
Module 2 — How researchers actually measure progress
Because static benchmarks get solved and replaced, researchers increasingly track trend lines instead of single scores. Two are widely cited:
METR's task-length time horizon
METR measures the length of a task (in typical human completion time) that a model can finish autonomously with 50% success. That horizon has been doubling roughly every 7 months across six years of models — from near-100% success on tasks under ~4 minutes to under 10% success on tasks over ~4 hours, as of METR's analysis.
Epoch AI's compute growth
Epoch AI tracks training compute across hundreds of notable models: frontier-model compute has grown roughly 4–5x per year since 2010, driven mostly by spending rather than hardware efficiency alone.
Both organizations are explicit that these are descriptions of a historical trend, not a guaranteed law. Extrapolating an exponential curve years into the future is exactly the kind of forecast that has been wrong before in the history of technology — which is precisely why Module 3 treats the resulting date estimates as a wide range, not a countdown.
Module 3 — Why expert timelines vary by decades
Three independent, verifiable data points show just how wide the real spread is — and none of them is obviously wrong.
Published researchers: 2047 (50% odds)
The 2023 AI Impacts survey of 2,778 researchers who had recently published at top ML venues (NeurIPS, ICML, ICLR and others) put 50% probability of “high-level machine intelligence” at 2047 — thirteen years sooner than the same survey found in 2022 (2060). The estimate itself moved by over a decade in twelve months.
Lab leaders: as early as 2026–2027
Anthropic's Dario Amodei has written that “powerful AI” could arrive “as early as 2026,” and Anthropic's policy filings pointed to late 2026 or early 2027. OpenAI's Sam Altman has written in “The Intelligence Age” that superintelligence is possible within “a few thousand days” — roughly a decade. Both are years to a couple of decades sooner than the researcher survey median.
Forecasters vs. domain experts: the widest gap in the whole study
The Forecasting Research Institute's Existential Risk Persuasion Tournament put 89 superforecasters and 80 domain experts through four months of structured, incentivized debate across several long-run risks. AI produced the largest, most persistent disagreement of any topic tested — domain experts stayed substantially more concerned about nearer-term risk, superforecasters stayed skeptical of long-horizon extrapolation, and neither side meaningfully moved the other.
None of these groups is obviously more credible than the others by default. Lab leaders have the closest technical view but also the strongest incentive (fundraising, recruiting, policy influence) to sound imminent. Surveyed researchers are less commercially exposed but were asked a single definitional question that may not match any one lab's roadmap. Superforecasters have the best calibration track record on short, resolvable questions, but that skill has never been tested on a question this far out. The honest conclusion is that this is a live, unresolved disagreement among serious people — not a case of some of them simply not knowing what they're talking about.
Module 4 — Reading an AGI headline or claim critically
You don't need to resolve the expert disagreement to evaluate a specific claim well. Four questions catch most overstated AGI headlines:
- Which definition are they using? Economic output (OpenAI), a performance level (DeepMind), or a capability list (Amodei-style)? Claims rarely state this, but the definition determines whether the claim is even coherent.
- Is this a benchmark score, a real deployment, or an announcement? A strong eval result, a supervised pilot, and a press release are three very different levels of evidence — and headlines routinely collapse them into one.
- What's the source's incentive? A funding round, a product launch, or a policy debate all reward sounding closer to (or further from) AGI than the evidence alone supports.
- Is the underlying trend continuing, plateauing, or being reinterpreted? Check the actual metric — task horizon, compute, benchmark saturation — rather than the framing around it.
Use the template below the next time an AGI claim crosses your feed — paste it into your notes, or into an AI chat tool alongside the claim, to force yourself through all four questions before you update your beliefs or your plans.
I'm evaluating this AGI-related claim: [paste the headline, quote, or announcement here]
Walk me through it against this checklist:
1. Which definition of AGI (or "powerful AI") does this claim implicitly use — economic output, a specific performance level, or a capability list? Is that stated or assumed?
2. Is the underlying evidence a benchmark score, a real-world deployment, or a company announcement/funding event? Be specific about which.
3. Who is making the claim, and do they have a financial, political, or reputational incentive for it to sound more or less imminent than the evidence supports?
4. What concrete, trackable metric (if any) does this claim relate to — and is that metric's trend continuing, plateauing, or being reinterpreted based on recent data?
Give me a one-paragraph honest assessment of how much this specific claim should move my beliefs, separate from how it was framed.Module 5 — Planning under real uncertainty
Put the three data points from Module 3 side by side and the honest range for “AGI or powerful-AI-equivalent capability” runs from roughly 2026 (the most aggressive lab forecasts) to 2047 and beyond (the researcher-survey median), with serious, calibrated forecasters on record doubting the shorter end. That is not a failure of the field to agree — it is the actual state of the evidence.
The practical implication: career, investment, product, and policy decisions that only make sense if a specific date is correct are fragile bets, not informed ones. Decisions that hold up across most of that 2026–2047+ range — investing in judgment and accountability-heavy skills, building governance now rather than waiting for a clear signal, tracking the underlying metrics instead of the headlines — are the ones worth prioritizing today.
From here, The Coming Wave covers how governance and containment frameworks are trying to prepare for that range of outcomes, and The Future of Work applies the same “don't bet on one date” logic to your own career.
Ready to Apply What You Learned?
AI Critical Thinking
The general evaluation habits this course applies specifically to AGI claims.
Start LearningRisks & Responsible Use
Know these before you go further.
Treating One Benchmark Score as "AGI Achieved"
A strong result on a single eval (a coding benchmark, a reasoning test) is routinely reported as if it settles the general question. By DeepMind's own "Levels of AGI" framework, a high score on one narrow benchmark says nothing about breadth — and current systems are still classified well below the "Competent" performance level even where they excel on specific tasks.
What this means for you
Ask which specific capability was tested, how broad it is, and whether independent replications exist — before treating any single result as evidence AGI has "arrived."
Confusing a Lab Prediction with a Peer-Reviewed Forecast
A CEO's essay, a funding-round pitch, and a 2,778-researcher survey are not the same kind of evidence, even when they get reported in the same tone. Lab leaders have real technical insight but also real incentives (fundraising, recruiting, policy influence) to sound imminent.
What this means for you
Note the source type explicitly — company statement, survey, or forecasting tournament — before letting any single AGI timeline claim move your plans.
Anchoring Career, Investment, or Policy Decisions to One Date
Verified expert estimates for AGI-equivalent capability span from roughly 2026 to 2047 and beyond. A decision that only makes sense if one specific date is correct is a fragile bet dressed up as a plan, not an informed one.
What this means for you
Stress-test any AGI-dependent decision against the full 2026–2047+ range, not just the date you find most persuasive.
Extrapolating a Trend Line Past What Its Own Authors Claim
METR and Epoch AI are explicit that their task-horizon and compute-growth trends are historical extrapolations, not physical laws — data availability, energy, chip supply, and diminishing returns are all open constraints on whether the curves continue.
What this means for you
Treat trend-line extrapolations as one input among several, and revisit the source's own caveats before repeating a multi-year projection as settled fact.
Test Your Knowledge
Complete this quiz to test your understanding of AGI definitions, progress metrics, and why expert timelines disagree.
Loading quiz...
Frequently asked questions
Key Insights: What You've Learned
AGI has at least three competing definitions — OpenAI's economic-output bar, DeepMind's five-level performance matrix, and Amodei's capability-list "powerful AI" — so "has AGI arrived?" isn't answerable without first asking which definition.
Researchers track concrete trend lines instead of vibes: METR's task-length time horizon (doubling ~7 months) and Epoch AI's compute growth (~4–5x/year) — both explicitly framed by their own authors as extrapolations, not guarantees.
The honest timeline is a range, not a date: a 2023 survey of 2,778 researchers put 50% odds at 2047, lab leaders have publicly forecast 2026–2027, and the Forecasting Research Institute found AI produced the largest, most persistent expert disagreement of any long-run risk it studied.