Artificial General Intelligence (AGI)

AGI is AI matching human ability across most cognitive work — and there is no agreed test for it, so the definition you pick decides the answer.

Published Updated

On this page

Definition

Artificial General Intelligence (AGI) is a hypothetical AI system that matches human ability across essentially the whole range of cognitive work rather than in one narrow domain — and there is no agreed test that would tell you when one had arrived. That second half is not a footnote: it is why two well-informed people can look at the same model and reach opposite conclusions about how close we are, and why the most useful thing you can learn about AGI is how to spot which definition a given claim is using.

The disagreement is not philosophical hair-splitting. It has been written into commercial contracts, into the safety policies that decide what a lab is allowed to deploy, and into regulation with fines attached. Three definitional families are in circulation, and they do not agree with each other about the present:

  • Economic. OpenAI's charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work." The threshold is about work displaced; it says nothing about how the system does it, so a very large collection of narrow competences could qualify.
  • Learning efficiency. François Chollet's 2019 paper On the Measure of Intelligence argues that intelligence is skill-acquisition efficiency — how little experience you need to handle a task you have never seen. Under this view a system that has memorised a domain has skill, not intelligence, and measuring it on tasks it was trained for measures nothing.
  • A threshold of tests. Pass some agreed battery — the Turing test, professional exams, a benchmark suite — and the label applies. This family is the most operational and has by far the worst track record, for reasons the next section is about.

Which family you pick determines the answer today. On the economic definition the argument is about degree, and reasonable people put frontier systems somewhere on the curve. On the learning-efficiency definition current systems are not close and scale is not obviously closing the gap: the BabyLM Challenge's estimate of a child's total language exposure by age 13 is about 100 million words — children hear 2 to 7 million words a year — while a frontier large language model is pretrained on tens of trillions of tokens, on the order of a hundred thousand times more input to reach comparable competence. On the test-threshold definition, several proposed AGI tests have already been passed and nobody changed their mind.

How It Works

AGI has no architecture to describe, because nothing has been built that anyone agrees qualifies. Inventing one here would be fiction. What does exist, and what a reader actually needs, is the apparatus the field uses to measure progress toward it — and the recurring reason that apparatus keeps failing.

Every proposed test has been passed without settling anything

Chess was supposed to be the marker until Deep Blue beat Kasparov in 1997, at which point chess became "just search". The Turing test held the position longer: in a 2025 pre-registered three-party study by Cameron Jones and Benjamin Bergen, GPT-4.5 prompted to adopt a persona was judged the human 73% of the time — more often than the actual human participants it was compared against. That is a formal pass of the test Turing described, and it moved almost no one. Winograd schemas, designed to need common-sense reference resolution, went the same way. So did professional exams.

This is the AI effect: once a capability is achieved it stops counting as intelligence. Framed as a complaint it sounds like bad faith, but it is a mechanism, and the mechanism is informative. Each test was a proxy for generality. Each proxy turned out to be passable by means its proposer had not anticipated — search, pattern statistics, a well-crafted persona — and passing it did not deliver the general capability the proxy stood in for. A test that a non-general system can pass is evidence the test was mis-specified, not evidence that generality arrived.

ARC-AGI: a benchmark designed against exactly that failure

Chollet's ARC-AGI, introduced alongside the 2019 paper, is worth understanding because its design follows directly from that history. Tasks are small coloured grids: you see a few input-output examples and must produce the output for a new input. Every task is novel, and the only background knowledge required is what ARC Prize calls core knowledge priors — objectness, counting, basic geometry, the things a young child already has. If the failure mode of every previous test was "passable by memorisation and domain-specific skill", then build a test where memorisation and domain skill are worth nothing. The specific transformation in each task appears nowhere in any training corpus, so a model must work it out during the test, which is precisely the generalization the other benchmarks failed to isolate.

ARC-AGI-2 was released in March 2025, and its results are best read as two numbers rather than one: how well a system scores, and how much compute it burns to get there. The benchmark reports both because either alone misleads. In the ARC Prize 2025 Kaggle competition — which ran from 26 March to 3 November 2025 under a hard efficiency cap of roughly $50 of compute per submission and no internet access — 1,455 teams filed 15,154 entries and the winning private-set score was 24.03%, far short of the 85% that unlocks the grand prize. Remove the compute cap and scores rise steeply, but only by spending: uncapped frontier systems reach well into the majority of the tasks, at a per-task compute cost orders of magnitude above the capped entries. That trade is the durable fact, and it is the one to carry away from any single leaderboard snapshot — the top line moves month to month, but the shape does not. Buying accuracy with test-time compute works, and it is exactly what a definition built on learning efficiency says should not count.

The human baseline is not what the headline says

ARC Prize reports that a human panel solves 100% of ARC-AGI-2 tasks, and that figure is routinely read as "humans score 100%, the best AI scores far less". It does not mean that. The 100% is a property of the task set, not a score any individual reached: in the benchmark's human study, 407 participants across 515 sessions attempted the tasks, and every task in the final set was solved by at least two independent testers within their first two attempts — which is the criterion for a task being included at all. Confirming each task is humanly solvable is a different claim from humans acing the set. The study's figure for an individual is much lower: the average test-taker solved about 66% of the tasks they attempted, and the curated test pairs were solved, on average, by roughly 75% of the people who tried them. The honest human-versus-machine comparison is against those figures, not against 100%. For the earlier ARC-AGI-1, where per-tester scores were published directly, human testers averaged 97-98% on the private evaluation set — a much harder standard that the ARC-AGI-2 headline is often mistaken for.

Check this before believing any human-versus-AI comparison. The human number is frequently the union of a committee, an expert subset, or a best-of-N, and the AI number is almost never any of those.

Levels instead of a line

Because a yes/no test keeps failing, DeepMind researchers proposed grading it. The Levels of AGI framework (Morris et al.) uses two axes: performance depth in six tiers — No AI, Emerging, Competent (50th percentile of skilled adults), Expert (90th), Virtuoso (99th), Superhuman — crossed with generality, narrow versus general. Deep Blue is Superhuman Narrow. Today's frontier models the paper places at Emerging AGI: broad in coverage, roughly unskilled-human in depth, with Expert-level performance in a few narrow places. The value of the framework is not the labels but the shape: it converts an unanswerable question into a set of claims that can each be checked against evidence.

The measurement problem, in one sentence

Contamination means a test cannot be trusted once it is public; generality means a test cannot be trusted once it is specific. Every AGI benchmark has to live between those two failures, and there is no stable place to stand.

Real-World Applications

There are no AGI deployments. Nothing has been built that satisfies any of the definitions above, and a list of things AGI would do if it existed would be fiction dressed as an application section. The concept nevertheless does concrete work today, in four places where somebody had to turn "is it AGI yet" into a decision with consequences.

Contracts, where the definition became money

The clearest case is the Microsoft-OpenAI partnership. The charter language — "most economically valuable work" — is unusable as a contract term, so the agreement reportedly substituted a commercial proxy: in December 2024, reporting on the agreement described AGI as treated as reached when OpenAI's systems generate profits on the order of $100 billion. In October 2025 the restructured partnership moved the call again: an AGI declaration by OpenAI must be verified by an independent expert panel, Microsoft's IP rights to research run until that verification or through 2030 whichever comes first, and model and product rights extend to 2032. By April 2026 the revenue-share terms had been decoupled from technology milestones altogether, draining most of the clause's force.

Follow the sequence, because it is the whole lesson of this page in one worked example: the moment the definition had to be enforceable, it stopped being cognitive and became financial, then procedural, then quietly irrelevant.

Safety frameworks, where the question got replaced

Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework all abandon the AGI question and substitute named capability thresholds that trigger specific obligations. In May 2025 Anthropic activated its ASL-3 protections alongside Claude Opus 4, explicitly as a precaution — it stated it had not determined that the model crossed the relevant threshold. What the activation triggered was not a press release but engineering: two-party authorisation, egress bandwidth controls, hardware security keys, cryptographically signed commits, and deployment measures aimed at chemical and biological uplift. The shape is what matters for AI safety generally. "Is it AGI" has no operational answer; "does it materially help someone acquire a weapon" is measurable and has a response attached.

Regulation, where a proxy replaced the concept

The EU AI Act does not use the word AGI at all. Article 51 presumes a general-purpose model has high-impact capabilities — and therefore systemic risk — when the cumulative training compute exceeds 10^25 FLOP. The threshold was chosen because it is auditable in a way that a capability claim is not; the presumption is rebuttable, and the Commission can move the number by delegated act as the frontier moves. Obligations for general-purpose models have been enforceable since 2 August 2025. This is what AI governance does when the underlying concept will not hold still: it regulates a measurable correlate and accepts that the correlate will drift.

Evaluation programmes, which exist because the question is unsettled

The ARC Prize Foundation is a nonprofit whose purpose is to keep a measure of progress honest, with $700,000 of prize money attached to a score nobody has reached. The International AI Safety Report 2026, published 3 February 2026 and chaired by Yoshua Bengio, was written by over 100 experts with an advisory panel nominated by 29 nations plus the UN, OECD and EU, and exists to give governments a shared evidence base about capabilities precisely because no single test provides one.

Challenges

Contamination. Once a benchmark is public it is a training-set candidate, and a strong score on it always has an alternative explanation that nobody can rule out, because frontier labs do not publish their training data. This is why serious AGI-adjacent evaluations now keep private or semi-private task sets, and why the scores you most want to trust are the ones you cannot independently reproduce. It is an uncomfortable trade and there is no version of it that is fully satisfying.

Specificity. The moment a test is precise enough to score reliably, it is precise enough for someone to build a system that does well on it and nothing else — which is how every previous milestone fell. The more sharply you operationalise generality, the less general the thing you are actually measuring.

No shared operationalisation, so evidence does not settle arguments. Two people can agree completely on what a model scored and disagree completely on what it means, because they entered with different definitions. That is not a failure of reasoning; it is what happens when a term has three incompatible referents in active use.

What this costs you if you ignore it. Two practical errors follow. The first is reading a benchmark number as a capability claim: a system's ARC-AGI-2 score has told you nothing about whether it can run your workflow, and a system that clears a professional licensing exam has told you nothing about whether it can do the job the licence covers. The second is taking an AGI timeline at face value without asking which definition it uses and who benefits from that definition — and the Microsoft-OpenAI clause is the demonstration that the answer to the second question is sometimes literally on a balance sheet.

Consciousness gets folded in, and should not be. None of the three working definitions requires subjective experience. A system could match humans across all economically valuable work with the question of whether it experiences anything entirely open, and the reverse is at least conceivable. Keeping consciousness separate from performance is not pedantry; the two claims need different evidence, and the word AGI routinely smuggles one in behind the other.

The most instructive forecasting result is not a date but a gap. In Thousands of AI Authors on the Future of AI, Katja Grace and colleagues surveyed 2,778 researchers who had published at top AI venues in October 2023. The aggregate forecast gave a 50% chance that unaided machines outperform humans at every possible task by 2047 — thirteen years earlier than the same survey's answer a year before. But the same respondents, in the same survey, put a 50% chance on all human occupations becoming fully automatable only by 2116. Sixty-nine years separate two questions that both sound like "when is AGI". Nothing about the world changed between them; only the wording did. That gap is the definitional problem, quantified by the people best placed to have an opinion.

The named disagreements are more useful when they are about mechanism than about dates. Demis Hassabis has publicly placed human-level AI around the end of this decade while maintaining that further research breakthroughs are still required. Yann LeCun has argued for years that scaling language models will not get there at all, and that the route runs through world models that learn how physical reality behaves and plan over it rather than predicting the next token. Those two positions make different predictions about what further scaling buys, which means evidence can eventually separate them — unlike a difference of opinion about a year.

The measurement to watch is efficiency rather than accuracy. ARC Prize's own reading of its 2025 results is that the accuracy gap is now mostly an engineering problem while the efficiency gap remains bottlenecked by ideas. A system that reaches high ARC-AGI-2 scores inside the Kaggle compute limit, rather than by spending heavily on compute per task, would be a substantively different event from one that buys the same score with test-time compute — and under Chollet's definition it is the only one of the two that counts. Beyond that threshold sit the questions this page deliberately does not answer: whether a system that reaches human level continues past it via self-improvement toward artificial superintelligence, and whether the values it acts on can be specified before rather than after it matters.

Frequently Asked Questions

There is no agreed test, and that is the honest answer. Every benchmark proposed as decisive — chess, the Turing test, professional exams — has been passed by systems nobody considered generally intelligent, which is why the field has largely moved from a single test to graded capability thresholds and to benchmarks like ARC-AGI that are designed so memorisation does not help.
Narrow AI is trained for a task family and fails outside it; AGI would handle intellectual work it was never prepared for, at human level. The disagreement is over how much breadth counts, which is why today's frontier models are called near-AGI by some definitions and nowhere near it by others.
Usually one of three: matching humans on most economically valuable work (OpenAI's charter), matching human learning efficiency on tasks never seen before (François Chollet's framing), or passing an agreed battery of tests. A claim that AGI is imminent is almost always using the first or third; a claim that it is distant is usually using the second.
In a 2023 survey of 2,778 published AI researchers, the aggregate forecast gave a 50% chance that unaided machines outperform humans at every task by 2047 — but the same respondents put a 50% chance on all occupations being fully automatable only by 2116. A 69-year gap between two phrasings of the same milestone is the best available measure of how unsettled the question is.
Current models are trained on roughly a hundred thousand times more language than a child hears before adolescence, and still fail novel puzzles that untrained adults solve. Whether that gap is closed by more scale or requires a different architecture is the central open disagreement, not a settled fact.
None of the working definitions require it. A system could match human performance across all economically valuable work while the question of whether it has any subjective experience stays open — and the reverse is conceivable too. Performance and experience are separate claims that the word AGI tends to blur together.

Continue Learning

Explore our use-case guides and prompts to deepen your AI knowledge.