---
source: 'https://howaiworks.ai/glossary/artificial-general-intelligence'
section: glossary
title: Artificial General Intelligence (AGI)
description: >-
  AGI is AI matching human ability across most cognitive work — and there is no
  agreed test for it, so the definition you pick decides the answer.
tags:
  - AGI
  - consciousness
  - AI Safety
  - Reasoning
  - Benchmarks
category: Artificial Intelligence
datePublished: '2025-07-19'
lastUpdated: '2026-07-24'
---

# Artificial General Intelligence (AGI)

> AGI is AI matching human ability across most cognitive work — and there is no agreed test for it, so the definition you pick decides the answer.

## Definition

**Artificial General Intelligence (AGI) is a hypothetical AI system that matches human ability
across essentially the whole range of cognitive work rather than in one narrow domain — and there
is no agreed test that would tell you when one had arrived.** That second half is not a footnote:
it is why two well-informed people can look at the same model and reach opposite conclusions about
how close we are, and why the most useful thing you can learn about AGI is how to spot which
definition a given claim is using.

The disagreement is not philosophical hair-splitting. It has been written into commercial
contracts, into the safety policies that decide what a lab is allowed to deploy, and into
regulation with fines attached. Three definitional families are in circulation, and they do not
agree with each other about the present:

- **Economic.** [OpenAI's charter](https://openai.com/charter/) defines AGI as "highly autonomous
  systems that outperform humans at most economically valuable work." The threshold is about work
  displaced; it says nothing about how the system does it, so a very large collection of narrow
  competences could qualify.
- **Learning efficiency.** François Chollet's 2019 paper
  [On the Measure of Intelligence](https://arxiv.org/abs/1911.01547) argues that intelligence is
  skill-acquisition efficiency — how little experience you need to handle a task you have never
  seen. Under this view a system that has memorised a domain has *skill*, not intelligence, and
  measuring it on tasks it was trained for measures nothing.
- **A threshold of tests.** Pass some agreed battery — the Turing test, professional exams, a
  benchmark suite — and the label applies. This family is the most operational and has by far the
  worst track record, for reasons the next section is about.

Which family you pick determines the answer today. On the economic definition the argument is
about degree, and reasonable people put frontier systems somewhere on the curve. On the
learning-efficiency definition current systems are not close and scale is not obviously closing
the gap: the BabyLM Challenge's estimate of a child's total language exposure by age 13 is
[about 100 million words](https://arxiv.org/abs/2504.08165) — children hear 2 to 7 million words a
year — while a frontier [large language model](https://howaiworks.ai/glossary/large-language-model) is pretrained on
tens of trillions of tokens, on the order of a hundred thousand times more input to reach
comparable competence. On the test-threshold definition, several proposed AGI tests have already
been passed and nobody changed their mind.

## How It Works

AGI has no architecture to describe, because nothing has been built that anyone agrees qualifies.
Inventing one here would be fiction. What *does* exist, and what a reader actually needs, is the
apparatus the field uses to measure progress toward it — and the recurring reason that apparatus
keeps failing.

### Every proposed test has been passed without settling anything

Chess was supposed to be the marker until Deep Blue beat Kasparov in 1997, at which point chess
became "just search". The Turing test held the position longer: in a 2025 pre-registered
three-party study by Cameron Jones and Benjamin Bergen,
[GPT-4.5 prompted to adopt a persona was judged the human 73% of the time](https://arxiv.org/abs/2503.23674)
— more often than the actual human participants it was compared against. That is a formal pass of
the test Turing described, and it moved almost no one. Winograd schemas, designed to need
common-sense reference resolution, went the same way. So did professional exams.

This is the AI effect: once a capability is achieved it stops counting as intelligence. Framed as
a complaint it sounds like bad faith, but it is a mechanism, and the mechanism is informative.
Each test was a *proxy* for generality. Each proxy turned out to be passable by means its
proposer had not anticipated — search, pattern statistics, a well-crafted persona — and passing it
did not deliver the general capability the proxy stood in for. A test that a non-general system
can pass is evidence the test was mis-specified, not evidence that generality arrived.

### ARC-AGI: a benchmark designed against exactly that failure

Chollet's ARC-AGI, introduced alongside the 2019 paper, is worth understanding because its design
follows directly from that history. Tasks are small coloured grids: you see a few input-output
examples and must produce the output for a new input. Every task is novel, and the only background
knowledge required is what ARC Prize calls *core knowledge priors* — objectness, counting, basic
geometry, the things a young child already has. If the failure mode of every previous test was
"passable by memorisation and domain-specific skill", then build a test where memorisation and
domain skill are worth nothing. The specific transformation in each task appears nowhere in any
training corpus, so a model must work it out during the test, which is precisely the
[generalization](https://howaiworks.ai/glossary/generalization) the other benchmarks failed to isolate.

ARC-AGI-2 was released in March 2025, and its results are best read as two numbers rather than
one: how well a system scores, and how much compute it burns to get there. The benchmark reports
both because either alone misleads. In the
[ARC Prize 2025](https://arcprize.org/blog/arc-prize-2025-results-analysis) Kaggle competition —
which ran from 26 March to 3 November 2025 under a hard efficiency cap of roughly $50 of compute per
submission and no internet access — 1,455 teams filed 15,154 entries and the winning private-set
score was **24.03%**, far short of the 85% that unlocks the grand prize. Remove the compute cap and
scores rise steeply, but only by spending: uncapped frontier systems reach well into the majority
of the tasks, at a per-task compute cost orders of magnitude above the capped entries. That trade
is the durable fact, and it is the one to carry away from any single leaderboard snapshot — the top
line moves month to month, but the shape does not. Buying accuracy with
[test-time compute](https://howaiworks.ai/glossary/test-time-compute) works, and it is exactly what a definition built on
learning *efficiency* says should not count.

### The human baseline is not what the headline says

ARC Prize reports that a human panel solves **100%** of ARC-AGI-2 tasks, and that figure is
routinely read as "humans score 100%, the best AI scores far less". It does not mean that. The
100% is a property of the *task set*, not a score any individual reached: in the benchmark's
[human study](https://arxiv.org/abs/2505.11831), 407 participants across 515 sessions attempted the
tasks, and every task in the final set was solved by at least two independent testers within their
first two attempts — which is the criterion for a task being *included* at all. Confirming each
task is humanly solvable is a different claim from humans acing the set. The study's figure for an
individual is much lower: the average test-taker solved about **66%** of the tasks they attempted,
and the curated test pairs were solved, on average, by roughly **75%** of the people who tried them.
The honest human-versus-machine comparison is against those figures, not against 100%. For the
earlier ARC-AGI-1, where per-tester scores were published directly, human testers averaged 97-98%
on the private evaluation set — a much harder standard that the ARC-AGI-2 headline is often mistaken
for.

Check this before believing any human-versus-AI comparison. The human number is frequently the
union of a committee, an expert subset, or a best-of-N, and the AI number is almost never any of
those.

### Levels instead of a line

Because a yes/no test keeps failing, DeepMind researchers proposed grading it. The
[Levels of AGI](https://arxiv.org/abs/2311.02462) framework (Morris et al.) uses two axes:
performance depth in six tiers — No AI, Emerging, Competent (50th percentile of skilled adults),
Expert (90th), Virtuoso (99th), Superhuman — crossed with generality, narrow versus general.
Deep Blue is Superhuman Narrow. Today's frontier models the paper places at *Emerging AGI*: broad
in coverage, roughly unskilled-human in depth, with Expert-level performance in a few narrow
places. The value of the framework is not the labels but the shape: it converts an unanswerable
question into a set of claims that can each be checked against evidence.

### The measurement problem, in one sentence

Contamination means a test cannot be trusted once it is public; generality means a test cannot be
trusted once it is specific. Every AGI [benchmark](https://howaiworks.ai/glossary/benchmark) has to live between those
two failures, and there is no stable place to stand.

## Real-World Applications

There are no AGI deployments. Nothing has been built that satisfies any of the definitions above,
and a list of things AGI *would* do if it existed would be fiction dressed as an application
section. The concept nevertheless does concrete work today, in four places where somebody had to
turn "is it AGI yet" into a decision with consequences.

### Contracts, where the definition became money

The clearest case is the Microsoft-OpenAI partnership. The charter language — "most economically
valuable work" — is unusable as a contract term, so the agreement reportedly substituted a
commercial proxy: in December 2024,
[reporting on the agreement](https://simonwillison.net/2026/Apr/27/now-deceased-agi-clause/)
described AGI as treated as reached when OpenAI's systems generate profits on the order of
**$100 billion**. In October 2025 the
[restructured partnership](https://blogs.microsoft.com/blog/2025/10/28/the-next-chapter-of-the-microsoft-openai-partnership/)
moved the call again: an AGI declaration by OpenAI must be verified by an independent expert panel,
Microsoft's IP rights to research run until that verification or through 2030 whichever comes
first, and model and product rights extend to 2032. By April 2026 the revenue-share terms had been
decoupled from technology milestones altogether, draining most of the clause's force.

Follow the sequence, because it is the whole lesson of this page in one worked example: the moment
the definition had to be *enforceable*, it stopped being cognitive and became financial, then
procedural, then quietly irrelevant.

### Safety frameworks, where the question got replaced

Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework and Google DeepMind's
Frontier Safety Framework all abandon the AGI question and substitute named capability thresholds
that trigger specific obligations. In May 2025 Anthropic
[activated its ASL-3 protections](https://www.anthropic.com/news/activating-asl3-protections)
alongside Claude Opus 4, explicitly as a precaution — it stated it had *not* determined that the
model crossed the relevant threshold. What the activation triggered was not a press release but
engineering: two-party authorisation, egress bandwidth controls, hardware security keys,
cryptographically signed commits, and deployment measures aimed at chemical and biological uplift.
The shape is what matters for [AI safety](https://howaiworks.ai/glossary/ai-safety) generally. "Is it AGI" has no
operational answer; "does it materially help someone acquire a weapon" is measurable and has a
response attached.

### Regulation, where a proxy replaced the concept

The EU AI Act does not use the word AGI at all. [Article 51](https://artificialintelligenceact.eu/article/51/)
presumes a general-purpose model has high-impact capabilities — and therefore systemic risk — when
the cumulative training compute exceeds **10^25 [FLOP](https://howaiworks.ai/glossary/flops)**. The threshold was chosen
because it is auditable in a way that a capability claim is not; the presumption is rebuttable, and
the Commission can move the number by delegated act as the frontier moves. Obligations for
general-purpose models have been enforceable since 2 August 2025. This is what
[AI governance](https://howaiworks.ai/glossary/ai-governance) does when the underlying concept will not hold still: it
regulates a measurable correlate and accepts that the correlate will drift.

### Evaluation programmes, which exist because the question is unsettled

The ARC Prize Foundation is a nonprofit whose purpose is to keep a measure of progress honest, with
$700,000 of prize money attached to a score nobody has reached. The
[International AI Safety Report 2026](https://arxiv.org/abs/2602.21012), published 3 February 2026
and chaired by Yoshua Bengio, was written by over 100 experts with an advisory panel nominated by
29 nations plus the UN, OECD and EU, and exists to give governments a shared evidence base about
capabilities precisely because no single test provides one.

## Challenges

**Contamination.** Once a benchmark is public it is a training-set candidate, and a strong score on
it always has an alternative explanation that nobody can rule out, because frontier labs do not
publish their training data. This is why serious AGI-adjacent evaluations now keep private or
semi-private task sets, and why the scores you most want to trust are the ones you cannot
independently reproduce. It is an uncomfortable trade and there is no version of it that is fully
satisfying.

**Specificity.** The moment a test is precise enough to score reliably, it is precise enough for
someone to build a system that does well on it and nothing else — which is how every previous
milestone fell. The more sharply you operationalise generality, the less general the thing you are
actually measuring.

**No shared operationalisation, so evidence does not settle arguments.** Two people can agree
completely on what a model scored and disagree completely on what it means, because they entered
with different definitions. That is not a failure of reasoning; it is what happens when a term has
three incompatible referents in active use.

**What this costs you if you ignore it.** Two practical errors follow. The first is reading a
benchmark number as a capability claim: a system's ARC-AGI-2 score has told you nothing about
whether it can run your workflow, and a system that clears a professional licensing exam has told
you nothing about whether it can do the job the licence covers. The second is taking an AGI
timeline at face value without asking which definition it uses and who benefits from that
definition — and the Microsoft-OpenAI clause is the demonstration that the answer to the second
question is sometimes literally on a balance sheet.

**Consciousness gets folded in, and should not be.** None of the three working definitions requires
subjective experience. A system could match humans across all economically valuable work with the
question of whether it experiences anything entirely open, and the reverse is at least conceivable.
Keeping [consciousness](https://howaiworks.ai/glossary/consciousness) separate from performance is not pedantry; the two
claims need different evidence, and the word AGI routinely smuggles one in behind the other.

## Future Trends

The most instructive forecasting result is not a date but a gap. In
[Thousands of AI Authors on the Future of AI](https://aiimpacts.org/wp-content/uploads/2023/04/Thousands_of_AI_authors_on_the_future_of_AI.pdf),
Katja Grace and colleagues surveyed **2,778 researchers** who had published at top AI venues in
October 2023. The aggregate forecast gave a 50% chance that unaided machines outperform humans at
every possible task by **2047** — thirteen years earlier than the same survey's answer a year
before. But the same respondents, in the same survey, put a 50% chance on all human occupations
becoming fully automatable only by **2116**. Sixty-nine years separate two questions that both
sound like "when is AGI". Nothing about the world changed between them; only the wording did.
That gap is the definitional problem, quantified by the people best placed to have an opinion.

The named disagreements are more useful when they are about mechanism than about dates. Demis
Hassabis has publicly placed human-level AI around the end of this decade while maintaining that
further research breakthroughs are still required. Yann LeCun has argued for years that scaling
language models will not get there at all, and that the route runs through world models that learn
how physical reality behaves and plan over it rather than predicting the next token. Those two
positions make different predictions about what further [scaling](https://howaiworks.ai/glossary/scaling-laws) buys,
which means evidence can eventually separate them — unlike a difference of opinion about a year.

The measurement to watch is efficiency rather than accuracy. ARC Prize's own reading of its 2025
results is that the accuracy gap is now mostly an engineering problem while the efficiency gap
remains bottlenecked by ideas. A system that reaches high ARC-AGI-2 scores inside the Kaggle
compute limit, rather than by spending heavily on compute per task, would be a substantively
different event from one that buys the same score with test-time compute — and under Chollet's
definition it is the only one of the two that counts. Beyond that threshold sit the questions this page deliberately does not
answer: whether a system that reaches human level continues past it via
[self-improvement](https://howaiworks.ai/glossary/self-improving-ai) toward
[artificial superintelligence](https://howaiworks.ai/glossary/artificial-superintelligence), and whether the
[values](https://howaiworks.ai/glossary/value-learning) it acts on can be specified before rather than after it matters.

## Frequently Asked Questions

### How would we know if we had AGI?

There is no agreed test, and that is the honest answer. Every benchmark proposed as decisive — chess, the Turing test, professional exams — has been passed by systems nobody considered generally intelligent, which is why the field has largely moved from a single test to graded capability thresholds and to benchmarks like ARC-AGI that are designed so memorisation does not help.

### What is the difference between AGI and narrow AI?

Narrow AI is trained for a task family and fails outside it; AGI would handle intellectual work it was never prepared for, at human level. The disagreement is over how much breadth counts, which is why today's frontier models are called near-AGI by some definitions and nowhere near it by others.

### Which definition is a given AGI claim using?

Usually one of three: matching humans on most economically valuable work (OpenAI's charter), matching human learning efficiency on tasks never seen before (François Chollet's framing), or passing an agreed battery of tests. A claim that AGI is imminent is almost always using the first or third; a claim that it is distant is usually using the second.

### When will AGI be achieved?

In a 2023 survey of 2,778 published AI researchers, the aggregate forecast gave a 50% chance that unaided machines outperform humans at every task by 2047 — but the same respondents put a 50% chance on all occupations being fully automatable only by 2116. A 69-year gap between two phrasings of the same milestone is the best available measure of how unsettled the question is.

### How is AGI different from current large language models?

Current models are trained on roughly a hundred thousand times more language than a child hears before adolescence, and still fail novel puzzles that untrained adults solve. Whether that gap is closed by more scale or requires a different architecture is the central open disagreement, not a settled fact.

### Does AGI require consciousness?

None of the working definitions require it. A system could match human performance across all economically valuable work while the question of whether it has any subjective experience stays open — and the reverse is conceivable too. Performance and experience are separate claims that the word AGI tends to blur together.

## Related

### Related terms

- [Artificial Intelligence (AI)](https://howaiworks.ai/glossary/artificial-intelligence)
- [Artificial Superintelligence (ASI)](https://howaiworks.ai/glossary/artificial-superintelligence)
- [Consciousness](https://howaiworks.ai/glossary/consciousness)
- [General Problem Solver (GPS)](https://howaiworks.ai/glossary/general-problem-solving)
- [Machine Learning (ML)](https://howaiworks.ai/glossary/machine-learning)
- [Neural Network](https://howaiworks.ai/glossary/neural-network)

---

Source: https://howaiworks.ai/glossary/artificial-general-intelligence — HowAIWorks.ai
