Definition
AI and employment is the question of how artificial intelligence changes what people are paid to do, and the evidence so far gives it a specific shape: exposure concentrates in tasks, not whole occupations, so the common outcome is a job whose composition changes rather than a job that disappears. A radiologist still has a job while software takes the first pass at the images; a support agent still has a job while a model drafts the reply they edit and send.
That distinction is not a diplomatic hedge; it is why two credible studies of the same labour market can differ by a factor of five. An occupation is a bundle of twenty or thirty tasks, and today's systems are good at some — summarising a document, drafting boilerplate, classifying an image — and poor at others, such as carrying legal responsibility or noticing that the client's real problem is not the one they described. Get the unit of analysis wrong and you will misread every headline about AI and jobs, starting with the most famous one.
How It Works
In 2013 Carl Benedikt Frey and Michael Osborne of the Oxford Martin School scored 702 US occupations and put 47% of total US employment in the high-risk category. That figure has been repeated for a decade as "AI will destroy half of all jobs", which is not what it says. It estimated the probability that an occupation was susceptible to computerisation over a horizon of perhaps a decade or two — not a prediction those jobs would go, with no date attached, and gross rather than net of work created.
The OECD then redid the exercise with a different unit of analysis. Arntz, Gregory and Zierahn (OECD working paper 189, 2016) used survey data on what workers in each job actually do hour by hour, and found that across 21 OECD countries only about 9% of jobs had more than 70% of their tasks automatable. Nedelkoska and Quintini (paper 202, 2018) extended it to 32 countries: 14% at high risk, plus a further 32% facing significant change to between half and 70% of their tasks.
47% and 9% are not a disagreement about capability but about what you count. Two accountants share a job title while one spends the week on reconciliation and the other on advising clients; scoring "accountant" as one unit forces both into the same bucket, and decomposing the role into tasks does not. This is the task framework David Autor and co-authors brought to labour economics, and it is why Autor's historical work describes automation as reallocating labour within occupations far more often than eliminating them.
Work the arithmetic on one team. Ten paralegals at 40 hours each is 400 hours a week. If document review is 12 of those 40 hours — 120 hours across the team — and an AI tool cuts review time by 70%, that removes 84 hours. The same caseload now takes 316 hours, about 7.9 full-time people: roughly 20% fewer paralegals for constant work, which is painful and still nothing like the occupation being abolished. Now let demand respond. If cheaper review attracts 25% more caseload, 316 × 1.25 is 395 hours and the team is back to ten people doing a differently-shaped job. Which branch you land on depends on the price elasticity of demand, not on how clever the model is.
That second branch has a precedent. American ATMs grew from roughly 100,000 in 1995 to about 400,000 by 2010, and bank teller employment did not fall — documented by James Bessen in Learning by Doing (2015) and cited in Autor's 2015 paper "Why Are There Still So Many Jobs?". Cash handling got cheap, so branches ran with fewer tellers, so branches got cheaper to open, so banks opened more.
Real-World Applications
A few studies of real deployments carry most of the useful evidence.
Customer support. Brynjolfsson, Li and Raymond studied 5,172 agents at a Fortune 500 software firm through a staged rollout of an LLM assistant ("Generative AI at Work", Quarterly Journal of Economics 140(2):889–942, 2025). Issues resolved per hour rose about 15% on average — but the average hides the finding. Gains ran near 34% for the newest and least-skilled agents and close to zero for the most experienced. Put numbers on it: a novice at 2.0 issues an hour and a veteran at 3.0 start 1.0 apart; afterwards the novice sits at roughly 2.68 and two-thirds of the measured gap is gone.
Professional writing. Noy and Zhang ran 453 college-educated professionals through mid-level writing tasks such as press releases and short reports (Science, 2023). Time per task fell 40%, evaluator-rated quality rose 18%, and again the weakest writers improved most.
Software. In Peng and colleagues' controlled trial of GitHub Copilot, 95 developers built an HTTP server in JavaScript; the assisted group finished 55.8% faster (1.66 hours against 2.41), again with the largest gains among less experienced developers — the dynamic underlying vibe coding and no-code tools.
Consulting. Dell'Acqua and co-authors gave GPT-4 to 758 Boston Consulting Group consultants ("Navigating the Jagged Technological Frontier", 2023). Inside the model's competence they completed 12.2% more tasks, 25.1% faster, at 40% higher quality, below-average performers gaining about 43% against 17% for above-average ones. On a task deliberately placed outside that competence, consultants using the model were 19 percentage points less likely to reach the right answer. Same tool, same people, opposite sign.
The economy-wide effect has not been measured well. Daron Acemoglu's "The Simple Macroeconomics of AI" (2024) argues only around 5% of tasks will be profitably automatable within ten years — an order of magnitude below industry projections, and contested. The microdata above is solid; aggregate claims in either direction are not yet.
Key Concepts
- Task exposure is not occupational risk. Exposure means a model can do part of the work; risk additionally requires that the rest not be worth a human's time and that demand not expand. Headline numbers report the first and are read as the second.
- Complementarity decides the sign. Where AI raises the value of the tasks a person keeps, headcount can rise; where it substitutes for the tasks that justified the headcount, it falls. That is human-AI collaboration as an economic claim.
- New work is real but slow. Autor, Chin, Salomons and Seegmiller found about 60% of US employment in 2018 sat in job titles that did not exist in 1940 — new categories arrive over decades, and not necessarily where the old ones were lost.
Challenges
The strongest reassurance available is a base rate: US agricultural employment fell from roughly 40% of the workforce in 1900 to under 2% today while total employment grew several-fold. But "aggregate employment held" is a statement about a country, not a person. Transition costs land on identifiable people in identifiable towns, and precedent is not a guarantee that this instance follows.
Every previous wave hit physical and routine work, and the reliable escape route was cognitive and language work. This wave points straight at that escape route, so "move up the skill ladder" assumes upper rungs that — for drafting, summarising, first-pass analysis and routine coding — are now the rungs under load.
The entry-level rung is the mechanism to watch. If the tasks juniors performed in order to become seniors are the tasks a model now does at a fraction of the cost, firms can buy the output and quietly stop producing seniors. Novices getting the biggest boost reads as good news and doubles as a warning: work a first-year employee can now match is work an employer has less reason to pay a first-year employee for. And measurement lags — layoffs get attributed to demand, restructuring or offshoring, and "AI" rarely appears as a stated cause in official statistics, so alarm and reassurance both run ahead of the evidence.
Future Trends
The variable to watch is not model capability but the size of the unit AI can take on. A model that does one task changes a job's composition; an AI agent running a multi-step agentic workflow reliably enough to be left alone takes a sequence of tasks, and a long enough sequence is a role. Reliability across chained steps, not benchmark scores, is the employment threshold.
The jagged frontier implies a durable category of work: verification. A 19-point accuracy penalty on out-of-frontier tasks means someone with domain judgement must check the output, and that person's value rises with the volume of generated material.
Expect the earliest visible effects in wage dispersion rather than headcount. If the compression seen inside firms holds at market scale, the return to experience falls before employment does — flat junior wages and slower promotion long before unemployment. Policy responses belong to AI governance and are mostly unwritten. All of this assumes AI stays a set of task-level capabilities; artificial general intelligence would be a different question, and nothing measured here speaks to it.