Definition
AI in science is the use of machine learning to do three quite different jobs that are usually spoken of as one: replacing an expensive calculation with a fast learned approximation, searching a candidate space too large for anyone to enumerate, and proposing hypotheses of its own. The first two have produced a great deal of verified output. The third — the one people mean when they say AI is "doing science" — has so far produced almost nothing that survived expert scrutiny.
So the honest answer to "what has AI actually discovered?" is that it has mostly not discovered things; it has made discovery cheaper, sometimes by four orders of magnitude. A learned weather model produces a 10-day global forecast in under a minute on one chip where the physics model needs hours on hundreds of machines. Structure prediction turned a months-long crystallography project into an overnight computation. But whenever a system has announced a count of new materials or new compounds, that count has shrunk under examination — in one case by way of a formal correction published in Nature. The gap between the headline number and the verified yield is the single most useful thing to understand about this field, and the rest of this page is organised around it.
How It Works
Mode 1: a surrogate for a calculation you already know how to do
Numerical weather prediction, molecular dynamics and computational fluid dynamics all solve known equations at enormous cost. A neural network trained on the simulator's own output learns to jump straight from input to answer, skipping the intermediate physics. The question and the kind of answer are unchanged; only the price moves.
GraphCast (Lam et al., Science, 2023) is the cleanest case. It predicts hundreds of weather variables 10 days ahead at 0.25° resolution — over a million grid points — in under a minute on a single Cloud TPU v4, and beat ECMWF's HRES on more than 90% of 1,380 verification targets, rising to 99.7% within the troposphere. Note what "more than 90% of 1,380" leaves behind: up to about 138 target-and-lead-time combinations where the physics model is still better, which is why operational centres run both.
The arithmetic is worth doing properly, because the speedup is often quoted without its cost. Training GraphCast took roughly four weeks on 32 TPU v4 devices — about 32 × 672 = 21,500 TPU-hours, paid once. Inference costs under 1/60 of a TPU-hour, so that training bill buys roughly 1.3 million forecasts' worth of inference. Against a conventional run that occupies hundreds of machines for hours, the one-off training is repaid within a few hundred forecast cycles. ECMWF acted on exactly this: its AIFS Single v1 became the first fully operational machine-learned forecast model on 25 February 2025, and ECMWF reports it runs more than ten times faster while using roughly a thousandth of the energy of the physics-based system.
There is a catch that decides how far this mode generalises. GraphCast was trained on 39 years (1979–2017) of ERA5 reanalysis — a dataset produced by the very physics model the network replaces. The surrogate cannot exist without the simulator it displaces, it inherits that simulator's biases, and it is weakest on events unlike anything in those 39 years, which are precisely the events a forecast matters most for.
Mode 2: search over a space too large to enumerate
Here the model does not reason about the science at all. It proposes candidates — structures, molecules, crystals, proofs — and something outside the model checks them. This is where nearly all of the field's real output has come from, and the external check is the part that gets left out of the summaries.
Mathematics shows the pattern in its purest form. AlphaGeometry (Trinh et al., Nature, 2024) pairs a language model that suggests auxiliary constructions with a symbolic engine that only accepts steps it can prove. On 30 olympiad geometry problems it solved 25 within the competition time limit, against 10 for the previous best system and 25.9 for the average human gold medallist. The neural half is a generator of guesses; the symbolic half is the reason you can trust the output. Remove the checker and the score is meaningless.
Biology works the same way — see protein folding for structure prediction and AI drug discovery for what happens to a shortlisted molecule afterwards. In both, the model narrows an intractable space to a list, and crystallography, assays or trials decide.
Materials science is where the check has been weakest and the lesson sharpest. DeepMind's GNoME (Nature, 2023) reported 2.2 million predicted crystal structures, of which about 380,000 — roughly 17% — were flagged as stable enough to be worth making. Cheetham and Seshadri reviewed the release in Chemistry of Materials in 2024 and reported "scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility", describing much of the set as trivial dopant substitutions, symmetry-broken variants of known phases, or chemically implausible compositions, alongside many radioactive entries with no plausible use. DeepMind rejected the criticism, saying it stands by the paper and that hundreds of its predictions have since been synthesised independently. Both statements can be true at once, and the distance between "380,000" and "hundreds" is the entire point.
Mode 3: hypothesis generation and autonomous experimentation
This is the mode everyone means by "AI is doing science", and it is by far the least established. Self-driving laboratories that plan and run their own experiments exist; systems that read the literature and propose mechanisms exist. What is thin is evidence that either produces insight a competent specialist would not have reached.
The strongest single anecdote is Google's AI co-scientist, given a short prompt in February 2025 about how capsid-forming phage-inducible chromosomal islands spread between bacterial species. In about two days it reached the same mechanism that José Penadés' group at Imperial College London had spent years establishing but had not yet published, and offered four further hypotheses, one of which the team had not considered. That is genuinely striking. It is also one case, ungated by any control, on a question whose answer was arguably derivable from published work on related systems, and where the humans had already got there. An anecdote of this shape is what you would expect from a very good literature synthesiser as well as from a discoverer, and nothing in the episode separates the two.
The counterweight is what happened to the field's most-cited quantitative claim about AI-assisted discovery. A November 2024 preprint reporting that an AI tool sharply increased materials discovery and patenting at a large laboratory was widely covered, and in May 2025 MIT announced it had "no confidence in the provenance, reliability or validity of the data" and asked for its withdrawal. The most quoted number about AI's effect on scientific productivity turned out not to be a number at all.
Real-World Applications
Operational weather forecasting is the clearest production deployment. ECMWF's AIFS Single v1 has run as a supported operational product alongside the physics-based IFS since February 2025 — not a demonstration, but a model whose output real forecasters use.
Real-time triage in physics and astronomy is unglamorous, load-bearing, and rarely mentioned in AI-for-science coverage. The LHC produces collisions at 40 MHz; the CMS Level-1 trigger, built on FPGAs, cuts that to about 100 kHz within 3.8 microseconds, and later stages reduce it to roughly 1,000 events per second written to disk. That is a 40,000:1 reduction — 39,999 of every 40,000 collisions are destroyed forever, in real time, on the basis of learned selection criteria. Get the classifier wrong and the physics is not recoverable, because the data no longer exists. The Vera C. Rubin Observatory faces the same problem from the other direction: roughly 10 million alerts and about 20 TB of raw imaging per night, routed to machine-learning community brokers that classify transients before any human looks.
Structural biology and drug discovery are covered in depth on protein folding and AI drug discovery; the short version is that structure prediction is now routine infrastructure and clinical validation is not.
Autonomous materials synthesis produced the field's most instructive correction. The A-Lab (Nature, 2023) reported synthesising 41 compounds from 58 attempted targets over 17 days of unattended operation. Robert Palgrave, Leslie Schoop and colleagues re-examined the diffraction data and argued that many products were already in the Inorganic Crystal Structure Database or were misidentified. An author correction published on 19 January 2026 clarified that "novel" had meant new to the prediction platform rather than new to science, removed one compound (Zn₂Cr₃FeO₈) that had mistakenly been in the training data, and reported that the platform reached the correct conclusion in 36 of its 40 remaining reported successes, with 4 inconclusive. The robot worked; the word "new" was doing more work than the chemistry supported.
Key Concepts
- Novel to the model is not novel to science. A generator's idea of novelty is "absent from my training set", which is a fact about the dataset. The A-Lab correction is the canonical illustration, and the same confusion underlies most inflated discovery counts.
- Generation is cheap; verification is the bottleneck. A model can emit 2.2 million candidate crystals in a compute run. Synthesising and characterising one takes days of instrument time. Increasing the numerator of that ratio without increasing the denominator does not increase discovery — it increases backlog.
- A surrogate is only as good as the simulator that trained it. Learned models trained on simulation output interpolate within the regime they saw and fail silently outside it, which is why they are least reliable on the rare, extreme cases that motivate the science.
- A hybrid with a checker is a different object from a bare model. AlphaGeometry's guarantee comes from its symbolic engine, not its neural half. When a result is reported from a neuro-symbolic system, ask which component is responsible for the correctness claim.
Challenges
The reproducibility problem is worse in ML-based science than in science generally, because the failure mode is invisible in the write-up. Kapoor and Narayanan surveyed the literature and identified 17 fields with 329 affected papers, all sharing one root cause: information from the test set leaking into training, which inflates reported accuracy without leaving a trace a reader could spot. Their civil-war-prediction case study is the sharpest illustration — every paper claiming that complex machine learning substantially outperformed decades-old logistic regression contained leakage, and once it was corrected the advantage disappeared entirely.
Peer review is poorly equipped for this. A reviewer can check a derivation and can, in principle, repeat an experiment, but cannot rerun a claim that depends on model weights the authors have not released. AlphaFold 3 is the well-documented case: inference code under an open licence, parameters granted only on request at the publisher's discretion and restricted to non-commercial use. A result that cannot be independently recomputed is being accepted partly on trust, which is a change in what publication means, not merely an inconvenience.
Then there is incentive asymmetry. Publishing a headline count of predicted discoveries costs a compute run; establishing that any of them are real costs years of laboratory work that a different group usually has to fund. The gap is not fraud — it is what happens when the cheap half of a claim can be published without the expensive half, and it is why the sensible response to any large discovery count is to wait for the verified yield.
Finally, the benchmark problem. Most scientific ML is scored on held-out prediction accuracy, which measures whether a model has captured the distribution it was trained on. Discovery is by construction about what lies outside that distribution, so the metric everyone optimises is not the quantity anyone cares about.
Future Trends
The direction that would matter most is the least discussed: raising verification throughput rather than candidate throughput. Cheap automated characterisation, standardised negative-result reporting and registries that track which predictions were tested and failed would do more for the field's credibility than a larger generative model, because they attack the ratio that actually binds.
Two structural changes are already visible. Journals and funders are beginning to require that model weights and full pipelines accompany claims, on the reasonable ground that a scientific result must be recomputable. And foundation models trained across many instruments and modalities are being proposed as shared substrates for a discipline rather than bespoke models per project — promising, but as yet without a discovery to their name that a specialist team would not have made.
The 2024 Nobel Prizes marked the field's arrival without settling the question this page opened with. The Physics prize went to John Hopfield and Geoffrey Hinton for the foundational work enabling machine learning with artificial neural networks — a prize for a method, not for a discovery it made. The Chemistry prize was split, half to David Baker for computational protein design that long predates deep learning, half to Demis Hassabis and John Jumper for structure prediction. Read carefully, the awards recognise tools that changed what is affordable to attempt. Whether AI generates scientific insight, or only accelerates the humans who do, remains genuinely open — and the honest position in 2026 is that the evidence supports acceleration and does not yet support generation.