Definition
AI drug discovery is the use of machine learning to carry out the parts of pharmaceutical research that are fundamentally searches: sifting millions of candidate proteins, molecules and existing medicines down to the handful worth synthesising and testing. Models nominate which protein a disease depends on, design molecules shaped to bind that protein, and forecast whether those molecules will dissolve, survive metabolism, reach the target tissue and avoid poisoning the liver — all before a chemist makes anything.
Then the question a sceptical reader actually arrives with: has any of this produced a real drug? Not an approved one. As of mid-2026 no regulator anywhere has approved a medicine discovered or designed by AI. Several AI-designed molecules have entered human trials since 2020, one has published a randomised Phase IIa result and started Phase III, and that is the whole record. Why it is thinner than the headlines — and why that is not, on its own, evidence of failure — comes down to one piece of arithmetic about where drug development spends its money.
How It Works
The industry pipeline runs target identification, hit finding, lead optimisation, preclinical development, then Phase I, II and III trials. AI enters at every stage, but it does very different amounts of work at each.
Target identification asks which protein or pathway to attack. Models trained on genomics, gene-expression, proteomics and the published literature rank candidate targets by their causal link to disease. Knowledge graphs are the workhorse here: BenevolentAI's system, searching approved drugs for one that could both block viral entry and damp the inflammatory response, surfaced the rheumatoid-arthritis drug baricitinib as a COVID-19 candidate in early 2020. The hypothesis went into The Lancet, was tested in the 4,000-patient RECOVERY trial, and baricitinib received FDA emergency use authorisation for hospitalised patients in November 2020.
Hit finding asks which molecules bind that target. Physical screening tops out around a few million compounds; docking against a predicted structure does not. Lyu et al. (2019, Nature) computationally docked 138 million make-on-demand molecules against the D4 dopamine receptor, synthesised 549 of them, and confirmed 122 as real ligands — a 22% hit rate, against the low single digits typical of high-throughput screening. That is the clearest demonstration of what virtual screening buys: not certainty, but a shortlist that is two orders of magnitude better enriched than a random one.
AlphaFold and its successors turned protein folding from a months-long crystallography project into an overnight computation, so structure-based design can now start on targets that never had an experimental structure at all.
Lead optimisation turns a hit into something drug-like. Generative chemistry models propose variants that keep the binding while fixing solubility, potency and half-life, constrained to molecules a chemist can actually synthesise. Graph neural networks, which read a molecule as a graph of atoms and bonds rather than a string, predict ADMET properties — absorption, distribution, metabolism, excretion, toxicity — so that bad candidates die on a workstation rather than in a rat.
Now the arithmetic. The Exscientia and Sumitomo Dainippon molecule DSP-1181 went from project start to clinical candidate in roughly 12 months, against an industry norm of four to five years. Take that at face value: it removes three to four years from a pipeline conventionally cited at 10 to 15 years, so about a quarter of the calendar in the best case. On cost the leverage is weaker still. DiMasi, Grabowski and Hansen (2016, Journal of Health Economics) — the standard reference, and a contested one, since Prasad and Mailankody's 2017 JAMA Internal Medicine study of ten cancer drugs put the median at $648 million — estimated $1,395 million of out-of-pocket spend per approved drug, of which $965 million, about 69%, was clinical. Discovery-stage chemistry is a fraction of the remaining 31%. Halving it moves the total by a few percent.
The reason is failure, not chemistry. Wong, Siah and Lo (2019, Biostatistics), working from 406,038 trial records, found that only 13.8% of molecules entering Phase I reach approval, with Phase II the narrowest gate at 30.7%. You therefore fund roughly seven clinical campaigns per approval and watch six of them fail, mostly on efficacy in Phase II. Suppose AI raised Phase II success from 30.7% to 45%: overall probability of success would rise to about 20%, cutting campaigns per approval from seven to five — a 30% reduction in the entire clinical budget. That single change would be worth more than eliminating discovery costs entirely. AI has so far accelerated the cheap, fast end of a pipeline whose expense lives at the expensive, slow end.
Real-World Applications
Insilico Medicine's rentosertib (ISM001-055) is the furthest-advanced case. Both the target — TNIK, for idiopathic pulmonary fibrosis — and the molecule came from generative models. A randomised, placebo-controlled Phase IIa in 71 patients across 22 sites in China published in Nature Medicine in June 2025 reported a mean forced vital capacity change of +98.4 mL in the 60 mg daily arm against −20.3 mL on placebo. That is a small, short, exploratory trial, and Insilico began a 320-patient, 52-week Phase III in July 2026 precisely because a Phase IIa signal is not proof.
Exscientia's DSP-1181, an OCD candidate designed with Sumitomo Dainippon, entered Phase I in January 2020 as the first AI-designed molecule in human trials. It passed safety and was discontinued afterwards without progressing. It is the field's most instructive result: reaching the clinic in a year proved nothing about whether the drug worked.
Recursion built a phenomics platform imaging millions of drug-treated cells, acquired Exscientia in a merger completed in November 2024, and in 2025 deprioritised several clinical-stage programmes including REC-2282 and REC-994. Consolidation and pipeline pruning are as much a part of this field's real history as the launches.
Isomorphic Labs, spun out of DeepMind on the back of AlphaFold, has signed drug-design partnerships with Eli Lilly and Novartis and has said its first candidates are approaching the clinic. It has not, at the time of writing, announced a molecule in human trials.
Alongside these sit uses that never produce a new molecule at all: repurposing approved drugs, and the precision medicine and AI healthcare work that decides who a drug is tested on.
Key Concepts
- Structure prediction is an input, not an answer: knowing a protein's shape tells you nothing about whether blocking it treats the disease. Target choice, not structure, is where most programmes die.
- Synthesisability is the binding constraint on generative chemistry: a model can propose molecules no laboratory can make, and an unmakeable molecule has a value of exactly zero.
- ADMET prediction is a filter, not an oracle: its job is to kill candidates cheaply. A false negative costs one discarded molecule; a false positive costs a failed animal study.
- Enrichment beats accuracy: virtual screening does not need to be right about any single molecule. It needs the top 500 of 138 million to be much better than random, which is a far easier target.
Challenges
The deepest problem is that the models are trained on what can be measured cheaply — binding affinity, solubility, cell-assay readouts — and none of these is clinical efficacy. Nothing in the current toolkit predicts whether a molecule will change disease course in a sick human, which is exactly what Phase II tests and exactly where two-thirds of programmes fail.
Data is the second constraint. Public bioactivity databases skew heavily toward well-studied targets, so a model's confidence is highest precisely where the medical need is lowest, and negative results mostly go unpublished — models learn from a corpus that has quietly deleted its own counterexamples.
Then there is the feedback loop. A prediction made in 2026 about whether a molecule survives Phase III is not settled until the mid-2030s. A field cannot iterate on a signal that arrives a decade late, which is why AI drug discovery has far more published methods than validated outcomes, and why Jayatunga et al.'s 2024 survey found just 67 AI-discovered molecules in clinical development out of roughly 6,147 — about 1%.
Future Trends
The most valuable direction is the least glamorous: moving AI from discovery into the clinical stages where the money burns. That means patient stratification so Phase II tests a drug on those most likely to respond, synthetic control arms built from historical data, and toxicity models good enough to shorten preclinical safety work.
Rentosertib's Phase III readout will be the first well-powered efficacy test of a fully AI-originated target-and-molecule pair, and its outcome — either way — will matter more than any benchmark.
Biological foundation models trained across sequence, structure, expression and imaging are the technical bet: one model that transfers across targets, rather than a bespoke model per programme. Commercially, the standalone AI-biotech is consolidating — as the Recursion-Exscientia merger showed — into partnerships where the AI company designs and a large pharmaceutical company pays for the trials. The economics above explain why.