AI Healthcare

Where AI is actually deployed in medicine — imaging, triage, ambient notes — and why 'beats doctors' headlines rarely survive contact with a clinic.

Published Updated

On this page

Definition

AI healthcare is the use of machine learning inside medical care itself — reading a scan, flagging a deteriorating patient, drafting the clinical note — as opposed to the research uses covered by AI drug discovery. On the question every reader actually arrives with: in one narrow, well-defined task there is now randomised evidence that AI beats the standard of care, and outside that shape of task the record is far thinner than the headlines suggest. The MASAI trial randomised 105,915 Swedish women to AI-supported mammography or to standard double reading by two radiologists, and the AI arm detected 81% of cancers against 74%, at essentially the same false-positive rate.

That is one task, one country, one screening programme. What "AI in medicine" means in practice is best read off the regulator's own list: the FDA had authorised more than 1,400 AI-enabled medical devices as of its March 2026 update, and roughly three-quarters of them are radiology products. The single most useful thing to know about this field is that it is mostly software that looks at a picture — and that the thing spreading fastest through clinics right now is not diagnostic at all.

How It Works

The clinical workflow runs triage, diagnosis, treatment planning, monitoring and documentation. AI has arrived at those five stages in wildly unequal amounts, and the reason is not clinical importance. It is which stage happened to have a large labelled archive lying around.

Why imaging is decades ahead. Radiology digitised in the 1990s under the DICOM standard, and every image it stored was paired with a radiologist's written report as a matter of routine record-keeping. By the time anyone wanted training data, hospitals were sitting on millions of images each already annotated by an expert — a labelled dataset assembled for billing and medico-legal reasons, entirely by accident. Imaging also happens to have the shape supervised learning handles best: a fixed-size input, one modality, and a small discrete answer ("cancer / no cancer"). Nothing else in medicine offers that combination. Electronic health records, by contrast, are sparse, irregularly sampled, coded differently at every institution, and record what was billed rather than what was true.

Triage is the easiest stage to automate because it does not require the model to be right about the diagnosis — only about the ordering. Viz.ai's LVO software received an FDA De Novo authorisation in February 2018 creating a new device class, computer-aided triage and notification. It scans CT angiography for a suspected large-vessel occlusion and pages the interventional team directly, in parallel with the radiologist's normal read. It replaces no judgement; it moves a scan up a worklist. The value is measured in minutes to treatment, not in accuracy, and in stroke those minutes are the whole intervention.

Diagnosis splits into two regulatory grades that are easy to confuse. Assistive devices mark or characterise a finding, and a clinician still reads the image and signs the result — this is nearly everything on the FDA list. Autonomous devices return a result on their own. IDx-DR, authorised in April 2018 for diabetic retinopathy screening, was the first, and it remains the cleanest example of what the category requires: a single well-framed retinal photograph as input, a binary referral decision as output, screening demand that vastly exceeds the supply of ophthalmologists, and a benign failure mode, since a false positive costs an unnecessary eye appointment rather than an unnecessary operation. Its pivotal trial at 10 primary-care sites in 900 patients reported 87.2% sensitivity and 90.7% specificity against a reading-centre reference standard.

Treatment planning and monitoring are thinner. The most routine use is automatic organ-at-risk contouring in radiotherapy planning, where FDA-cleared segmentation software produces a first pass that a dosimetrist corrects — a labour saving on a task that is tedious rather than difficult. Risk prediction from the record — deterioration, sepsis, readmission — is the most widely deployed and the most weakly evidenced category in the whole field, for reasons the Epic Sepsis Model below makes concrete.

Documentation is where the actual boom is. Ambient clinical documentation puts a microphone in the consultation room, transcribes it, and has a large language model draft the structured note, which the clinician edits and signs. It is administrative rather than diagnostic, and that is precisely why it spread: software that summarises a conversation for a human to approve does not diagnose or recommend treatment, so it generally sits outside the medical-device pathway altogether. No De Novo, no pivotal trial, no 18-month review. The fastest-moving clinical AI of the 2020s got there by not being a medical device.

Real-World Applications

MASAI (Mammography Screening with Artificial Intelligence, Sweden) is the strongest evidence in the field, because it is prospective, randomised and population-scale rather than a retrospective reader study. 105,915 women were randomised to AI-supported screening or to standard double reading. Final results published in The Lancet in January 2026 reported sensitivity of 81% against 74%, interval cancers at 1.55 per 1,000 versus 1.76, and false-positive rates of 1.5% versus 1.4%. An earlier interim analysis reported a 44% reduction in screen-reading workload, since AI-triaged low-risk exams needed only one human reader instead of two.

IDx-DR (now marketed as LumineticsCore) runs autonomous diabetic retinopathy screening in primary care, endocrinology and pharmacy settings, where the point is that the patient gets a result during a visit they were already making rather than a referral they may never attend.

Viz.ai LVO is deployed across stroke networks in the United States and was granted a Medicare New Technology Add-on Payment in 2020 — notable because most clinical AI has no billing code at all, and reimbursement, not accuracy, is usually what decides whether a hospital buys.

Kaiser Permanente's ambient scribe deployment is the largest published measurement of clinical AI use of any kind: 7,260 Permanente physicians across 2,576,627 patient encounters, saving an estimated 15,791 documentation hours — about 22 seconds per encounter. That number is simultaneously unimpressive and the reason the technology is winning, because the gains landed in after-hours charting and burnout rather than in shorter visits. See ambient clinical documentation for the randomised evidence, the consent and liability questions, and why two vendors in the same trial produced significantly different results.

The Epic Sepsis Model is the counterexample, and the most instructive deployment in the field. A proprietary sepsis-prediction model shipped inside the most widely used US electronic health record, it was adopted by hundreds of hospitals on the vendor's internal validation. When Michigan Medicine externally validated it on 38,455 hospitalisations, the AUC was 0.63, sensitivity 33% and positive predictive value 12%. It generated alerts on 18% of all hospitalisations and missed 1,709 of the 2,552 sepsis cases. It was deployed at national scale for years before anyone outside the vendor measured it.

Key Concepts

The base rate is the concept that decides whether a screening deployment is useful, and it is arithmetic rather than machine learning. Work it on MASAI's own numbers. In the intervention arm, 53,043 women were screened and 338 screen-detected cancers were found — a prevalence of roughly 0.64%. The published false-positive rate of 1.5% applies to the ~52,600 women who do not have cancer, which is about 790 women recalled for nothing. Against 338 who genuinely have cancer, that is roughly 1,130 recalls of which 30% are true. Seven in ten women called back are called back for no reason — in the better arm. (Run the same sum on the control arm: 1.4% of 52,600 is about 737 false positives against 262 cancers, a precision of 26%.) No achievable improvement in the model fixes this, because it is a statement about a disease that occurs in six women per thousand, not about the classifier. Any headline of the form "99% accurate" that does not tell you the prevalence has told you nothing.

That is the first reason "better than a doctor" claims mislead. The second is study design. Nagendran and colleagues reviewed 81 non-randomised studies comparing deep learning against clinicians in medical imaging: 61 of them asserted in the abstract that AI was at least comparable to clinicians, only 9 were prospective, and only 6 were tested in a real clinical environment. The median comparison was against 4 experts. A retrospective result on a curated dataset, scored against a handful of readers who knew they were in a study, is not a measurement of clinical performance — it is an upper bound on it.

  • Distribution shift is documented and repeated, not theoretical. Zech et al. trained pneumonia detectors on chest radiographs from three hospital systems: a jointly trained model scored 0.931 AUC internally and 0.815 on an external site, and internal beat external in 3 of 5 natural comparisons. The networks could tell which hospital an image came from, and since pneumonia prevalence differed between the sites, identifying the hospital was a genuinely predictive shortcut — right up to the moment the model met a fourth hospital. This is a generalization failure that a validation split cannot catch, because the split shares the shortcut.
  • The data has to arrive in the shape the model was trained on. When a retinopathy screening model was deployed across 11 clinics in Thailand, it rejected 21% of the retinal photographs nurses captured as too low-quality to grade — largely because the screening rooms could not be darkened. Laboratory-grade accuracy quietly assumed laboratory-grade inputs, and patients whose images were rejected were sent to a hospital appointment they had come to the clinic specifically to avoid.
  • Calibration matters more than discrimination for anything that triggers an alert. A model that ranks patients correctly but assigns miscalibrated probabilities will fire at the wrong threshold, and in an inpatient setting the threshold — not the AUC — determines how many nurses stop what they are doing.

Challenges

The regulatory pathway creates a problem with no clean solution: a model authorised as a medical device is authorised as it was tested, so retraining it on new data is a modification to the device rather than routine maintenance. The FDA finalised guidance on Predetermined Change Control Plans on 4 December 2024, which lets a manufacturer pre-authorise a bounded set of future changes and their validation protocol. But the default is still frozen weights, and a frozen model inside a hospital that keeps changing scanners, coding practice and case mix degrades silently, with no mechanism that notifies anybody. Continuous learning, the thing that makes machine learning valuable elsewhere, is the thing device regulation is structurally built to prevent.

Liability lands on the clinician, and this shapes behaviour more than any technical property. The authorisation covers the manufacturer; the signature on the chart is the doctor's, and the malpractice exposure follows the signature whether they accept the model's output or override it. The rational response to a tool you are liable for but cannot inspect is to use it when it agrees with you and disregard it when it does not — which, either way, changes no decision. This is why explainable AI is a deployment requirement in medicine rather than a research nicety, and why accountability questions are not downstream of adoption but a precondition for it.

Alert burden is the failure mode of every model wired to the record. Take the Epic arithmetic: 18% of 38,455 hospitalisations over roughly ten months is about 6,970 alerts, of which a 12% precision makes around 840 true — leaving some 6,100 false alarms, on the order of 19 a day on a single hospital's inpatient service. Clinicians learn to dismiss such an alert within weeks, and once dismissal is habitual the model's sensitivity becomes irrelevant, because nothing downstream of the alert happens.

Then the test that most pilots fail and few measure: a model that changes no decision has no value however accurate it is. If it predicts an outcome the team already acts on, or produces its answer after the decision point, or lives behind a separate login outside the record system, its accuracy is real and its clinical benefit is zero. Integration is not a deployment detail; it is most of the work.

Finally, bias. Medical training data records who received care, not who needed it, so a model fit to historical records inherits historical access patterns as though they were biology — and a diagnostic trained on a cohort that underrepresents a group will be least reliable exactly where verification is hardest. Privacy constraints compound it: the datasets that would fix representation are the hardest to assemble and share.

The signal worth watching is not benchmark scores but whether autonomous authorisation spreads beyond retinopathy. Eight years after IDx-DR, autonomous diagnosis is still close to a category of one, and every fresh clinical area that clears it would say more about the field's maturity than any leaderboard.

The second is whether Predetermined Change Control Plans work in practice. The first devices shipping with pre-authorised retraining will be the test of whether a medical AI can legally improve after approval, or whether "locked model" remains the permanent condition of regulated clinical software.

Ambient documentation is moving upstream from drafting the note toward drafting orders, billing codes and referrals. The moment a generated draft influences a clinical order, the software starts to resemble decision support, and the regulatory question that ambient scribes bypassed reopens on much less favourable terms.

Large multimodal and foundation models pass medical licensing examinations comfortably, and there is at present no randomised evidence that they improve patient outcomes. That gap is the honest summary of clinical AI in 2026: capability is not the constraint, evidence is. MASAI took years and 105,915 women to answer one question about one cancer. Until comparable trials exist for conversational diagnostic models, the correct thing to say about them is that we do not yet know — which is different from saying they do not work, and different again from the claim that they already do.

The adjacent research uses continue on their own track: precision medicine for deciding which patient gets which treatment, protein folding and AI drug discovery for producing treatments in the first place. They share techniques with clinical AI and almost none of its constraints, which is exactly why their headlines travel further.

Frequently Asked Questions

In one narrow task, on the best evidence available, yes. The MASAI randomised trial of 105,915 Swedish women found AI-supported mammography screening detected 81% of cancers against 74% for standard double reading by two radiologists, with essentially the same false-positive rate. That is a single task, in one screening programme, with a well-defined image and a binary answer — it does not generalise to diagnosis in general.
Overwhelmingly in medical imaging — roughly three-quarters of the more than 1,400 AI-enabled devices the FDA has authorised are radiology products. The fastest-spreading use is not diagnostic at all: ambient documentation software that drafts the clinical note from the consultation audio.
IDx-DR, authorised by the FDA in April 2018 to screen for diabetic retinopathy. It was the first device permitted to return a diagnostic result without a clinician reading the image. Its pivotal trial in 900 patients at 10 primary-care sites reported 87.2% sensitivity and 90.7% specificity.
Two reasons. Distribution shift: a model trained on one hospital's scanners and patient mix degrades on another's — Zech et al. saw an AUC fall from 0.931 internally to 0.815 externally. And base rates: in screening, disease is rare, so even excellent specificity produces mostly false positives.
Nothing currently authorised is built to. Almost all imaging AI is assistive — a second reader, a triage flag, a pre-populated measurement — and even in MASAI, where AI cut screen-reading workload by 44%, radiologists still made every recall decision. Autonomous authorisation remains close to a category of one.
Because it is authorised as it was tested, so retraining it is a modification to the device rather than routine maintenance. The FDA finalised guidance on Predetermined Change Control Plans on 4 December 2024, which lets a manufacturer pre-authorise a bounded set of future changes — but the default is still frozen weights inside a hospital whose scanners, coding practice and case mix keep changing, degrading silently with no mechanism that notifies anybody.

Continue Learning

Explore our use-case guides and prompts to deepen your AI knowledge.