AI for Good

AI for Good is an aspiration, not a technique, and nothing certifies the label. Here it is tested against named projects with published, measured results.

Published Updated

On this page

Definition

AI for Good is a label for applying machine learning to problems that markets do not pay to solve: flood warnings for people without river gauges, speech recognition for people whose speech no product was trained on, protein structures for parasites no pharmaceutical company is funding. It is not a technique, a standard or a certification. The phrase entered wide use as the name of a conference — the ITU's AI for Good Global Summit, first held in Geneva in June 2017 — and nothing since has attached a testable definition to it. Any organisation may describe any project this way, and none is audited for it.

That makes the label worthless as a filter and the evidence indispensable. The useful question about a project is never what it aims at; it is what was measured, on how many people, and who published the number. Most of the field cannot answer that. This page is built from the ones that can, and it says plainly where a project has a deployment count instead of a result.

How It Works

There is no distinct technology here. The models are ordinary supervised learners — a recurrent network over rainfall series, a convolutional classifier over camera-trap frames, a fine-tuned speech recogniser. What is unusual is the data situation, and it takes three recurring shapes.

The first is prediction where measurement is absent. A classical hydrological model must be calibrated against a long streamflow record in each watershed, so a basin with no gauge gets no forecast — and gauges are sparsest in the countries floods hurt most. A single model trained across many gauged basins can transfer to ungauged ones: the 2024 Nature paper behind Google's flood forecasts trained on 5,860 gauges and evaluated on 4,089.

The second is triage of a firehose. A camera trap fires on moving grass as readily as on a leopard, and a rainforest microphone records hours of insects between chainsaws. The model's value is not clever recognition; it is discarding the overwhelming majority of frames that contain nothing, so that scarce expert hours land on the remainder.

The third is personalisation to a sample of one, for populations so heterogeneous that a general model is useless and only per-person fitting works. Disordered speech is the clearest case: word error rates that sit under 10% for typical speakers can reach 90% for severe dysarthria, and no amount of averaging across such speakers produces a model that serves any of them.

Then comes the part that decides whether any of it matters, and it is not machine learning. In 2019 a pilot study in Bihar, India tested precisely the deployment everyone imagines — an accurate flood forecast delivered as a smartphone push notification. Despite widespread flooding, fewer than half of households received an alert, and the study measured no effect on behaviour. The model was never the bottleneck. Three subsequent years of work on how a warning physically reaches a village were.

Real-World Applications

Six projects with published numbers. Each entry gives the measurement, then what it does not show.

Flood forecasting in ungauged basins

The March 2024 Nature paper reported that an AI model matched or beat the Copernicus Emergency Management Service's Global Flood Awareness System at a five-day lead time against that system's zero-day nowcast, and achieved at five-year return period events the accuracy the incumbent achieved at one-year events. It was operational in over 80 countries at publication. Google's flood forecasting site now claims more than 150 countries and 2 billion people, with up to 7 days of lead for riverine flooding and 24 hours for urban flash floods — figures published without a date, which is reason to trust them less than the paper.

The evaluation is contested. A 2026 paper in Journal of Hydrology X, "When are AI models ready for deployment? Reassessing Google's global AI flood forecasting system through the lens of responsible modelling", argues that the headline reliability depends on defensible-but-flattering choices — how a flood event is defined, how wide a time window still counts a forecast as a hit, which return periods qualify as extreme — and that global deployment ran ahead of the evidence.

Whether a warning changes anything

The follow-up nobody usually funds. Jagnani and Pande (working paper, April 2024) randomised community-based dissemination of those same alerts across 319 villages in 12 Bihar districts, treating 160 communities reaching roughly 1.8 million people over the 2022 and 2023 flood seasons. Treated communities received more alerts and more false positives, but far fewer missed floods, so overall accuracy as experienced by households improved — and reported trust in alerts rose with it. Households in severely flooded treated areas scored 0.18 standard deviations higher on indices of proactive adaptation and physical health, and reported 30% lower medical expenditure.

This is the strongest causal evidence in the field, and it is a single unrefereed working paper covering one intervention in one Indian state. An effect of 0.18 SD is real and modest. Both statements are true at once, and the second is the one usually dropped.

Camera-trap wildlife monitoring

Wildlife Insights' SpeciesNet classifier was trained on over 35 million labelled camera-trap images spanning 1,295 species and 237 higher taxonomic classes, and is released openly. Its most valuable number is the dullest one: blank-frame detection at 99.89% precision and 93.44% recall. That is the labour saving — a survey that would have consumed weeks of volunteer scrolling becomes a day of checking animals.

Per-species performance is far more uneven, which is why the project publishes a per-species table rather than one headline accuracy. Rare animals sit in the long tail with the worst scores, and rare animals are usually the reason the survey exists.

Acoustic detection of illegal logging

Rainforest Connection's solar-powered Guardian units listen for chainsaws, engines and gunshots and relay alerts to rangers' phones, where a ranger confirms or rejects the detection. The organisation reports 587 devices across 37 countries, covering about 736,200 hectares.

Read what that is: a deployment count, not an outcome. There is no published evaluation showing that logging or poaching fell where the units were installed. Hectares monitored is exactly the kind of figure that makes a programme sound larger than anything it has demonstrated, and the honest summary is that the detection works and the conservation effect is unmeasured.

Speech recognition for disordered speech

Google's Project Euphonia collected over 1 million utterances — roughly 1,400 hours from 1,330 speakers with impaired speech — and fine-tuned individual models for about 430 of them. Reported at Interspeech 2021, median word error rate on home-automation phrases for speakers with severe impairment fell from around 89% to 13%, outperforming human listeners on short phrases.

The catch is in the method. Each personalised model required at least 300 phrases recorded by that specific person. That is the opposite of scalable, and it is why this stayed research for years rather than shipping as a feature anyone could switch on.

Protein structures for neglected diseases

AlphaFold's predicted structures — see Protein Folding for how the prediction works — were released openly, including proteomes of organisms on the WHO's neglected-tropical-disease list. That matters because these parasites attract almost no commercial structural biology: the structures did not exist and nobody was going to pay to solve them.

What has come of it so far is small and specific. A 2025 study in Computational and Structural Biotechnology Journal docked roughly 30,000 compounds against AlphaFold-predicted Trypanosoma cruzi structures, tested 24 experimentally, and found two already-approved drugs, pimecrolimus and ledipasvir, with antiparasitic activity in vitro. Two in-vitro hits are a starting point for Chagas disease research, not a treatment; nearly all the cost and years lie between there and a dosed patient. AI Drug Discovery covers that pipeline.

Challenges

The gap is deployment, not accuracy. Bihar is the pattern in miniature: a technically excellent forecast produced no measurable benefit until someone spent three years on volunteers, loudspeakers and flag-planting. Model accuracy is the metric every project reports because it is cheap to compute from data already on hand. The metric that decides whether the project mattered — what fraction of at-risk households received a warning and acted on it — requires a household survey, a control group and years, and is therefore reported by almost nobody.

Capacity sits where the problems are not. Africa holds about 0.6% of global data-centre capacity against roughly 18% of world population (Africa Data Centres Association, 2026 report). The consequence is not merely that models are trained elsewhere. The people deciding which problem to attack, which categories to label and which failure mode is acceptable are elsewhere too. The flood model needs gauge data, and the gauge network is thinnest exactly where the paper says the benefit would be largest — which is why that paper closes by asking for more hydrological data rather than for a better model.

"For good" is an unregulated claim. Overstating AI is punishable when it is a securities claim: in March 2024 the SEC fined two investment advisers $225,000 and $175,000 for "AI washing". No equivalent mechanism covers social-impact claims. Nothing stops a device count standing in for an outcome, or a 200-user pilot being described as transformation. AI Governance covers what is and is not actually enforceable.

Field-level evidence does not exist, only opinion. The most cited attempt to assess AI against the UN Sustainable Development Goals (Vinuesa et al., Nature Communications, 2020) polled experts and concluded AI could enable 134 of the 169 targets while inhibiting 59. The first number is quoted constantly and the second almost never — and both are elicited judgement, not measurement. When the field's headline evidence is a survey of expectations, individual projects with control groups are worth more than any aggregate claim.

The compute is billed to the same budget as the bed nets. A flood model is small; the general-purpose models increasingly proposed for humanitarian and educational work are not, and their inference cost is recurring. See AI Energy Consumption for the underlying numbers.

Two shifts are worth watching, and both are about evidence rather than capability.

The first is the move from proprietary demonstrations to released artefacts. SpeciesNet is open, the AlphaFold database is open, and Google has open-sourced the hydrology framework behind its flood models. This changes who can audit a claim: a released model can be run against someone else's data by someone with no stake in the answer, which is how the Journal of Hydrology X critique became possible at all.

The second is that last-mile delivery is finally being funded as its own problem rather than as an afterthought. The UN's Early Warnings for All initiative, launched in 2022, proposed US$3.1 billion over five years, explicitly covering dissemination and communication alongside observation and forecasting. The test for the next decade is simple: more randomised evaluations of the Bihar kind, and fewer coverage announcements. If that ratio moves, AI for Good will have become a claim someone can check rather than a label anyone can apply.

Frequently Asked Questions

It is a label with no definition, no certification and no audit behind it. The phrase entered wide use as the name of the ITU's AI for Good Global Summit, first held in Geneva in June 2017, and any organisation may apply it to any project without meeting a criterion.
A randomised evaluation of AI-based flood alerts across 319 villages in Bihar, India (Jagnani and Pande, working paper, April 2024). Treated households in severely flooded areas scored 0.18 standard deviations higher on adaptation and health indices and reported 30% lower medical spending. It is one working paper about one intervention in one state — which is itself telling about how thin the evidence base is.
Because the hard part is delivery, not modelling. A 2019 pilot in Bihar pushed accurate flood forecasts as smartphone notifications and found no effect on behaviour: fewer than half of households received an alert at all. Fixing that took three more years of volunteers, loudspeakers and WhatsApp groups, work that attracts neither research funding nor engineering headcount.
Ask what was measured and who published it. A deployment count — devices installed, hectares covered, countries reached — is not an outcome. An outcome is a number describing what changed for people, ideally with a control group, and most projects do not have one.
Loosely, and the mapping cuts both ways. The most cited attempt (Vinuesa et al., Nature Communications, 2020) used expert elicitation to conclude AI could enable 134 of the 169 SDG targets while inhibiting 59 of them. Both figures are expert judgement rather than measurement.

Continue Learning

Explore our use-case guides and prompts to deepen your AI knowledge.