Introduction
Google made Gemini 3.8 Flash generally available on September 2, 2026, at exactly the same rate card as the model it replaces: $0.75 per million input tokens and $3.75 per million output tokens. Google's own framing is "our best reasoning and coding model yet, at the same speed and low cost of 3.7."
The rate card is not the bill. Buried in the same announcement is a sentence that undoes the pricing claim for anyone running agents: on complex tasks 3.8 Flash "exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively," and "at times, the model might use more tokens to maximize performance."
Artificial Analysis put a number on that. Running its Intelligence Index v4.2 suite at high reasoning effort, Gemini 3.8 Flash consumed about 30% more output tokens per task — 48,000 — and cost $0.58 per task against $0.40 for Gemini 3.7 Flash. Same price per token, roughly 40% more money.
That is the story of this launch, and it has a second half: the $0.75/$3.75 rate is introductory and expires on December 31, 2026.
What shipped
- Model ID:
gemini-3.8-flash, stable, GA September 2, 2026. It is the third Flash release in six weeks — 3.6 Flash on July 21 and 3.7 Flash on August 13. - Context: 1,048,576 input tokens, 65,536 output tokens. Unchanged across the Flash line. See context window for what those two limits mean in practice.
- Knowledge cutoff: March 2026, per Google's model card — unchanged since 3.6 Flash.
- Modalities: text, image, video, audio and PDF in; text out. No image generation, no audio generation, no Live API.
- Thinking levels: low, medium and high. The
minimalsetting available on some earlier Google models is not offered here, which is part of why 3.8 Flash spends more tokens than its predecessor at the floor. - Alias warning:
gemini-flash-latestis repointed on every Flash launch. Three swaps in six weeks is enough to change the model under a running application without a deploy. Pingemini-3.8-flashexplicitly.
A second model shipped alongside it. Gemini 3.8 Flash Cyber, tuned for autonomous vulnerability discovery and patching, is not on the public API and cannot be selected as a model ID. Access runs only through the Fairwind Program, Google's vetted channel for government authorities, critical infrastructure operators and software maintainers.
The introductory rate and what follows it
Every price on the Gemini 3.8 Flash row of Google's pricing page carries a "through Dec 31, 2026" qualifier, and every one of them doubles the next day.
| Tier (per 1M tokens) | Through Dec 31, 2026 | From Jan 1, 2027 |
|---|---|---|
| Standard input / output | $0.75 / $3.75 | $1.50 / $7.50 |
| Batch and Flex | $0.375 / $1.875 | $0.75 / $3.75 |
| Priority | $1.35 / $6.75 | $2.70 / $13.50 |
| Context caching | $0.075 + $0.50/hr storage | $0.15 + $1.00/hr storage |
There is no cheaper corner to hide in. The discount is applied uniformly, so prompt caching and Batch scale with it and revert with it.
Worth noting because it is easy to misread as a price cut: Gemini 3.6 Flash was repriced onto this same window. It launched on July 21 at $1.50 / $7.50 — the post-introductory number — and now bills at $0.75 / $3.75. The entire current Flash line shares one expiry date. Nothing in the catalog gives you a hedge against January 1.
The bill is not the price
Put the two facts together and the arithmetic gets uncomfortable. Take Artificial Analysis's measured $0.58 per Intelligence Index v4.2 task at high reasoning effort. Hold the token count constant and apply the January 1 rate card, and the same task costs roughly $1.16 — about 2.9x what Gemini 3.7 Flash costs today.
That is a projection, not a quote: the token count is one benchmark suite at one effort level, and your workload's ratio of input to output will move it. But the direction is not in doubt, and it comes from two independent mechanisms — more tokens per task, and a higher price per token — that compound rather than offset.
Three things follow for anyone doing capacity planning:
- Measure tokens per task, not price per token. For agentic work the two have now visibly decoupled. A migration that looks free on the rate card is not free in inference spend.
- Effort level is a budget control. Google recommends lower thinking levels where token overhead matters, and the index scores fall with them. That is the dial to test first.
- 3.7 Flash is a supported destination, not a legacy one. Google says it "remains fully supported for efficiency-first workloads." For an agentic workflow whose accuracy is already adequate, staying put is the cheaper answer and Google is telling you so.
Benchmarks, including one that disagrees with itself
Google's launch table, 3.8 Flash against 3.7 Flash. These are vendor-run and vendor-selected figures; independent evaluation may differ.
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| DeepSWE v1.1 | 73.7% | 65.3% |
| Terminal-Bench 2.1 | 89.4% | 85.8% |
| OSWorld-2.0 | 59.0% | 50.6% |
| Vals Finance Agent v2 | 61.4% | 59.0% |
| HLE-Verified | 54.9% | 53.6% |
Real gains, and DeepSWE v1.1 at 73.7% lands within a third of a point of Claude Opus 5 at 74.0% — from a model at a fraction of the price. But note the shape of the table: the large jumps are on agentic benchmarks with long tool-calling loops, which is precisely where the extra tokens are being spent. The 1.3-point gain on HLE-Verified, a single-shot reasoning test, is what 3.8 Flash buys you when it cannot work harder.
One number does not agree with itself. Google's announcement table gives Terminal-Bench 2.1 as 89.4% / 85.8%; Google's own developer guide reports 90.8% / 81.6% for the same pair — a different run configuration, and a 4.2-point swing in the gap between the two models. The announcement figures are used above. When a vendor publishes two versions of one benchmark three days apart, treat the gap between models as the soft number, not the score.
On the independent side, Gemini 3.8 Flash scores 59 on the Artificial Analysis Intelligence Index v4.2 at high reasoning effort, three points above 3.7 Flash on the same index version. Three points of measured intelligence for roughly 40% more cost per task is the trade in one line. (v4.1 scores are not comparable to v4.2 — the index was reweighted on September 4, 2026.)
The Pro tier that never arrived
The other thing this launch says is what it does not contain. Gemini 3.5 Pro was announced at Google I/O on May 19, 2026 with a June rollout. As of September 6, 2026 it has no API model ID, no pricing entry, and no changelog entry; Google's position is that it is testing with enterprise partners. The newest callable Pro-tier model is still Gemini 3.1 Pro Preview from February 2026, still labelled Preview, and priced several times above the Flash model that outscores it.
So the Pro tier of this generation shipped zero times while Flash shipped three, and the Flash brand quietly stopped meaning "the distilled cheap one." Google now calls Flash its most intelligent model for agentic and coding tasks. That is a reasonable thing for the Gemini line to do — but it is also why the cost story matters. There is no Pro model above this one to migrate to, and no cheaper Flash below it that is not on the same expiry date.
Conclusion
Gemini 3.8 Flash is a genuine improvement, and on agentic coding it is close to models costing an order of magnitude more. It is not, however, a free upgrade from 3.7 Flash, and the identical rate card is the reason that is easy to miss.
Two dates decide the economics. The first is your next billing cycle, where a model that "works harder" spends about 30% more output tokens on the same work. The second is January 1, 2027, when the introductory rate across the whole Flash line doubles. A team that budgeted 3.8 Flash from today's sticker price has underestimated it on both axes at once.
Benchmark the effort levels, meter tokens per completed task rather than per call, pin gemini-3.8-flash instead of the alias, and treat staying on 3.7 Flash as a live option rather than technical debt. Google does.
Sources
- Gemini 3.8 Flash and 3.8 Flash Cyber — Google's September 2, 2026 announcement
- Gemini 3.8 Flash model docs — limits, modalities, supported features
- Gemini API pricing — introductory rates and the January 1, 2027 reversion
- Gemini API changelog — GA dates for 3.6, 3.7 and 3.8 Flash
- Gemini 3.8 Flash model card — knowledge cutoff, evaluations, safety
- Google has released Gemini 3.8 Flash — Artificial Analysis's independent token and cost-per-task measurements
- Announcing Artificial Analysis Intelligence Index v4.2 — why v4.1 scores are not comparable
- Gemini 3.8 Flash — our model page
- Google Ships Gemini 3.6 Flash at a Lower Price Than 3.5 — the July 2026 launch, before the repricing
- Google's Fairwind Program — how gated access to Flash Cyber works