Gemini 3.8 Flash: Same Price Per Token, 40% More Per Task

Google's Gemini 3.8 Flash shipped September 2, 2026 at $0.75/$3.75 per million. Artificial Analysis measured cost per task rising 40% anyway.

by HowAIWorks Team
On this page

Introduction

Google made Gemini 3.8 Flash generally available on September 2, 2026, at exactly the same rate card as the model it replaces: $0.75 per million input tokens and $3.75 per million output tokens. Google's own framing is "our best reasoning and coding model yet, at the same speed and low cost of 3.7."

The rate card is not the bill. Buried in the same announcement is a sentence that undoes the pricing claim for anyone running agents: on complex tasks 3.8 Flash "exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively," and "at times, the model might use more tokens to maximize performance."

Artificial Analysis put a number on that. Running its Intelligence Index v4.2 suite at high reasoning effort, Gemini 3.8 Flash consumed about 30% more output tokens per task — 48,000 — and cost $0.58 per task against $0.40 for Gemini 3.7 Flash. Same price per token, roughly 40% more money.

That is the story of this launch, and it has a second half: the $0.75/$3.75 rate is introductory and expires on December 31, 2026.

What shipped

  • Model ID: gemini-3.8-flash, stable, GA September 2, 2026. It is the third Flash release in six weeks — 3.6 Flash on July 21 and 3.7 Flash on August 13.
  • Context: 1,048,576 input tokens, 65,536 output tokens. Unchanged across the Flash line. See context window for what those two limits mean in practice.
  • Knowledge cutoff: March 2026, per Google's model card — unchanged since 3.6 Flash.
  • Modalities: text, image, video, audio and PDF in; text out. No image generation, no audio generation, no Live API.
  • Thinking levels: low, medium and high. The minimal setting available on some earlier Google models is not offered here, which is part of why 3.8 Flash spends more tokens than its predecessor at the floor.
  • Alias warning: gemini-flash-latest is repointed on every Flash launch. Three swaps in six weeks is enough to change the model under a running application without a deploy. Pin gemini-3.8-flash explicitly.

A second model shipped alongside it. Gemini 3.8 Flash Cyber, tuned for autonomous vulnerability discovery and patching, is not on the public API and cannot be selected as a model ID. Access runs only through the Fairwind Program, Google's vetted channel for government authorities, critical infrastructure operators and software maintainers.

The introductory rate and what follows it

Every price on the Gemini 3.8 Flash row of Google's pricing page carries a "through Dec 31, 2026" qualifier, and every one of them doubles the next day.

Tier (per 1M tokens)Through Dec 31, 2026From Jan 1, 2027
Standard input / output$0.75 / $3.75$1.50 / $7.50
Batch and Flex$0.375 / $1.875$0.75 / $3.75
Priority$1.35 / $6.75$2.70 / $13.50
Context caching$0.075 + $0.50/hr storage$0.15 + $1.00/hr storage

There is no cheaper corner to hide in. The discount is applied uniformly, so prompt caching and Batch scale with it and revert with it.

Worth noting because it is easy to misread as a price cut: Gemini 3.6 Flash was repriced onto this same window. It launched on July 21 at $1.50 / $7.50 — the post-introductory number — and now bills at $0.75 / $3.75. The entire current Flash line shares one expiry date. Nothing in the catalog gives you a hedge against January 1.

The bill is not the price

Put the two facts together and the arithmetic gets uncomfortable. Take Artificial Analysis's measured $0.58 per Intelligence Index v4.2 task at high reasoning effort. Hold the token count constant and apply the January 1 rate card, and the same task costs roughly $1.16 — about 2.9x what Gemini 3.7 Flash costs today.

That is a projection, not a quote: the token count is one benchmark suite at one effort level, and your workload's ratio of input to output will move it. But the direction is not in doubt, and it comes from two independent mechanisms — more tokens per task, and a higher price per token — that compound rather than offset.

Three things follow for anyone doing capacity planning:

  • Measure tokens per task, not price per token. For agentic work the two have now visibly decoupled. A migration that looks free on the rate card is not free in inference spend.
  • Effort level is a budget control. Google recommends lower thinking levels where token overhead matters, and the index scores fall with them. That is the dial to test first.
  • 3.7 Flash is a supported destination, not a legacy one. Google says it "remains fully supported for efficiency-first workloads." For an agentic workflow whose accuracy is already adequate, staying put is the cheaper answer and Google is telling you so.

Benchmarks, including one that disagrees with itself

Google's launch table, 3.8 Flash against 3.7 Flash. These are vendor-run and vendor-selected figures; independent evaluation may differ.

BenchmarkGemini 3.8 FlashGemini 3.7 Flash
DeepSWE v1.173.7%65.3%
Terminal-Bench 2.189.4%85.8%
OSWorld-2.059.0%50.6%
Vals Finance Agent v261.4%59.0%
HLE-Verified54.9%53.6%

Real gains, and DeepSWE v1.1 at 73.7% lands within a third of a point of Claude Opus 5 at 74.0% — from a model at a fraction of the price. But note the shape of the table: the large jumps are on agentic benchmarks with long tool-calling loops, which is precisely where the extra tokens are being spent. The 1.3-point gain on HLE-Verified, a single-shot reasoning test, is what 3.8 Flash buys you when it cannot work harder.

One number does not agree with itself. Google's announcement table gives Terminal-Bench 2.1 as 89.4% / 85.8%; Google's own developer guide reports 90.8% / 81.6% for the same pair — a different run configuration, and a 4.2-point swing in the gap between the two models. The announcement figures are used above. When a vendor publishes two versions of one benchmark three days apart, treat the gap between models as the soft number, not the score.

On the independent side, Gemini 3.8 Flash scores 59 on the Artificial Analysis Intelligence Index v4.2 at high reasoning effort, three points above 3.7 Flash on the same index version. Three points of measured intelligence for roughly 40% more cost per task is the trade in one line. (v4.1 scores are not comparable to v4.2 — the index was reweighted on September 4, 2026.)

The Pro tier that never arrived

The other thing this launch says is what it does not contain. Gemini 3.5 Pro was announced at Google I/O on May 19, 2026 with a June rollout. As of September 6, 2026 it has no API model ID, no pricing entry, and no changelog entry; Google's position is that it is testing with enterprise partners. The newest callable Pro-tier model is still Gemini 3.1 Pro Preview from February 2026, still labelled Preview, and priced several times above the Flash model that outscores it.

So the Pro tier of this generation shipped zero times while Flash shipped three, and the Flash brand quietly stopped meaning "the distilled cheap one." Google now calls Flash its most intelligent model for agentic and coding tasks. That is a reasonable thing for the Gemini line to do — but it is also why the cost story matters. There is no Pro model above this one to migrate to, and no cheaper Flash below it that is not on the same expiry date.

Conclusion

Gemini 3.8 Flash is a genuine improvement, and on agentic coding it is close to models costing an order of magnitude more. It is not, however, a free upgrade from 3.7 Flash, and the identical rate card is the reason that is easy to miss.

Two dates decide the economics. The first is your next billing cycle, where a model that "works harder" spends about 30% more output tokens on the same work. The second is January 1, 2027, when the introductory rate across the whole Flash line doubles. A team that budgeted 3.8 Flash from today's sticker price has underestimated it on both axes at once.

Benchmark the effort levels, meter tokens per completed task rather than per call, pin gemini-3.8-flash instead of the alias, and treat staying on 3.7 Flash as a live option rather than technical debt. Google does.

Sources

Frequently Asked Questions

$0.75 per million input tokens and $3.75 per million output tokens. That is an introductory rate through December 31, 2026; both figures double to $1.50 and $7.50 on January 1, 2027. Batch and Flex run at half the standard rate, Priority at 1.8x, and every one of those tiers doubles on the same date.
The rate card is identical, but the bill is not. Google says 3.8 executes extra reasoning steps and calls tools iteratively and 'might use more tokens to maximize performance.' Artificial Analysis measured roughly 30% more output tokens per task and a cost per task on its Intelligence Index v4.2 suite of $0.58 against $0.40 for 3.7 Flash.
Only if the accuracy gain is worth roughly 40% more spend on comparable work. Google explicitly keeps 3.7 Flash fully supported and recommends it for efficiency-first workloads, so this is a real choice rather than a deprecation.
gemini-3.8-flash, stable and generally available since September 2, 2026. Do not rely on the gemini-flash-latest alias in production: it has been hot-swapped three times in six weeks.
No. The security-specialised variant is not on the public API at all. It is distributed only through Google's Fairwind Program, a vetted channel for government authorities, critical infrastructure operators and software maintainers.
No. Announced at Google I/O on May 19, 2026 with a June rollout, it still has no API model ID, no pricing entry and no changelog entry as of September 6, 2026. The newest callable Pro-tier model remains Gemini 3.1 Pro Preview from February 2026.

Continue Your AI Journey

Explore our glossary and model catalog to deepen your understanding.