---
source: 'https://howaiworks.ai/models/gemini'
section: models
title: Gemini 3.8 Flash
description: >-
  Google's Gemini 3.8 Flash, GA September 2, 2026. 1M-token context, $0.75/$3.75
  per 1M introductory, plus a gated Flash Cyber variant.
tags:
  - Gemini
  - Google
  - Google DeepMind
  - Language Model
  - Large Language Model
  - AI Assistant
  - Multimodal
  - Agentic AI
  - Coding AI
  - Latest
category: Language Models
developer: Google
developerWebsite: 'https://ai.google/'
modelType: Multimodal Language Model
releaseDate: '2026-09-02'
lastUpdated: '2026-09-06'
license: Proprietary
contextWindow: 1M tokens
inputPrice: $0.75
outputPrice: $3.75
knowledgeCutoff: March 2026
officialWebsite: >-
  https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/
docsPage: 'https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash'
pricingPage: 'https://ai.google.dev/gemini-api/docs/pricing'
---

# Gemini 3.8 Flash

> Google's Gemini 3.8 Flash, GA September 2, 2026. 1M-token context, $0.75/$3.75 per 1M introductory, plus a gated Flash Cyber variant.

## Overview

**Gemini 3.8 Flash** has been generally available since **September 2, 2026**. It is Google's third Flash release in six weeks — after 3.6 Flash on July 21 and 3.7 Flash in mid-August — and Google pitches it as its "best reasoning & coding model yet, at the same speed and low cost of 3.7." Google DeepMind's own framing is blunter: "our most intelligent workhorse model yet for coding and agents."

That positioning inverts the usual Flash-is-the-cheap-one story. Flash is the frontier model of this line in practice, not a distillation of one: the Pro tier of the generation, Gemini 3.5 Pro, still has not shipped, so Flash is what Google actually competes with.

The line now also has a **gated variant**. Gemini 3.8 Flash Cyber, a security-specialised model for finding and patching vulnerabilities, is not on the public API at all: it is distributed only through the [Fairwind Program](https://howaiworks.ai/blog/google-fairwind-program-2026), Google's vetted-access channel for government authorities, critical infrastructure operators and software maintainers.

Pricing is aggressive and temporary. The whole current Flash line — 3.6, 3.7 and 3.8 — sits at an **introductory $0.75 in / $3.75 out per million through December 31, 2026**, and on **January 1, 2027 both numbers double** to $1.50 / $7.50. Unit economics built on today's rate have four months to run.

One behavioural caveat comes with 3.8: Google says the model "works harder," executing extra reasoning steps and calling tools iteratively, and "might use more tokens to maximize performance, especially at higher effort levels." Google explicitly keeps 3.7 Flash supported for efficiency-first workloads, and recommends lower effort levels where token overhead matters. Identical per-token pricing does not mean an identical bill.

The generation's knowledge cutoff has not moved. Gemini 3.5 Flash launched with a **January 2025** cutoff — already 16 months stale on release day. Gemini 3.6 Flash moved it to **March 2026**, and 3.7 and 3.8 Flash both stay there; Google's model card warns that some domains are still only covered to January 2025.

The rest of the line has not moved. **Gemini 3.5 Pro** still has no model ID and no price; **Gemini 3.1 Pro**, from February 2026, is still Preview and still several times dearer per token than the Flash model that outscores it.

## Capabilities

- **Agentic loops**: Google tunes Flash for long-horizon agentic coding. 3.8's gain over 3.7 is diligence rather than speed — more reasoning steps, more iterative tool calls, and the token cost that implies.
- **Native multimodal input**: Text, image, video, audio, and PDF in a single request. Output is text.
- **Thinking levels**: Low, medium and high. The `minimal` option available on some earlier models is not offered on 3.8 Flash, which is part of why it spends more tokens.
- **Computer use**: A built-in client-side tool in the Gemini API and Gemini Enterprise rather than a separate model path, though the feature remains preview-labelled.
- **Grounding**: Google Search grounding and Google Maps grounding, plus URL context and file search.
- **Code execution**: Runs code as a tool during a response.
- **Long-context retrieval**: Google's last published MRCR figure for the line is 54.0% on GDM-MRCR v2 at 1M tokens, for 3.6 Flash, against under 27% for 3.5 Flash and 3.1 Pro. It has not republished the metric for 3.8.

## Technical Specifications

- **Model ID**: `gemini-3.8-flash` (stable). Predecessors `gemini-3.7-flash`, `gemini-3.6-flash` and `gemini-3.5-flash` remain callable.
- **Alias**: `gemini-flash-latest` tracks the newest Flash release and is hot-swapped on each launch — pin an explicit ID in production.
- **Input limit**: 1,048,576 tokens
- **Output limit**: 65,536 tokens
- **Knowledge cutoff**: March 2026 (some domains only to January 2025, per Google's model card)
- **Input modalities**: Text, image, video, audio, PDF
- **Output modalities**: Text
- **Supported**: Thinking (low/medium/high), structured outputs, function calling, Search grounding, Google Maps grounding, code execution, Batch API, flex inference, priority inference, caching, file search, URL context, computer use (preview)
- **Not supported**: Audio generation, image generation, Live API, `minimal` thinking

## Model Family

### Gemini 3.8 Flash — generally available
Released September 2, 2026. `gemini-3.8-flash`. The current default for agentic and coding work. **$0.75 in / $3.75 out** per million introductory through December 31, 2026, then $1.50 / $7.50. Available in Google AI Studio, the Gemini API, [Google Antigravity](https://howaiworks.ai/ai-tools/google-antigravity), Android Studio, Stitch, Vertex AI, Gemini Enterprise, the Gemini app for Pro and Ultra subscribers, AI Mode in Google Search, and Google Sheets.

### Gemini 3.8 Flash Cyber — Fairwind Program only
A cybersecurity variant for autonomous vulnerability discovery and patching. **Not on the public API.** Access runs through the [Fairwind Program](https://howaiworks.ai/blog/google-fairwind-program-2026): an application, background checks on the organization, phishing-resistant MFA, and access confined to internal security teams. Google's stated reason is worth quoting — the model ships with a "more permissive" set of mitigations for cybersecurity work than public 3.8 Flash under the Frontier Safety Framework. It runs standalone or inside **CodeMender**, Google's patching harness.

### Gemini 3.7 Flash — superseded, still supported
Released mid-August 2026, three weeks before 3.8, at the same $0.75 / $3.75 introductory pricing. Google keeps it recommended for efficiency-first workloads where 3.8's extra reasoning steps and tool calls cost more tokens than the added accuracy is worth.

### Gemini 3.6 Flash — superseded
Released July 21, 2026. Same context limits, March 2026 cutoff. It launched at $1.50 / $7.50 per million, but Google has folded it into the same introductory window as 3.7 and 3.8: it currently bills at $0.75 / $3.75 and reverts on January 1, 2027.

### Gemini 3.5 Flash — superseded, still callable
Released May 19, 2026. Same context limits, January 2025 cutoff. It scores below every later Flash release on the benchmarks Google publishes for both — there is no reason to start new work on it.

### Gemini 3.5 Flash-Lite — generally available
Released July 21, 2026. `gemini-3.5-flash-lite`. Roughly **350 output tokens per second** at **$0.30 in / $2.50 out** per million, for document processing, high-volume automation and subagent roles. It scores 54.2% on SWE-Bench Pro against 49.6% for Gemini 3 Flash. No context caching.

### Gemini 3.5 Flash Cyber — superseded by 3.8 Flash Cyber
A limited-access CodeMender pilot for governments and trusted partners from July 21, 2026, never on the public API. The pilot became the Fairwind Program on September 2, 2026.

### Gemini 3.5 Pro — announced, not shipped
Announced May 19, 2026. No API model ID, no pricing entry and no changelog entry as of September 6, 2026, more than three months after announcement and several missed rollout targets. Google's position is that it is testing with enterprise partners and will launch when ready.

### Gemini 3.1 Pro Preview — available
Released February 19, 2026. Still labeled Preview. Tiered pricing: $2.00 in / $12.00 out per million for prompts up to 200K tokens, $4.00 in / $18.00 out above 200K. Context caching at $0.20 (≤200K) or $0.40 (>200K) per million plus $4.50 per hour.

### Gemini Omni Flash Preview — public preview
Released to developers June 30, 2026. **Video generation** — 3-10 second clips at 720p from text or still images, with conversational editing across turns. Input $1.50 per million; output $9.00 (text) or $17.50 (video) per million, roughly $0.10 per second at 720p.

### Gemini 3.5 Live Translate — preview
Real-time speech-to-speech translation across 70+ languages. Listed in the model catalog and on the pricing page, still Preview.

## Use Cases

- **Long-horizon agentic coding**: The workload Google names first and the one 3.8 was tuned for — 73.7% on DeepSWE v1.1 and 89.4% on Terminal-Bench 2.1, both the best in Google's own comparison set.
- **Domain agent work**: Vals Finance Agent v2 (61.4%) and Harvey's Legal Agent Benchmark (10.0%) are the two places where 3.8 Flash beats every model Google charted, including Claude Opus 5. If you are building a finance or legal agent, this is the specific claim to test.
- **Long-document analysis**: A 1M-token input window covers large codebases, contract sets, and research corpora, with PDFs accepted natively.
- **Grounded question answering**: Search and Maps grounding for questions past the March 2026 cutoff.
- **Multimodal understanding**: Video, audio, and image inputs into a text answer — chart reading (86.2% CharXiv), long-video comprehension (87.8% LVBench), screen understanding.
- **High-throughput batch jobs**: Batch or Flex at 50% of interactive pricing, or Flash-Lite at 40% of the input price when quality demands are lower.
- **Video generation and editing**: Via `gemini-omni-flash-preview`, with conversational refinement instead of full re-prompting.

## Performance / Benchmarks

Google's launch table for **Gemini 3.8 Flash**, September 2, 2026. All figures are vendor-run and vendor-selected; independent evaluations may differ.

| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| DeepSWE v1.1 | 73.7% | 65.3% | 74.0% | 72.7% |
| Terminal-Bench 2.1 | 89.4% | 85.8% | 89.1% | 88.8% |
| Terminal-Bench 4.0 | 19.1% | 11.2% | 51.8% | 37.3% |
| OSWorld-2.0 | 59.0% | 50.6% | 75.4% | 62.6% |
| Vals Finance Agent v2 | 61.4% | 59.0% | 58.6% | 53.8% |
| Harvey Legal Agent | 10.0% | 8.8% | 6.7% | 2.5% |
| HLE-Verified | 54.9% | 53.6% | 54.4% | 54.5% |
| GDPval-AA v2 | 1545 Elo | 1482 Elo | 1824 Elo | 1710 Elo |
| BioMysteryBench (Hard) | 56.5% | 43.5% | 49.4% | 44.7% |
| Gray Swan IPI (lower is better) | 5.5% | 9.2% | 4.8% | 27.0% |

One number to treat carefully: **Terminal-Bench 2.1**. Google's announcement table gives 89.4% for 3.8 Flash and 85.8% for 3.7 Flash, but Google's own developer guide has been reported as showing 90.8% and 81.6% for the same pair — a different run configuration, presumably. The announcement figures are used above.

**It does not lead everywhere.** Terminal-Bench 4.0 is the tell: 19.1% against Claude Opus 5's 51.8%, on the harder successor to the benchmark 3.8 Flash tops. Opus 5 also leads OSWorld-2.0 and GDPval-AA v2 by wide margins. The argument for 3.8 Flash is price-performance in a specific band — agentic coding and domain knowledge work — not the crown.

For **Gemini 3.8 Flash Cyber**, Google reports 47.2% pass@1 on CWE-Bench against 47.8% for the leading frontier model at far lower cost, over 70% success on an internal 20-language vulnerability-discovery benchmark, and — per Google's Chrome Security team — 2.6x more correct Chrome patches than larger commercial models. Its CyberGym result is called frontier-level, but **no score was published**.

## Limitations

- **Weak on the hardest agentic terminal work.** 19.1% on Terminal-Bench 4.0 is a third of Claude Opus 5's score, and a warning against extrapolating from the 2.1 result.
- **Prices double on January 1, 2027.** The $0.75 / $3.75 rate is introductory across the whole Flash line, not a permanent repricing.
- **It may cost more than 3.7 Flash at the same rate card**, because Google says it spends more tokens at higher effort levels.
- **Text output only.** Image and audio generation and the Live API are unsupported; use Omni Flash for video.
- **No `minimal` thinking level**, so the cheapest reasoning setting available on some earlier models is not an option here.
- **Gemini 3.5 Pro still does not exist** as a callable model, more than three months after announcement and several missed rollout dates.
- **Output cap of 65,536 tokens** is one-sixteenth of the 1,048,576-token input window; very long generations must be chunked.
- **`gemini-flash-latest` is hot-swapped** on every Flash release — three times in six weeks — so an unpinned alias silently changes model under a running application.
- **Grounding costs extra** beyond 5,000 prompts per month (shared across the Gemini 3 line), at $14 per 1,000 search queries.
- **Flash Cyber is unavailable to almost everyone.** The most differentiated model in the line cannot be bought.

## Pricing & Access

### Gemini 3.8 Flash (per 1M tokens)

| Tier | Input | Output |
|---|---|---|
| Standard, introductory (through Dec 31, 2026) | $0.75 | $3.75 |
| Standard, from Jan 1, 2027 | $1.50 | $7.50 |
| Batch / Flex (introductory) | $0.375 | $1.875 |
| Priority (introductory) | $1.35 | $6.75 |

Context caching: $0.075 per million tokens plus $0.50 per million tokens per hour of storage, both doubling on January 1, 2027. Gemini 3.7 Flash and Gemini 3.6 Flash carry the same rate card. **Gemini 3.8 Flash Cyber has no published price** — it is distributed through the Fairwind Program rather than sold per token.

### Gemini 3.5 Flash-Lite (per 1M tokens)

| Tier | Input | Output |
|---|---|---|
| Standard | $0.30 | $2.50 |
| Batch / Flex | $0.15 | $1.25 |
| Priority | $0.54 | $4.50 |

Context caching is not available.

### Gemini 3.1 Pro Preview (per 1M tokens)

| Prompt size | Input | Output |
|---|---|---|
| ≤ 200K tokens | $2.00 | $12.00 |
| > 200K tokens | $4.00 | $18.00 |

Batch and Flex: $1.00 / $6.00 (≤200K) and $2.00 / $9.00 (>200K). Context caching: $0.20 or $0.40 per million plus $4.50 per hour.

### Grounding

5,000 prompts per month free, shared across the Gemini 3 line, then $14 per 1,000 search queries.

### Where the flagship Flash model is available

As of Gemini 3.8 Flash: Google AI Studio, the Gemini API, Google Antigravity, Android Studio, Stitch, Vertex AI, the Gemini Enterprise Agent Platform, Gemini Enterprise, the Gemini app for Pro and Ultra subscribers, AI Mode in Google Search, and Google Sheets.

## Ecosystem & Tools

**Google platforms**
- **[Google AI Studio](https://aistudio.google.com/)** - Prototyping and API keys
- **[Google Antigravity](https://antigravity.google/)** - Agentic development platform
- **[Vertex AI](https://cloud.google.com/vertex-ai)** - Enterprise deployment and MLOps
- **Gemini Enterprise Agent Platform** - Agent hosting for enterprises, with zero data retention
- **Android Studio** and **Stitch** - In-IDE and UI-generation access

**Open-weight sibling**
- **[Gemma 4](https://howaiworks.ai/models/gemma)** - Google's open-weight family, Apache 2.0 licensed, for local and on-device deployment

**Developer**
- **[Gemini API docs](https://ai.google.dev/gemini-api/docs)** and official SDKs

## Community & Resources

- [Gemini 3.8 Flash and 3.8 Flash Cyber announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/) - Google's September 2, 2026 launch post
- [Gemini 3.8 Flash model card](https://deepmind.google/models/model-cards/gemini-3-8-flash/) - Knowledge cutoff, evaluations, safety
- [Gemini 3.8 Flash model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash) - Limits, modalities, capabilities
- [Gemini 3.6 Flash announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/) - Google's July 2026 launch post
- [Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing)
- [Gemini API changelog](https://ai.google.dev/gemini-api/docs/changelog) - Release dates for every model version
- [Google DeepMind — Gemini Flash](https://deepmind.google/models/gemini/flash/)
- [Google Ships Gemini 3.6 Flash at a Lower Price Than 3.5](https://howaiworks.ai/blog/google-gemini-3-6-flash-launch-2026) - Our launch write-up
- [Google's Fairwind Program](https://howaiworks.ai/blog/google-fairwind-program-2026) - How gated access to Flash Cyber works

## Frequently Asked Questions

### What is the current Gemini Flash model?

Gemini 3.8 Flash, generally available since September 2, 2026. It is Google's third Flash release in six weeks, after 3.6 Flash on July 21 and 3.7 Flash in mid-August. Both predecessors remain callable.

### What is the context window of Gemini 3.8 Flash?

1,048,576 input tokens and 65,536 output tokens — unchanged across the Flash line since Gemini 3.5 Flash.

### What is Gemini 3.8 Flash's knowledge cutoff?

March 2026, the same cutoff as 3.6 and 3.7 Flash. Google's model card adds the caveat that coverage in some domains only reaches January 2025.

### How much does Gemini 3.8 Flash cost?

$0.75 per million input tokens and $3.75 per million output tokens, an introductory rate through December 31, 2026. Both figures double on January 1, 2027, to $1.50 and $7.50. Batch and Flex run at half the standard rate; Priority is 1.8x.

### Is Gemini 3.8 Flash cheaper to run than 3.7 Flash?

The per-token price is identical, but the bill may not be. Google says 3.8 "works harder" — extra reasoning steps and iterative tool calls — and "might use more tokens to maximize performance, especially at higher effort levels." Google keeps 3.7 Flash recommended for efficiency-first workloads.

### What modalities does Gemini 3.8 Flash accept?

Text, image, video, audio, and PDF as input; text as output. It does not generate images or audio, and does not support the Live API.

### How does Gemini 3.8 Flash score on benchmarks?

On Google's own announcement table: 73.7% on DeepSWE v1.1, 89.4% on Terminal-Bench 2.1, 59.0% on OSWorld-2.0, 61.4% on Vals Finance Agent v2, and 54.9% on HLE-Verified. All are vendor-run figures.

### Is Gemini 3.5 Pro available?

No. Announced at Google I/O on May 19, 2026, it still has no API model ID and no pricing entry as of September 6, 2026, and Google describes it as testing with enterprise partners. The newest Pro-tier model you can call is `gemini-3.1-pro-preview`.

### What is Gemini 3.8 Flash Cyber?

A security-specialised variant released alongside 3.8 Flash for autonomous vulnerability discovery and patching. It is not on the public API: access runs only through the [Fairwind Program](https://howaiworks.ai/blog/google-fairwind-program-2026), Google's vetted channel for government authorities, critical infrastructure operators and software maintainers.

### What is Gemini 3.5 Flash-Lite?

A high-throughput, low-latency model released July 21, 2026 at $0.30 in / $2.50 out per million tokens, running at roughly 350 output tokens per second. It remains the cheapest model in the line and does not support context caching.

### Where can I use Gemini 3.8 Flash?

Google AI Studio, the Gemini API, Google Antigravity, Android Studio, Stitch, Vertex AI, Gemini Enterprise, the Gemini app for Pro and Ultra subscribers, AI Mode in Google Search, and Google Sheets.

## Related

### Related models

- [GPT-6 Astra](https://howaiworks.ai/models/gpt)
- [Claude Opus 5](https://howaiworks.ai/models/claude-opus)
- [Claude Sonnet 5](https://howaiworks.ai/models/claude-sonnet)
- [Grok 4.6](https://howaiworks.ai/models/grok)
- [Gemma 4](https://howaiworks.ai/models/gemma)
- [Llama 4](https://howaiworks.ai/models/llama)

---

Source: https://howaiworks.ai/models/gemini — HowAIWorks.ai
