Overview
DeepSeek V4 is the fourth generation of DeepSeek's open-weight large language model series. It is not one model but a family of three callable checkpoints, and the family grew again in August 2026:
| Date | What shipped |
|---|---|
| 2026-04-24 | V4-Pro and V4-Flash preview, on the OpenAI and Anthropic API surfaces |
| 2026-07-31 | deepseek-v4-flash official release (DeepSeek-V4-Flash-0731), re-post-trained on the same architecture |
| 2026-08-13 | deepseek-v4-pro GA (DeepSeek-V4-Pro-0813): agent upgrades, three thinking-effort levels, native Responses API |
| 2026-08-16 | Peak / off-peak pricing replaces the flat rate โ a substantial price rise |
| 2026-08-21 | deepseek-v4-flash-vision-exp on the API: the family's first multimodal model |
| 2026-08-31 | Vision-Exp weights published on Hugging Face under MIT |
The generation's defining engineering work is attention. V4 replaces dense attention with a hybrid stack of Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), layered on Manifold-Constrained Hyper-Connections (mHC) and trained with the Muon optimizer. DeepSeek reports that at a 1M-token context, V4-Pro "requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2." Long context stops being a premium feature and becomes the default operating mode.
The August addition is eyes. deepseek-v4-flash-vision-exp takes the V4-Flash backbone, adds a vision encoder and aligner, and continues training to unlock image understanding โ the payoff to the "Now, We See You" teaser DeepSeek posted in May. It is explicitly experimental. DeepSeek's model card calls it "our first experimental multimodal model in the DeepSeek-V4 family," and the -exp suffix carries no stability, versioning or deprecation guarantee. It is a model to prototype against and self-host, not one to put behind a customer-facing SLA.
Capabilities
- Frontier open-weight coding: 80.6 on SWE-bench Verified and a 3206 Codeforces rating put V4-Pro at the top of the open-weight field for software engineering.
- Image understanding at text prices: Vision-Exp reads screenshots, charts, documents and UI states, billed at V4-Flash token rates with a 384-token ceiling per image.
- Multimodal agents: DeepSeek's headline claim for Vision-Exp is a large jump on multimodal agent benchmarks with text-only agent performance held level against V4-Flash-0731.
- Graded thinking effort: V4-Pro's GA build exposes low / high / max effort levels โ low for simple calls, high for everyday agent loops, max for hard problems.
- Million-token working memory: The hybrid attention stack is what makes a 1M window economically usable rather than a specification-sheet number.
- 384K output: Long enough to emit a large refactor, a full technical report or an extended agent trajectory in one response.
- Three API surfaces: OpenAI ChatCompletions, Anthropic Messages, and โ since the August GA โ the OpenAI Responses API.
- Self-hostable: MIT-licensed weights for all three checkpoints, with vLLM support including an optional speculative decoding module.
Technical Specifications
deepseek-v4-pro | deepseek-v4-flash | deepseek-v4-flash-vision-exp | |
|---|---|---|---|
| Checkpoint | V4-Pro-0813 (GA) | V4-Flash-0731 | V4-Flash-Vision-Exp (experimental) |
| Total parameters | 1.6T | 284B | 284B backbone + vision modules |
| Activated per token | 49B | 13B | 13B |
| Modalities | text | text | text + image in, text out |
| Context window | 1M tokens | 1M tokens | 1M tokens |
| Max output | 384K tokens | 384K tokens | 384K tokens |
| License | MIT | MIT | MIT |
| Thinking modes | non-thinking / thinking, effort lowยทhighยทmax | non-thinking / thinking (default) | non-thinking / thinking (default) |
- Architecture: Mixture-of-Experts Transformer with a Hybrid Attention Architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA)
- Residual structure: Manifold-Constrained Hyper-Connections (mHC)
- Optimizer: Muon
- Pre-training corpus: more than 32T tokens
- Efficiency: at 1M-token context, V4-Pro uses 27% of the single-token inference FLOPs and 10% of the KV cache of DeepSeek-V3.2
The vision stack
Vision-Exp is described by DeepSeek as building on the V4-Flash architecture "by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities." Community serving recipes put the added tower at a 32-layer, 1024-dimension Vision Transformer with a two-layer aligner, roughly 0.5B parameters on top of the 284B MoE backbone. Note that the Hugging Face repository metadata reports 305B parameters, against DeepSeek's own 284B for the backbone; the backbone figure is the one that matches the serving recipes, so treat 284B / 13B active as load-bearing and 305B as repo bookkeeping.
What the multimodal API actually accepts, per DeepSeek's vision guide:
- Formats: JPEG, PNG, GIF, WebP โ detected from file content, not the filename
- Delivery: base64 data URL, HTTP(S) URL (max 8192 characters), or a
file_idfrom the new free Files API - Volume: up to 600 images per request; 64 MiB total inline, 200 MiB including Files API images
- Resolution: 8192 px per side, dropping to 4096 px once a request carries 15 or more images
- Billing: images are resized automatically and cost up to 384 tokens each, at V4-Flash text rates
- Placement: images are accepted in
usermessages only โ a system or assistant message carrying an image is an error
DeepSeek-V4-Pro-DSpark
DeepSeek-V4-Pro-DSpark appears as a separate Hugging Face repository, but it is not a separate model. DeepSeek states plainly: "DeepSeek-V4-Pro-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached." In vLLM it is enabled with a single flag. Treat it as a serving optimization, not a capability upgrade.
Legacy model IDs
The two legacy API names, deepseek-chat and deepseek-reasoner, were discontinued on July 24, 2026. During the transition window they pointed at the non-thinking and thinking modes of deepseek-v4-flash. Applications still referencing them are already broken.
Serving hardware
DeepSeek trains on Nvidia and is moving its serving fleet elsewhere. Bloomberg reported in September 2026 that the company plans to deploy at least 160,000 Huawei Ascend 950DT accelerators at a gigawatt-scale site in Inner Mongolia โ for inference only, with training staying on Nvidia hardware. See DeepSeek's 160,000-chip Huawei cluster for the detail; the practical read for an API customer is that DeepSeek's serving cost structure, and therefore its pricing, is being decoupled from Nvidia supply.
Use Cases
Which variant to call, and why:
deepseek-v4-proโ repository-scale software engineering. An 80.6 SWE-bench Verified score, a 1M context window and a 384K output ceiling suit whole-repo refactors and migrations that must read broadly before writing. Usemaxeffort for the hard steps andlowfor the mechanical ones; the effort dial is the main cost lever now that the price is three times Flash's.deepseek-v4-proโ competitive and algorithmic programming. A 3206 Codeforces rating and 93.5 on LiveCodeBench make it a credible assistant for algorithm design and contest-style problems.deepseek-v4-flashโ high-volume text work. Same context and output limits as Pro at a third of the price, and per Artificial Analysis' Intelligence Index v4.1 the 0731 refresh scores above V4-Pro. For most text pipelines Flash is the default and Pro is the escalation.deepseek-v4-flash-vision-expโ document, chart and screenshot understanding. Reading tables out of scanned filings, extracting figures from charts, or turning UI screenshots into structured state. The 384-token image cap makes bulk document ingestion unusually cheap.deepseek-v4-flash-vision-expโ multimodal agents, in prototype. Browser and desktop agents that must look at what they are operating. This is the model's stated strength and its least production-safe use; gate it behind an evaluation set of your own.- Long-document analysis. Legal discovery, regulatory filings, multi-paper literature review โ a million tokens of context removes most chunking and retrieval plumbing.
- Private and air-gapped deployment. MIT-licensed weights make on-premises hosting viable for regulated industries. V4-Flash and Vision-Exp are the realistic self-hosting targets; V4-Pro at 1.6T is not.
Performance / Benchmarks
All figures below are vendor-run, from DeepSeek's own model cards on Hugging Face, unless attributed otherwise. Vendor benchmarks are self-selected; treat them as an upper bound and validate on your own evaluation set.
DeepSeek-V4-Pro (April preview card)
| Benchmark | Score |
|---|---|
| SWE-bench Verified | 80.6 |
| GPQA Diamond | 90.1 |
| LiveCodeBench | 93.5 |
| Codeforces (rating) | 3206 |
| MMLU-Pro | 87.5 |
| Humanity's Last Exam (HLE) | 37.7 |
DeepSeek-V4-Flash (April preview card, Think Max mode)
| Benchmark | Score |
|---|---|
| LiveCodeBench | 91.6 |
| MMLU-Pro | 86.2 |
| SimpleQA-Verified | 34.1 |
| MRCR (1M long context) | 78.7 |
DeepSeek has not republished a full comparable table for the 0731 and 0813 refreshes, so the two tables above describe the April preview checkpoints, not the current GA builds.
DeepSeek-V4-Flash-Vision-Exp (August model card)
| Text agent benchmark | Score |
|---|---|
| Terminal Bench 2.1 | 83.9 |
| NL2Repo | 57.7 |
| Cybergym | 75.3 |
| DeepSWE | 59.3 |
| Toolathlon-Verified | 75.9 |
| DSBench-Hard | 63.6 |
| AutomationBench | 25.7 |
| Multimodal agent benchmark | Score |
|---|---|
| ApexBench (Pass@1) | 36.5 |
| Agents' Last Exam | 27.3 |
| Chartography | 64.3 |
| ZeroBench (Pass@5) | 35.0 |
DeepSeek's claim is that Vision-Exp "achieves substantial improvements on its multimodal agent capabilities, while maintaining comparable performance on text-only agent tasks" relative to V4-Flash-0731, and that it lands close to Anthropic's Opus-4.8 on multimodal agent work (see Claude Opus). That comparison is DeepSeek's own and has not been independently replicated.
Third-party
Artificial Analysis, running its Intelligence Index v4.1 on July 31, 2026, scored DeepSeek-V4-Flash-0731 at 50 โ ten points above the April V4-Flash (40) and six above V4-Pro (44). Artificial Analysis lists an Intelligence Index figure for the vision model as well, but flags it as an estimate and on a later index revision, so it is not comparable to the numbers above and is omitted here.
Limitations
- Vision-Exp is experimental, and DeepSeek says so. No stability guarantee, no announced GA successor, no promised migration path. Pin the model ID and keep a text-only fallback.
- Prices rose sharply in August 2026. The flat $0.435 / $0.87 V4-Pro rate is gone; peak input is now roughly 3x that and peak output about 4.5x. Cost models built before August 16 are wrong.
- Peak/off-peak billing is a scheduling problem. Rates double during 01:00-04:00 and 06:00-10:00 UTC on weekdays. Latency-sensitive interactive traffic cannot be shifted, so it pays the peak rate.
- No single "V4" endpoint. You must choose an explicit model ID. Code assuming a generic
deepseek-v4will fail. - Vision serving is Nvidia-only in practice. Community vLLM recipes report no vision implementation in the ROCm or XPU builds, and roughly 202 GB of VRAM before KV cache at FP4+FP8 mixed precision.
- Serving cost at 1.6T. Even at 49B active parameters, self-hosting V4-Pro needs a substantial multi-GPU cluster.
- Weights, not full openness. MIT covers the weights. Training data and the full training pipeline are not published.
- No published knowledge cutoff. DeepSeek does not state a reliable cutoff for V4. Ground time-sensitive queries with retrieval or tool use.
- Benchmarks are vendor-reported. Independent replication of the headline scores, and of the Opus-4.8 multimodal comparison in particular, is limited.
Pricing & Access
Prices are from DeepSeek's models and pricing page, per 1M tokens in USD. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; all other hours are off-peak, at half the peak rate. This structure took effect on August 16, 2026.
| Model | Input, cache hit | Input, cache miss | Output |
|---|---|---|---|
deepseek-v4-pro (peak) | $0.044 | $1.32 | $3.96 |
deepseek-v4-pro (off-peak) | $0.022 | $0.66 | $1.98 |
deepseek-v4-flash (peak) | $0.014 | $0.44 | $1.32 |
deepseek-v4-flash (off-peak) | $0.007 | $0.22 | $0.66 |
deepseek-v4-flash-vision-exp | same as deepseek-v4-flash |
Two things follow. First, the cache-hit price is about 1/30th of the cache-miss price, so prompt caching remains the highest-leverage optimization for any repeated-context workload. Second, DeepSeek is no longer the outlier-cheap option it was in the spring: the permanent 75% discount announced in May was retired with the GA rate card, and V4-Pro output at peak is more than four times its former price.
Access options:
- DeepSeek API โ OpenAI ChatCompletions, Anthropic Messages, and OpenAI Responses interfaces
- DeepSeek Chat โ official web and mobile applications, with V4-Pro behind "Expert Mode"
- Hugging Face and ModelScope โ MIT-licensed weights for private deployment
Ecosystem & Tools
- DeepSeek API documentation โ reference, changelog and migration guides
- Vision guide โ image formats, size limits and tokenization rules for Vision-Exp
- Files API โ free image upload; reference a
file_idinstead of resending base64 on every request - DeepSeek-V4-Pro on Hugging Face โ frontier-tier weights
- DeepSeek-V4-Flash on Hugging Face โ efficient-tier weights
- DeepSeek-V4-Flash-Vision-Exp on Hugging Face โ MIT vision weights, published August 31, 2026
- DeepSeek-V4-Pro-DSpark โ the same V4-Pro checkpoint with a speculative decoding module for vLLM
- GitHub โ inference code and research releases
- vLLM โ supported serving path, including the vision encoder and DSpark speculative decoding behind a single flag
- DeepSeek Harness โ DeepSeek's agent harness, updated for Vision-Exp on release
Community & Resources
- DeepSeek API changelog โ the release entries for V4, the 0731 and 0813 refreshes, and Vision-Exp
- Models and pricing โ the current peak / off-peak rate card
- DeepSeek-V4-Flash-Vision-Exp model card
- DeepSeek on GitHub
- Our coverage: the V4-Pro and V4-Flash launch, the price cut that became permanent, and the 160,000-chip Huawei cluster
- Compare with Kimi K3, GLM-5.3, Qwen3.8-Max, and Ling-2.6-1T