---
source: 'https://howaiworks.ai/models/llama'
section: models
title: Llama 4
description: >-
  The last release in Meta's Llama line: Scout and Maverick, natively multimodal
  Mixture-of-Experts models with 17B active and up to a 10M token context.
tags:
  - Llama
  - Meta
  - Open Weights
  - Language Model
  - Large Language Model
  - Multimodal
  - Mixture of Experts
  - Scout
  - Maverick
category: Language Models
developer: Meta
developerWebsite: 'https://ai.meta.com/'
modelType: Multimodal Language Model
releaseDate: '2025-04-05'
lastUpdated: '2026-09-06'
license: Llama 4 Community License Agreement
contextWindow: 10M tokens (Scout)
knowledgeCutoff: August 2024
officialWebsite: 'https://ai.meta.com/blog/llama-4-multimodal-intelligence/'
docsPage: 'https://developer.meta.com/ai/'
pricingPage: ''
---

# Llama 4

> The last release in Meta's Llama line: Scout and Maverick, natively multimodal Mixture-of-Experts models with 17B active and up to a 10M token context.

## Overview

Llama 4 is the **last release in the Llama line**. Published on April 5, 2025, it introduced Scout and Maverick: two natively multimodal [Mixture-of-Experts](https://howaiworks.ai/glossary/mixture-of-experts) models that each activate 17 billion parameters per token, and pushed open-weight [context windows](https://howaiworks.ai/glossary/context-window) to 10 million tokens.

Seventeen months later, no new Llama foundation model has shipped. There is no Llama 4.5 and no Llama 5. The line was extended once, at LlamaCon on April 29, 2025, with safety components rather than a base model: Llama Guard 4, Llama Prompt Guard 2, and LlamaFirewall. Llama 4 Behemoth — announced in the same April 2025 post at 288B active parameters and roughly 2 trillion total — was still training and was never publicly released.

Meta's model work has moved to the **Muse** line, built by Meta Superintelligence Labs, and it went in two directions at once. **[Muse Spark](https://howaiworks.ai/models/muse-spark)** — April 8, 2026, updated to 1.1 on July 9, 2026 and to 1.3 on September 2, 2026 — is the proprietary frontier model, with no downloadable weights and a paid endpoint on the Meta Model API. **[Muse Glimmer](https://howaiworks.ai/models/muse-glimmer)** — August 10, 2026 — is the open one: a 30B dense multimodal model published under **Apache 2.0**, a more permissive licence than any Llama ever carried, and small enough to run offline on a single 24 GB consumer GPU.

So the accurate statement is narrower than "Meta left open weights." Meta stopped shipping *Llama*, and it stopped publishing open weights at frontier scale — but it did not stop publishing open weights. What it releases openly now is a small local-agent model rather than a frontier MoE family.

Meta has published no statement discontinuing Llama, and none announcing a successor. Llama 4 is simply the most recent release in a line that has not been extended for seventeen months. Meta shut down its own hosted Llama API on July 6, 2026, and `llama.com` now redirects to `developer.meta.com/ai/`, where Llama 3 and Llama 4 both remain listed and downloadable.

## Model Family

### Llama 4 Scout
- **Active parameters**: 17B
- **Total parameters**: 109B
- **Experts**: 16
- **Context window**: 10,000,000 tokens
- **Distinguishing feature**: iRoPE — interleaved attention layers carrying *no* positional embeddings, with rotary position embeddings (RoPE) in most other layers and inference-time temperature scaling of attention; the mechanism behind the 10M window

Meta positioned Scout against Gemma 3, Gemini 2.0 Flash-Lite, and Mistral 3.1 at launch.

### Llama 4 Maverick
- **Active parameters**: 17B
- **Total parameters**: 400B
- **Experts**: 128
- **Context window**: 1,000,000 tokens
- **Distinguishing feature**: the general-purpose workhorse of the pair, stronger on reasoning and coding

Meta's April 2025 claim was that Maverick beat GPT-4o and Gemini 2.0 Flash, with results comparable to DeepSeek v3 on reasoning and coding.

### Llama 4 Behemoth — never released
- **Active parameters**: 288B
- **Total parameters**: ~2 trillion
- **Experts**: 16

Previewed as still in training. Meta claimed it outperformed GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. It served as an internal teacher model for distillation and never shipped as public weights.

## Capabilities

- **Native multimodality via early fusion**: Text and vision tokens are integrated in a unified backbone from the start, rather than through a separately trained vision encoder.
- **Mixture-of-Experts efficiency**: Alternating dense and MoE layers mean only 17B parameters activate per token, on models totalling 109B and 400B.
- **10M token context (Scout)**: The longest context window in any widely deployed open-weight model, enabled by iRoPE.
- **Massively multilingual pre-training**: 200 languages, including more than 100 with over a billion tokens each.
- **Open weights**: Downloadable and self-hostable under the Llama 4 Community License Agreement.
- **Text and code output**: From text and image input.

## Technical Specifications

| | Scout | Maverick | Behemoth (unreleased) |
|---|---|---|---|
| Active parameters | 17B | 17B | 288B |
| Total parameters | 109B | 400B | ~2T |
| Experts | 16 | 128 | 16 |
| Context window | 10M tokens | 1M tokens | not disclosed |

- **Architecture**: [Mixture-of-Experts](https://howaiworks.ai/glossary/mixture-of-experts) with alternating dense and MoE layers, on a [Transformer](https://howaiworks.ai/glossary/transformer) backbone
- **Multimodality**: Early fusion of text and vision tokens
- **Long context (Scout)**: iRoPE — interleaved attention layers with *no* positional embeddings, rotary position embeddings (RoPE) in most other layers, and inference-time temperature scaling of attention
- **Training corpus**: More than 30 trillion tokens across 200 languages
- **Input modalities**: Multilingual text and images
- **Output modalities**: Multilingual text and code
- **License**: Llama 4 Community License Agreement (custom commercial license)
- **Officially supported languages** (Maverick Instruct model card): Arabic, English, French, German, Hindi, Indonesian, Italian, Portuguese, Spanish, Tagalog, Thai, Vietnamese
- **Knowledge cutoff**: August 2024, per Meta's Maverick Instruct model card

## Where Llama 4 Stands Today

Llama 4 is still a capable, freely downloadable multimodal MoE family, and Scout's 10M token window has no open-weight equal. But three things have changed since April 2025:

1. **Meta's frontier moved to closed weights — its open-weight programme did not stop.** Muse Spark (April 8, 2026, now at 1.3) has no public checkpoint. Muse Glimmer (August 10, 2026) does, under Apache 2.0 — but at 30B dense with a 128K context it is a local-agent model, not a frontier-scale replacement for Maverick.
2. **The open-weight field caught up and moved on.** [Gemma 4](https://howaiworks.ai/models/gemma) shipped under Apache 2.0 on March 31, 2026 with a 31B dense model Google places at #3 among *open models* on the Arena text leaderboard, and [DeepSeek V4](https://howaiworks.ai/models/deepseek) and the [Qwen](https://howaiworks.ai/models/qwen) Max line both post newer checkpoints.
3. **Llama 4's benchmark claims are dated.** Meta's comparisons were drawn against GPT-4o, Gemini 2.0 Flash, and Claude Sonnet 3.7 — models that have themselves been superseded twice over.

Choose Llama 4 today for its 10M token context and its maturity in the local-inference tooling ecosystem — not because it is the strongest open-weight model available, and not for its licence, which Muse Glimmer's Apache 2.0 now beats outright.

On the question a reader actually arrives with — *should I still build on Llama 4?* — the supported answer is that it is available but not maintained. The weights are still published by Meta, the licence is unchanged, and Scout and Maverick are still served by the major cloud and inference providers. What Llama 4 does not have is a roadmap: no successor has been announced, nothing will be fixed, and the August 2024 knowledge cutoff will not move. That is fine for a frozen fine-tuning base or a long-context retrieval workload, and a poor bet for anything that needs the model to improve.

## Use Cases

- **Extreme-length retrieval and analysis**: Scout's 10M token window handles entire document archives or repository histories without chunking.
- **Self-hosted multimodal inference**: Text and image understanding on private infrastructure with no per-token cost.
- **Fine-tuning base**: A well-supported, widely documented starting point for domain adaptation.
- **Cost-controlled serving**: 17B active parameters keeps inference cost far below what the 109B and 400B totals suggest.
- **Multilingual applications**: 200 languages in pre-training, 12 officially supported in the instruction-tuned checkpoints.
- **Research on MoE and early fusion**: Open weights make Llama 4 one of the few frontier-scale MoE models available for architectural study.

## Performance / Benchmarks

Meta's April 2025 announcement leads with comparative claims. **These comparisons are against 2025-era models and should not be read as current standings:**

- **Scout** outperforms Gemma 3, Gemini 2.0 Flash-Lite, and Mistral 3.1.
- **Maverick** beats GPT-4o and Gemini 2.0 Flash, with results comparable to DeepSeek v3 on reasoning and coding.
- **Behemoth** (unreleased) outperforms GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks.

One score appears in the announcement itself: Maverick reached an **ELO of 1417 on LMArena**. Read the qualifier carefully — Meta attributes that number to "an experimental chat version," not to the weights it published. The two are not the same artifact.

Fuller benchmark tables do exist, on Meta's own instruction-tuned model cards rather than in the blog post. See the [Maverick Instruct card](https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct) and the [Scout Instruct card](https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct) for the current figures.

## Limitations

- **No new foundation model in seventeen months, and no successor announced.** Nothing since April 5, 2025 but safety components. Meta has not said the line is over, which also means it has not committed to maintaining it.
- **The licence is no longer competitive within Meta's own catalog.** Muse Glimmer ships under Apache 2.0; Llama 4 does not.
- **Behemoth was never released.** The strongest model in the announcement does not exist as downloadable weights.
- **The license is not open source.** The Llama 4 Community License Agreement is a custom commercial license, not Apache 2.0 or MIT. Among its restrictions: entities with more than **700 million monthly active users** must obtain separate permission from Meta before using the model at all.
- **Meta's own hosted API is gone.** The Llama API — free, and in public preview its whole life — shut down on July 6, 2026. Requests to it now return a sunset response.
- **Knowledge cutoff of August 2024.** Nearly two years stale as of mid-2026.
- **10M tokens is a ceiling, not a promise.** Scout's window is architecturally supported; quality and cost at extreme lengths are separate questions from whether the model accepts the input.
- **Text output only.** Llama 4 understands images; it does not generate them.
- **Twelve officially supported languages** in the instruct checkpoints, despite 200 in pre-training.
- **Serving cost of Maverick.** 400B total parameters must be resident in memory even though only 17B activate per token.

## Pricing & Access

Llama 4 weights are **free to download and self-host** under the Llama 4 Community License Agreement.

Meta no longer hosts Llama itself. The first-party **Llama API**, announced at LlamaCon in April 2025 and never taken out of free public preview, **shut down on July 6, 2026**; requests now return a sunset response, and Meta directs developers to third-party hosts. There is no first-party paid API to migrate to.

**Where to get the weights**
- [developer.meta.com/ai](https://developer.meta.com/ai/) - Official download portal (`llama.com` redirects here)
- [Hugging Face](https://huggingface.co/meta-llama) - Base and instruction-tuned checkpoints for Scout and Maverick

**Hosted inference** is available through major cloud providers and third-party model-hosting platforms; pricing varies by provider and Meta publishes no reference rate.

**The Muse models** are separate products, not upgrades on this page's terms. [Muse Spark](https://howaiworks.ai/models/muse-spark) is proprietary and, since the 1.3 release of September 2, 2026, sold through the Meta Model API rather than confined to a private preview; it is also served through meta.ai and the Meta AI app. See its own page for the rate card. [Muse Glimmer](https://howaiworks.ai/models/muse-glimmer) is free to download under Apache 2.0 and is the model to weigh against Llama 4 Scout for local deployment: 30B dense on a single 24 GB GPU, against Scout's 109B total and 10M token window.

## Ecosystem & Tools

- **[developer.meta.com/ai](https://developer.meta.com/ai/)** - Downloads, model cards, and documentation
- **[Hugging Face](https://huggingface.co/meta-llama)** - `Llama-4-Scout-17B-16E-Instruct` and `Llama-4-Maverick-17B-128E-Instruct`
- **[Llama Cookbook](https://github.com/meta-llama/llama-cookbook)** - Official recipes and fine-tuning examples
- **Local runtimes** - llama.cpp, Ollama, vLLM, and LM Studio all support Llama 4 checkpoints
- **Cloud hosting** - Available through major cloud providers and dedicated inference platforms

## Community & Resources

- [The Llama 4 herd](https://ai.meta.com/blog/llama-4-multimodal-intelligence/) - Meta's April 5, 2025 announcement
- [Introducing Muse Spark](https://ai.meta.com/blog/introducing-muse-spark-msl/) - Meta Superintelligence Labs, April 8, 2026
- [Muse Glimmer model page](https://developer.meta.com/ai/models/muse-glimmer/) - Meta's open-weight 30B model, August 10, 2026
- [Llama 4 Community License Agreement](https://developer.meta.com/ai/llama4/license/) - The full licence text
- [Llama 4 Maverick model card](https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct)
- [Llama 4 Scout model card](https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct)
- [Llama 4 product page](https://developer.meta.com/ai/models/llama-4/)
- [Llama Cookbook on GitHub](https://github.com/meta-llama/llama-cookbook)

## Frequently Asked Questions

### When was Llama 4 released?

April 5, 2025. The release comprised Llama 4 Scout and Llama 4 Maverick, both natively multimodal Mixture-of-Experts models with open weights.

### Is Llama 4 still Meta's flagship model?

No. Meta's current frontier model is Muse Spark, released April 8, 2026 by Meta Superintelligence Labs and updated to 1.3 on September 2, 2026. Muse Spark is proprietary, with no downloadable weights. Llama 4 is the last release in the Llama line — it is not Meta's last open-weight model, because Meta released the open-weight Muse Glimmer on August 10, 2026.

### Did Meta stop releasing open-weight models after Llama 4?

No. Meta stopped shipping open Llama models, and its frontier models are now proprietary — but on August 10, 2026 it released Muse Glimmer, a 30B dense multimodal model under an Apache 2.0 licence, small enough to run offline on a single 24 GB consumer GPU. The open-weight effort moved to the Muse line; it did not end. What ended is open weights at frontier scale.

### Is Llama 4 still supported and downloadable?

Yes. Scout and Maverick are still published by Meta at developer.meta.com (llama.com now redirects there) and on Hugging Face, under the unchanged Llama 4 Community License Agreement, and are still served by the major cloud and inference providers. What is gone is Meta's own hosted Llama API, shut down on July 6, 2026, and any roadmap: no successor has been announced, so expect no fixes and no refresh of the August 2024 knowledge cutoff.

### Is there a Llama 5 or Llama 4.5?

No. Meta has shipped no new Llama foundation model since April 5, 2025 — seventeen months as of September 2026. The line was extended with safety components at LlamaCon on April 29, 2025 (Llama Guard 4, Llama Prompt Guard 2, LlamaFirewall), but no new base model. Meta has published no statement discontinuing Llama and has announced no successor; Llama 3 and Llama 4 are both still listed for download. Its model work has moved to the Muse line — the proprietary Muse Spark and the open-weight Muse Glimmer.

### What is the context window of Llama 4 Scout?

10 million tokens. Meta calls the mechanism iRoPE: the 'i' is for interleaved attention layers, which carry no positional embeddings at all, while rotary position embeddings (RoPE) are used in most other layers. Inference-time temperature scaling of attention then extends length generalization.

### What is the difference between Scout and Maverick?

Both activate 17B parameters per token. Scout has 109B total parameters across 16 experts and a 10M token context window. Maverick has 400B total parameters across 128 experts and a 1M token context window, and is the stronger general-purpose model.

### Was Llama 4 Behemoth ever released?

No. Behemoth — 288B active parameters, roughly 2 trillion total, 16 experts — was previewed as still in training in April 2025 and never shipped as public weights.

### What license does Llama 4 use?

The Llama 4 Community License Agreement, a custom commercial license — not a standard open-source license like Apache 2.0 or MIT.

### Is Llama 4 multimodal?

Yes, natively. Meta used early fusion to integrate text and vision tokens in a unified backbone rather than bolting a vision encoder onto a text model. Input is text and images; output is text and code.

### How many languages does Llama 4 support?

Pre-training covered 200 languages, including over 100 with more than a billion tokens each. The instruction-tuned Maverick model card lists 12 officially supported languages: Arabic, English, French, German, Hindi, Indonesian, Italian, Portuguese, Spanish, Tagalog, Thai, and Vietnamese.

### Where can I download Llama 4?

From developer.meta.com/ai (llama.com now redirects there) and from Hugging Face, under the Llama 4 Community License Agreement. Both Scout and Maverick are available as base and instruction-tuned checkpoints.

## Related

### Related models

- [Muse Spark 1.3](https://howaiworks.ai/models/muse-spark)
- [Muse Glimmer](https://howaiworks.ai/models/muse-glimmer)
- [Gemma 4](https://howaiworks.ai/models/gemma)
- [Inkling](https://howaiworks.ai/models/inkling)
- [DeepSeek V4](https://howaiworks.ai/models/deepseek)
- [Qwen3.8-Max](https://howaiworks.ai/models/qwen)

---

Source: https://howaiworks.ai/models/llama — HowAIWorks.ai
