Composer 2.5

Cursor's in-house agentic coding model, released May 18, 2026. Built by post-training Moonshot's Kimi K2.5, delivering frontier-adjacent coding at low cost.

Updated

Released
May 18, 2026
Type
Code Model
Pricing
$0.50 / $2.50 per Mtok
License
Proprietary
On this page

Overview

Composer 2.5 is Cursor's in-house agentic coding model, released on May 18, 2026. It succeeded Composer 2 (March 19, 2026) and, as of September 6, 2026, remains the only shipping model in the Composer line.

The successor Cursor announced has still not appeared. At its Compile conference in June 2026, Cursor announced "a new model — our first model trained from scratch," giving it no name, no version number, and no date. Nearly three months later it appears in neither Cursor's model list, its pricing table, nor its changelog. Treat the name "Composer 3" and the parameter counts and training-cluster details quoted alongside it as press invention until Cursor says otherwise.

What arrived instead was a change of owner. SpaceX completed its acquisition of Cursor on August 14, 2026 — "This completes the acquisition process that started in April, when we announced our partnership with SpaceXAI to accelerate our model training efforts" — and two days before that, Cursor shipped Grok 4.6 "together with SpaceXAI," pointing to it as "an early look at what we can now build together." Grok 4.6 now sits in Cursor's own model list beside Composer 2.5, described in Cursor's documentation as jointly trained by Cursor and SpaceXAI, at $2.00 / $6.00 per million tokens. Cursor's newest model therefore shipped under the Grok name, not the Composer one, and Composer 2.5 has been carrying the Composer line alone since May.

The single most important thing to understand about Composer, and the thing most coverage gets wrong, is stated plainly by Cursor itself:

"Composer 2.5 is built on the same open-source checkpoint as Composer 2, Moonshot's Kimi K2.5."

Composer is not a from-scratch Cursor model. It is a post-trained derivative of an open-source base — Moonshot AI's Kimi K2.5 — refined through continued training, reinforcement learning on long-horizon coding tasks, and synthetic task generation. Cursor's value-add is the post-training and the harness, not the pretraining.

That framing explains the model's economics. Cursor describes Composer 2 as "a new, optimal combination of intelligence and cost," and when Artificial Analysis first measured Composer 2.5 in May 2026 it landed third on the Coding Agent Index at roughly 1/60th the per-task cost of the two models above it. That ranking has not been refreshed since; the price has not moved either.

Cursor characterises Composer 2.5 as "a substantial improvement in intelligence and behavior over Composer 2" — "better at sustained work on long-running tasks, follows complex instructions more reliably, and is more pleasant to collaborate with."

Capabilities

  • Long-horizon agentic execution — Cursor's documentation says the model "excels at long-horizon tasks via reinforcement learning on long-horizon coding tasks." Composer 2 was described as "able to solve challenging tasks requiring hundreds of actions."
  • Tool use, file edits, and terminal operations — the model is "tuned for tool use, file edits, and terminal operations inside Cursor." It has access to all agent tools when used in Cursor.
  • Sustained work — the specific improvement Cursor claims for 2.5 over 2: holding coherence across long-running tasks.
  • Complex instruction following — Cursor cites more reliable adherence to multi-part instructions.
  • Frontier-level coding — Cursor's own characterisation of Composer 2 was "frontier-level at coding."
  • Low-latency variant — Composer 2.5 Fast is the default for interactive sessions, trading price for speed at the same intelligence tier.

Technical Specifications

  • Base checkpoint: Moonshot AI's Kimi K2.5, open-source. Same base as Composer 2.
  • Post-training: continued training and fine-tuning via reinforcement learning and synthetic task generation, applied by Cursor.
  • Variants: Composer 2.5 (standard) and Composer 2.5 Fast (default for interactive sessions).
  • Harness: full access to Cursor's agent tools — file editing, search, terminal.
  • Weights: closed. Cursor's post-trained model is proprietary, even though its base is open.
  • Availability: inside Cursor only. No standalone API.
  • Context window: Not published by Cursor. Neither the launch post nor the model documentation states one.
  • Maximum output tokens: Not published by Cursor.
  • Architecture: Not disclosed beyond the base checkpoint identification. Cursor's Composer 2 announcement disclosed no architecture details.

The base-model disclosure matters for reasons beyond credit. Kimi K2.5's architecture, tokenizer, and pretraining data are Composer 2.5's architecture, tokenizer, and pretraining data. If you want to understand what Composer knows, read about the Kimi line.

One caveat on that lineage: the base checkpoint is no longer callable at Moonshot. Moonshot retired the kimi-k2.5 and moonshot-v1 series from its own API on August 31, 2026, and requests for them now return a 404; its current line is Kimi K3. This does not affect Composer 2.5 — Cursor serves its own post-trained weights and never depended on Moonshot's endpoint — but you can no longer call K2.5 at Moonshot to benchmark base against derivative, and Composer's foundation is now a model its author has moved on from.

Use Cases

  • Long-running autonomous coding tasks — the workload Cursor explicitly reinforcement-trained for, spanning hundreds of sequential actions.
  • In-editor agentic development — the model's native context, with full tool access inside Cursor.
  • Terminal-driven work — Composer is tuned for terminal operations, and Terminal-Bench is one of the three benchmarks Cursor reports.
  • Multilingual codebases — SWE-bench Multilingual is Cursor's third headline benchmark, at 73.7 for Composer 2.
  • Cost-sensitive agent workloads — at $0.07 per task on Artificial Analysis's harness, Composer 2.5 makes high-volume agentic coding economically different from frontier alternatives.
  • Interactive pair programming — the Fast variant exists for this, and is the default for interactive sessions.
  • Mobile agent supervision — via Cursor's iOS app, in public beta since June 2026, which lets you "launch and manage always-on agents from anywhere."

Performance / Benchmarks

Cursor's own benchmarks (primary source)

From Cursor's Composer 2 announcement, comparing Composer 2 against Composer 1:

ModelCursorBenchTerminal-Bench 2.0SWE-bench Multilingual
Composer 261.361.773.7
Composer 138.040.056.9

Cursor's Composer 2.5 launch post presents its benchmark tables and effort curves as charts rather than in text; the figures above are the ones Cursor states numerically.

Artificial Analysis (third-party evaluator)

Artificial Analysis independently evaluated Composer 2.5 on its Coding Agent Index and published the results on May 20, 2026. All figures below are Artificial Analysis's, not Cursor's, and all of them are that May 2026 snapshot — see the note underneath for why they no longer describe the current field.

Model + harnessCoding Agent IndexCost per task
Claude Opus 4.7 (max) in Claude Code66$4.10
GPT-5.5 (xhigh reasoning) in Codex65$4.82
Composer 2.562$0.07
Composer 248—
Composer 2.5 Fast—$0.44

Composer 2.5 ranks third, a 14-point gain over Composer 2's 48. Component movements from Composer 2 to Composer 2.5:

  • SWE-Bench-Pro-Hard-AA: 12% → 47% (+35 points). Artificial Analysis notes this score is "comparable to Claude Opus 4.7 (max)."
  • Terminal-Bench v2: 64% → 66% (+2 points)
  • SWE-Atlas-QnA: 69% → 72% (+3 points)

Artificial Analysis also measured Composer 2.5 Fast at an average wall time of 6.7 minutes per task.

Why those figures are a May 2026 snapshot

Two things have moved since, and both matter for how you read the table above.

The index was recomposed. The Coding Agent Index is now at v1.4, an equal-weight average of DeepSWE, Terminal-Bench v2.1 and SWE-Atlas-QnA. The SWE-Bench-Pro-Hard-AA component that produced Composer 2.5's 35-point jump is no longer in it, so 62 is not comparable to a score published under the current version.

The field was replaced. In Artificial Analysis's September 3, 2026 write-up of GPT-6 Astra, the Coding Agent Index is led by Claude Fable 5.1 in Claude Code at 70, with GPT-6 Astra in Codex at 67, roughly level with Claude Opus 5. Claude Opus 4.7 and GPT-5.5 — the two models Composer 2.5 was measured against in May — are no longer the top of the board. Artificial Analysis has not published a Composer 2.5 result at the current index version, so its standing against today's frontier is unmeasured. Third place is a May 2026 fact and should not be quoted as a present-tense one.

What has not changed is the cost column: at $0.07 per task, Composer 2.5 ran at roughly 1.7% the per-task cost of the leading coding agent, and Cursor has not raised its token prices since.

Limitations

  • Cursor-only. No standalone API, no weights, no third-party hosting. Using Composer means using Cursor.
  • Not a from-scratch model. Composer inherits the capabilities, knowledge, and blind spots of Moonshot's Kimi K2.5. Cursor's post-training reshapes behaviour, not pretraining knowledge.
  • No published context window. Neither the launch post nor the model docs state one, so long-context planning is guesswork.
  • No published architecture or parameter count. Cursor discloses the base checkpoint and nothing further.
  • Specialised for coding. Composer is tuned for tool use, file edits, and terminal operations inside Cursor. It is not a general-purpose assistant.
  • Its independent ranking is stale. Artificial Analysis put Composer 2.5 third on the Coding Agent Index in May 2026, behind Claude Opus 4.7 and GPT-5.5. The index has since been recomposed at v1.4 and its leaders replaced, and Composer 2.5 has not been re-published against the current field — so where it now ranks is simply unmeasured. It was never the capability leader; whether it is still close is unknown.
  • No successor timeline. Cursor announced a from-scratch model in June 2026 and has said nothing further about it in the three months since, while its newest model shipped under the Grok name instead. Composer's roadmap is not something you can plan against.
  • The base checkpoint is end-of-life upstream. Moonshot retired kimi-k2.5 from its own API on August 31, 2026. Composer 2.5 keeps running on Cursor's own weights, but its pretraining lineage now traces to a model its author no longer hosts.
  • The Fast variant carries a real premium. At $3.00/$15.00 it is 6x the standard price, and it is the default for interactive sessions — worth knowing before your bill arrives.
  • Knowledge cutoff not published. Recency of framework and library knowledge is unstated.

Pricing & Access

API pricing (per million tokens, from Cursor)

VariantInputOutput
Composer 2.5$0.50$2.50
Composer 2.5 Fast$3.00$15.00

Cursor notes the Fast tier is "a lower cost than the fast tiers of other frontier models." Fast is the default for interactive sessions in Cursor.

Plan access

  • Individual plans — Composer usage draws from the standalone usage pool.
  • Team and enterprise plans — billed at direct API pricing.
  • Cursor offered double usage for the first week after the Composer 2.5 launch.

Surfaces

  • Cursor editor — the primary interface.
  • Cursor for iOS — public beta on all paid plans since June 29, 2026. Launch and manage always-on agents remotely, with live activity tracking and artifact review.
  • Cursor's new interface — Composer has been available in its early alpha since the Composer 2 release.

There is no way to access Composer 2.5 outside Cursor.

Ecosystem & Tools

  • Cursor — the editor, and the only place Composer runs.
  • Kimi — Moonshot AI's open-source model line. Kimi K2.5 is Composer 2.5's base checkpoint, retired from Moonshot's own API on August 31, 2026; understanding the K2 generation is understanding Composer's foundation.
  • Grok 4.6 — released by Cursor with SpaceXAI on August 12, 2026 and listed among Cursor's own models. The other model you can pick inside Cursor that Cursor had a hand in training.
  • Cursor docs: Composer 2.5 — official model documentation.
  • Cursor agent tools — file editing, semantic search, terminal execution. Composer has full access to all of them in Cursor.
  • Cursor for iOS — mobile supervision of always-on agents, public beta since June 2026.
  • Artificial Analysis Coding Agent Index — the leading independent evaluation of Composer 2.5.

Community & Resources

Frequently Asked Questions

May 18, 2026. It succeeded Composer 2, which shipped on March 19, 2026. Release dates in April 2026 that appear in some coverage are incorrect.
No. At its Compile conference in June 2026, Cursor announced "a new model — our first model trained from scratch," giving no name, version number, or release date. Nearly three months later, on September 6, 2026, it still does not appear in Cursor's model list, pricing table, or changelog. What did ship in the interval is Grok 4.6, released on August 12, 2026 "together with SpaceXAI" and now listed among Cursor's own models. The name "Composer 3" and the parameter counts and training-cluster details circulating alongside it come from press coverage, not from Cursor.
No, and Cursor says so directly: "Composer 2.5 is built on the same open-source checkpoint as Composer 2, Moonshot's Kimi K2.5." Cursor's contribution is continued training and fine-tuning on that checkpoint through reinforcement learning and synthetic task generation. The base model is open-source and comes from Moonshot AI.
Moonshot AI's Kimi K2.5 open-source checkpoint — the same base as Composer 2. Cursor applies continued training, reinforcement learning on long-horizon coding tasks, and synthetic task generation on top of it.
$0.50 per million input tokens and $2.50 per million output tokens for the standard variant. The Fast variant, which is the default for interactive sessions in Cursor, is $3.00 per million input tokens and $15.00 per million output tokens — which Cursor notes is "a lower cost than the fast tiers of other frontier models."
Third, as of May 20, 2026, and unmeasured since. Artificial Analysis scored Composer 2.5 at 62 on its Coding Agent Index that day — behind Claude Opus 4.7 (max) in Claude Code at 66 and GPT-5.5 (xhigh reasoning) in Codex at 65, at $0.07 per task against $4.10 and $4.82. That index has since been recomposed to v1.4 and its leaders replaced; Artificial Analysis has not published a Composer 2.5 result against the current field, so the third-place claim is a May 2026 fact, not a current one.
Artificial Analysis measured a 14-point gain on its Coding Agent Index (48 to 62). The largest single jump was SWE-Bench-Pro-Hard-AA, from 12% to 47% — a 35-point improvement that Artificial Analysis describes as "comparable to Claude Opus 4.7 (max)." Terminal-Bench v2 rose from 64% to 66%, and SWE-Atlas-QnA from 69% to 72%.
Cursor has not published one. Neither the launch post nor the model documentation states a context window or a maximum output length, so we do not report a figure.
No. Composer 2.5 is available only inside Cursor. There is no standalone API, no weights release, and no third-party hosting. It can be used through the Cursor editor, Cursor's iOS app (public beta since June 2026), and the early alpha of Cursor's new interface.
Speed and price. Fast is the default for interactive sessions in Cursor and costs $3.00/$15.00 per million tokens against the standard variant's $0.50/$2.50. Artificial Analysis measured Fast at an average wall time of 6.7 minutes per task, at $0.44 per task versus $0.07 for standard.
No. Moonshot retired the kimi-k2.5 and moonshot-v1 series from its own API on August 31, 2026, and calls to them now return a 404. Composer 2.5 is unaffected: Cursor serves its own post-trained weights rather than calling Moonshot. The practical consequence is that the base checkpoint is no longer available to call at Moonshot for comparison, and Composer's lineage now runs from a model its original author no longer hosts.
No. Although Composer 2.5 is post-trained from Moonshot's open-source Kimi K2.5 checkpoint, Cursor's resulting model is proprietary and is not released. The base checkpoint is open; Cursor's derivative is not.

Explore More Models

Discover other AI models and compare their capabilities.