Inworld Ships Realtime TTS-2 Voice Model Family

Inworld completes its Realtime TTS-2 rollout with a Flash variant and per-character pricing. Where the model actually ranks on the Artificial Analysis arena.

by HowAIWorks Team
On this page

Introduction

Inworld AI released Realtime TTS-2 Flash and published pricing for its Realtime TTS-2 voice models on 2 September 2026, completing a rollout that began four months earlier. The sequence: a research preview of Realtime TTS-2 on 5 May, general availability of that model on 31 August, and then the September announcement, which added the Flash variant and put per-character rates on both.

The result is a two-model family. Realtime TTS-2 is the quality model; Realtime TTS-2 Flash trades some of that quality for speed and unit cost. Both are text-to-speech models built for live conversation rather than offline narration, and both are available through Inworld's API and its Realtime API, including a native integration with LiveKit Agents.

What the Model Does

The headline feature is that TTS-2 conditions on the audio of previous conversation turns, not only on the text it is asked to read. Inworld calls this a closed loop: the model builds a profile of the person it is talking to — pacing, tone, emotional state — and adjusts its own delivery as the conversation goes on. For a conversational AI system this is the difference between a narrator and a participant.

Delivery is also steerable from the text itself. Inline natural-language tags such as [whispering] or [speak tired but warm, like she just got home] change how a line is read, and non-verbal markers — [laugh], [sigh], [breathe], [cough], [clear_throat] — are placed inline. Three stability modes decide how far the model may improvise: Expressive, Balanced (the default) and Stable, the last for deployments where consistency matters more than performance.

Voice cloning works from a 5–15 second reference sample, and the cloned identity carries across languages, including a switch made mid-utterance. Inworld's documentation claims 200+ languages and locales, with 16 Tier 1 languages that get the most natural pronunciation — Russian among them — and a long tail of 180+ Tier 2 languages whose quality it describes as variable.

On latency, the firmest number is third-party: evaluation company Coval measured 25 ms time-to-first-byte for TTS-2 Flash. Inworld's own figures for the larger model do not agree with each other — its launch post cites a sub-200 ms median time-to-first-audio, the press release under 100 ms at the 99th percentile — and neither has been independently checked.

Where It Ranks

On the Artificial Analysis provider-voice Speech Arena — blind pairwise listening tests scored with Elo — the order at the time of publication is Cartesia Sonic 3.6 first at 1282, Inworld Realtime TTS-2 second at 1252, Alibaba's Qwen-Audio-3.0-TTS-Plus third at 1241, TTS-2 Flash sixth at 1222 and ElevenLabs' v3 Conversational eighth at 1210. Inworld's "#1 on Artificial Analysis" marketing refers to the arena view filtered to realtime models, where Sonic 3.6 is not in the comparison; second overall and first among streaming voices is the accurate reading.

Worth knowing what that benchmark does not cover. Arena Elo aggregates listener preference on short clips into a single number — it says nothing about how a voice holds up over a ten-minute call, how it handles interruptions, whether pronunciation survives outside the 16 Tier 1 languages, or how stable a cloned voice stays across turns. Those are the properties that decide whether a voice agent works in production, and no public leaderboard measures them yet.

Pricing

Realtime TTS-2 is metered per character, roughly 1,000 characters per minute of audio. On demand it is $25 per million characters, dropping through the paid plans to $20, $17.50, $15 and $12.50 on Growth; Inworld's pricing page ties the $15 rate to a $300/month plan and $12.50 to $1,500/month, with enterprise rates quoted as low as $5. TTS-2 Flash runs from $15 down to $7. The monthly dollar figures are plan commitments, not the price of the model — a distinction worth making, since at roughly 1,000 characters per minute the $12.50 tier works out to about 1.25 cents per minute of speech.

Conclusion

Second place on the arena puts Inworld behind Cartesia, which shipped Sonic 3.6 in August, and alongside ElevenLabs as an option teams will actually evaluate. What separates this release is control rather than raw audio quality: audio-conditioned delivery, text-directed style, stability modes and cross-lingual voice identity are the features that matter when the output is a live AI agent instead of a rendered file. Whether the adaptation to a speaker's tone holds up over long conversations is the open question, and it is exactly the thing current benchmarks do not test.

Sources

Frequently Asked Questions

It is Inworld's text-to-speech model family. Realtime TTS-2 is the quality model, generally available since 31 August 2026; Realtime TTS-2 Flash, released on 2 September, trades quality for speed and lower cost. Both stream speech in real time and clone voices across languages.
No. On the Artificial Analysis provider-voice leaderboard Cartesia's Sonic 3.6 holds first place at 1282 Elo, with Realtime TTS-2 second at 1252 and TTS-2 Flash sixth at 1222. Inworld's own '#1' claim refers to the arena view filtered to realtime models.
Yes. Inworld's documentation lists Russian among the 16 Tier 1 languages that receive the highest and most consistent quality, out of 200+ supported languages and locales.
Realtime TTS-2 starts at $25 per million characters on demand and falls to $12.50 on the Growth plan; TTS-2 Flash starts at $15 and falls to $7. The lower rates come with monthly plan commitments, not with the character price alone.
Delivery is steered with inline natural-language tags in the text, such as [whispering] or [speak tired but warm], plus non-verbal markers like [laugh], [sigh] and [breathe]. Three stability modes — Expressive, Balanced and Stable — set how much variation the model is allowed.

Continue Your AI Journey

Explore our glossary and model catalog to deepen your understanding.