Introduction
Inworld AI released Realtime TTS-2 Flash and published pricing for its Realtime TTS-2 voice models on 2 September 2026, completing a rollout that began four months earlier. The sequence: a research preview of Realtime TTS-2 on 5 May, general availability of that model on 31 August, and then the September announcement, which added the Flash variant and put per-character rates on both.
The result is a two-model family. Realtime TTS-2 is the quality model; Realtime TTS-2 Flash trades some of that quality for speed and unit cost. Both are text-to-speech models built for live conversation rather than offline narration, and both are available through Inworld's API and its Realtime API, including a native integration with LiveKit Agents.
What the Model Does
The headline feature is that TTS-2 conditions on the audio of previous conversation turns, not only on the text it is asked to read. Inworld calls this a closed loop: the model builds a profile of the person it is talking to — pacing, tone, emotional state — and adjusts its own delivery as the conversation goes on. For a conversational AI system this is the difference between a narrator and a participant.
Delivery is also steerable from the text itself. Inline natural-language tags such as [whispering] or [speak tired but warm, like she just got home] change how a line is read, and non-verbal markers — [laugh], [sigh], [breathe], [cough], [clear_throat] — are placed inline. Three stability modes decide how far the model may improvise: Expressive, Balanced (the default) and Stable, the last for deployments where consistency matters more than performance.
Voice cloning works from a 5–15 second reference sample, and the cloned identity carries across languages, including a switch made mid-utterance. Inworld's documentation claims 200+ languages and locales, with 16 Tier 1 languages that get the most natural pronunciation — Russian among them — and a long tail of 180+ Tier 2 languages whose quality it describes as variable.
On latency, the firmest number is third-party: evaluation company Coval measured 25 ms time-to-first-byte for TTS-2 Flash. Inworld's own figures for the larger model do not agree with each other — its launch post cites a sub-200 ms median time-to-first-audio, the press release under 100 ms at the 99th percentile — and neither has been independently checked.
Where It Ranks
On the Artificial Analysis provider-voice Speech Arena — blind pairwise listening tests scored with Elo — the order at the time of publication is Cartesia Sonic 3.6 first at 1282, Inworld Realtime TTS-2 second at 1252, Alibaba's Qwen-Audio-3.0-TTS-Plus third at 1241, TTS-2 Flash sixth at 1222 and ElevenLabs' v3 Conversational eighth at 1210. Inworld's "#1 on Artificial Analysis" marketing refers to the arena view filtered to realtime models, where Sonic 3.6 is not in the comparison; second overall and first among streaming voices is the accurate reading.
Worth knowing what that benchmark does not cover. Arena Elo aggregates listener preference on short clips into a single number — it says nothing about how a voice holds up over a ten-minute call, how it handles interruptions, whether pronunciation survives outside the 16 Tier 1 languages, or how stable a cloned voice stays across turns. Those are the properties that decide whether a voice agent works in production, and no public leaderboard measures them yet.
Pricing
Realtime TTS-2 is metered per character, roughly 1,000 characters per minute of audio. On demand it is $25 per million characters, dropping through the paid plans to $20, $17.50, $15 and $12.50 on Growth; Inworld's pricing page ties the $15 rate to a $300/month plan and $12.50 to $1,500/month, with enterprise rates quoted as low as $5. TTS-2 Flash runs from $15 down to $7. The monthly dollar figures are plan commitments, not the price of the model — a distinction worth making, since at roughly 1,000 characters per minute the $12.50 tier works out to about 1.25 cents per minute of speech.
Conclusion
Second place on the arena puts Inworld behind Cartesia, which shipped Sonic 3.6 in August, and alongside ElevenLabs as an option teams will actually evaluate. What separates this release is control rather than raw audio quality: audio-conditioned delivery, text-directed style, stability modes and cross-lingual voice identity are the features that matter when the output is a live AI agent instead of a rendered file. Whether the adaptation to a speaker's tone holds up over long conversations is the open question, and it is exactly the thing current benchmarks do not test.
Sources
- Realtime TTS-2: A new frontier voice model that feels as human as it sounds — Inworld AI, 31 August 2026
- Inworld Launches Realtime TTS-2 Voice Model Family for Controllable, Realtime Speech — press release, 2 September 2026
- Realtime TTS-2 multilingual capabilities — Inworld documentation
- Text to Speech Leaderboard: Provider Voice — Artificial Analysis
- Inworld pricing — Inworld AI