rkj dev

ElevenLabs Launches Eleven v4 Audio Model with In-Text Voice Direction

The new architecture introduces inline natural-language voice directing, low-latency streaming for AI agents, and multilingual native accent adoption.

Illustration of a professional sound design studio with digital audio waveform visualizations
Illustration: Modern digital audio engineering environment optimized for voice synthesis and directional audio control.AI-generated illustration

Key takeaways

  • ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28, 2026, built on an entirely new architecture for expressive text-to-speech.
  • The model introduces inline natural-language audio tags, letting creators steer emotion, pacing, and sound effects directly in the script.
  • ElevenLabs reports that Eleven v4 Turbo achieves a median time to first speech of approximately 150ms based on internal testing against identical scripts for conversational agent workflows.
  • Support extends across more than 90 languages, automatically matching native accents while preserving speaker identity from source voices.

ElevenLabs announced the launch of Eleven v4 and Eleven v4 Turbo on September 28, 2026, introducing an overhauled architecture designed to interpret tone, pacing, and context from written text. According to ElevenLabs' official announcement, the new generation delivers natural multi-speaker dynamics and in-text directional control, while reducing voice generation latency for real-time applications.

Alongside the flagship model, ElevenLabs launched Eleven v4 Turbo to serve low-latency interactive workflows, such as customer service tools and conversational characters. Both model tiers are available immediately through the company's web platforms and API, including access on the free tier.

Natural Language Voice Direction and Audio Tags

The central architectural shift in Eleven v4 is the ability to interpret performance directions written directly into script text. Rather than relying solely on sliders for speed or style, users can embed natural language instructions and bracketed tags to control emotional delivery, pacing, and non-speech sounds.

According to ElevenLabs, the model processes descriptive inline tags such as [laughs] and [said angrily in French accent], alongside sound effects like [phone buzzing] or [light rain]. Developer documentation also indicates improved support for the International Phonetic Alphabet (IPA) to handle custom pronunciations.

Illustration of script directing with visual indicators of emotion and pacing
Illustration: Natural-language script direction allows creators to guide emotional tone and timing directly within written text.AI-generated illustration

AI audio generation and workflow platform Morphic noted that previous manual style and speed sliders have been replaced by this context-driven approach. Instead of using SSML break tags, creators shape rhythm and timing by adjusting punctuation, capitalization, or inserting natural cues such as [rushed] or [slowly].

Analysis from MindStudio highlighted that natural language direction lowers friction for creators building narrative audiobooks, games, and AI video sound design, comparing the feature trajectory to capabilities previously demonstrated by ByteDance's Seed Audio.

Speed, Benchmarks, and the Turbo Architecture

To address interactive voice applications, ElevenLabs deployed Eleven v4 Turbo alongside the primary model. High-fidelity voice models historically traded response speed for expressive performance, but ElevenLabs reports that Eleven v4 Turbo reaches a median inference latency of approximately 100 milliseconds and a median time to first speech of roughly 150 milliseconds over WebSocket streaming.

In benchmark testing conducted in September 2026, ElevenLabs stated that Eleven v4 ranked number one on the Artificial Analysis Provider Voice Arena Leaderboard. In blind head-to-head preference evaluations against competing models—including Cartesia Sonic 3.6, Inworld TTS-2, Google Gemini 3.8 Flash TTS, and Google Gemini 3.8 Flash-Lite TTS—graders preferred Eleven v4 approximately 75% of the time for expressiveness and natural tone.

Illustration representing low-latency streaming audio data transferring through server infrastructure
Illustration: Low-latency audio streaming enables real-time conversational responses in AI voice agent platforms.AI-generated illustration

For real-time evaluation, ElevenLabs benchmarked time-to-first-speech against systems including OpenAI GPT-4o mini TTS and xAI TTS, measuring performance with network latency removed.

Multilingual Native Accents and Voice Cloning

Eleven v4 expands language coverage to more than 90 languages, an increase from over 70 supported in Eleven v3. A key change in the model's multilingual behavior is native accent adoption: when an existing voice clone speaks a target language, it automatically adopts the accent of a native speaker in that language while preserving the original vocal identity.

According to ElevenLabs' feature documentation, voice cloning fidelity has also been updated:

  • Instant Voice Clones: Can capture speaker characteristics from 10 seconds of source audio.
  • Professional Voice Clones (PVC): Supported for high-fidelity studio productions.
  • Request Stitching: Improved consistency across long-form multi-line generations in ElevenLabs Studio and the Reader App.
  • Accent Adherence: Voice clones maintain consistent pronunciation across generations without drifting back to their source language accent.

Interactive demonstrations published on the ElevenLabs Emotional Range showcase illustrate the model generating more than 35 discrete emotional states, including variations for anxious, excited, and vulnerable states across repeated phrases.

Illustration of a connected globe representing multilingual speech synthesis
Illustration: Multilingual speech synthesis enables cross-language communication with native regional accents.AI-generated illustration

Availability and Platform Access

Eleven v4 and Eleven v4 Turbo are available immediately through ElevenLabs within ElevenAgents, ElevenCreative, and via the developer API. Free account holders can test the models directly without entering a paid tier, though standard platform usage limits apply.

Third-party workflow platforms have also integrated the models. As detailed by Morphic, Eleven v4 serves as the default engine for text generation up to 3,000 characters per request and supports multi-speaker dialogue scenes with two to ten distinct voices.

Frequently asked questions

What is new in ElevenLabs Eleven v4?

Eleven v4 introduces an updated architecture with natural-language voice direction, inline sound effect tags, native accent adaptation across 90+ languages, and improved speaker identity retention.

What is the latency of Eleven v4 Turbo?

According to ElevenLabs, Eleven v4 Turbo achieves a median inference latency of approximately 100ms and a median time to first speech of around 150ms over WebSocket streaming.

Is Eleven v4 free to use?

Yes, Eleven v4 and Eleven v4 Turbo are available on ElevenLabs' free tier as well as on third-party platforms like Morphic's free tier, subject to standard credit limits.

Sources

  1. Introducing Eleven v4, our most emotive modelElevenLabs · Sep 28, 2026 · Official
  2. Eleven v4: Meet the fastest and most emotive model we've ever built | ElevenLabselevenlabs.io · Official
  3. Eleven v4: ElevenLabs' most expressive AI voice modelMorphic
  4. ElevenLabs v4: Natural Language Voice Direction Now on the Free TierMindStudio · Oct 2, 2026

How this story was made: the newsroom picked it up from elevenlabs.io, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (33 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.

#ElevenLabs #Voice AI #Text to Speech #Artificial Intelligence #Audio Generation

Published October 5, 2026 at 01:05 UTC