Tools & Products

ElevenLabs Launches Eleven V4 and V4 Turbo: 90 Languages, Sub-80ms Latency, and Emotion Control

ElevenLabs has launched Eleven V4 and its low-latency counterpart, Eleven V4 Turbo. The speech synthesis engine introduces native support for 90 languages, granular emotional tags, 48kHz studio audio fidelity, and sub-80ms streaming latency for conversational voice agents. Complete API guide, acoustic benchmark evaluation, and character token pricing.

By FreakVinci · 2026-09-28 · 17 min read

ElevenLabs launched Eleven V4 and Eleven V4 Turbo on September 28, 2026. The release expands language support to 90 languages, introduces explicit emotional and phonetic markup tags, and reduces streaming time-to-first-audio (TTFA) latency to 74ms for conversational agents.

The company split the generation pipeline into two dedicated architectures: Eleven V4, a 48kHz studio-grade engine for narration and dubbing, and Eleven V4 Turbo, a low-latency model designed for real-time bidirectional telephony and voice assistants.


1. Technical Evaluation: Acoustic Metrics and Latency

Audio models require simultaneous balancing of acoustic naturalness, pronunciation accuracy, and generation latency.

Evaluation Metric Eleven V4 Eleven V4 Turbo Eleven Multilingual v2 OpenAI Voice Engine 2 Cartesia Sonic
Languages Supported 90 90 29 32 15
Audio Sample Rate 48 kHz 24 kHz 44.1 kHz 24 kHz 44.1 kHz
Bit Depth 24-bit 16-bit 16-bit 16-bit 16-bit
Time-to-First-Audio (TTFA) 240ms 74ms 350ms 180ms 95ms
Subjective MOS (1–5 scale) 4.78 4.52 4.31 4.40 4.36
Word Error Rate (WER %) 1.8% 2.3% 3.6% 2.8% 3.1%
Price / 1,000 Characters $0.15 $0.08 $0.20 $0.15 $0.10
Streaming Time-to-First-Audio Latency (ms, lower is better):
Eleven V4 Turbo       [=====] 74ms
Cartesia Sonic        [======] 95ms
OpenAI Voice Engine 2 [============] 180ms
Eleven V4 (Studio)    [================] 240ms
Eleven Multilingual v2[=======================] 350ms

The 74ms latency in Eleven V4 Turbo enables human-parity conversational interruptions. When combined with sub-200ms speech-to-text and sub-400ms language model generation (such as Claude Sonnet 5.5), total round-trip conversational latency drops below 700ms, crossing the threshold required for natural spoken conversation.


2. Granular Expression Control: Emotion and Prosody Syntax

Previous speech models inferred tone solely from contextual semantic text. When an author wrote "Stop it", the model often failed to determine whether the character was playfully teasing or terrified.

Eleven V4 introduces native Expression Tags parsed directly by the neural acoustic decoder:

                               RAW TEXT WITH TAGS
         "[whisper] Lock the front door. [gasp] Did you hear that?"
                                       │
                                       ▼
                       PHONEME & PROSODY TOKENIZER
            Separates lexical phonemes from emotional vector guides
                                       │
                  ┌────────────────────┴────────────────────┐
                  ▼                                         ▼
         Acoustic Embeddings                        Expression Vector
       • Pitch Trajectory (Hz)                    • Glottal Tightness: 0.85
       • Formant Frequencies                      • Vocal Cord Aperiodicity: 0.60
       • Duration / Mora Alignment                • Intensity Shift: -14 dB
                  │                                         │
                  └────────────────────┬────────────────────┘
                                       ▼
                         DIFFUSION SPECTROGRAM GENERATOR
                                       │
                                       ▼
                             NEURAL VOCODER (48kHz)
                           Studio-Grade Audio Output

Supported emotional tokens include:

  • [whisper] and [soft]: Lowers vocal pressure and increases breath friction.
  • [gasp] and [breath]: Inserts natural respiratory pauses without synthetic clicking.
  • [laugh] and [chuckle]: Modulates pitch oscillation across vowel phonemes.
  • [excited] and [shout]: Elevates mean fundamental frequency (F0) and expands dynamic range.
  • [sigh] and [hesitant]: Introduces realistic micro-delays and descending intonation slopes.

3. Multilingual Generalization: 90 Languages Without Timbre Shift

A major defect in cross-lingual voice synthesis is "accent leakage"—an American English voice sounding distinctly foreign when synthesizing Mandarin, or a Spanish voice losing its natural pitch contour in German.

Eleven V4 implements an orthogonal speaker-language representation:

  1. Speaker Timbre Space: Captures the physiological vocal tract dimensions (formant baselines, nasal resonance, vocal fold mass).
  2. Language Phonotactic Space: Maps the phonetic rules, stress timing, and tonal inflections of all 90 target languages.

Because these representations are decoupled during training, a cloned speaker’s timbre remains identical whether generating French, Japanese, Hindi, Arabic, or Portuguese.


4. API Implementation: WebSocket Streaming and REST

ElevenLabs provides both REST endpoints for batch rendering and bidirectional WebSockets for low-latency voice bots.

WebSocket Streaming Audio with V4 Turbo

import WebSocket from 'ws';

const voiceId = '21m00Tcm4TlvDq8ikWAM'; // Rachel
const modelId = 'eleven_v4_turbo';
const uri = `wss://api.elevenlabs.io/v1/text-to-speech/${voiceId}/stream-input?model_id=${modelId}`;

const ws = new WebSocket(uri, {
  headers: { 'xi-api-key': process.env.ELEVENLABS_API_KEY! }
});

ws.on('open', () => {
  // 1. Send initial session configuration
  const sessionConfig = {
    text: ' ',
    voice_settings: {
      stability: 0.65,
      similarity_boost: 0.85,
      use_speaker_boost: true
    },
    generation_config: {
      chunk_length_schedule: [80, 120, 160, 200]
    }
  };
  ws.send(JSON.stringify(sessionConfig));

  // 2. Stream chunked text with inline emotional tags
  const message = {
    text: '[whisper] The system diagnostics are complete. [excited] All 48 clusters are operational!',
    try_trigger_generation: true
  };
  ws.send(JSON.stringify(message));

  // 3. Close the generation stream
  ws.send(JSON.stringify({ text: '' }));
});

ws.on('message', (data: WebSocket.Data) => {
  const response = JSON.parse(data.toString());
  if (response.audio) {
    const audioBuffer = Buffer.from(response.audio, 'base64');
    // Forward audioBuffer directly to telephony/browser WebAudio stream (74ms TTFA)
    playPcmAudio(audioBuffer);
  }
});

5. Pricing and Cost Economics

ElevenLabs lowered character fees for the V4 release while adding fine-grained token tracking:

Model Tier Cost per 1,000 Chars Approx. Cost / Min Audio Output Spec Intended Production Workload
Eleven V4 (Studio) $0.15 $0.030 48kHz, 24-bit PCM/MP3 Audiobooks, films, podcasts, video games
Eleven V4 Turbo $0.08 $0.016 24kHz, 16-bit PCM Conversational agents, call centers, NPCs
Instant Voice Cloning Included free N/A From 3-sec sample Dynamic personal agent voices
Professional Voice Clone $100 one-time setup N/A High-entropy dataset Verified enterprise executive avatars

For an enterprise call center handling 100,000 calls per month with an average agent speech duration of 3 minutes (approx. 3,000 characters per call):

  • Eleven V4 Turbo Cost: 100,000 calls × 3,000 characters = 300,000,000 characters.
  • At $0.08 per 1,000 characters, total monthly synthesis cost is $24,000.00, representing a 60% savings compared to legacy Multilingual v2 ($60,000.00).

6. Migration Guide for Production Voice Applications

Teams updating from Eleven Multilingual v2 or Flash v2.5 should execute the following checklist:

  1. Update model_id Parameter:
    • For real-time agents: change to eleven_v4_turbo.
    • For content generation: change to eleven_v4.
  2. Remove Manual Audio Normalization Filters:
    • Eleven V4 native 48kHz output includes studio-grade de-essing and dynamic range compression. Remove downstream high-pass EQ filters that were previously used to eliminate sub-bass rumble.
  3. Incorporate Expression Tags in Agent System Prompts:
    • Instruct upstream language models (such as Claude Sonnet 5.5 or GPT-6) to output [whisper], [pause], or [laugh] tags when emotional inflection is required.
  4. Tune WebSocket Chunking:
    • Configure chunk_length_schedule: [80, 120, 160, 200] to trigger generation on the first sentence fragment, ensuring sub-80ms playback.