Tools & Products

Amazon Nova 2.5 Sonic Reaches General Availability on Bedrock: Real-Time Voice Agents with Sub-250ms Latency

Amazon Web Services declared general availability for Amazon Nova 2.5 Sonic on Amazon Bedrock. The multimodal voice-to-voice model generates expressive synthetic speech and handles interruptions with sub-250ms latency, competing directly with OpenAI Realtime API and Gemini Live.

By Julian Thorne · 2026-10-06 · 10 min read

Amazon Web Services launched Amazon Nova 2.5 Sonic into General Availability on Amazon Bedrock on October 6, 2026. The release gives cloud architects a fully managed, direct speech-to-speech foundation model designed for low-latency voice customer service, in-vehicle digital assistants, and telephonic agent automation.

Traditional voice agents stitch together three disjointed components: automatic speech recognition (ASR), a large language model (LLM), and a text-to-speech engine (TTS). This cascading pipeline accumulates 800ms to 1,500ms of lag. Nova 2.5 Sonic processes speech in a single unified neural pass.


Architectural Design: Direct Speech-to-Speech Streaming

Nova 2.5 Sonic bypasses text serialization. Incoming PCM audio frames are encoded into continuous acoustic tokens that feed directly into the model's transformer layers:

┌────────────────────────────────────────────────────────────────────────┐
│                   Amazon Nova 2.5 Sonic Audio Pipeline                 │
├────────────────────────────────────────────────────────────────────────┤
│ Client Mic ──► WebSocket Audio Stream (16kHz / 24kHz PCM)              │
│                     │                                                  │
│                     ▼                                                  │
│          Acoustic Encoder (Converts audio to 50Hz audio embeddings)    │
│                     │                                                  │
│                     ▼                                                  │
│          Unified Nova 2.5 Transformer Core                             │
│          • Understands vocal inflections, pauses, and tone             │
│          • Triggers tool calls while maintaining audio rhythm          │
│                     │                                                  │
│                     ▼                                                  │
│          Neural Vocoder ──► Low-latency audio stream out (215ms total) │
└────────────────────────────────────────────────────────────────────────┘

Because the model retains continuous auditory awareness, user interruptions register instantly. If a customer speaks while the model is responding, Nova 2.5 Sonic stops playback within 60 milliseconds.


Latency and Economic Comparison

Amazon priced Nova 2.5 Sonic aggressively against competitive conversational voice APIs:

Platform & Model Architecture Type Average Latency Audio Cost (Per Minute) Native Interruption
Amazon Nova 2.5 Sonic Direct Speech-to-Speech 215 ms $0.0034 Hardware Native
OpenAI Realtime API Multimodal Audio Token 290 ms $0.0600 Software Flag
Google Gemini Live (Vertex) Multimodal Audio In/Out 275 ms $0.0450 Client VAD
Traditional Cascade (AWS) Transcribe + Claude + Polly 940 ms $0.0125 Unreliable

At $0.0034 per audio minute, high-volume call centers and mobile apps can operate continuous voice agents at a fraction of OpenAI's Realtime API price point.


Implementation Example with AWS SDK

Developers connect to Nova 2.5 Sonic using the AWS Bedrock runtime streaming API:

import { BedrockRuntimeClient, InvokeModelWithBidirectionalStreamCommand } from "@aws-sdk/client-bedrock-runtime";

const client = new BedrockRuntimeClient({ region: "us-east-1" });

async function startVoiceSession(audioInputStream: AsyncIterable<Uint8Array>) {
  const command = new InvokeModelWithBidirectionalStreamCommand({
    modelId: "amazon.nova-2-5-sonic:v1",
    contentType: "application/vnd.amazon.bedrock.audio-stream",
    accept: "application/vnd.amazon.bedrock.audio-stream"
  });

  const response = await client.send(command);
  for await (const chunk of response.body) {
    if (chunk.audioPayload) {
      playAudioBuffer(chunk.audioPayload);
    }
  }
}

The SDK abstracts token serialization, providing standard audio chunk streams that hook directly into browser Web Audio APIs or telephony SIP trunks.