Amazon Nova 2.5 Sonic Reaches General Availability on Bedrock: Real-Time Voice Agents with Sub-250ms Latency
Amazon Web Services declared general availability for Amazon Nova 2.5 Sonic on Amazon Bedrock. The multimodal voice-to-voice model generates expressive synthetic speech and handles interruptions with sub-250ms latency, competing directly with OpenAI Realtime API and Gemini Live.
Amazon Web Services launched Amazon Nova 2.5 Sonic into General Availability on Amazon Bedrock on October 6, 2026. The release gives cloud architects a fully managed, direct speech-to-speech foundation model designed for low-latency voice customer service, in-vehicle digital assistants, and telephonic agent automation.
Traditional voice agents stitch together three disjointed components: automatic speech recognition (ASR), a large language model (LLM), and a text-to-speech engine (TTS). This cascading pipeline accumulates 800ms to 1,500ms of lag. Nova 2.5 Sonic processes speech in a single unified neural pass.
Architectural Design: Direct Speech-to-Speech Streaming
Nova 2.5 Sonic bypasses text serialization. Incoming PCM audio frames are encoded into continuous acoustic tokens that feed directly into the model's transformer layers:
┌────────────────────────────────────────────────────────────────────────┐
│ Amazon Nova 2.5 Sonic Audio Pipeline │
├────────────────────────────────────────────────────────────────────────┤
│ Client Mic ──► WebSocket Audio Stream (16kHz / 24kHz PCM) │
│ │ │
│ ▼ │
│ Acoustic Encoder (Converts audio to 50Hz audio embeddings) │
│ │ │
│ ▼ │
│ Unified Nova 2.5 Transformer Core │
│ • Understands vocal inflections, pauses, and tone │
│ • Triggers tool calls while maintaining audio rhythm │
│ │ │
│ ▼ │
│ Neural Vocoder ──► Low-latency audio stream out (215ms total) │
└────────────────────────────────────────────────────────────────────────┘
Because the model retains continuous auditory awareness, user interruptions register instantly. If a customer speaks while the model is responding, Nova 2.5 Sonic stops playback within 60 milliseconds.
Latency and Economic Comparison
Amazon priced Nova 2.5 Sonic aggressively against competitive conversational voice APIs:
| Platform & Model | Architecture Type | Average Latency | Audio Cost (Per Minute) | Native Interruption |
|---|---|---|---|---|
| Amazon Nova 2.5 Sonic | Direct Speech-to-Speech | 215 ms | $0.0034 | Hardware Native |
| OpenAI Realtime API | Multimodal Audio Token | 290 ms | $0.0600 | Software Flag |
| Google Gemini Live (Vertex) | Multimodal Audio In/Out | 275 ms | $0.0450 | Client VAD |
| Traditional Cascade (AWS) | Transcribe + Claude + Polly | 940 ms | $0.0125 | Unreliable |
At $0.0034 per audio minute, high-volume call centers and mobile apps can operate continuous voice agents at a fraction of OpenAI's Realtime API price point.
Implementation Example with AWS SDK
Developers connect to Nova 2.5 Sonic using the AWS Bedrock runtime streaming API:
import { BedrockRuntimeClient, InvokeModelWithBidirectionalStreamCommand } from "@aws-sdk/client-bedrock-runtime";
const client = new BedrockRuntimeClient({ region: "us-east-1" });
async function startVoiceSession(audioInputStream: AsyncIterable<Uint8Array>) {
const command = new InvokeModelWithBidirectionalStreamCommand({
modelId: "amazon.nova-2-5-sonic:v1",
contentType: "application/vnd.amazon.bedrock.audio-stream",
accept: "application/vnd.amazon.bedrock.audio-stream"
});
const response = await client.send(command);
for await (const chunk of response.body) {
if (chunk.audioPayload) {
playAudioBuffer(chunk.audioPayload);
}
}
}
The SDK abstracts token serialization, providing standard audio chunk streams that hook directly into browser Web Audio APIs or telephony SIP trunks.