How to Build Multi-Speaker Expressive Voice Applications with Gemini 3.8 Flash TTS API
Step-by-step developer tutorial for integrating Google Gemini 3.8 Flash TTS and Flash-Lite TTS APIs with prompt-steered vocal design, multi-character dialogue scripts, and 24 kHz PCM streaming.
Step 1: Install the Google GenAI SDK and Audio Utilities
Add the official @google/genai TypeScript SDK to your project. Ensure your Node.js runtime has buffer and filesystem capabilities.
npm install @google/genai # For Python environments: pip install google-genai soundfile
Step 2: Configure API Key and Client Initialization
Initialize the GoogleGenAI client with your API key from Google AI Studio. Set the environment variable GEMINI_API_KEY in your deployment environment.
import { GoogleGenAI } from '@google/genai';
// Automatically reads process.env.GEMINI_API_KEY
const ai = new GoogleGenAI({});
Step 3: Define Voice Design Prompt and Single-Speaker Synthesis
Configure responseModalities to AUDIO and specify voiceDesign prompts in speechConfig to control pacing, emotional tension, whispering, or regional accents without providing reference audio.
export async function generateNarratorAudio(scriptText: string) {
const response = await ai.models.generateContent({
model: 'gemini-3.8-flash-tts',
contents: scriptText,
config: {
responseModalities: ['AUDIO'],
speechConfig: {
voiceConfig: {
prebuiltVoiceConfig: {
voiceName: 'Puck' // One of 2,048 available voice presets
}
},
voiceDesign: {
prompt: 'Thoughtful investigative journalist, mid-30s, speaking in calm, measured cadence with distinct pauses between evidentiary statements.'
},
audioFormat: 'audio/pcm;rate=24000'
}
}
});
const part = response.candidates?.[0]?.content?.parts?.find(p => p.inlineData?.mimeType?.startsWith('audio/'));
if (!part?.inlineData?.data) {
throw new Error('No audio data returned by Gemini TTS engine');
}
return Buffer.from(part.inlineData.data, 'base64');
}
Step 4: Configure Multi-Speaker Turn-Taking Scripts
Format character dialogues with structured speaker tags. Gemini 3.8 Flash TTS synthesizes natural conversational turn-taking, pauses, and overlapping speech in a single generative pass.
export async function generateDialoguePodcast(dialogueScript: string) {
// Format script with explicit character names and styles
const prompt = `
[Speaker: Maya | Voice: Aoede | Style: Animated, fast-paced tech analyst]
Have you seen the latency numbers on the new flash models? We are looking at under ninety milliseconds!
[Speaker: David | Voice: Fenrir | Style: Skeptical, grounded engineer]
Sub-ninety is solid for raw generation, but what about the network roundtrip across cross-region clusters?
[Speaker: Maya | Voice: Aoede | Style: Chuckling, dismissive of doubts]
They ran test nodes in Tokyo and Frankfurt. The time-to-first-chunk holds steady under 120ms globally.
`;
const response = await ai.models.generateContent({
model: 'gemini-3.8-flash-tts',
contents: prompt,
config: {
responseModalities: ['AUDIO'],
speechConfig: {
multiSpeaker: {
enabled: true,
turnDetection: true
},
audioFormat: 'audio/wav'
}
}
});
return response.candidates?.[0]?.content?.parts?.[0]?.inlineData?.data;
}
Step 5: Stream Low-Latency Chunks with Flash-Lite TTS
For interactive conversational agents, switch the model target to gemini-3.8-flash-lite-tts to cut initial buffer wait time to 85 milliseconds.
export async function streamAgentSpeech(userPrompt: string, onChunkReceived: (chunk: Buffer) => void) {
const stream = await ai.models.generateContentStream({
model: 'gemini-3.8-flash-lite-tts',
contents: userPrompt,
config: {
responseModalities: ['AUDIO'],
speechConfig: {
voiceConfig: {
prebuiltVoiceConfig: { voiceName: 'Charon' }
},
audioFormat: 'audio/pcm;rate=24000'
}
}
});
for await (const chunk of stream) {
const audioPart = chunk.candidates?.[0]?.content?.parts?.find(p => p.inlineData?.mimeType?.startsWith('audio/'));
if (audioPart?.inlineData?.data) {
const buffer = Buffer.from(audioPart.inlineData.data, 'base64');
onChunkReceived(buffer);
}
}
}