Intermediate · 15 minutes

How to Build Multi-Speaker Expressive Voice Applications with Gemini 3.8 Flash TTS API

Step-by-step developer tutorial for integrating Google Gemini 3.8 Flash TTS and Flash-Lite TTS APIs with prompt-steered vocal design, multi-character dialogue scripts, and 24 kHz PCM streaming.

Step 1: Install the Google GenAI SDK and Audio Utilities

Add the official @google/genai TypeScript SDK to your project. Ensure your Node.js runtime has buffer and filesystem capabilities.

npm install @google/genai
# For Python environments:
pip install google-genai soundfile

Step 2: Configure API Key and Client Initialization

Initialize the GoogleGenAI client with your API key from Google AI Studio. Set the environment variable GEMINI_API_KEY in your deployment environment.

import { GoogleGenAI } from '@google/genai';

// Automatically reads process.env.GEMINI_API_KEY
const ai = new GoogleGenAI({});

Step 3: Define Voice Design Prompt and Single-Speaker Synthesis

Configure responseModalities to AUDIO and specify voiceDesign prompts in speechConfig to control pacing, emotional tension, whispering, or regional accents without providing reference audio.

export async function generateNarratorAudio(scriptText: string) {
  const response = await ai.models.generateContent({
    model: 'gemini-3.8-flash-tts',
    contents: scriptText,
    config: {
      responseModalities: ['AUDIO'],
      speechConfig: {
        voiceConfig: {
          prebuiltVoiceConfig: {
            voiceName: 'Puck' // One of 2,048 available voice presets
          }
        },
        voiceDesign: {
          prompt: 'Thoughtful investigative journalist, mid-30s, speaking in calm, measured cadence with distinct pauses between evidentiary statements.'
        },
        audioFormat: 'audio/pcm;rate=24000'
      }
    }
  });

  const part = response.candidates?.[0]?.content?.parts?.find(p => p.inlineData?.mimeType?.startsWith('audio/'));
  if (!part?.inlineData?.data) {
    throw new Error('No audio data returned by Gemini TTS engine');
  }

  return Buffer.from(part.inlineData.data, 'base64');
}

Step 4: Configure Multi-Speaker Turn-Taking Scripts

Format character dialogues with structured speaker tags. Gemini 3.8 Flash TTS synthesizes natural conversational turn-taking, pauses, and overlapping speech in a single generative pass.

export async function generateDialoguePodcast(dialogueScript: string) {
  // Format script with explicit character names and styles
  const prompt = `
[Speaker: Maya | Voice: Aoede | Style: Animated, fast-paced tech analyst]
Have you seen the latency numbers on the new flash models? We are looking at under ninety milliseconds!

[Speaker: David | Voice: Fenrir | Style: Skeptical, grounded engineer]
Sub-ninety is solid for raw generation, but what about the network roundtrip across cross-region clusters?

[Speaker: Maya | Voice: Aoede | Style: Chuckling, dismissive of doubts]
They ran test nodes in Tokyo and Frankfurt. The time-to-first-chunk holds steady under 120ms globally.
`;

  const response = await ai.models.generateContent({
    model: 'gemini-3.8-flash-tts',
    contents: prompt,
    config: {
      responseModalities: ['AUDIO'],
      speechConfig: {
        multiSpeaker: {
          enabled: true,
          turnDetection: true
        },
        audioFormat: 'audio/wav'
      }
    }
  });

  return response.candidates?.[0]?.content?.parts?.[0]?.inlineData?.data;
}

Step 5: Stream Low-Latency Chunks with Flash-Lite TTS

For interactive conversational agents, switch the model target to gemini-3.8-flash-lite-tts to cut initial buffer wait time to 85 milliseconds.

export async function streamAgentSpeech(userPrompt: string, onChunkReceived: (chunk: Buffer) => void) {
  const stream = await ai.models.generateContentStream({
    model: 'gemini-3.8-flash-lite-tts',
    contents: userPrompt,
    config: {
      responseModalities: ['AUDIO'],
      speechConfig: {
        voiceConfig: {
          prebuiltVoiceConfig: { voiceName: 'Charon' }
        },
        audioFormat: 'audio/pcm;rate=24000'
      }
    }
  });

  for await (const chunk of stream) {
    const audioPart = chunk.candidates?.[0]?.content?.parts?.find(p => p.inlineData?.mimeType?.startsWith('audio/'));
    if (audioPart?.inlineData?.data) {
      const buffer = Buffer.from(audioPart.inlineData.data, 'base64');
      onChunkReceived(buffer);
    }
  }
}