OpenAI Python SDK 3.24.0 Adds Voices API: Custom Voice Creation from Prompts and Audio Samples
OpenAI released version 3.24.0 of its official Python library, introducing client.voices.create to generate persistent synthetic voices either by providing reference WAV audio or describing acoustic traits in plain text.
OpenAI updated its official Python SDK to version 3.24.0, introducing an unannounced client.voices resource namespace.
The API gives developers programmatic control over creating, storing, and addressing custom voice personas for text-to-speech runs and Realtime API audio sessions.
Previously, developers using OpenAI's voice models were restricted to pre-set stock personas (alloy, echo, fable, onyx, nova, shimmer, sage, and coral). Version 3.24.0 removes this constraint by supporting both descriptive acoustic prompting and reference sample synthesis.
Installation and Version Check
The endpoint requires updating the official library to at least 3.24.0:
pip install --upgrade "openai>=3.24.0"
Verify the module exports the voices sub-client:
import openai
print(openai.__version__)
client = openai.OpenAI()
assert hasattr(client, "voices"), "Voices API requires openai>=3.24.0"
Modality 1: Prompt-Driven Acoustic Generation
Developers who lack pristine studio audio recordings can synthesize new voices by providing text instructions specifying vocal timbre, cadence, pitch, and accent:
from openai import OpenAI
client = OpenAI()
# Generate a custom voice via natural language acoustic prompt
voice = client.voices.create(
name="technical-instructor-en-uk",
description="A calm, articulate Scottish speaker in his mid-40s with a measured cadence, moderate baritone pitch, and crisp diction suitable for technical engineering audiobooks.",
modality="prompt"
)
print(f"Created Voice ID: {voice.id}")
# Output: voice_94k2ms820fa923
The backend translates the descriptive adjectives into latent speaker conditioning vectors. The operation completes in 3.2 to 4.5 seconds and returns a permanent identifier usable across all audio-capable endpoints.
Modality 2: Reference Audio Sample Cloning
For organizations that hold explicit intellectual property rights to specific narrator voices, the API accepts short WAV or MP3 reference files:
with open("reference_sample_12s.wav", "rb") as audio_file:
cloned_voice = client.voices.create(
name="company-brand-voice",
audio_sample=audio_file,
consent_verification_token="usr_verif_99812",
modality="sample"
)
print(f"Cloned Voice ID: {cloned_voice.id}")
OpenAI enforces safety guardrails on uploaded samples:
- Minimum audio duration of 8 seconds with high signal-to-noise ratio (>20 dB SNR).
- Automated biometric watermarking that embeds cryptographic tags in output audio streams.
- Real-time vocal fingerprint filtering against public politician and celebrity registries.
Using Custom Voice IDs in Realtime and TTS Calls
Once generated, the custom voice_id replaces stock voice strings in Chat Completions and WebSockets:
# Standard TTS request with custom voice
response = client.audio.speech.create(
model="tts-1-hd",
voice=voice.id,
input="The database migration executed successfully with zero table locks."
)
response.stream_to_file("output.mp3")
# Realtime Session Configuration
realtime_session_payload = {
"type": "session.update",
"session": {
"modalities": ["text", "audio"],
"voice": voice.id,
"turn_detection": {"type": "server_vad"}
}
}
Operational Limits and Pricing
| Dimension | Prompt-Generated Voices | Sample-Cloned Voices |
|---|---|---|
| Generation Fee | $0.15 per generated voice profile | $1.20 per reference audio ingest |
| Storage Fee | Free for active accounts | Free up to 50 active voice profiles |
| Latency to First Audio Chunk | 180 ms (TTS-1) / 310 ms (TTS-1-HD) | 195 ms (TTS-1) / 325 ms (TTS-1-HD) |
| Max Voices per Organization | 100 profiles (standard tier) | 20 profiles (requires verified KYC) |
Summary
The addition of client.voices.create in Python SDK 3.24.0 bridges the gap between static stock voices and costly external voice cloning services, letting developers produce unified sonic brand identities directly inside their OpenAI pipelines.