Google Gemini 3.8 Flash TTS Brings Prompt-Driven Voice Design
Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS, enabling teams to design voices from text prompts and direct dialogue line by line via API.

Google has expanded its Gemini Audio family with the launch of Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The two text-to-speech models transition speech synthesis away from static presets toward prompt-steered vocal production. Google describes the release as its most expressive audio generation system to date, allowing engineers and creators to generate bespoke voices and control performances using natural language instructions. Both models are accessible immediately through the Gemini API and Google AI Studio under the model identifiers gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Google confirmed that access remains strictly API-based, with no open weights provided for local hosting, while an enterprise release within Gemini Enterprise is marked as coming soon.
The release introduces a two-tier architecture designed to balance creative depth against operational cost and latency. Gemini 3.8 Flash TTS is engineered specifically for deep creative direction, character design, and rich audio experiences across video games, podcasts, audiobooks, and interactive media. In contrast, Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient scale, targeting large-scale media dubbing, content generation, and conversational voice agents. These models complement Google’s existing Gemini Audio lineup, which includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.
Generative Voice Design and Granular Performance Direction
A major technical shift in the 3.8 release is the expansion beyond the previous library of 30 original voices. Developers can now utilize generative voice design to construct custom voices from scratch using descriptive natural language prompts covering role, accent, and vocal timbre across more than 100 languages and dialects. In its technical release materials, Google showcased examples ranging from a Melbourne radio DJ and a tinny monotone robot to a dramatic Japanese dragon. For projects requiring pre-built assets, the platform supplies a library of more than 2,000 production-ready voices, including regional linguistic varieties such as Scots English, Quebec French, and Mexican Spanish. Custom-designed voices can be saved and reused to prevent acoustic drift across long-term production pipelines.
- Line-by-line script directing: Developers can insert written stage directions or rely on natural textual context to steer delivery from whispers to authoritative dialogue.
- Native two-speaker scene staging: Multi-turn conversations can be orchestrated from a single screenplay script while maintaining distinct vocal separation and realistic turn-taking.
- Long-form stability: The architecture preserves timbre, speech pacing, and audio fidelity continuously across hours of generated material.
- Conversational non-verbal bursts: Scripts can embed direct tags such as « laughs », « sigh », and « gasp » to add human texture.
- Backchanneling interjections: Active listening cues like « mhm » and « yeah » allow fine-grained calibration of comedic timing and reaction pauses.
- Voice remixing: A forthcoming feature designed to let creators adjust the timbre, pitch, pace, and accent of library voices via text prompts.
The generation workflow also supports voice replication from a 30-second reference audio sample. To mitigate unauthorized cloning, Google requires an accompanying verbal consent recording from the voice owner, which is verified against the reference speaker before generation is authorized. Replicated voice tools within Google AI Studio are currently restricted and unavailable in several jurisdictions, including the European Economic Area, the United Kingdom, and India.
Benchmark Performance and Content Provenance
According to benchmark figures published by Google, Gemini 3.8 Flash TTS secured first place overall on Hume AI’s Voice Design Benchmark with a score of 71.4, alongside a leading mark of 60.8 in accent modeling. On Hume AI’s Overall Quality Index, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS captured the number one and number two positions, respectively, showing notable performance gains over Gemini 3.1 Flash TTS. Furthermore, Google reported that blind human preference evaluations on Voice Arena placed both models at top rankings across several major languages, including Modern Standard Arabic, Hindi, Brazilian Portuguese, Japanese, Vietnamese, and Mexican Spanish.
To address media provenance and misuse, Google embeds an imperceptible SynthID watermark directly into the audio output generated by its Gemini Audio models. Cloned voices generated through the replication pipeline also incorporate C2PA content credentials, providing verifiable metadata regarding the synthetic nature of the media.
Deployment Architecture and Operational Impact
The models are being integrated natively into Google products such as Google Vids and Gemini Notebook. For external software integration, developer infrastructure platforms including Agora, LiveKit, Pipecat, and Vercel have enabled support via the Gemini API. Commercial media and software companies such as Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang are deploying the models to handle automated dubbing, regional audio localization, and real-time interactive agents.
For enterprises and technical organizations, the release alters the economics and engineering of synthetic speech deployment. Teams are no longer restricted to rigid voice catalogs or separate, fragile pipelines for multi-character dialogue generation. The dual-tier structure allows system architects to route latency-sensitive, high-throughput applications like customer support agents to Flash-Lite TTS, while reserving Flash TTS for brand-aligned narration and complex scripted media, all managed under standardized API controls and embedded provenance tracking.
Sources
- Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTSMarkTechPost · September 23, 2026
- Gemini 3.8 text-to-speech says helloGoogle · September 23, 2026



