Google Ships Gemini 3.8 TTS Models With Voice Design And Line By Line Direction

Gemini 3.8
Image credit: Google
Google has released two new text to speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash‑Lite TTS, calling them the most expressive audio generation models it has built to date and pitching them as a shift from static preset voices toward what the company describes as a dynamic creative studio for audio.
The models were introduced on September 23 by Leland Rechis, group product manager, and Alan Cowen, director of research science on behalf of the Gemini Audio team. They join a fast growing Gemini Audio lineup that already includes 3.5 Live Translate, 3.5 Transcribe, and 3.8 Live, positioning speech generation as an increasingly central piece of Google's broader Gemini 3.8 rollout.
The two models are built for different jobs. Gemini 3.8 Flash TTS is designed for deep creative direction and character design, aimed at gaming, immersive audiobooks, podcasts, and interactive media where a developer wants to shape a distinct voice and control delivery line by line. Gemini 3.8 Flash‑Lite TTS is built for high volume, cost efficient scale, optimised for dubbing large batches of content and powering expressive voice agents that need to run continuously rather than for one off creative projects. According to Artificial Analysis benchmark data, Flash TTS processes audio at roughly 44.1 characters per second and Flash‑Lite at about 40.2 characters per second, translating to close to 2.7 times and 2.4 times faster than real time playback respectively.
Central to the release is what Google calls generative voice design. Rather than choosing from a fixed list of presets, developers using Flash TTS can describe a voice's role, accent, and characteristics in plain language across more than 100 languages and dialects, whether that means a high energy radio host from Melbourne or a fire breathing dragon with a specific regional cadence. Alongside that capability, both models ship with a library of more than 2,000 production ready voices covering regional varieties such as Mexican Spanish, Quebec French, and Scots English.
The models can also replicate a specific voice from a short reference clip. Google says the system can recreate a consistent vocal profile from a 30 second audio sample, but only after a verbal consent recording from the voice's owner is checked against the reference speaker, a safeguard the company says is intended to protect voice talent and prevent unauthorised cloning. Voice replication through AI Studio is not currently available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, or India, according to Google's own disclosure. A voice remixing feature that will let users fine tune timbre, pitch, pace, and accent of an existing library voice through simple prompts is described as coming soon rather than available at launch.
Beyond generating a voice, both models give developers granular control over how a line is actually performed. Users can write their own stage directions or let the model interpret natural script cues, adjusting pacing, dialect shifts, and backchanneling on a line by line basis. The models are built to maintain consistent voice quality and timbre across long form content such as full length podcasts and audiobooks with minimal speaker drift, and support native two speaker scene staging that keeps multiple voices distinctly separated through natural conversational turn taking. Google has also added support for non verbal vocal bursts such as laughing, sighing, or gasping, along with short active listening interjections, aimed at making multi speaker dialogue sound less mechanically scripted.
On third party benchmarks, Google says Gemini 3.8 Flash TTS holds the top overall spot on Hume AI's Voice Design Benchmark with a score of 71.4, and also leads that benchmark's accent modelling category. Flash TTS and Flash‑Lite TTS take the first and second positions respectively on Hume's Overall Quality Index, and Google reports that both models rank among the top performers in blind human preference testing on Voice Arena across languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.
Safety measures are built directly into the output rather than offered as an optional layer. Every audio clip generated by the models carries a SynthID watermark, an imperceptible signal woven into the audio itself that is designed to keep AI generated speech detectable even after the clip has been shared, edited, or re‑encoded, alongside C2PA content credentials intended to support broader transparency about how a piece of audio was produced.
Both models are rolling out starting today. Developers can access Gemini 3.8 Flash TTS and Flash‑Lite TTS through the Gemini API and in Google AI Studio's new speech generation playground, where voices can be prompted from scratch and then moved into a dual speaker screenplay editor for line by line direction. Enterprise access through Gemini Enterprise is described as coming soon, while consumer facing integration is already live through Gemini Notebook for Flash TTS and Google Vids for Flash‑Lite TTS. Google says it is also partnering with developer platforms including Agora, LiveKit, Pipecat, and Vercel, along with companies such as Figma, HeyGen, Linguana, and Wondercraft, who are integrating the new models into dubbing, localisation, and conversational voice agent products.
The launch lands as competition in AI powered speech generation intensifies, with rival providers racing to improve naturalness, multilingual coverage, and low latency performance for real time voice agents. Google's own benchmark data shows Gemini's TTS models still trailing some specialised providers on raw generation speed, suggesting the company's near term pitch rests more on expressiveness, voice design flexibility, and safety tooling than on being the fastest option available.
Topics
Sources
Stay informed
Startup news in your inbox
Get important funding rounds, founder stories, and startup updates.
No spam - only important startup updates.





