Trending: On-device modelsSearch
iHeartGeek
iTECH

Gemini 3.8 Flash TTS takes voice work into create-a-voice

Google has added Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two speech models that generate original character voices from a prompt instead of selecting from a preset list.

Google's pale blue Gemini announcement artwork reading Introducing Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS above the multicoloured Gemini spark logo

Google is done with voice presets. The company has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two speech models that build a voice from a written prompt rather than asking a developer to pick one from a catalogue, and both are available from today.

From thirty voices to an open library

The older approach capped creators at around thirty recorded voices. Google says the 3.8 generation removes that ceiling: Flash TTS can generate an entirely new character voice, or a consistent brand voice, and then perform scripts in it. Both models accept direction at the line level, covering accent and emotional tone, and both support dual-speaker screenplay editing for dialogue scenes rather than single-voice narration.

Where it is available

Rollout starts today across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids. In AI Studio the company has set up something closer to a voice design desk, where a developer can prompt a new vocal identity, or replicate their own, and move it into a screenplay editor for line-by-line delivery. Agora, LiveKit, Pipecat and Vercel are named as developer platforms carrying the models, with Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang using them for dubbing, localisation and conversational agents.

The benchmark claims and the safety controls

Google says Flash TTS took first place overall on Hume AI's Voice Design Benchmark with a score of 71.4, and first in accent modelling at 60.8, while the two models hold the top two positions on Hume's Overall Quality Index. In blind listening tests on Voice Arena the company reports leading positions in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi, with support for more than 100 languages.

Voice replication comes with a consent check: a user must supply a verbal consent recording from the voice owner that matches the reference speaker before a cloned voice can be created. Every clip generated by the Gemini audio models also carries a SynthID watermark woven into the audio itself, so synthetic speech stays detectable.

Our opinion

The interesting shift here is not quality, it is the unit of work. Benchmarks get quoted every release cycle, but moving from thirty voices to an unlimited generated library changes what a small studio can attempt: an audiobook, a game or a podcast no longer needs a casting session before the first line is written.

That is also where the pressure lands. The consent recording and the SynthID watermark are sensible defaults, and Google deserves credit for shipping them on day one rather than in a later patch, but they shift the burden of diligence onto whoever holds the microphone. Voices are the part of a performer's livelihood that is least protected by copyright, and a system that makes a convincing replica in one prompt raises the value of proving who actually agreed to it.