Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, adding tools for designing voices and directing spoken dialogue. The company describes Flash TTS as the more expressive creative option, while Flash-Lite targets higher-volume speech generation. Source: Google
The release is aimed at applications such as narrated content, dubbing, games and voice interfaces. Its practical change is the ability to specify more of the performance rather than choosing only a preset voice and submitting text.
Voice design and dialogue
Google says Flash TTS can create vocal identities from natural-language descriptions, with controls for characteristics such as accent and delivery. The announcement also describes a library of more than 2,000 voices and support across more than 100 languages and dialects.
Both models include line-by-line direction and two-speaker scene staging. The company also claims improved consistency for long-form audio. Those quality statements come from Google’s announcement; they should be tested with the language, script and production conditions of the intended application.
Voice remixing is described as coming later, so it should not be treated as part of every currently available workflow.
Replication has a separate access boundary
Google’s voice-replication feature uses a short sample and a matching verbal-consent recording. The company says generated audio includes SynthID watermarking and describes additional provenance measures for replication.
The announcement states that voice replication through AI Studio is unavailable in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India. That restriction is narrower than a statement about the availability of all text-to-speech features.
A product team therefore needs to distinguish designing a new synthetic voice from reproducing a particular person’s voice.
Where the models are offered
Google identifies the Gemini API and Google AI Studio among the developer access routes and describes rollout into other products. Availability and feature support can differ between those surfaces.
The announcement does not provide a single price covering every model, product and usage pattern. A production decision still requires checking the current billing terms for the chosen route.
For creators, the release offers more control over how a script is performed. The remaining evaluation is familiar to anyone working with audio: pronunciation, consistency, editing effort and whether the result remains usable across a full project rather than one short demo.





