←Back to AI News
AI News

Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS: A New Era of AI Voice Generation

Google's latest text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, challenge the AI voice industry with advanced voice cloning capabilities from just 30 seconds of audio, multi-speaker scripting, and an expansive library of over 2,000 voices.

V
Vicky Rathore
Google Gemini 3.8 Flash TTS & Flash-Lite TTS Launched: Features & Specs

Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS: A New Era of AI Voice Generation

Google has officially launched the Gemini 3.8 Flash TTS voice models, introducing powerful new speech generation systems designed for both creative scripting and high-volume audio production. With this release, developers are gaining unprecedented control over how AI voices are directed, cast, and generated.

The Dual-Model Approach: Flash and Flash-Lite

To accommodate diverse development needs, Google has bifurcated its new text-to-speech offering into two distinct models:

  • Gemini 3.8 Flash TTS: This model is built for studio-grade voice quality and highly expressive acting. It supports 130 languages and is designed for nuanced emotional delivery, custom voice design, regional dialects, and long-form narrative stability. It currently ranks #1 on Hume AI's Voice Design Benchmark with a score of 71.4.

  • Gemini 3.8 Flash-Lite TTS: Acting as the high-throughput, cost-efficient workhorse, this model supports 101 languages. It is positioned as the recommended drop-in upgrade from earlier preview models, perfect for real-time voice agents, customer support bots, and high-volume dubbing. It ranks #2 on Hume AI's Overall Quality Index.

Expanded Voice Library and Voice Cloning

Historically, developers had to choose from a limited pool of generic AI voices. Google has drastically changed this by expanding its voice library from just 30 standard studio voices to over 2,000 regional voices.

Furthermore, Gemini 3.8 Flash TTS supports rapid voice replication. Developers can clone a voice using just a 30-second audio clip, or alternatively, generate a completely brand-new voice from scratch using simple text-based descriptions via Voice Design.

A New Scripting Format for Director-Level Control

A major issue with previous text-to-speech models was the blending of stage directions with actual dialogue, which often caused the AI to accidentally read its own instructions out loud. To solve this, Google has separated the prompt structure into three distinct parts:

  1. Cast (Voice): Developers select who is speaking by choosing from the vast voice library, replicating a voice, or generating a new one.

  2. Direct (Style): This field handles how the line is delivered, allowing developers to set emotions, pacing, and tone (e.g., "whispered urgently" or "bored and monotone").

  3. Speech (Text): This is the verbatim transcript of what the AI will say.

Bringing Realism with Acoustic Tokens and Multi-Speaker Scenes

To make conversations sound truly human, the text field now supports inline non-verbal sounds. Developers can place specific acoustic tokens like <laughs>, <coughs>, and <sighs> directly into the text right where they want them to happen.

The models also support realistic two-speaker conversations. Instead of awkward, unnatural pauses between different voices, the system allows for dynamic interactions and overlapping reactions. This includes the use of conversational backchanneling tokens like |mhm| and |yeah| to mimic how humans listen and respond during a dialogue.

By providing granular control over emotion, pacing, and natural acoustic textures, Google's Gemini 3.8 Flash TTS models are setting a new standard for interactive voice AI and automated dubbing pipelines.

Comments

Log in to leave a comment.