Search results for

All search results
Best daily deals

Affiliate links on Android Authority may earn us a commission. Learn more.

Gemini can now clone your voice and perform scripts like an actor

The upgrade adds 2,000+ voices, two-speaker scenes, and finer delivery controls.
By

Sep 24, 2026 — 4:02 AM ET

Google-Gemini-3.8-Flash-text-to-speech
Google
Add Android Authority on Google:
TL;DR
  • Google is rolling out Gemini 3.8 Flash TTS and Flash-Lite TTS, its latest text-to-speech models for more expressive AI-generated audio.
  • Gemini 3.8 Flash TTS can create custom voices, replicate an authorized voice from a 30-second sample, and control accents, pacing, emotion, and two-speaker dialogue.
  • The new models are reaching Gemini Notebook and Google Vids, with both also available to developers through Google AI Studio and the Gemini API.

AI voiceovers can read a script. But getting them to actually perform it with the right accent, pacing, emotion, pauses, and even the occasional sigh is another story. To address this, Google has spent weeks growing Gemini 3.8 from the original Gemini 3.8 Flash into new Live and Live Extended Thinking models. And now it wants the same family to act the part, with Gemini 3.8 Flash Text-To-Speech (TTS) and Flash-Lite TTS rolling out.

The biggest change for users is control. In a blog post, Google says Gemini 3.8 Flash TTS can create an original voice from a natural-language description, with support for more than 100 languages and dialects. You also get over 2,000 production-ready voices, and voice replication can recreate a consistent voice from a 30-second sample, provided you have permission to use it.

This goes well beyond “paste text, get audio.” Both models can follow line-by-line directions for tone, pacing, dialect changes, whispers, laughs, sighs, and other conversational cues. They can also stage two-speaker conversations from one script and keep voices consistent across hours of audio. That could mean podcasts, audiobooks, dubbing, voice agents, and narrated videos that all sound less robotic.

Google’s benchmark results are strong, too. Gemini 3.8 Flash TTS ranked first overall on Hume AI’s Voice Design Benchmark, including a 60.8 score for accents, while Flash and Flash-Lite took the top two spots on Hume’s overall quality index. Voice Arena tests also put them at or near the top across several languages. ElevenLabs still scored higher in Hume’s individual voice quality and human-like variation categories.

Regular users won’t have to wait for every piece to reach developer tools. Flash TTS is rolling out in Gemini Notebook, while Flash-Lite TTS is rolling out in Google Vids, which just opened its latest Omni video generator to more users. Developers can access both through the Gemini API and AI Studio.

Google is also putting guardrails in place for voice replication. It requires consent verification, adds SynthID watermarking and C2PA credentials, and blocks AI Studio voice replication in several regions, including the UK, EEA (European Economic Area), India, Texas, and Illinois.

Follow

Thank you for being part of our community. Read our Comment Policy before posting.