Skip to content
AZ Labs
Product Updates23 September 2026•3 min read

Gemini 3.8 Flash TTS and Flash-Lite TTS add custom voices and script direction

AZ Labs editorial illustration for Gemini 3.8 text-to-speech
Inspect
Original AZ Labs editorial illustration Original AZ Labs editorial illustration • © 2026 AZ Labs

Google has released two speech generation models: Flash TTS for creative voice direction and Flash-Lite TTS for high-volume audio workflows.

AI Neural Narration

48kHz Studio

Fish Audio Neural Engine · Natural editorial narration

0:000:00
verified

Key Takeaways

  • check_circleFlash TTS and Flash-Lite TTS serve different production needs.
  • check_circleEvaluate a voice on the actual script and language your audience hears.
  • check_circleKeep approval records for voices used in customer-facing content.

Two new speech generation options

Google announced both models on 23 September 2026. Flash TTS focuses on voice design and performance direction; Flash-Lite TTS targets efficient volume production. Both support multilingual audio.

Developer rollout is through the Gemini API and AI Studio; enterprise API access is coming soon. Consumer surfaces include Gemini Notebook and Google Vids. Replication requires consent verification and has regional restrictions. Remixing is a future feature.

Choose for the finished recording

Our assessment: an audio team should compare the models using a finished script, rather than a few flattering sample sentences. Include abbreviations, numbers, customer names and words that are easy to mispronounce. Listen for changes in character or delivery across a longer recording. Ask a listener who knows the target language to assess the result.

For a support agent, clarity and response time may matter more than dramatic range. For a narrated lesson, pacing and consistent pronunciation may determine whether the audio is useful. Write down the acceptance criteria before selecting a voice, then score both models against those criteria. Keep the reference script stable so revisions can be compared fairly.

Make the production workflow reviewable

Save the approved script, voice settings and final recording together. That gives an editor a way to reproduce a clip when a price, date or product name changes. Record which version was actually published. A regenerated sentence should be checked in context, because a correct standalone clip can still sound out of place in the complete recording.

For large dubbing jobs, calculate the cost per approved minute after regeneration and review. Build a pronunciation list for recurring terms and sample each language before expanding the batch. Where a voice is based on a person, document the permission and permitted uses before recording. These are editorial recommendations, not a claim that a model's safety tools replace your own approval process.

Frequently Asked Questions

Is Gemini Enterprise API access already included?

Google's launch announcement describes enterprise API access for both TTS models as coming soon.

How should I compare creative and high-volume TTS?

Use the same representative script and assess pronunciation, pacing, review effort and cost per approved minute.

Primary Sources

Share this articlePost on X
arrow_backBack to all news