Gemini 3.8 Flash TTS
Direct every line of your script with Gemini 3.8 Flash TTS — set tone turn by turn, stage two speakers and produce 130-language audio.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Gemini 3.8 Flash TTS: The Creative Tier of Google's Voice Family
Launched on September 23, 2026, the lineup pairs an expressive flagship with a lean, high-volume engine built for mass audio production.
- A Single Release, Two Distinct JobsGemini 3.8 Flash TTS handles nuanced, character-driven reads, while Flash-Lite TTS keeps the cost per minute low for large batches.
- Direction Instead of PresetsStyle notes applied per turn, structured speech metadata and inline vocal events govern tone, tempo, feeling and accent.
- Designed Voices and Consented ReplicationDescribe the voice you want in plain language, or mirror a real speaker using a reference clip plus a matching consent recording.
How to Prompt Gemini 3.8 Flash TTS for Clean Output
Four habits that keep the spoken text untouched and let the performance metadata do the work.
Capability Deep Dive: Gemini 3.8 Flash TTS
Performance control, voice design and multilingual reach — a closer look at what the flagship Gemini TTS tier actually delivers.
Turn-by-Turn Performance Control
Per-line styling plus inline laughs, sighs, coughs, breaths and pauses — the feel of coaching a performer rather than loading a preset.
Describe a Voice in Plain Words
Prompts can set age range, personality, accent, vocal texture and role, drawing on 2,000+ production voices through the Voices endpoint.
Cloning Guarded by Consent
A clean reference recording plus a matching consent take from the same adult speaker, with SynthID watermarking and C2PA credentials attached.
Two-Voice Scene Staging
Scripts carry the exchange for podcasts, teaching dialogues, product demos and game scenes with no manual line stitching.
Steady Voice Across Long Form
Google documents consistent identity, timbre, volume and room tone across multi-minute narration and extended dialogue.
130 Languages, Regional Accents Included
Flash TTS covers 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.
Questions People Ask About Gemini 3.8 Flash TTS
Straight answers on cost per minute, which tier to pick, benchmark standing and the safeguards that apply.
What does it cost per minute of audio?
Roughly 1.35 cents per audio minute at launch rates of $0.50 for input and $9 for output per million tokens.
Which tier should I choose?
Take the flagship when acting nuance and long-form quality matter; take Flash-Lite when volume and low latency matter more.
How does it stack up against rival voice models?
Google reports 71.4 on Hume's Voice Design Benchmark; Voice Arena places it second at 1,260 Elo.
What changed from the Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS takes over from the 3.1 preview and brings audio output pricing down from $20 to $6 per million tokens.
Can I clone a voice, and what rules apply?
Replication requires a reference clip plus a matching consent recording from the same adult speaker.
Why does the model read my stage directions aloud?
Everything in the input is treated as transcript, so move lasting directions into the speech metadata instead.
Take Gemini 3.8 Flash TTS for a Spin on Your Own Scripts
Try both tiers inside the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS means swapping a single model identifier. Weigh batch against priority inference before you lock in a production budget.
