Gemini 3.8 Flash TTS

Direct every line of your script with Gemini 3.8 Flash TTS — set tone turn by turn, stage two speakers and produce 130-language audio.

Gemini 3.8 Flash TTS
Steer each line's delivery with the flagship tier, or switch to Flash-Lite for budget-friendly bulk runs
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Gemini 3.8 Flash TTS: The Creative Tier of Google's Voice Family

Launched on September 23, 2026, the lineup pairs an expressive flagship with a lean, high-volume engine built for mass audio production.

  • A Single Release, Two Distinct Jobs
    Gemini 3.8 Flash TTS handles nuanced, character-driven reads, while Flash-Lite TTS keeps the cost per minute low for large batches.
  • Direction Instead of Presets
    Style notes applied per turn, structured speech metadata and inline vocal events govern tone, tempo, feeling and accent.
  • Designed Voices and Consented Replication
    Describe the voice you want in plain language, or mirror a real speaker using a reference clip plus a matching consent recording.

How to Prompt Gemini 3.8 Flash TTS for Clean Output

Four habits that keep the spoken text untouched and let the performance metadata do the work.

Capability Deep Dive: Gemini 3.8 Flash TTS

Performance control, voice design and multilingual reach — a closer look at what the flagship Gemini TTS tier actually delivers.

Turn-by-Turn Performance Control

Per-line styling plus inline laughs, sighs, coughs, breaths and pauses — the feel of coaching a performer rather than loading a preset.

Describe a Voice in Plain Words

Prompts can set age range, personality, accent, vocal texture and role, drawing on 2,000+ production voices through the Voices endpoint.

Cloning Guarded by Consent

A clean reference recording plus a matching consent take from the same adult speaker, with SynthID watermarking and C2PA credentials attached.

Two-Voice Scene Staging

Scripts carry the exchange for podcasts, teaching dialogues, product demos and game scenes with no manual line stitching.

Steady Voice Across Long Form

Google documents consistent identity, timbre, volume and room tone across multi-minute narration and extended dialogue.

130 Languages, Regional Accents Included

Flash TTS covers 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Questions People Ask About Gemini 3.8 Flash TTS

Straight answers on cost per minute, which tier to pick, benchmark standing and the safeguards that apply.

1

What does it cost per minute of audio?

Roughly 1.35 cents per audio minute at launch rates of $0.50 for input and $9 for output per million tokens.

2

Which tier should I choose?

Take the flagship when acting nuance and long-form quality matter; take Flash-Lite when volume and low latency matter more.

3

How does it stack up against rival voice models?

Google reports 71.4 on Hume's Voice Design Benchmark; Voice Arena places it second at 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS takes over from the 3.1 preview and brings audio output pricing down from $20 to $6 per million tokens.

5

Can I clone a voice, and what rules apply?

Replication requires a reference clip plus a matching consent recording from the same adult speaker.

6

Why does the model read my stage directions aloud?

Everything in the input is treated as transcript, so move lasting directions into the speech metadata instead.

Take Gemini 3.8 Flash TTS for a Spin on Your Own Scripts

Try both tiers inside the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS means swapping a single model identifier. Weigh batch against priority inference before you lock in a production budget.