Try AI Narrator Free
Sign up free — no credit card required. Turn your documents into natural-sounding audio in minutes.
Sign Up FreeTry AI Narrator Free
Experience the voices mentioned in this article. Install the Google Docs add-on and generate audio from any document in seconds.
Browse All VoicesBrowse other categories
Gemini TTS Best Practices: Generate Consistent, Expressive Audio Every Time
Quick Summary: Gemini TTS models generate audio from language model prompts rather than simple text-to-waveform mapping. That means your prompts, voice choice, and inline bracket tags heavily shape the output. This guide covers five pillars of consistent audio generation: presets, voice selection, emotion tags, system prompts, and guard rails.
Style Prompt (System Instruction)
The primary driver of overall emotional tone and delivery style. Sets the context for the whole segment.Text Content (Transcript)
The meaning of the words themselves. Emotionally rich text gives the model more to work with than neutral text.Inline Tags (Markup)
Bracketed modifiers like [whispers] or [laughs] for localized control over specific phrases, working with the style prompt.
Key Principle: For maximum predictability, keep Style Prompt, Text Content, and Inline Tags pointing at the same emotional goal. A tense prompt paired with neutral, meeting-style text will produce muddled results.
- Reproducibility: The same preset applied to different transcripts produces audio that sounds like the same speaker in the same session.
- Multi-speaker workflows: Define a preset per speaker (interviewer, guest, narrator) to keep distinct voices across episodes.
- Reduced prompt drift: Without a preset, the model infers style from text alone, which drifts across a long project.
- Faster iteration: Once a preset sounds right, you only adjust the transcript, not reinvent the voice each time.
- Give the character a name. Grounding the model with a name ("Nova", "Dr. Claire", "Marcus") ties the performance together.
- Define identity, not just tone. "Radio DJ" or "documentary narrator" gives the model a performance archetype to anchor to.
- Be specific about environment. "Busy morning coffee shop" vs "quiet studio" produces different acoustic profiles.
- Include director's notes. Address style, pace, accent, and tone explicitly rather than leaving them to chance.
- Keep presets concise. A focused preset is easier for the model to follow consistently.
Want to hear these voices yourself?
Try our AI Narrator for free right now. No sign-up required.
| Voice Name | Gender | Best For | Characteristics |
|---|---|---|---|
| Kore | Female | Warm narration, podcasts | Friendly, warm, approachable |
| Puck | Male | Upbeat content, promos | Energetic, confident |
| Enceladus | Male | Calm narration, audiobooks | Breathy, introspective |
| Aoede | Female | Professional presentations | Clear, polished |
| Fenrir | Male | Dramatic narration | Powerful, intense |
| Orus | Male | Technical content | Precise, measured |
- Match voice to preset style. If your preset describes a tired, introspective character, pick a breathier voice like Enceladus rather than an upbeat one.
- Test voices in the playground. Listen to how each voice handles a variety of prompts before committing.
- Use voice + prompt synergy. A deep, authoritative voice paired with an authoritative style prompt reinforces the effect.
- Avoid mismatched age or tone prompts. Asking a deep voice to sound like a young child produces uncanny results.
- Lock one voice per character. In multi-speaker content, assign one voice per speaker and don't change it mid-project.
Pro Tip: When using the Google Cloud TTS API (Vertex AI), you can combine voice selection with inline audio tags and multi-speaker dialogue mode for automatic speaker separation in interviews and podcasts.
The tag is replaced by an audible, non-speech vocalization rather than spoken words.
| Tag | Effect | Reliability |
|---|---|---|
| [sigh] | Inserts a sigh sound | High |
| [laughing] | Inserts a laugh | High |
| [uhm] | Inserts a hesitation | High |
| [gasp] | Inserts a gasp | High |
| [cough] | Inserts a cough sound | Medium |
The tag changes how the following text is delivered — tone, pace, or emphasis.
| Tag | Effect | Reliability |
|---|---|---|
| [whispers] | Whispered delivery | High |
| [excitedly] | Excited, energetic tone | High |
| [bored] | Flat, disinterested delivery | High |
| [calmly] | Relaxed, measured delivery | High |
| [slowly] | Slower pace | High |
| [quickly] | Faster pace | High |
| [pause] | Brief pause before continuing | High |
Tags that control speed and emphasis of delivery.
// Default — no tags Hey there, I'm a new text to speech model. How can I help you today? // Excited [excitedly] Hey there, I'm a new text to speech model! How can I help you today? // Bored [bored] Hey there, I'm a new text to speech model... How can I help you today? // Whispered with laugh [whispers] Hey there... [laughing] I'm a new text to speech model. How can I help you today?
Ready to Try AI Narrator?
Start converting your Google Docs into professional audio today. Free forever!
Add to Doc - It's free- Use tags for localized moments, not overall tone. Set the overall tone with the system prompt; use tags for a laugh here, a whisper there.
- Don't overuse tags. Too many tags in a short passage sounds unnatural. Aim for one or two per paragraph.
- Keep tags aligned with the prompt and text. A [cheerful] tag in a scene the prompt describes as somber will fight the output.
- Test new tags first. A tag you expect to be a style modifier might be spoken as literal text. Test it in a playground before using it in production.
You are a scriptwriter and audio director. I have a context but no transcript. TASK: 1. Write an engaging script based on the context. 2. Format the output as a structured TTS prompt. OUTPUT FORMAT: # AUDIO PROFILE: [Name] ## "[Title]" ## THE SCENE: [Scene] [Description] ### DIRECTOR'S NOTES Style: [Style] Pace: [Pace] Accent: [Accent]
- Be specific, not generic. "Speak like a 1940s radio announcer" beats "speak in an old-fashioned way."
- Set the scene. Environment context (a busy airport, a quiet studio) guides the acoustic interpretation.
- Define the character's emotional state. "Nova is relaxed and confident" gives the model a baseline for every line.
- Include delivery details. Mention pacing, pauses, and vocal texture. "Slightly breathy, measured pace" is more actionable than "calm."
- Keep prompts consistent across a series. For a podcast, use the same system prompt for every episode.
| Consideration | Impact | Mitigation |
|---|---|---|
| Long outputs | Quality can drift over very long single generations | Split transcripts at natural break points and stitch audio in post |
| Voice mismatch with prompt | Audio may not match the intended speaker | Keep preset, voice, and prompt aligned |
| Tag vocalization | Some tags can be spoken as literal text | Test tags in the playground before production |
| Rate limits & quotas | API throttling on high-volume use | Use the right model tier and queue requests |
- Presets — Define who, where, and how (reusable per speaker)
- Voice selection — Match voice characteristics to your character
- Inline tags — Granular, per-phrase control (use sparingly, test first)
- System prompts — Set the overall tone, scene, and performance style
- Guard rails — Split long content, test tags, align voice to prompt
Skip the Code: If you want Gemini-powered voices inside your Google Docs without writing API integrations, try the free AI Narrator Add-on — it handles the voice generation pipeline, including emotion and pacing control.
- Best AI Voice APIs in 2026 — Full comparison of every TTS API provider
- AI Voice Cloning Guide — Clone any voice with step-by-step instructions
- AI Narrator vs ElevenLabs — Head-to-head quality and pricing comparison
Disclosure
Some of the tools compared on this site — including AI Narrator — are our own products. We review them on the same footing as competitors, on current pricing and hands-on testing, and always tell you when a product is ours.