A natural voice is only one part of a useful narration. The speaker must say the correct words, pronounce names properly, leave room for the picture, and survive a revision without sounding like a different person. Google AI Studio’s text-to-speech workspace lets you direct those choices. This guide follows the Gemini 3.8 interface checked on September 24, 2026, and uses a short production exercise you can repeat with your own script.
Download the narration audition and revision sheet · Try a separate narration in Voice Lab
Choose scripted speech for a locked narration
Google’s speech-generation documentation distinguishes TTS from the Live API. TTS is intended for reciting supplied text with delivery control; Live is designed for interactive conversation. For an approved tutorial script, begin with the speech-generation workflow rather than expecting a live conversation to produce a repeatable narration master.
Write the words before choosing the voice. A model cannot solve an unclear explanation by sounding more expressive. Read the script yourself, remove phrases you would not naturally say aloud, and check that the listener can follow it without seeing the paragraph. If an instruction depends on a screen label, use the exact current label.
Decide whether the result is one narrator or a dialogue. Start with one speaker for this exercise. Multiple speakers introduce questions about who says each line, turn-taking, and whether the voices remain distinguishable in the final mix. Solve those after one short narration works.
Find the current speech controls in AI Studio
Open Google AI Studio’s speech workspace and choose Create new dialog. In the interface checked for this guide, the editor contains a speech block, a speaker selector, Style, + Expression, and an Add speech block control. The run-settings panel shows the selected model and speaker settings.
The observed model was Gemini 3.8 Flash TTS. Confirm your own selection before following model-specific syntax, because saved projects and older tutorials can refer to earlier models. The inspected workspace also showed “No API key selected.” If Run is unavailable, check the project and access requirements shown in your account; opening the workspace alone does not establish a free generation entitlement.
You do not need a custom voice for this exercise. Begin with an available voice that suits the language and delivery. Keep the model name and voice identifier in your notes. A voice label without the model makes it harder to explain why a later attempt sounds different.
Keep spoken words separate from performance direction
Paste only the intended narration into the speech block. Open Style to describe the delivery. In the current interface, this opens a custom-style field and preset examples. A restrained instruction such as “clear, conversational explanation with room between steps” is a better first test than several conflicting emotions.
For Gemini 3.8, Google documents transcript text as verbatim speech and separate style metadata for sustained delivery. Point-in-time expressions use supported inline tags. A stage direction accidentally placed in ordinary text may become something the voice says. Do not copy another provider’s bracket syntax and assume it means the same thing.
The script contains a name, a time, a number, and a three-step instruction. Those give you concrete things to check. Substitute a real name only after confirming its pronunciation with the person or an approved reference. This is an audition plan, not a recording we have benchmarked.
Make one neutral audition before adding expression
Use a short passage first. Listen for intelligibility and fit: does the voice suit the audience, can you understand every word without looking at the script, and does the sentence rhythm leave room to absorb the instructions? Save the neutral version before changing the style.
Then make a directed version using the same words, model, and voice. Change only the delivery instruction. If you alter the script and switch voices together, you cannot tell whether the improvement came from clearer writing or better performance. Name the files so the difference remains obvious, such as menu-neutral-v01 and menu-directed-v01.
Listen at comparable loudness. A louder take can seem more convincing even when its phrasing is worse. Check on headphones for artifacts and on a small speaker for everyday clarity. A short list of pass/fail requirements keeps the decision grounded.
Use a listening checklist with actual acceptance criteria
| Check | Pass condition | What to revise |
|---|---|---|
| Words | Every approved word is present and in order | Transcript or failed generation |
| Names and numbers | The intended pronunciation and reading are clear | Spelling, written-out form, or supported pronunciation control |
| Pacing | Instructions can be followed without rushing | Sentence length, pauses, or style |
| Performance | Tone matches the audience and stays consistent | One delivery instruction at a time |
| Picture fit | Key words land near the intended visual cue | Script length or edit timing |
| Technical quality | No audible clipping, abrupt joins, or truncated ending | Take selection or finishing edit |
For this sample, “twelve dishes” should sound like a count, and “nine thirty” should sound like the intended time. Writing the spoken form explicitly removes some ambiguity. It also makes your approved transcript a better record of what the listener should hear.
Do not mark a take successful only because it sounds human. A beautifully spoken wrong number is a failed narration. If the voice adds a word, drops a phrase, or reads a direction aloud, correct the input or regenerate the affected passage and listen again.
Repair one line without losing the surrounding performance
Suppose the first two sentences are good but the instruction about the plate is rushed. Keep the accepted file. Revise that sentence with enough neighboring context to preserve its rhythm, using the same model and voice. A separate speech block can organize the work, but it is not a guarantee of seamless matching between generations.
Compare the replacement with the line before and after it. Listen for a sudden change in energy, pitch, room character, and pace. If the join is distracting, regenerate a larger phrase rather than cutting in the middle of a tightly connected thought. A clean sentence boundary is usually easier to manage than a splice inside a word.
Assemble the final version in an audio or video editor. Keep a little silence around accepted passages, then trim deliberately. Avoid extreme time compression to force an overlong script into a short scene. Often the better correction is removing a redundant sentence while preserving the important instruction.
Export, reopen, and check the real deliverable
Use the download or export control available for the accepted result, then reopen the saved file in the application that will finish the project. Verify its duration, beginning, ending, channel layout, and playback. A successful preview does not prove that you downloaded the intended version.
For API users, Google’s current documentation distinguishes complete WAV output for ordinary requests from raw PCM chunks for streaming by default. Follow the current output-format documentation if you are building an integration; this browser exercise does not require API code.
Place the narration under the real picture before adding music. Confirm that a pause does not leave a visual step unexplained and that an emphasized word lands where it helps. Then add a restrained sound bed and check that the words remain clear. Keep the narration master separate from the mixed export so one music change does not require another voice generation.
How QuestStudio fits as a separate option
If you want to try the same production brief in another workspace, use QuestStudio Voice Lab with its available voice workflow. It is a separate option; this guide does not claim that Voice Lab exposes Gemini 3.8 or mirrors Google AI Studio’s controls.
Keep the words and acceptance criteria consistent across trials while adapting delivery syntax to the selected tool. For timing problems, use the voiceover timing workflow before generating the entire project. The useful result is an approved file that fits your edit, not a growing folder of auditions.
About this guide
This is an editorial production method, not a claim that every AI model can perform every step. Examples and worksheets are illustrative; no measured generation results are implied. The hero is an original AI-generated illustration. Google’s official documentation and visible AI Studio editor were checked September 24, 2026. No paid generation or comparative listening benchmark was run for this article.
Frequently asked questions
Where is text to speech in Google AI Studio?
Open the speech workspace linked in this guide and choose Create new dialog. The checked interface uses speech blocks with a speaker selector and separate Style control.
Which Gemini model does this tutorial use?
The interface checked September 24, 2026 showed Gemini 3.8 Flash TTS. Confirm the model selected in your own workspace before using version-specific instructions.
Why are my style instructions being spoken aloud?
They may have been entered as ordinary narration. Keep spoken text in the speech block and sustained delivery instructions in the separate Style control.
Is QuestStudio Voice Lab the same as Google AI Studio?
No. Voice Lab is a separate narration option with its own available models and controls. Use the same script and acceptance checklist when comparing workflows.

