The voice says “warm and reassuring” before it reads your opening line. Then it announces “pause” in the middle of the sentence. Your script contains instructions for a performer, but the speech tool has treated them as words to say. The first fix is to separate the recording script from the production brief.

Quick answer: Make a clean version containing only the words the audience should hear. Keep stage directions in separate notes or documented controls. Add model-specific tags only after confirming that the selected model and interface support them. Test one short passage and check every unexpected word before rendering the full script.

Download the spoken script and direction preflight · Test a clean passage in Voice Lab

Separate words from instructions before changing voices

A script written for a human actor often mixes dialogue, timing, speaker labels and notes. A plain text-to-speech field may have no reliable way to know which parts are instructions. Parentheses, brackets and capital letters are not universal control languages.

Create two versions of the script. The production version keeps everything the team needs: speaker assignment, timing, pronunciation notes and performance direction. The spoken version contains the actual narration. Copy from the spoken version into a plain speech field.

This does not mean every parenthetical phrase should be removed. “The blue cable, which is optional, connects here” contains real information for the listener. “Pause while the cable is connected” may be an editing note. Decide based on meaning rather than deleting every phrase between punctuation marks.

Sort the script with a three-column check

For this proposed exercise, imagine a short assembly narration. The draft says: “NARRATOR: [calm] Place the tray on the shelf. (Wait for the close-up.) Use the shorter screw, not the longer one.” The clean spoken version should retain the assembly instruction and the important contrast.

Draft elementWhere it belongsWhat the listener should hear
NARRATOR:Speaker assignment outside a plain speech fieldNo speaker label unless it is intentionally part of the narration
[calm]Production note or a documented supported controlThe intended delivery, not the word calm
Place the tray on the shelf.Spoken scriptThe complete instruction
(Wait for the close-up.)Edit timing or a documented pause controlA pause if needed, not the production sentence
Use the shorter screw, not the longer one.Spoken scriptThe contrast that helps the viewer choose correctly

The last row matters. An automatic cleanup that removes everything after “not” or strips explanatory asides can change the message. Keep a human-readable comparison between the approved script and the cleaned version. Script cleanup is also a meaning check.

Mark the expected first and last spoken words for each passage. If the result starts with “narrator” or ends with “end scene,” you can identify the failure immediately without debating whether the voice sounds natural.

Confirm the control language for the actual model

Different systems interpret directions differently. Narakeet’s narration documentation says its stage directions need their own paragraph with blank lines around them; otherwise the instruction can be read aloud. That is a specific parser rule for that workflow, not a general rule for all AI speech tools.

ElevenLabs’ current audio-tag help documents square-bracket tags for specified supported models. It does not make arbitrary bracketed prose a portable instruction format. Confirm the selected model and the interface that sends the text, not merely the provider name.

Likewise, SSML support is not automatic. Some tools parse a defined set of XML speech tags; others accept plain text or a different direction field. Pasting a break tag into an unsupported input can fail, be ignored or become unwanted spoken material.

If a wrapper, template or automation sits between your script and the voice model, inspect the visible text it is about to send where the interface allows that. A correct production note can become a problem if the wrapper concatenates it with the narration. Fix the input boundary before searching for a more expressive voice.

Start with a clean baseline passage

Place the tray on the shelf. Use the shorter screw, not the longer one. Turn it gently until the tray stays level.

The block above is spoken copy only. It intentionally contains no performance notes, speaker headings or hidden instructions. Adapt the words to your project, then generate a short test with the voice and model you intend to use for the full recording.

Listen for literal correctness first. Did every intended sentence appear? Were there added introductions, spoken labels or missing words? Compare against the script instead of judging from memory. Only after that pass should you assess pace and delivery.

If the plain passage works, add one documented control if the project needs it. Compare that version with the baseline. When a direction suddenly becomes spoken, you have a smaller compatibility question to investigate than a full page of mixed instructions.

If the clean baseline still introduces words absent from the submitted text, retain the exact passage, model and result. Check visible presets or templates, try a shorter passage and consult the tool’s current documentation. Do not assume script punctuation alone explains an addition that was never in the input.

Handle pauses and speaker labels deliberately

A request such as “two-second pause” should not appear in plain narration unless the audience is meant to hear those words. Use a documented pause mechanism where supported, or place the pause in the audio editor. Splitting paragraphs may affect phrasing, but it does not guarantee an exact duration.

For dialogue, keep speaker names in the interface’s speaker assignment controls when available. If the tool has only a single plain text field, consider generating separate clean passages for each speaker, then assembling them in an editor. Check the joins for abrupt changes in level, pacing and room character.

Some projects intentionally say names aloud, such as an interview introduction or an educational reading of a play. Preserve those names in the spoken version. The goal is not to remove all labels mechanically; it is to make every audible word intentional.

Keep timing notes attached to the correct passage in the production document. Removing them from speech input should not erase the editor’s instructions. A clean recording script and a useful production brief serve different jobs.

Avoid cleanup that silently changes the message

Do not strip parentheses across an entire document without reviewing what they contain. A pronunciation expansion, an alternate product name or an accessibility description may be intended speech. Rewrite it into a natural sentence if necessary rather than dropping the information.

Also inspect Markdown headings, bullet labels, code fences and quotation wrappers copied from a writing assistant. “Here is your final script” is useful conversational framing in a chat, but it usually does not belong at the start of a voiceover. Copy the approved words, not the surrounding response.

For longer projects, keep a small list of recurring contaminants: scene headings, camera notes, pronunciation reminders and export labels. Search the draft for those patterns and review each match. A pattern can flag a likely mistake; it cannot decide the intended meaning on its own.

After cleaning, read the spoken version once from beginning to end. This catches missing transitions that were previously carried by stage directions. If the listener needs to know the scene changed, add an actual spoken transition instead of hoping a silent production note will communicate it.

Approve wording before polishing performance

Use two listening passes. In the first, follow the script word by word and mark additions, omissions and literal control text. In the second, hide the script and judge whether the delivery is clear for a listener. This keeps a pleasant voice from masking a wording error.

When a short passage passes, apply the same preparation to the next section. Do not insert a new tag format halfway through a long recording without testing it. Save the clean script version alongside the accepted audio so later revisions begin from a known input.

If the words are now right but the meaning still sounds flat, move to emphasis or pacing work. That is a separate problem from stage directions being spoken. Keeping those decisions separate makes each retry easier to evaluate.

How QuestStudio helps

Use Voice Lab for a short clean-script test with the selected model and voice. Start with spoken words only. Check the controls available for that model before adding any markup; this guide does not claim that every Voice Lab model shares a tag parser or automatically removes production notes.

Keep the worksheet with the script and accepted audio. Once literal wording passes, refine one delivery choice at a time. You should be able to explain why every word in the final recording is there.

The examples are proposed production exercises, not measured tool tests. The original hero image is an AI-generated editorial illustration.

Frequently asked questions

Why does an AI voice read words inside brackets?

The selected model or input field may treat them as ordinary text. Brackets are only controls when the specific system documents and supports that syntax.

Should I delete every parenthetical phrase?

No. Some parenthetical material is intended information. Separate production instructions from spoken meaning and review each change.

Do blank lines make stage directions work everywhere?

No. Some tools define whitespace-sensitive syntax, but that behavior is specific to the tool. Check the actual model and interface documentation.

Can I use SSML in any text-to-speech field?

No. SSML support and supported tags vary. Use plain spoken text unless the selected tool explicitly supports the markup you need.

What is the fastest useful test before a long voiceover?

Generate a short clean passage with the intended model and voice, compare every spoken word with the script, then add at most one documented control for a second comparison.