The voiceover is thirty seconds long, but the video only has twenty-four seconds available for speech. The title, reveal, and closing hold need time too. A timing budget solves that mismatch before you start accelerating the narrator.

Quick answer: Subtract silent visual holds from the total video length, estimate a draft word budget, then generate and measure the actual take. Cut redundant wording before removing meaningful pauses. Use small pitch-preserving tempo changes only after the script and performance are right.

Download the voiceover timing budget · Time a short narration take

Separate total duration from available speech time

A fifteen-second ad may need an opening visual, a product reveal, and a final card. If all fifteen seconds contain speech, the audience has no room to process those elements. Begin by reserving the moments that should remain quiet or carry only music.

For an illustrative thirty-second edit, reserve two seconds at the opening and four at the end. That leaves twenty-four seconds for narration. If a mid-video demonstration needs another two seconds without speech, the available voice budget falls to twenty-two.

These are editorial choices, not universal rules. A fast announcement and a reflective story have different needs. Write the allocations in a small table so everyone reviewing the cut can see why a technically thirty-second take still does not fit.

Use word count as a draft estimate

The basic estimate is words = available seconds × words per minute ÷ 60. At a hypothetical 135 words per minute, twenty-four seconds holds about fifty-four words. Actual speech varies with sentence shape, names, emphasis, pauses, voice, and language, so the estimate only tells you where to start.

Total editIllustrative visual holdsSpeech budgetDraft words at 135 WPM
15 seconds3 seconds12 seconds27 words
30 seconds6 seconds24 seconds54 words
60 seconds10 seconds50 secondsAbout 112 words

Do not write exactly to the estimated limit and assume the generation will land there. A short phrase containing a long product name may take longer than a longer phrase of simple words. Generate the real wording in the intended voice and use that measured take for the next decision.

Build a script with one job per sentence

For a short product video, one sentence can identify the problem, one can show the useful action, and one can tell the viewer what to do next. Repeating the same benefit in three different ways consumes time without adding information.

Here is an original sample for a fictional image-workflow demonstration. It is spoken text only, so it can be auditioned without accidentally reading performance notes:

Start with one clear product photo. Choose a background that fits the story, then check the edges and shadow. Keep the real label and color. Save the clean version before making another variation. One approved image is a better starting point than a folder of almost-right drafts.

Put tone and pace in the voice controls or a separate supported direction field. Read the script aloud once yourself. Any phrase you naturally stumble over deserves attention before you pay for repeated synthetic takes.

Measure the take you will actually use

Import the generated file into your editor and note the first audible word, the last audible word, and the full file duration. Leading silence, trailing silence, and a long final breath can make those measurements differ.

Listen without looking at the waveform first. Mark the phrase that feels rushed, the word that is unclear, and any pause that carries meaning. Then use the waveform to locate the relevant edit. A quiet region is not automatically disposable; it may separate two ideas or give the viewer time to see a result.

Record the voice, model, script revision, and measured duration. Keep the same conditions for the next take if you are testing a script cut. Changing the voice and wording simultaneously makes it harder to know what solved the timing problem.

Choose the right repair for the size of the mismatch

If the take is slightly long because of empty head or tail space, trim that space while protecting the first consonant and final decay. If it is long because the script repeats itself, edit the wording. If every phrase sounds slow despite a concise script, adjust the delivery or voice speed where supported.

Observed mismatchUseful first changeWhat to preserve
Long silence before the first wordTrim the unused headInitial breath or consonant when it belongs
Repeated explanationRemove a sentence or clauseThe actual benefit and required details
One awkward phraseRewrite and regenerate that passageMeaning and pronunciation
Small overall overrunCompare a modest tempo adjustmentNatural consonants and expression
Large overrunRewrite or change the edit lengthComprehension rather than every original word

Do not make the entire take faster to solve one slow sentence. A local replacement may fit better and preserve the comfortable pace elsewhere.

Understand speed percentage before changing it

To fit a twenty-seven-second passage into twenty-four seconds, the playback-speed factor is 27 ÷ 24 = 1.125. That means 12.5% faster playback. The duration becomes about 11.1% shorter; those percentages are different because they use different starting quantities.

Use a pitch-preserving tempo control when you want shorter speech without the pitch shift of simple speed-up. Audacity’s Change Tempo documentation describes changing duration without changing pitch and warns that more extreme processing can introduce audible distortion. Listen to the processed result; a correct formula is not a quality guarantee.

If a small speed change makes sibilants watery or consonants clipped, return to the script or regenerate with a different delivery. Do not keep processing an already processed file. Preserve the original take for comparison.

Keep provider controls and editing controls separate

A text-to-speech speed setting influences the generated performance. A tempo effect changes an existing recording. Those are different operations, and a numeric value in one system may not behave like the same number in another.

ElevenLabs’ pace documentation describes a Speed control and notes that extreme values can affect quality. That does not mean every QuestStudio voice model or interface exposes the same control. Use the options actually available for the selected model.

Likewise, punctuation and pause tags are model-dependent. If exact silence matters, place it explicitly in the editor after approving the speech. That makes the timeline reviewable and avoids asking a language instruction to serve as an exact clock.

Align picture and captions after the voice is approved

Once the narration fits, mark the meaningful words that the picture should support. If the line says “check the shadow,” show the shadow during that phrase. Cutting every three seconds regardless of meaning can make a technically synchronized video feel disconnected.

Generate or update captions from the final audio. If you change the tempo afterward, old timestamps may no longer match. The same applies to translated captions, on-screen highlights, and sound effects synchronized to a particular word.

Review the closing card with the final sentence in place. The voice should not finish so late that the viewer loses the next action, nor should the screen remain empty for several seconds because the script was cut without revisiting the visual plan.

Check the exported file, including its edges

Listen to the first word after export. A trim that looks precise can remove the attack of a consonant. Listen to the last word and any intended tail. Then verify the total duration in the exported file, not only in the project timeline.

Play the result on a phone at a normal listening level. If the sentence only makes sense because you already know the script, it needs more room or clearer wording. A fixed-duration deliverable still has to be understandable.

Keep one full-quality voice master, the accepted script, and the final mixed video. Name versions by purpose or revision so a future editor does not accidentally restore the overlong take.

How QuestStudio helps

Use Voice Lab to audition a short clean-text take and revise the words before generating the full narration. Keep the timing budget beside the script. Exact trimming, tempo processing, and frame-level synchronization happen in your audio or video editor.

The timing table and sample script are original editorial examples. They are not measured delivery rates for a particular voice model. The useful result is a final take that fits the real speech window and remains clear after mixing and export.

Frequently asked questions

How many words fit in a thirty-second AI voiceover?

It depends on the voice and delivery. At an illustrative 135 words per minute, thirty seconds is about 67 words, but a video with visual holds may have much less time available for speech.

Why does the same script have different durations?

Voice, model, punctuation, emphasis, and generation variability can change pacing. Measure the selected take instead of relying only on word count.

Should I speed up a long voiceover?

First remove redundant wording and unused silence. A modest pitch-preserving tempo change can help a small overrun, but listen for artifacts.

How do I calculate the required speed change?

Divide current duration by target duration. A 27-second take fitted to 24 seconds needs a 1.125 speed factor, or 12.5% faster playback.

When should I generate captions?

After the voice and its timing are approved. Recheck captions whenever you replace, trim, or change the speed of the narration.

To apply that direction in Google’s current speech workspace, use the Google AI Studio narration guide with separate spoken text, style controls, and a short revision test.