The narration sounds clear on its own. The music sounds good on its own. Together, the words disappear. This is a mixing problem, and the answer is rarely a universal “set music to ten percent” rule. Start with the voice, then decide where the music belongs around it.

Quick answer: Keep voice and music on separate tracks, establish a clear narration level, and add the music from silence upward. Lower the bed across spoken phrases and let it recover deliberately in larger gaps. Check the exported mix on headphones and an ordinary speaker at modest volume.

Download the voice and music mix map · Draft a sparse music bed

Keep the voice and music separate for as long as possible

Start with a clean narration file and a separate music file. If the only source is a finished mix, turning down the whole file also turns down the voice. Source separation may help, but it can introduce artifacts and does not guarantee a clean reconstruction.

Preserve the original files before processing. Name the narration by script revision and the music by cue or take. A folder containing several files called “final” makes it easy to mix the wrong voice version and discover the mismatch after captions have been prepared.

Our worked example is a thirty-second product explainer with two spoken passages, a short demonstration gap, and a closing card. The music should connect those parts and leave the words in front. It does not need to behave like a full song with a dominant chorus.

Choose a bed that leaves space before touching a fader

A dense lead melody, a vocal hook, or bright percussion can compete with narration even at a low level. Choose an arrangement whose musical role fits the scene. Sparse accompaniment is easier to place under speech than a track that demands attention on every beat.

For the explainer, use a soft rhythmic texture with a restrained bass part and no sung words. Keep the pulse steady and avoid dramatic impacts at arbitrary points. If a product reveal needs an accent, plan that accent against the picture rather than accepting whatever arrived in the generated music.

Create a restrained instrumental background bed for a clear spoken product explainer. Soft muted percussion, warm simple chords, a light steady bass pulse, and plenty of space between musical gestures. No vocals, no spoken words, no dominant lead melody, no sudden impacts or dramatic build. Keep the arrangement even and supportive so narration remains the main focus.

This describes a musical intention. Inspect the actual output for vocals, sudden entrances, and changes in density before using it under a voice.

Establish the narration level first

Listen to the voice alone. Correct a much quieter sentence, a distracting click, or a clipped word before adding music. A background bed can conceal defects during casual review, but it cannot restore missing speech.

Use your editor’s meters to watch for clipping and your ears to judge intelligibility. Peak level, perceived loudness, and clarity are different things. A file can avoid clipping while still sounding uneven or difficult to understand.

Once the voice is comfortable, leave it stable while introducing the bed. Start with the music muted and raise it until it supports the mood. Then back it down if you begin following the melody instead of the sentence. This establishes a useful relationship for the actual material.

Why a fixed percentage is not a mix recipe

Two music files at the same fader percentage can have very different loudness, instrumentation, and frequency balance. A quiet ambient track and a heavily limited pop track do not become equivalent because both are set to ten percent.

A numerical offset can be a starting comparison within one project, but the pass condition is intelligible speech in context. Keep a note of your starting settings, then adjust after listening to the quietest phrase and the densest part of the music.

Do not solve every collision by boosting the voice. That can make loud words harsh while the background still masks softer consonants. A sparser bed, a local music reduction, or a modest tonal adjustment may solve the actual conflict more cleanly.

Draw a phrase-based music map

Music does not need to jump up during every breath. Mark larger spoken passages and meaningful gaps. For the fictional explainer, the demonstration gap is a planned opening for the bed; the half-second pause inside a sentence is not.

IntervalPicture and voiceMusic decision
0–2 secondsOpening image, no speechGentle introduction
2–11 secondsFirst explanationStable lower bed across the whole passage
11–14 secondsSilent product demonstrationSmall deliberate recovery
14–25 secondsSecond explanation and next stepReturn below the voice before speech begins
25–30 secondsClosing cardControlled finish and fade

These timestamps are an original planning example. Build your map from the approved narration and picture. If the voice changes, revisit the music automation rather than assuming the old dips still land correctly.

Use manual envelopes for a short edit

For a thirty-second piece, a few volume keyframes can be easier to control than an automatic detector. Place a point before the spoken passage, lower the bed by the time the first word arrives, hold it through the phrase, then recover gradually after the last word.

Audacity’s narration-and-music tutorial describes manual envelope editing and automatic ducking as two options. Manual automation is useful when you want to decide exactly which pauses allow the music to return.

Listen across each move instead of previewing only the low section. A sudden dip can draw attention to itself. A very slow dip can allow the music to cover the first words. Shape the transition around the actual voice entrance.

Use automatic ducking as a first pass for longer narration

Automatic ducking reduces the music in response to another track. In Audacity’s documented setup, select the music and leave the synchronized voice control track unselected underneath it. The effect uses that control signal to decide when to reduce the selected material.

Adjust the amount of reduction, detection threshold, pause behavior, and fade timing for the recording. If the bed rises between every few words, extend the behavior that keeps it down across short pauses. If quiet words fail to trigger a reduction, inspect the voice level and detection setting.

Adobe’s Premiere ducking workflow uses assigned audio types and generated keyframes. Those keyframes can be refined, but generating them again replaces manual changes. Save a version before recalculating an edit you have already adjusted.

Diagnose pumping, masking, and sudden accents separately

Pumping is the distracting rise and fall of the bed around small speech gaps. Reduce unnecessary level changes or lengthen the recovery behavior. The music should follow the structure of the narration, not advertise every breath.

Masking happens when competing sound makes words harder to hear. Try a quieter or less busy arrangement first. If you use EQ, make a focused adjustment and compare at a similar overall level. A broad aggressive cut can make the music thin without solving the most difficult word.

Sudden accents may need local editing. A loud cymbal hit under a product name can remain distracting even when the rest of the bed sits well. Move, lower, or replace that event instead of turning down the entire soundtrack until it loses its purpose.

Review with ordinary listening conditions

Play the mix at a modest level through a phone or ordinary speaker. Can you understand the narration without reading along? Then use headphones to inspect clicks, rough edits, stereo distractions, and exaggerated breaths.

Compare voice-only and mixed versions. The mixed version should add mood without requiring more effort to follow the sentence. Ask another person to repeat the key message if the wording is especially important. A listener unfamiliar with the script can reveal a problem you have learned to ignore.

Check the final combined output for clipping and the delivery requirements of the destination. There is no single loudness number that this article can prescribe for every platform or client. Master the actual combined mix to the required specification, preserving intelligibility.

How QuestStudio helps

Draft a restrained instrumental in Music Lab and generate or revise the narration in Voice Lab. Download accepted files separately, then perform the mix and automation in an audio or video editor. Automatic ducking described here belongs to those external editors.

Keep the final cue map, source files, and exported mix together. If the script changes next week, you can update the affected passage without reconstructing every music decision. The cue map and prompts are original editorial examples; the hero is an AI-generated studio illustration, not a screenshot or measured audio test.

Frequently asked questions

How loud should background music be under AI voiceover?

Low enough that every word remains easy to understand in the intended listening conditions. The right setting depends on the actual voice and music, so a fixed percentage is not reliable.

What is audio ducking?

It is a reduction in background level while a foreground signal such as speech is present. It can be drawn manually or generated automatically.

Why does the music rise between every word?

The recovery behavior may be too fast or the detector may treat short gaps as the end of a passage. Keep the bed down through coherent phrases.

Can I duck music after it has been mixed into the voice?

Separate tracks are much easier to control. A finished combined file may require source separation, which can add artifacts and may not fully isolate the voice.

Does QuestStudio automatically mix and duck these files?

This guide uses QuestStudio to create source audio and an external editor for the detailed mixing and ducking workflow.