A key turns, a door opens, and the soundtrack tells us none of it. Audio description fills that information gap. The challenge is choosing the visual facts a listener needs and delivering them before the next important sound.
Download the description timing sheet · Voice an approved description line
Editorial guide checked September 13, 2026. Examples and worksheets are original planning material; the hero is an AI-generated illustration, not a model benchmark.
Listen before you describe
Audio description makes important visual information available through sound. It is different from captions, a video summary, or a transcript of the dialogue. Before writing, listen to the existing soundtrack without watching. Note what the words and sounds already explain, and what remains impossible to understand from audio alone.
Then watch the video and compare. A spoken instruction might already explain that a person unlocks the door. In that case, repeating the same action could add clutter. But if the soundtrack contains only music, the key, lock, and opening door may need description. The decision depends on the content, not on whether the frame contains many objects.
The W3C description guidance distinguishes integrated description, a separate described video, and supported separate tracks or files. Choose a delivery approach early. A voice file sitting in a project folder does not help a viewer who cannot find or play it.
Build a small visual-information ledger
Use this original fictional scene as a planning exercise. A person named Mara stands at a blue door. She turns a brass key marked by red thread. The door opens to an empty room, and she places the key on a shelf. A later line refers to “the marked key,” making the thread relevant to the story.
The scene runs thirty seconds in this example. Its exact timing is invented for the exercise; it is not an analyzed production clip. Separate what the soundtrack provides from what the description needs to add. That separation keeps the writing grounded in the scene instead of becoming a list of everything visible.
| Example interval | Existing sound | Visual information to consider |
|---|---|---|
| 0–4 seconds | “This is the place,” Mara says | Blue door and key in hand |
| 4–8 seconds | Music, then a lock click | Red thread marks the key |
| 8–12 seconds | Door creak | Room beyond is empty |
| 12–18 seconds | Dialogue about waiting | No essential new action |
| 18–23 seconds | Quiet room tone | Key placed on shelf |
| 23–30 seconds | “Leave the marked key here” | Earlier action explains the reference |
The ledger reveals why red thread matters while the precise grain of the door may not. It also reveals a possible production issue: if the line about leaving the key comes after the action, the sequence should make that relationship understandable. Description should communicate the edit faithfully rather than silently rewriting the story.
Write a line that earns its time
For the opening listening window, try the nine-word line below. It supplies the person, the object, and the identifying feature. Whether it fits depends on the actual generated delivery and the timing of the lock click. Word count is a planning aid, not a substitute for listening.
If the take runs into the lock click, try “Mara turns the key with red thread.” The shorter wording preserves the identifying feature. If it still crowds the sound, revise the timing or introduce the key earlier. The better line is the one that communicates naturally inside the real window, not the one with the most compressed wording.
For the next window, “The door opens onto an empty room” adds information the creak alone cannot provide. For the later quiet interval, “She leaves the key on a narrow shelf” locates the object. If the shelf matters later, preserve that fact; if its width does not, remove “narrow.” Each revision should answer what a listener needs at this moment.
Avoid describing Mara as frightened merely because her hand pauses. The image may support a visible hesitation, but the emotion is an interpretation. Choose observable action or information established by the story. The DCMP Audio Description Tip Sheet emphasizes observable information, intelligible delivery, and protecting essential audio. Also avoid explaining what will happen next before the scene reveals it.
Audition the approved words
Use Voice Lab to generate a reviewed line, then play it against the scene. The tool supplies speech from text; it does not inspect the video or decide which visual facts belong in a description. Keep the script and the generated take linked so a later edit cannot silently change the meaning.
For this example, audition a calm, clear voice that can be distinguished from Mara’s dialogue. Listen to the proper name, the phrase “red thread,” and the final consonant of “key.” A smooth voice that blurs those words is less useful than a plainer delivery that remains understandable.
Measure the spoken duration after generation. Leave room for the line to finish without colliding with the lock click or the next dialogue. If it does not fit, revise the sentence or the editorial timing. Increasing speed until every word fits may defeat the purpose by making the information difficult to follow.
Mix for a listener following the scene
Place each line on the timeline and listen through the surrounding section. The description should be understandable, but the original dialogue and meaningful effects still need space. In our example, the lock click confirms the action and the door creak establishes the opening. Do not bury both beneath an unnecessarily long sentence.
If you lower background music during the description, use transitions that do not call attention to the volume changes. Check the spoken words on ordinary speakers as well as headphones. Keep a separate copy of the original mix so revisions remain reversible and the described version can be compared against it.
When the scene has no usable listening window, reconsider the delivery method. For a new production, the main narration may be able to explain the action naturally. For an existing dense sequence, a separately edited described version may need additional time. Do not claim that simply adding an AI voice establishes compliance with an accessibility standard.
Review information, timing, and delivery separately
First, verify every described fact against the picture. Second, listen without looking and note whether the key’s identity and location remain clear. Third, inspect the final player. These passes catch different failures: an invented fact, an unintelligible mix, or a track that exists but cannot be selected.
Where possible, include feedback from blind or low-vision listeners. Ask about specific comprehension points rather than whether the voice sounds professional. For this scene, useful questions are which key Mara uses, what is beyond the door, and where she leaves the key. Record confusion before deciding whether the fix belongs in wording, timing, or delivery.
Make the described version discoverable
YouTube’s audio-description instructions say availability depends on access to multi-language audio. They describe adding descriptive audio under Languages after an original or dubbed track exists in that language. The uploaded file should use a supported audio format and roughly match the video’s duration. Check your account’s current options.
Test what the selected track actually plays. A viewer needs the scene’s audio experience, not isolated description lines with silence where dialogue should be. If you provide a separate described video, label and link it clearly, and verify the exact exported version. The worksheet preserves the script, windows, review notes, and delivery location so the useful work survives the final upload.
Frequently asked questions
Is audio description the same as captions?
No. Captions convey speech and relevant sounds visually; audio description conveys important visual information through sound.
Should I describe every object on screen?
Describe information needed to understand the content. Avoid crowding dialogue with decorative details that do not help the listener.
What if the description cannot fit between lines?
Revise the wording, integrate the information into a new production, or make an extended described version. Do not simply rush the voice.
Can every YouTube creator upload a separate description track?
YouTube says access is tied to multi-language audio availability. Check your account and provide a clearly labeled described version when a separate track is unavailable.

