A voice nails one dramatic sentence, then sounds exhausted reading a simple instruction. Voice Design can create a fictional voice from a description, but the preview is only the beginning of casting. The better question is whether the voice works across the material you actually publish.

Quick answer: Describe the narrator's language, vocal texture, pacing, and role. Keep the design brief separate from the preview script. Compare candidates on the same neutral, emotional, and practical passages before saving one. QuestStudio offers a separate voice-selection workflow and does not currently expose ElevenLabs Voice Design.

Download the narrator audition worksheet · Audition a script in Voice Lab

Editorial guide checked September 13, 2026. Examples and worksheets are original planning material; the hero is an AI-generated illustration, not a model benchmark.

Cast a role before describing a sound

A narrator has an assignment. A museum guide should make unfamiliar objects understandable. A product demonstration needs clear instructions. A fictional travel diary may need intimacy and hesitation. Begin with that job, because a list of attractive vocal adjectives can produce a memorable voice that is wrong for the script.

ElevenLabs Voice Design creates a fictional voice from a description. Its current guide separates the voice description from preview text and returns three candidates to audition. Describe an original role instead of requesting an imitation of a recognizable person. If the assignment requires someone's actual voice, that is a different, authorized voice-cloning workflow.

Our example is a narrator for short videos about repairing and appreciating everyday objects. The voice should explain a practical step clearly, tell a brief personal story without melodrama, and sound consistent when the script contains a number or unfamiliar name. Those are three different demands; one polished preview cannot settle all of them.

Separate identity from performance

Voice identity includes qualities you want to carry across videos: language, regional accent, perceived age, resonance, texture, and general pace. Performance is what the speaker does in a particular sentence: a pause, a brighter greeting, a quieter admission, or emphasis on an instruction. Keep the stable characteristics in the design brief.

DecisionExample for this narratorWhere it belongs
Language and accentEnglish with a general American accentVoice description
Vocal characterAdult, warm midrange, lightly textured, clear consonantsVoice description
RolePatient host explaining familiar objectsVoice description
One emotional turnA restrained admission in a personal storyThat passage's delivery direction
Exact wordsThe sentence viewers should hearPreview or speech text

Avoid defining the whole identity as “whispering dramatically” unless every intended use really requires that behavior. It may make the first sample compelling and the ordinary instructions tiring. Similarly, asking for a telephone or tape sound can confuse a vocal quality with a recording effect. Cast the voice cleanly; add production effects later when the project needs them.

Write a brief you can evaluate

The description below deliberately avoids a celebrity reference and a long list of incompatible traits. “Clear consonants” and “unhurried pace” are useful because the audition can reveal whether they are present. “The best voice ever” gives you no acceptance criterion.

An original adult narrator speaking English with a general American accent. Warm midrange resonance, a lightly textured natural tone, clear consonants, and an unhurried conversational pace. The character is a patient host who explains how everyday objects work, with understated curiosity rather than a sales delivery. Clean, close studio recording without room echo, background noise, or artificial distortion.

Enter this in the voice-description field, not as spoken copy. In the separate preview field, use a passage from the kind of work you plan to publish. ElevenLabs' documentation notes that preview text influences the resulting performance. A theatrical monologue is therefore a poor sole audition for a voice intended to read calm instructions.

Record the description, preview text, and available settings for each generation. If you change the accent request and the preview script together, you will not know which change caused a different impression. Keep one stable audition while comparing revisions to the brief.

Use three passages with different demands

These original passages form a small audition pack. They are not examples of audio already generated or scored. Use the same words across candidates, then listen without looking at the candidate names. For a longer production, add a representative paragraph from the real script before making the final selection.

Practical instruction

Place the empty frame on a soft cloth. Count the four tabs along the back, then turn them gently outward. Keep the small pieces together so they are easy to find when you put the frame back.

Listen for a natural separation between the actions. “Four tabs” should be understandable without checking the transcript. The host should sound helpful, not impatient, and the final sentence should not accelerate as if the model is trying to finish within an imaginary time limit.

Quiet personal story

I nearly gave the little radio away. Then I found my grandfather's pencil mark underneath it, beside the date he repaired the speaker. I decided it could stay on the shelf a little longer.

This tests whether the voice can suggest feeling without turning the story into a trailer. A small pause may help the discovery land. A sob, dramatic whisper, or exaggerated tremble would change the character of this particular assignment.

Names, numbers, and contrast

For the Luma Vale display, we compared twelve small frames and selected three. The widest was not the winner. The narrow oak frame gave the photograph more room to breathe.

This passage tests a fictional two-word name, a number, and a contrast. Decide how “Luma Vale” should be pronounced before judging the audio. A candidate that handles the name differently is not automatically worse; it needs to match the project's approved pronunciation.

Choose with evidence from the script

Label the candidates A, B, and C. Match their approximate listening level so the loudest sample does not win by default. Rate intelligibility, role fit, consistency across passages, and distracting artifacts on a simple one-to-five scale. Add a timestamped reason for every unusually high or low rating.

Do not average away a critical failure. A voice that sounds excellent in the story but makes the key instruction unclear is not ready for this job. Treat essential-word intelligibility as a gate, then compare style among the candidates that pass.

Ask a second listener what the instruction told them to do without showing the script first. This is a practical comprehension check, not a scientific benchmark. If both listeners need the transcript to understand an ordinary sentence, investigate the voice, pronunciation, or script before committing to it.

For recurring narration, revisit the preferred candidate after a break and play a longer passage. A texture that is interesting for ten seconds may become distracting over several minutes. The relevant question is whether the voice supports attention through the intended runtime.

Revise the right layer

All passages sound too promotional: revise the stable role description and remove advertising language. Test the same audition again. Do not try to counter a permanently sales-oriented identity with a new negative instruction on every sentence.

Only the story is overacted: adjust that passage's delivery or punctuation before abandoning the voice. The issue may be performance rather than identity. The Eleven v3 audio-tag guide addresses direction for a compatible model; tags are not a universal control across all speech engines.

One name is wrong: use the provider's supported pronunciation approach and verify the result in context. Do not distort every sentence in the design brief to fix one proper noun. Keep the approved spoken form with the production script.

The voice changes across scripts: check synthesis model and settings as well as text. A saved voice does not make every synthesis configuration equivalent. Record the combination used for the accepted audition before rendering the rest of the project.

Run a separate audition in Voice Lab

QuestStudio does not currently expose ElevenLabs Voice Design or an import for externally designed voices. You can use the same three-passage method with the available voices in Voice Lab. Choose a suitable language and model, then test the script before creating the full narration.

Save the selected voice, model, settings, approved pronunciation, and accepted sample together. The downloadable audition sheet provides a place for those decisions. A successful casting session ends with a narrator that works on the real material and a clear reason for choosing it.

Frequently asked questions

Is Voice Design the same as voice cloning?

No. Voice Design creates a voice from a description. Cloning uses authorized recordings of a particular voice.

Why does my designed voice sound different on another script?

The words, emotion, sentence structure, and synthesis settings can change the performance. Audition several representative passages.

Should the voice description be read aloud?

No. Keep design instructions in the description field and the spoken passage in the preview or speech field.

Can I import an ElevenLabs designed voice into QuestStudio?

The reviewed QuestStudio interface does not expose that import. Use the available Voice Lab voices or the external provider's own workflow.