If an emotion tag seems to do nothing, the first problem may be the model version. Fish Audio’s older S1 examples use parentheses, while its S2 family uses square brackets and natural-language descriptions. Copying a tag list without that distinction makes voice direction unnecessarily confusing.

Quick answer: Confirm the selected model, use its matching syntax, and compare a plain reading with one emotional cue. Judge whether the listener understands the intended feeling without losing the words or timing.

Download the emotion audition sheet · Audition your plain narration

Model documentation checked September 14, 2026. The packing-scene script and audition plan are original examples, not recorded Fish Audio test results. The hero is an AI-generated illustration.

Resolve the syntax before directing the voice

The current models overview identifies s2.1-pro as the recommended production model. S2.1-Pro and S2-Pro accept square-bracket descriptions; S1 uses parenthesized controls. The S2 descriptions are interpreted as natural language rather than a fixed catalog of dedicated control tokens.

The emotion best-practices page explicitly says its parenthesis examples apply to S1. Its S1 rules should not be copied indiscriminately into an S2 project. Record the model name beside the script so the next person editing it knows which convention is intentional.

Model familySyntax to verifyPractical consequence
S1Parentheses, such as an approved S1 emotion markerUse the S1 guide’s documented controls and placement rules
S2-Pro / S2.1-ProSquare brackets with a delivery descriptionWrite a clear cue and audition how the model interprets it

Keep the spoken transcript separate from the performance script. The transcript contains only words the audience should hear. The performance script can include supported cues. This prevents direction such as “tired but relieved” from leaking into captions or a later translation.

Direct a reason for the emotion

Our original scene is simple: two friends have finished packing a studio after a long day. One narrator says, “That’s the last box. We can go home.” The desired feeling is tired relief. The line should not sound like a high-energy advertisement or a distressed emergency.

Write the intention in ordinary language before selecting a tag. What changed for the speaker? The work is complete. What do they want the listener to do? Relax and leave. This gives the delivery a purpose that a broad label such as “happy” may not express.

You can draft and audition a plain narration in Voice Lab while refining the words. Fish Audio’s model-specific bracket syntax belongs in the Fish workflow; this article does not imply that those controls are implemented in QuestStudio.

Make three small audition candidates

Keep the voice, model, words, and available generation settings stable. First produce the untagged line as a baseline. Then try one cue describing relief. Finally try one cue that adds a restrained sense of fatigue. These are proposed experiments; the results need listening review.

Plain transcript:
That's the last box. We can go home.

S2-family audition A:
[relieved] That's the last box. We can go home.

S2-family audition B:
[tired but relieved] That's the last box. We can go home.

The second cue is an original natural-language direction, not a named guaranteed preset. If it produces a useful delivery, save the exact wording and model version. If it does not, simplify the cue before stacking more instructions around it.

Avoid asking for relief, excitement, sadness, whispering, laughter, and a gasp in the same short line. Even when a system accepts the text, the performance may become crowded. One clear emotional turn is enough for this scene.

Listen for meaning before decoration

Compare the candidates at similar playback levels. Listen once without reading the cues and write the feeling you hear. If you have a second listener, give them neutral filenames so the label does not tell them what they are supposed to perceive.

Then check the words. The final consonant in “last” and the beginning of “box” should remain understandable. A beautifully weary delivery that swallows the key noun is not the right take for a short explanatory video.

Finally inspect the phrase shape. The first sentence resolves the task; the second releases the tension. If both sentences rise as questions, the speaker may sound uncertain that the work is finished. That can be an interesting performance for another scene, but it does not meet this brief.

Change wording only when the script causes the problem

Sometimes the emotional cue is clear and the sentence is awkward. Read the line aloud yourself. Does “We can go home” naturally follow the setup, or does the scene need a name, a pause, or a shorter phrase? Fix the script for a specific reason, then begin a new comparison.

Do not quietly rewrite one candidate and keep calling the exercise a tag comparison. A different sentence can produce a different rhythm regardless of the cue. Save the earlier script and label the revision so you know what changed.

Use punctuation to clarify the intended sentence structure. Do not fill the text with repeated ellipses merely to force slowness. A readable script is easier to direct, caption, translate, and revise later.

Diagnose the common failures

The cue seems ignored: confirm the model family and syntax, then try a shorter unambiguous description. Compare against the plain baseline before concluding that no change occurred. A subtle shift may be appropriate for the scene.

The delivery becomes theatrical: reduce the cue’s intensity or remove the extra emotional adjective. The narrator has finished packing a room; they do not need to sound as if they have survived a disaster.

A cue appears in the spoken output: check that the selected tool and model actually interpret that syntax. Remove the cue, generate a plain line, and consult the matching interface guidance. Do not keep adding punctuation around an unsupported instruction.

The emotional change damages timing: decide whether the picture can allow a longer delivery. If not, simplify the wording or choose a less extended performance. Compressing the whole take aggressively may remove the natural rhythm you were trying to create.

Approve the line inside the scene

Put the selected take under the actual picture and any background sound. Relief that is obvious in headphones may disappear under loud music. A breath that sounds natural on its own may land awkwardly across an edit.

Keep the spoken words, emotional intention, duration, and final mix as separate approval notes. This makes a later pickup easier: you can ask for the same restrained relief without guessing which of several unrelated tags created it.

For a series, keep a compact direction sheet for the recurring narrator. Describe the ordinary delivery and the few emotional departures the project uses. A consistent voice with meaningful variations usually serves the audience better than a different extreme performance in every sentence.

Frequently asked questions

Should Fish Audio emotion tags use brackets or parentheses?

Use the syntax for the selected model. Fish documents parentheses for S1 and square-bracket natural-language descriptions for the S2 family.

Is “tired but relieved” a guaranteed preset?

No. It is an example of a natural-language direction for an S2-family audition. Listen to the result and simplify or revise it as needed.

Should I put the cues into my captions?

No. Build captions from the approved spoken transcript. Keep performance direction in a separate script unless an accessibility caption intentionally describes an audible event.

Can I paste Fish tags into QuestStudio Voice Lab?

Do not assume they will be interpreted. Use Voice Lab for its supported narration controls and Fish Audio for the model-specific syntax described here.