The sentence “I thought you had the key” can sound like a practical question, a worried realization, or a quiet accusation. An Inworld TTS-2 tutorial is most useful when it helps you direct that difference without losing the words. This guide builds three readings of the same line and explains which choices belong to authored narration versus live conversation.

Quick answer: Keep the voice and spoken text fixed, place a concise direction at the start of the input, and compare the resulting intention, clarity, and timing. Test conversational adaptation separately; a standalone text-to-speech request does not automatically contain the preceding audio exchange.

Download the three-read voice direction sheet · Test a short narration script

Documentation checked September 16, 2026. The scripts are original proposed exercises, not measured Inworld outputs or latency benchmarks. The hero is an AI-generated illustration.

What is new in Realtime TTS-2

Inworld’s August 31, 2026 release announcement describes Realtime TTS-2 and natural-language voice direction. It distinguishes authored performance from a realtime conversational path that can use preceding audio context. Those are related capabilities, but they require different tests.

The current TTS documentation is the source for supported models and request options. Use its current model identifiers and parameters if implementing an API integration; older tutorials can refer to a different release or access stage.

For a recorded explainer, your main concern may be a repeatable reading of approved text. For an interactive character, the concern includes whether the reply responds appropriately to the person who just spoke. Do not use one polished narration clip as evidence that a conversational system works.

Define three intentions for one sentence

Our scene takes place outside a locked rehearsal room. Two fictional performers have arrived early. The spoken text stays exactly the same: “I thought you had the key.” The difference is what the speaker wants from the listener.

ReadingIntentionWhat to avoid
PracticalCheck who can open the doorAn accusation that starts an unintended argument
WorriedRecognize that rehearsal may be delayedPanic too large for the situation
DisappointedExpress a small breach of expectationSarcasm that changes the character relationship

Keep the selected voice, text, and other settings unchanged across the initial comparison. If you change the voice between readings, you cannot tell whether the difference comes from the direction or the casting.

Write the intended listener response too. The practical reading should invite an answer. The worried reading should invite reassurance or a solution. The disappointed reading should make the other character notice the expectation. That gives reviewers a concrete way to describe what they heard.

Put concise direction before the spoken text

Inworld’s release documentation describes bracketed natural-language direction at the beginning of the input. Use a short, actionable instruction. Avoid a paragraph of competing emotions that asks the model to be calm, furious, amused, and exhausted at once.

[Speak matter-of-factly, as a practical question to a colleague. Keep the tone friendly and the pace natural.] I thought you had the key.
[Sound quietly worried as you realize the door may stay locked. Keep the concern restrained, without panic.] I thought you had the key.
[Express mild disappointment to someone you trust. Keep the line soft and direct, without sarcasm.] I thought you had the key.

These are proposed inputs for the supported Inworld workflow. They are not universal syntax for every speech provider. If you use another tool, translate the intention into its documented controls instead of assuming bracketed text will be interpreted as direction.

Listen for whether the instruction itself is spoken, whether the sentence changes, and whether the intended emotion is audible. A text field accepting the input does not prove the model followed it.

Review intention, words, and timing separately

First, listen without looking at the labels and name the intention you hear. If the practical take sounds accusatory, record that observation before adjusting anything. The purpose is to hear the output, not to confirm the prompt you wrote.

Second, transcribe the sentence as heard. Check every word, especially short function words that can disappear in soft speech. For longer production scripts, include names, dates, and numbers in this pass. Expressiveness does not compensate for incorrect information.

Third, measure the usable duration from the first audible sound to the end of the final word. Include meaningful breaths when judging the edit. A worried pause can help the scene but may not fit a tightly timed interface response.

Finally, listen inside the actual mix or scene. Background music can mask a soft ending. An isolated reading that feels subtle may become inaudible when placed over a busy soundtrack. Preserve clarity without automatically making the character more dramatic.

Add nonverbal sounds only when they do a job

Inworld documents nonverbal controls such as laugh, sigh, and breath. Use them sparingly and verify their behavior with the current model. A sigh can establish frustration, but it also adds duration and can make a neutral line sound judgmental.

Try the scene without an added nonverbal cue first. If the intention already reads clearly, another breath or laugh may only draw attention to the performance. Add one cue when you can explain the story information it contributes.

Keep a clean version without the cue. During editing, the director may decide that the pause works but the sigh does not. A simpler alternate is useful production material, not a failed take.

Test realtime context as its own experiment

Now imagine the preceding speaker says, “I checked every pocket. It’s gone.” In one version, they sound calm; in another, they sound close to tears. A context-aware response may need to react differently even if the reply text remains similar.

To evaluate that behavior, use the documented realtime path that actually supplies the prior audio context. An ordinary request containing only the reply text cannot reveal whether the system understood an audio turn it never received.

Keep the words of the preceding line fixed while changing its delivery, and record both input audio and output. Compare whether the response remains intelligible, fits the situation, and avoids an inappropriate emotional jump. Do not infer success solely from a vendor’s aggregate benchmark.

For an application, also examine turn endings and interruptions in its real environment. A well-acted response that begins at the wrong moment can still create a poor conversation. This guide does not claim measured latency or interruption performance for your integration.

Turn the winning direction into a production note

Save the approved text, voice identifier, model version, direction, settings, and rendered audio together. Describe the intended behavior in ordinary language so another editor can understand it without reconstructing the entire audition.

Use a new sentence to check whether the direction transfers: “Let’s call the caretaker before we move anything.” A practical voice should remain helpful; a worried voice should not become breathless on every clause. If the effect fails outside the first sentence, refine the instruction.

For general script preparation, test a short narration in Voice Lab and use our voiceover prompting guide. The Inworld-specific direction and realtime context tests happen separately in Inworld. If you instead need to change the reusable identity of a voice, see the voice-remix audition workflow.

Frequently asked questions

Where should I put voice direction?

The documented Inworld workflow places a bracketed natural-language direction at the start of the input. Check the current documentation for your model and integration.

Does every TTS request understand earlier audio?

No. Conversational context must be supplied through a supported path. A request containing only new text does not automatically include the previous audio exchange.

Should I add a breath or sigh to every emotional line?

No. Add a nonverbal cue only when it communicates something useful, and check its effect on intention and timing.

Does this article include measured performance results?

No. It provides original scripts and a repeatable evaluation method. Test the actual outputs and application behavior in your own workflow.