One explorer wears a teal coat and carries a lantern. The other wears ochre and has a satchel. Halfway across the bridge, the lantern changes hands and both coats match. A plausible frame can hide a failed reference transfer. This guide makes that failure easy to identify.
Download the two-subject continuity sheet · Prepare a reference in Image Lab
Editorial guide checked September 13, 2026. Examples and worksheets are original planning material; the hero is an AI-generated illustration, not a model benchmark.
Give the two subjects identities that can be checked
For this exercise, create two fictional stop-motion explorers. Mira wears a teal raincoat and carries a copper lantern. Niko wears an ochre coat and has a square canvas satchel. Their task is simple: cross a short wooden bridge and stop beside each other. There is no chase, costume change, or prop exchange.
The distinction is deliberate. Names alone are weak visual evidence. Coat color, silhouette, and prop ownership give you three separate checks. If the faces remain different but the satchel appears on Mira, the scene has still failed one requirement. Write those requirements before generating, when you are less likely to excuse an attractive mistake.
Use original or authorized reference material. The example characters and scene here are a planning exercise. They are not supplied Vidu benchmark outputs, and the hero illustration does not establish what that model can produce.
Separate references from first-frame composition
A reference identifies what should appear. A first frame also fixes where it already appears at the beginning of a clip. Confusing the two can lead to disappointment: you provide a studio character portrait, expect the output to begin on a bridge, then judge it against an opening composition you never specified.
Vidu's Q3 overview describes native audio-video output and clips up to sixteen seconds. That does not mean every reference mode accepts every kind of source. The reference-to-video API documentation distinguishes model-specific subject inputs. In the reviewed Q3 subjects mode, subjects are represented with images and text, and named subjects can be referenced in prompts. Do not copy a video-reference recipe from a different mode and assume it applies.
In the interface you use, record the selected model and reference mode. If the integration provides only a single image input, it may be offering image to video instead. The upload control, not a broad model headline, determines the inputs you can actually use.
Build a small reference pack with no competing story
Give each explorer a clear full-body image on a simple background. Add a second angle only when it resolves an important hidden feature. A second image that introduces different boots or a different lantern creates a conflict rather than useful coverage.
Keep each subject's pictures together and name the files plainly. For example, mira-front and mira-side are easier to audit than an unordered set of exports. Put the identity notes in the same folder. If you later replace a reference, update its version so a collaborator can reproduce the accepted setup.
| Subject | Appearance to retain | Exclusive prop | Must not inherit |
|---|---|---|---|
| Mira | Teal coat, rounded hood, short boots | Copper lantern in left hand | Niko's satchel or ochre coat |
| Niko | Ochre coat, narrow collar, tall boots | Canvas satchel at right hip | Mira's lantern or hood |
A location image can help establish the bridge, but avoid one that already contains a similar person. That extra figure can compete with your intended cast. First test the characters against a simply described bridge. Add a location reference only if the environment itself needs tighter control.
Run the single-subject check before the pair
Begin with Mira walking two steps and stopping. Keep the camera fixed and show enough of her body to inspect the coat, boots, and lantern. If her own reference cannot survive this action, combining her with Niko will make the failure harder to diagnose.
Repeat with Niko under comparable conditions. You are not trying to select a winner between characters. You are checking whether each reference group gives the model a coherent subject. If one works and the other does not, inspect the weaker reference for occlusion, tiny details, inconsistent costume, and confusing background objects.
Only then generate the pair. Retain the same general scene, lighting, and restrained motion. This staged exercise costs more than a single lucky attempt, but it tells you where a problem begins. Keep the test set small and review the displayed generation cost before each run.
Write ownership into the interaction
A two-person prompt should specify more than “Mira and Niko walk together.” State their starting positions, direction, props, and final arrangement. For the first trial, avoid crossing paths or handing objects between them. Those are useful later tests, but they introduce extra identity and contact problems.
If your provider exposes named subject references, bind those names using its documented interface. The plain-language names in this example are not a universal command syntax. Before spending on a render, confirm that both reference groups are attached to the intended subjects.
Once the pair works, try one change: a closer view, a small head turn, or a deliberate cut. Do not add all three at once. You want to know whether the extra request preserved the identities you already established.
Score the beginning, interaction, and ending separately
Take a frame near the beginning, one during the walk, and one after the stop. Compare all three against the references. A model can recover a character by the last frame while still showing a distracting swap halfway through.
| Check | What passes | What requires another look |
|---|---|---|
| Identity | Distinct face and silhouette remain | Both acquire the same head or proportions |
| Wardrobe | Colors and garment shapes stay assigned | Coats merge or exchange |
| Ownership | Lantern and satchel stay with their subjects | Props duplicate or teleport |
| Space | Feet remain on the bridge; paths make sense | Subjects pass through each other or railings |
Do not average these into a flattering overall score. A beautiful environment cannot compensate for the wrong person holding the lantern if that prop matters to the story. Keep aesthetic preference separate from the pass requirements.
Listen to the footsteps as well. If the characters stop and the walking continues, the sound needs correction even when the picture is usable. Native audio is a starting point for review, not an exemption from it.
Fix the reference before rewriting the whole scene
When both coats drift toward teal, make the references more distinct and shorten competing style directions. When only the lantern fails, increase its visibility in Mira's source image or simplify her hand movement. When identities fail only during overlap, adjust the staging so the subjects remain readable.
Keep the last successful setup. If the pair worked before you introduced a camera move, return to that checkpoint and test a smaller move. Replacing the prompt, references, duration, and model simultaneously loses the evidence that one part was already working.
For QuestStudio, the useful next step is preparing and checking the reference stills. Vidu Q3 is not currently listed in Video Lab. You can use available models to generate separate character shots, but you should not expect the same multi-subject binding controls. Export an accepted source pack with the ownership table so the next generation begins from a clear brief.
Frequently asked questions
Is reference to video the same as image to video?
No. Image to video commonly starts from a particular frame. Reference to video uses supplied subjects to guide a newly staged scene.
Can every Vidu Q3 integration take video references?
No. Verify the exact mode. The reviewed Q3 subjects API describes image and text subjects; do not import capabilities from another model's reference mode.
Why do characters swap clothes or props?
Ambiguous inputs, similar silhouettes, and complicated interactions can make subject ownership unclear. Isolate each subject and simplify the action to locate the cause.
Can I generate Vidu Q3 in QuestStudio?
Vidu Q3 is not currently listed. You can prepare reference images and use supported Video Lab models for separate shots.

