One explorer wears a teal coat and carries a lantern. The other wears ochre and has a satchel. Halfway across the bridge, the lantern changes hands and both coats match. A plausible frame can hide a failed reference transfer. This guide makes that failure easy to identify.

Quick answer: Prepare separate, clean reference groups for each subject, give them stable names, and specify ownership of distinguishing props. Test one subject, then both together, before adding a cut. Vidu Q3 is not currently listed in QuestStudio; Image Lab can help prepare the still references.

Download the two-subject continuity sheet · Prepare a reference in Image Lab

Editorial guide checked September 13, 2026. Examples and worksheets are original planning material; the hero is an AI-generated illustration, not a model benchmark.

Give the two subjects identities that can be checked

For this exercise, create two fictional stop-motion explorers. Mira wears a teal raincoat and carries a copper lantern. Niko wears an ochre coat and has a square canvas satchel. Their task is simple: cross a short wooden bridge and stop beside each other. There is no chase, costume change, or prop exchange.

The distinction is deliberate. Names alone are weak visual evidence. Coat color, silhouette, and prop ownership give you three separate checks. If the faces remain different but the satchel appears on Mira, the scene has still failed one requirement. Write those requirements before generating, when you are less likely to excuse an attractive mistake.

Use original or authorized reference material. The example characters and scene here are a planning exercise. They are not supplied Vidu benchmark outputs, and the hero illustration does not establish what that model can produce.

Separate references from first-frame composition

A reference identifies what should appear. A first frame also fixes where it already appears at the beginning of a clip. Confusing the two can lead to disappointment: you provide a studio character portrait, expect the output to begin on a bridge, then judge it against an opening composition you never specified.

Vidu's Q3 overview describes native audio-video output and clips up to sixteen seconds. That does not mean every reference mode accepts every kind of source. The reference-to-video API documentation distinguishes model-specific subject inputs. In the reviewed Q3 subjects mode, subjects are represented with images and text, and named subjects can be referenced in prompts. Do not copy a video-reference recipe from a different mode and assume it applies.

In the interface you use, record the selected model and reference mode. If the integration provides only a single image input, it may be offering image to video instead. The upload control, not a broad model headline, determines the inputs you can actually use.

Build a small reference pack with no competing story

Give each explorer a clear full-body image on a simple background. Add a second angle only when it resolves an important hidden feature. A second image that introduces different boots or a different lantern creates a conflict rather than useful coverage.

Keep each subject's pictures together and name the files plainly. For example, mira-front and mira-side are easier to audit than an unordered set of exports. Put the identity notes in the same folder. If you later replace a reference, update its version so a collaborator can reproduce the accepted setup.

SubjectAppearance to retainExclusive propMust not inherit
MiraTeal coat, rounded hood, short bootsCopper lantern in left handNiko's satchel or ochre coat
NikoOchre coat, narrow collar, tall bootsCanvas satchel at right hipMira's lantern or hood

A location image can help establish the bridge, but avoid one that already contains a similar person. That extra figure can compete with your intended cast. First test the characters against a simply described bridge. Add a location reference only if the environment itself needs tighter control.

Run the single-subject check before the pair

Begin with Mira walking two steps and stopping. Keep the camera fixed and show enough of her body to inspect the coat, boots, and lantern. If her own reference cannot survive this action, combining her with Niko will make the failure harder to diagnose.

Repeat with Niko under comparable conditions. You are not trying to select a winner between characters. You are checking whether each reference group gives the model a coherent subject. If one works and the other does not, inspect the weaker reference for occlusion, tiny details, inconsistent costume, and confusing background objects.

Only then generate the pair. Retain the same general scene, lighting, and restrained motion. This staged exercise costs more than a single lucky attempt, but it tells you where a problem begins. Keep the test set small and review the displayed generation cost before each run.

Write ownership into the interaction

A two-person prompt should specify more than “Mira and Niko walk together.” State their starting positions, direction, props, and final arrangement. For the first trial, avoid crossing paths or handing objects between them. Those are useful later tests, but they introduce extra identity and contact problems.

Mira, the explorer in the teal raincoat carrying her copper lantern in her left hand, starts on the left side of a narrow wooden bridge. Niko, the explorer in the ochre coat with his square canvas satchel at his right hip, starts beside her on the right. In one steady full-body view, both walk two slow steps toward the far end of the bridge and stop. Mira keeps the lantern; Niko keeps the satchel. Their paths remain side by side. Soft morning mist sits behind the bridge, with warm light coming from camera left. Quiet footsteps on wood and distant running water; no dialogue or music.

If your provider exposes named subject references, bind those names using its documented interface. The plain-language names in this example are not a universal command syntax. Before spending on a render, confirm that both reference groups are attached to the intended subjects.

Once the pair works, try one change: a closer view, a small head turn, or a deliberate cut. Do not add all three at once. You want to know whether the extra request preserved the identities you already established.

Score the beginning, interaction, and ending separately

Take a frame near the beginning, one during the walk, and one after the stop. Compare all three against the references. A model can recover a character by the last frame while still showing a distracting swap halfway through.

CheckWhat passesWhat requires another look
IdentityDistinct face and silhouette remainBoth acquire the same head or proportions
WardrobeColors and garment shapes stay assignedCoats merge or exchange
OwnershipLantern and satchel stay with their subjectsProps duplicate or teleport
SpaceFeet remain on the bridge; paths make senseSubjects pass through each other or railings

Do not average these into a flattering overall score. A beautiful environment cannot compensate for the wrong person holding the lantern if that prop matters to the story. Keep aesthetic preference separate from the pass requirements.

Listen to the footsteps as well. If the characters stop and the walking continues, the sound needs correction even when the picture is usable. Native audio is a starting point for review, not an exemption from it.

Fix the reference before rewriting the whole scene

When both coats drift toward teal, make the references more distinct and shorten competing style directions. When only the lantern fails, increase its visibility in Mira's source image or simplify her hand movement. When identities fail only during overlap, adjust the staging so the subjects remain readable.

Keep the last successful setup. If the pair worked before you introduced a camera move, return to that checkpoint and test a smaller move. Replacing the prompt, references, duration, and model simultaneously loses the evidence that one part was already working.

For QuestStudio, the useful next step is preparing and checking the reference stills. Vidu Q3 is not currently listed in Video Lab. You can use available models to generate separate character shots, but you should not expect the same multi-subject binding controls. Export an accepted source pack with the ownership table so the next generation begins from a clear brief.

Frequently asked questions

Is reference to video the same as image to video?

No. Image to video commonly starts from a particular frame. Reference to video uses supplied subjects to guide a newly staged scene.

Can every Vidu Q3 integration take video references?

No. Verify the exact mode. The reviewed Q3 subjects API describes image and text subjects; do not import capabilities from another model's reference mode.

Why do characters swap clothes or props?

Ambiguous inputs, similar silhouettes, and complicated interactions can make subject ownership unclear. Isolate each subject and simplify the action to locate the cause.

Can I generate Vidu Q3 in QuestStudio?

Vidu Q3 is not currently listed. You can prepare reference images and use supported Video Lab models for separate shots.