Reference image
Identity, wardrobe, product shape, composition, palette, or location design.
Use images for identity, video for movement, and audio for timing. The completed generation appears here.
MiniMax H3 Reference to Video is a reference-guided AI video workflow. Text-to-video starts with words and asks the model to construct the complete shot. Reference-to-video starts with visual or audio evidence: a face that should remain recognizable, a product that must keep its shape, a movement pattern, a camera rhythm, or an audio cue that should influence timing.
The prompt still matters, but it has a different job. Instead of redescribing the source, the prompt tells the model what to do with it. You can ask a character to turn toward the camera, direct a slow orbit around a product, preserve a wardrobe detail, or match the pace of a reference clip. This makes the mode useful for creators who need more continuity than an open-ended prompt usually provides.
Reference control is especially practical for campaign teams, product marketers, character-led content, previsualization, and social production. It does not guarantee a pixel-identical copy of every source. It gives the generation clearer constraints, which makes the output easier to brief, review, and refine.
Identity, wardrobe, product shape, composition, palette, or location design.
Gesture, movement, camera pace, choreography, transitions, or shot rhythm.
Beat timing, emphasis, speech rhythm, or the pace of a sequence.
A strong request gives every source one clear responsibility. Follow these four steps in the live tool above.
Add at least one source. Signed-in users can upload a JPEG, PNG, GIF, or WebP image, while image, video, and audio references can also be supplied as public HTTPS URLs. Choose clean assets with an obvious subject. Avoid several nearly identical files unless each one contributes a different angle or detail.
Describe the new action and assign roles to the sources. For example: “Use the image for identity, the video for walking rhythm, and the audio for beat timing.” Add the camera path, framing, environment, and anything that must remain stable. The prompt should resolve ambiguity rather than narrate the source.
Select duration, resolution, and aspect ratio for the intended placement. Use a shorter 768P generation to test reference roles or motion. Move to a longer duration or 2K when the direction is already working and the extra detail has a clear production purpose.
Review the visible credit estimate, sign in, and submit the task. The page monitors generation status and displays the completed result when it is ready. Download the video or open your creation history, then change one prompt clause or source at a time if another pass is needed.
Reference-guided generation works best when the project already has an approved visual decision that should survive into the new shot.
Use a clean portrait or full-body image to establish identity, clothing, and color. Add a motion reference when the performance requires a particular gesture or walking style. In the prompt, state which facial and wardrobe details should remain recognizable while the character acts.
Use product photographs to protect silhouette, material, logo placement, and surface details. A reference video can suggest an orbit, push-in, turntable motion, or lighting reveal. Keep the requested action simple enough for the product to remain readable throughout the clip.
Combine a style frame with a movement reference to direct atmosphere and camera language separately. Name the beginning and end of the shot, keep the horizon or subject position stable when needed, and avoid stacking several incompatible camera moves into one short generation.
Create vertical or square variations from an approved character, campaign image, or product asset. Use concise movement, a clear focal point, and timing that fits the intended placement. Test the composition in the final aspect ratio instead of relying on a later crop.
Reuse a controlled set of identity and style references across related shots. Keep the reference roles and core prompt constraints consistent, then vary one action, camera move, or environment at a time so the series remains connected without becoming repetitive.
Test how an approved frame might move before a full production or edit. Reference-guided drafts can help a team compare camera directions, action beats, and scene pacing while keeping the original creative brief visible.
People searching for MiniMax H3 Ref2VA are generally looking for reference-guided video creation that can combine visual and audio direction. Ref2VA is treated here as a useful description of the workflow, not as a separate model claim. The live route is MiniMax H3 Reference to Video, and the accepted inputs shown in the tool are the source of truth.
A reference image usually carries appearance or composition. A reference video is better suited to movement and pacing. Audio can influence timing. The prompt connects those sources by explaining what should be transferred into the new output and what should remain unchanged.
Do not add a source simply because the interface accepts it. If two references compete for identity, style, or camera direction, the request becomes harder to interpret. Start with the minimum useful set, review the result, and add another reference only when it solves a specific missing constraint.
Choose sharp, well-lit sources where the subject is easy to separate from the background. Compression artifacts, heavy motion blur, and tiny subjects make identity and geometry harder to interpret.
Write “use the image for face and wardrobe” instead of “make it look like this.” A named role helps separate identity, movement, composition, style, and timing.
State the starting position, action, camera path, and finishing frame. Chronological instructions are easier to review than a list of unrelated cinematic terms.
Call out the details that must not drift: face, logo, product shape, clothing, palette, or camera angle. Keep that constraint wording stable across retries.
A short, focused movement often produces a clearer test than several events compressed into one clip. Increase duration after the action and camera path are working.
When a result misses, adjust one reference, setting, or prompt clause. Controlled retries reveal whether the problem came from the source, motion, framing, or wording.
Need help assigning camera and reference roles? Use the MiniMax H3 Prompt Guide, then return to the generator with a clearer direction.
Reference generations use credits. The live generator calculates the current estimate from duration and resolution before you submit, so the displayed cost—not a static article example—is the source of truth.
Start with a short 768P pass when you are validating source roles. Use 2K or a longer duration once the identity, motion, and camera direction are stable. The same account balance works across text, image, and reference workflows.
It is a generation mode that uses one or more source assets to guide a new video. A reference image can define identity or composition, a reference video can guide motion or camera pace, and reference audio can help establish timing.
Upload a supported image or paste a public HTTPS image URL, then explain what the image should control. Clear instructions such as “preserve the character face and jacket” are more useful than repeating every visible detail.
Yes, this workspace accepts public HTTPS video URLs for the reference-to-video route. Use a video reference when movement, gesture, choreography, or camera rhythm matters more than a single still frame can express.
Ref2VA is a search term used for reference-guided video and audio workflows. On this site it points to the same MiniMax H3 reference-to-video process rather than a separate model or product page.
The visible estimate changes with the selected duration and resolution. The generator shows the calculated amount before submission, so you can compare a shorter 768P test with a longer or 2K pass.
Yes. Add your prompt and references in the browser, sign in, review the credit estimate, and generate. Completed results can be opened and downloaded from the same workflow.
Move from research to a relevant generator, prompt resource, or pricing page without restarting your workflow.
Write and test prompts for specific video goals.
Follow the complete generation workflow.
Understand the speed-focused search intent and availability.
Camera movement, shot composition, and motion control prompting.
Transform an existing clip with a video-first reference generator.
Reapply movement from a video reference onto a new subject.
See every online input path in one workflow map.
Understand LoRA availability and the closest online option.
See what a ComfyUI-based path involves and the online alternative.
What API access means here and the online generator alternative.
VRAM, GPU, Mac, AMD, and model size questions in one place.
What running MiniMax H3 locally would involve, mapped out.
Where to look for repositories, workflows, and implementations.
What to look for on Hugging Face: weights, variants, and LoRA.
Open the generator with MiniMax H3 selected.
Compare plans and one-time credit packs.