Reference-to-Video (R2V)

Reference-to-video (R2V) uses reference images to keep characters, products, and styles consistent in AI video — and how it differs from image-to-video.

Open the full generator

See it in action

One real prompt and the result it produced on Molyin. Remix it to start from here.

  • Seedance 2.5
  • Reference to video
  • 30s
  • 720P

A premium, highly cinematic 30-second 3D motion-graphics sequence in an intricate steampunk and vintage-miniature diorama style, with continuous, fluid orbiting and fly-through camera moves. [0-10s]: Macro close-up of an antique brass clock face that miraculously unfolds, layer by layer, into interlocking rings of rotating gears and volumetric fog. The camera plunges down through the gears to reveal a mechanical ornithopter spiraling up out of a miniature canyon stacked from weathered antique books. [10-20s]: The camera glides forward along the ornithopter's flight path, passing seamlessly into a fast-spinning ornate brass zoetrope whose interior projects the dynamic light and shadow of galloping mechanical horses. The projection leaps out of the drum and the scene instantly transforms into a brass-textured suspended cable car threading through a forest of mechanical gears along faintly glowing copper rails, bathed in cinematic golden-hour light. [20-30s]: The camera pans down elegantly to reveal an exquisite clockwork wooden sailing ship below the cable car, cutting through rolling waves of deep-blue glass. At the far end of the waves the sea seamlessly evolves into a huge glowing moon, and a group of silhouetted explorers holding swaying lanterns treks laboriously along a crystal-vein ridge under the starry sky. The camera pulls back in a smooth spiral through ethereal clouds, returning to the grand ticking brass clock face. In the final second the logo appears, per @Image1. Technical specs: Ultra-detailed mechanical textures, rich brass and gold tones, cinematic shallow depth of field. Smooth, coherent, seamless fly-through camerawork with an overwhelming sense of epic fantasy adventure.

Reference-to-video (R2V) is a generative AI technique that uses one or more reference images to guide video generation — not as the literal first frame, but as an identity anchor. The model studies your references (a character's face, a product's design, an art style) and generates new scenes where that identity stays consistent.

Reference vs. first frame

This is the key distinction from image-to-video:

  • Image-to-video treats your image as the opening frame. The video starts exactly there.
  • Reference-to-video treats your images as a definition of what things look like. The model can then place that character or product into an entirely new scene, angle, or action described by your prompt.

In short: I2V continues a picture; R2V casts it.

("R2V" is simply shorthand for reference-to-video — you'll also see it written as reference2video or ref-to-video.)

T2V vs I2V vs R2V at a glance

ModeYour inputThe model treats it asBest for
Text-to-video (T2V)a written prompt onlythe full creative briefnet-new scenes from scratch
Image-to-video (I2V)one image (plus an optional last frame)the literal opening frameanimating a still you already have
Reference-to-video (R2V)a set of images, clips, and audioidentity and style anchorsconsistent characters, products, and series

Why it matters

Consistency is the hardest problem in AI video. Generate the same "red-haired girl in a yellow raincoat" twice from text alone and you'll get two different girls. Reference-to-video solves this:

  • Character consistency — keep the same protagonist across shots, scenes, and episodes.
  • Product fidelity — show your actual product from new angles, in new environments.
  • Style continuity — carry an illustration style or brand look through a whole series.

Prompting tips

Give the model clean, well-lit references that show the subject clearly. Then let the prompt do the directing: new setting, new action, new camera. Say what should change — the references already say what should stay. More in our prompt writing guide.

Reference-to-video on Molyin

Reference-to-video is a standard mode across Molyin's video lineup rather than a single-model feature. Reference budgets per generation:

ModelReference budgetNotes
Seedance 2.5 / 2.0up to 9 images + 3 video clips + 3 audio tracksaudio references can drive voice or score
Wan 3.0up to 10 images + 5 videos + 5 audioall-in-one reference with up to 30-second output
Wan 2.7up to 5 images and videos combinedno audio reference slot
MiniMax H3up to 9 images + 3 videos + 3 audioreference clips capped at 15 seconds total
Kling 3.02–4 element imagesnamed-element casting
FLUX 33–10 keyframe imagesstoryboard semantics — keyframes pin moments in time rather than identity

Pick a model, drop in your references, and mention them inline ("@Image1 walks through the alley from @Video1") to direct the shot.

Questions about Reference-to-Video (R2V)

How is reference-to-video different from image-to-video?

Image-to-video animates one picture as the literal first frame. Reference-to-video takes several images, clips, or audio as references for identity, style, or motion, and composes a new scene that is not locked to any of them.

How many references can I use?

It depends on the model; the generator shows the limit as you add them. Reference video seconds are billed at a reduced rate on most models, and the estimate updates live before you generate.

Go deeper

What is Reference-to-Video (R2V)? | Molyin