Verified against Veo 3.1 · 2026-08-02
Keep a character's face consistent across shots using a Flow reference image
A prompt structured for Flow's "ingredients to video" reference-image feature, which anchors a character's appearance from an actual uploaded image rather than a text description — written to deliberately avoid re-describing the face in text, since a conflicting text description competes with the image the model is supposed to be locking onto.
The prompt
Ready to copy — highlighted parts are example details you can swap.
Write a Veo 3.1 prompt intended to be used together with an uploaded reference image in Flow's ingredients-to-video mode, where the reference image — not the text description — is what should govern the character's actual facial appearance across every generated shot. The discipline here is describing action, setting, and camera in full detail while deliberately not re-describing the face, hair, or body type the reference image already establishes, since a text description that conflicts even slightly with the reference image gives the model two different sources of truth about what the character looks like, and it will not reliably favor the image over the text. REFERENCE IMAGE a headshot of a man in his 40s with short greying hair, uploaded as the character anchor — a note describing what the uploaded image actually shows, for your own tracking, not a substitute for the image itself; the actual appearance comes from the file, not from this line. WHAT TO DESCRIBE IN TEXT the man, now wearing a navy suit jacket over the same build shown in the reference, walking into a glass-walled office lobby. Describe what the character is doing, wearing in this specific shot (if different from the reference image — a costume change is fine to describe in text), and where they are, in full detail — this is exactly the information the reference image cannot supply on its own. WHAT NOT TO RE-DESCRIBE IN TEXT Do not restate the character's face, hair color, hair style, body type, or any other physical trait already visible in the reference image, even in passing — every one of those restated details is a chance for the text to say something subtly different from the image, and Flow does not reliably resolve that conflict in the image's favor. CAMERA a slow tracking shot following him from a three-quarter angle as he crosses the lobby for this specific shot. LIGHTING AND ENVIRONMENT bright, even daylight through floor-to-ceiling glass, a modern corporate lobby, matched to whatever look this shot needs, independent of the reference image's own lighting, since the reference image is only anchoring appearance, not the lighting or setting of the new shot. STYLE AND MOOD clean corporate commercial look, neutral color grade. CONSISTENCY CHECK ACROSS MULTIPLE SHOTS a second shot of him sitting down at a boardroom table, and a third of him shaking hands with a colleague. For each additional shot using the same reference image, repeat this same discipline — describe the new action and setting, never the face — so the character reads as the same person across every shot for the reasons stated above, not because the text happened to describe them identically each time. WHAT TO AVOID Do not swap in a different reference image mid-sequence for the same character without expecting a visible discontinuity — a character's appearance is only as consistent as the single reference image anchoring it, and two different photos of even the same real person can anchor subtly different appearances across a sequence. OUTPUT The finished prompt for this shot, followed by one line confirming no facial or body description was restated in the text, and a running list of which reference image is anchoring which character across the full sequence.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Flow's ingredients-to-video mode anchors a character's appearance from the pixels of an actual uploaded image rather than from a text description, which is a categorically different mechanism than describing a face in words and hoping the same words produce the same face twice — text-to-video generation has no persistent memory of "the same character" between separate prompts, so two independently written text descriptions of "a man in his 40s with short greying hair" will produce two different faces, while the same reference image anchoring two separate generations produces a recognizably consistent one. This is exactly why re-describing the face in text alongside the reference image is actively counterproductive rather than merely redundant: doing so gives the model two sources of truth about the same attribute, and when the text's wording drifts even slightly from what the image actually shows — a slightly different hair length, an age description that does not quite match — the model has no reliable rule for which source to trust, and the resulting face is often neither fully the reference image nor fully the text description, but some inconsistent blend of both. Restricting the text to action, setting, and camera — the information the reference image genuinely cannot supply — respects the actual division of labor between the two inputs: the image's job is appearance, the text's job is everything the image is a single still frame and cannot express, which is what the character is doing, wearing differently, and where they are in this specific shot. Explicitly warning against swapping reference images mid-sequence addresses a specific and easy mistake: because the anchor is the image itself, not a persistent abstract "character," two different photographs of even the same real person carry subtly different lighting, angle, and expression information, and switching the anchor mid-sequence introduces exactly the kind of visible discontinuity the reference-image workflow exists to prevent in the first place.
Verified against
Veo 3.1 3.1 · 2026-08-02
Google Flow 1.5 · 2026-08-02
Changelog
- 2026-08-02 — Initial publish, verified against Veo 3.1 via Flow ingredients-to-video.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
