Veo

Verified against Veo 3.1 · 2026-08-02

Keep a character's face consistent across shots using a Flow reference image

A prompt structured for Flow's "ingredients to video" reference-image feature, which anchors a character's appearance from an actual uploaded image rather than a text description — written to deliberately avoid re-describing the face in text, since a conflicting text description competes with the image the model is supposed to be locking onto.

Veo 3.1Google Flow6 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

Write a Veo 3.1 prompt intended to be used together with an uploaded reference image in Flow's ingredients-to-video mode, where the reference image — not the text description — is what should govern the character's actual facial appearance across every generated shot. The discipline here is describing action, setting, and camera in full detail while deliberately not re-describing the face, hair, or body type the reference image already establishes, since a text description that conflicts even slightly with the reference image gives the model two different sources of truth about what the character looks like, and it will not reliably favor the image over the text.

REFERENCE IMAGE
a headshot of a man in his 40s with short greying hair, uploaded as the character anchor — a note describing what the uploaded image actually shows, for your own tracking, not a substitute for the image itself; the actual appearance comes from the file, not from this line.

WHAT TO DESCRIBE IN TEXT
the man, now wearing a navy suit jacket over the same build shown in the reference, walking into a glass-walled office lobby. Describe what the character is doing, wearing in this specific shot (if different from the reference image — a costume change is fine to describe in text), and where they are, in full detail — this is exactly the information the reference image cannot supply on its own.

WHAT NOT TO RE-DESCRIBE IN TEXT
Do not restate the character's face, hair color, hair style, body type, or any other physical trait already visible in the reference image, even in passing — every one of those restated details is a chance for the text to say something subtly different from the image, and Flow does not reliably resolve that conflict in the image's favor.

CAMERA
a slow tracking shot following him from a three-quarter angle as he crosses the lobby for this specific shot.

LIGHTING AND ENVIRONMENT
bright, even daylight through floor-to-ceiling glass, a modern corporate lobby, matched to whatever look this shot needs, independent of the reference image's own lighting, since the reference image is only anchoring appearance, not the lighting or setting of the new shot.

STYLE AND MOOD
clean corporate commercial look, neutral color grade.

CONSISTENCY CHECK ACROSS MULTIPLE SHOTS
a second shot of him sitting down at a boardroom table, and a third of him shaking hands with a colleague. For each additional shot using the same reference image, repeat this same discipline — describe the new action and setting, never the face — so the character reads as the same person across every shot for the reasons stated above, not because the text happened to describe them identically each time.

WHAT TO AVOID
Do not swap in a different reference image mid-sequence for the same character without expecting a visible discontinuity — a character's appearance is only as consistent as the single reference image anchoring it, and two different photos of even the same real person can anchor subtly different appearances across a sequence.

OUTPUT
The finished prompt for this shot, followed by one line confirming no facial or body description was restated in the text, and a running list of which reference image is anchoring which character across the full sequence.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

Flow's ingredients-to-video mode anchors a character's appearance from the pixels of an actual uploaded image rather than from a text description, which is a categorically different mechanism than describing a face in words and hoping the same words produce the same face twice — text-to-video generation has no persistent memory of "the same character" between separate prompts, so two independently written text descriptions of "a man in his 40s with short greying hair" will produce two different faces, while the same reference image anchoring two separate generations produces a recognizably consistent one. This is exactly why re-describing the face in text alongside the reference image is actively counterproductive rather than merely redundant: doing so gives the model two sources of truth about the same attribute, and when the text's wording drifts even slightly from what the image actually shows — a slightly different hair length, an age description that does not quite match — the model has no reliable rule for which source to trust, and the resulting face is often neither fully the reference image nor fully the text description, but some inconsistent blend of both. Restricting the text to action, setting, and camera — the information the reference image genuinely cannot supply — respects the actual division of labor between the two inputs: the image's job is appearance, the text's job is everything the image is a single still frame and cannot express, which is what the character is doing, wearing differently, and where they are in this specific shot. Explicitly warning against swapping reference images mid-sequence addresses a specific and easy mistake: because the anchor is the image itself, not a persistent abstract "character," two different photographs of even the same real person carry subtly different lighting, angle, and expression information, and switching the anchor mid-sequence introduces exactly the kind of visible discontinuity the reference-image workflow exists to prevent in the first place.

Verified against

Veo 3.1 3.1 · 2026-08-02

Google Flow 1.5 · 2026-08-02

Changelog

  • 2026-08-02 Initial publish, verified against Veo 3.1 via Flow ingredients-to-video.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All Veo prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY