Veo

Verified against Veo 3.1 · 2026-07-25

Write a lip-synced two-character dialogue scene for Veo

A Veo 3.1 prompt structured for its native dialogue generation — quoted lines attributed per speaker, one speaking turn at a time, so the model actually lip-syncs the audio to the correct face instead of blending two voices into one mouth.

Veo 3.17 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

Write a Veo 3.1 prompt for an 8-second dialogue scene between two characters, using Veo's native ability to generate synchronized spoken audio tied to whichever character's mouth is on camera. This only works if exactly one character is speaking at a time and the quoted line is attributed to a named, described character — an unattributed line, or two lines that overlap in the same beat, is how Veo ends up syncing the wrong mouth or blending two voices into one.

CHARACTERS
a woman in her 50s, silver-streaked hair in a low bun, wearing a grey wool cardigan
a man in his 20s, unshaven, wearing a rumpled delivery-uniform jacket

SCENE AND SETTING
the doorway of a small apartment; the woman stands just inside, the man on the step outside holding a package. State where each character is positioned relative to the other and relative to camera, since the model needs to know whose face is closer to lens when only one of them is speaking in a given beat.

DIALOGUE, ONE TURN AT A TIME
Mrs. Alvarez says: "You're two hours late, and I already told the building manager."
Danny says: "I know, I'm sorry — the truck broke down on the highway."
Keep lines short enough to plausibly fit their share of the eight seconds — a rushed, over-long line is what produces audio that outruns the mouth movement generated for it. If a natural pause, a reaction, or a beat of silence belongs between the two lines, say so explicitly rather than leaving the gap unstated.

CAMERA
{{camera_direction}}. If the shot cuts from one character to the other mid-scene, say exactly when the cut happens relative to the two lines, since a cut landing mid-word is a common and avoidable source of desynced audio.

ENVIRONMENT AND LIGHTING
{{lighting_and_mood}}. Note any acoustic detail worth carrying into the audio — a small room versus an open outdoor space changes what a "natural" voice should sound like, and naming it steers the generated voice's tonal quality even though it is not a literal audio-mixing instruction.

STYLE AND MOOD
{{visual_style}}. State the emotional register of the delivery for each line — flat, warm, sarcastic — since Veo uses that description to shape both the performance and the voice, not just the picture.

WHAT TO AVOID
Do not write both characters' lines as happening "at the same time" or as overlapping dialogue — Veo's lip-sync model is built around one active speaker per beat, and simultaneous dialogue is the single most common cause of a scene where neither character's mouth matches what is heard. Do not leave a character's voice or delivery style undescribed if it matters to the scene; an unstated voice defaults to whatever the model infers from the visual description alone, which is frequently a generic register that flattens a character meant to sound distinct.

OUTPUT
The finished prompt exactly as it should be pasted into Veo, followed by one line confirming that only one character speaks per beat and stating which beat, if any, includes a silent reaction shot.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

Veo's dialogue generation ties synthesized speech to visible mouth movement on whichever character is framed as the active speaker, which is a fundamentally different mechanism from adding a voiceover track after the fact — the model has to decide, frame by frame, whose face is producing sound, and it can only make that decision correctly if the prompt states unambiguously who is speaking in each beat. This is exactly why the one-turn-at-a-time structure is not a stylistic preference but a hard constraint: a prompt describing two lines as simultaneous gives the model two competing claims about which mouth should be moving in the same window, and the observable failure mode is audio that syncs to neither face convincingly, or a blended, garbled attempt at both voices at once. Naming each character's position relative to camera and to the other character matters for the same underlying reason — lip-sync accuracy depends on the model correctly identifying which face is close enough to lens to read as the current speaker, and an unstated spatial arrangement leaves that judgment call entirely to the model's own defaults, which skew toward whichever face is largest in frame regardless of who the text says is talking. Keeping each line short enough to plausibly fit its share of the eight-second window addresses a specific and measurable failure: when a written line is too long for the time available, the generated audio has to compress or the mouth movement has to rush to keep pace, and the visible mismatch between spoken cadence and lip movement is one of the fastest ways a viewer clocks a clip as synthetic. Finally, describing each line's emotional register (flat, warm, sarcastic) does real work beyond flavoring the picture — Veo's voice generation reads that description as a performance instruction, not just a visual mood cue, so an unstated delivery style defaults to a flat, generic register that erases whatever distinct character voice the visual description on its own was trying to establish.

Verified against

Veo 3.1 3.1 · 2026-07-25

Changelog

  • 2026-07-25 Initial publish, verified against Veo 3.1 native dialogue and lip-sync generation.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All Veo prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY