Veo

Verified against Veo 3.1 · 2026-07-26

Write a voiceover-narrated explainer clip without an on-screen talking mouth

A Veo 3.1 prompt for generating a clip with off-screen narration audio synced to on-screen visuals rather than to a speaking mouth, structured to keep narration length matched to the fixed clip duration instead of running long or short.

Veo 3.17 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

Write a Veo 3.1 prompt for an 8-second explainer clip carried by off-screen voiceover narration, not on-camera dialogue. This is a different generation mode from a talking-head shot: there is no mouth on screen for the audio to sync to, so the narration timing has to be matched to the visuals by pacing description alone, and the prompt must say explicitly that the speaker is off-screen so Veo does not invent an on-camera narrator whose mouth then fails to match the words.

SUBJECT AND VISUALS
a small business owner sorting invoices at a cluttered desk, shown as switching from a messy paper stack to typing on a laptop, visibly relieved across the clip. Since no mouth is being synced, the visuals can carry their own pacing independent of the narration's rhythm — use that freedom to show, not just illustrate, what the narration is describing.

NARRATION
Off-screen voiceover, no speaker visible in frame. The narrator says: "Stop chasing invoices across three different apps. One dashboard, every client, paid on time."
Count roughly two to three spoken words per second of clip — an 8-second clip comfortably holds 16-24 words at a natural pace; a longer line either gets rushed into an unnatural cadence or gets clipped before it finishes, and a shorter line leaves dead air the model will fill with something unintended, often ambient noise that competes with the narration instead of supporting it.

VOICE AND TONE
calm and reassuring, like a founder explaining a fix they wish they had sooner. State the tone in performance terms — measured, urgent, warm — since this is a real instruction to the voice generation, not decoration on a visual description.

ENVIRONMENT AND LIGHTING
a small home office, late-afternoon light through a window blind, warm and slightly cluttered.

STYLE AND MOOD
clean, natural, unforced — not a glossy studio ad look.

AMBIENT AUDIO UNDER THE NARRATION
a faint keyboard clatter and the soft hum of a desk fan, kept quiet enough that it sits under the narration rather than competing with it. Naming an ambient layer explicitly, even a subtle one, prevents Veo from generating an unintentionally silent or unintentionally noisy environment around the voice track.

WHAT TO AVOID
Do not describe a person visible in frame as the one delivering the line unless the intent is genuinely a talking-head shot with lip-sync — an ambiguous prompt that shows a person's face while also requesting "voiceover" risks the model treating the visible face as the speaker and attempting lip-sync anyway, producing a mismatched mouth. Do not write a narration line by word count alone without reading it aloud first at a natural pace; a line that looks short on the page but is dense with hard consonants or long words will not fit the seconds allotted to it.

OUTPUT
The finished prompt, followed by a rough word count for the narration line and the target seconds-per-word ratio it produces, so pacing can be checked before generating.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

Off-screen narration and on-camera dialogue are two genuinely different generation problems for Veo, not two ways of phrasing the same request — dialogue asks the model to tie synthesized speech to a specific visible mouth, while voiceover asks it to generate speech with no mouth to sync at all, and stating "off-screen, no speaker visible" explicitly is what tells the model which problem it is solving. Leaving that ambiguous is the actual failure mode this prompt is built to avoid: a prompt that shows a person's face on screen while also asking for "voiceover narration" gives Veo two conflicting signals, and it will frequently default to treating the visible face as the speaker and attempting lip-sync anyway, producing a mouth that moves out of sync with words the prompt intended to be disembodied. The two-to-three-words-per-second pacing rule exists because the clip's duration is fixed before the narration is written, not the other way around — a script written first and then fitted into eight seconds routinely runs long, and when it does, Veo either compresses the delivery into an unnaturally fast cadence or truncates the line mid-thought, both of which are more noticeable and more damaging to a short ad than a slightly shorter line would have been. Explicitly naming a quiet ambient bed under the narration, rather than leaving the sound environment unstated, prevents a specific and common artifact: an unstated soundscape does not reliably default to clean silence, and depending on the described visual scene, the model may generate an ambient layer at a volume or character that fights the narration for attention, which a single named, deliberately quiet ambient cue heads off. Because the visuals here do not need to be paced to a mouth's movement the way a dialogue shot does, the visual layer is genuinely freer to carry its own rhythm — a fact worth stating in the prompt itself, since it is the reason a voiceover-driven clip can afford a visual beat (a gesture, a cut, a reveal) that would be impossible to time correctly against a synced speaking face.

Verified against

Veo 3.1 3.1 · 2026-07-26

Changelog

  • 2026-07-26 Initial publish, verified against Veo 3.1 voiceover audio generation.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All Veo prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY