Verified against Veo 3.1 · 2026-07-21
Turn a product photo into a cinematic hero showcase clip
A layered Veo 3.1 brief for an 8-second hero shot of a physical product, written in the fixed subject/camera/lighting/style/audio order the model conditions on best, with the native audio described in the same brief instead of added afterward.
The prompt
Ready to copy — highlighted parts are example details you can swap.
Write this as a single Veo 3.1 text-to-video prompt for an 8-second hero shot of a physical product. Keep the five layers below in this exact order, as short separate sentences rather than one run-on paragraph — Veo conditions on each layer somewhat independently, and merging them into one dense sentence makes it harder for the model to tell which adjective belongs to which layer. SUBJECT AND ACTION matte-black wireless earbuds case, resting on a brushed concrete pedestal, rotating slowly and evenly through roughly 180 degrees over the full clip. Name only one action. A product that is simultaneously rotating, being splashed, and catching falling confetti is three separate ideas competing for the same eight seconds, and Veo will blend them into something that reads as neither. CAMERA Exactly one camera movement for the whole clip: a slow dolly-in from a medium shot to a close-up, locked-off tripod smoothness, no handheld shake, no simultaneous orbit or crane. State the shot type at the start (medium) and the shot type at the end (close-up) explicitly, since Veo uses the named start and end framing as an anchor for how far the move should travel across the eight seconds. ENVIRONMENT AND LIGHTING A dark, uncluttered studio backdrop. One soft key light from the upper left, roughly 45 degrees above the product and a subtle rim light separating the product's silhouette from the background. Describe what the light is doing, not just where it is positioned — "catching the brushed concrete pedestal edge as it turns" gives Veo a moving highlight to render across the rotation, instead of a static lighting setup that ignores the motion already asked for. STYLE AND MOOD warm minimalist commercial, soft 35mm film grain. Shallow depth of field with the background genuinely soft, not merely described as "blurry." Premium and quiet, not busy or cluttered. AUDIO a low mechanical hum with a faint metallic resonance, rising subtly the instant the camera reaches the close-up, no dialogue, no music track. Native audio is generated from this exact sentence, so write it as a real sound description rather than a mood word — "tense" is not a sound; "a low mechanical hum climbing half a tone" is. DURATION AND FORMAT 8 seconds, 16:9 for the website hero and a separate 9:16 pass for Instagram Stories. If the output is needed for more than one platform, generate the primary aspect ratio first and treat every other ratio as its own separate generation with the same five layers re-described for the new frame, not a crop of this one — Veo composes lighting position and camera-move endpoints relative to the frame it is told to fill, so cropping after the fact clips the rim light and cuts the dolly's intended endpoint short. WHAT TO AVOID Do not add a second product, a hand entering frame, or on-screen text — each is a separate subject Veo has to reconcile with the single rotating product already described, and product clips are the category most likely to show warped geometry when two subjects are asked to share a frame. Do not describe the surface as "shiny" without naming what it reflects; an unspecified reflection is where Veo most often renders smeared, illegible detail instead of a clean highlight. OUTPUT The finished five-layer prompt, ready to paste into Veo as-is, followed by one line naming which single camera move and which single action were chosen, so it is checkable against the "exactly one of each" rule above before generating.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Veo 3.1 generates native audio directly from the text of the prompt, which is why the audio layer here is written as a literal sound description rather than a mood word — "a low mechanical hum climbing half a tone" gives the model an actual waveform-shaped instruction, while a word like "tense" gives it nothing concrete to render and it will either default to silence or invent a generic tone that has no relationship to the visual beat it is meant to land on. Keeping the five layers in a fixed order as short separate sentences rather than one dense paragraph matters because Veo appears to condition on each descriptive clause somewhat independently rather than parsing full grammatical dependency — when "soft" and "warm" and "shiny" are all crammed into one sentence describing three different things at once, the model has to guess which adjective belongs to the light, the material, or the mood, and it guesses wrong often enough that separating the layers measurably reduces that ambiguity. Naming exactly one camera movement is the single biggest lever for output quality on 5-10 second clips — a prompt that asks for a dolly-in while also implying an orbit or a crane is the most common cause of warped product geometry and inconsistent shape across the frames, because the model is trying to satisfy two spatial instructions that pull the virtual camera in incompatible directions at once, and the product itself absorbs the resulting error as visible distortion. Naming the start shot type and end shot type explicitly (medium to close-up) rather than just "dolly in" gives Veo two concrete anchor points to travel between across the fixed eight-second duration, which is a meaningfully different instruction than an unanchored "camera moves closer" that leaves how far and how fast entirely to the model's own defaults. Finally, treating each aspect ratio as its own generation rather than a crop is a real constraint of how the model composes a shot: the rim light position, the pedestal's negative space, and the dolly's calculated endpoint are all composed relative to the frame boundaries it was told to fill, so a 16:9 clip cropped down to 9:16 after generation loses the composed edges rather than reframing them, which is different from how a still photograph crops.
What you get back
An 8-second clip: the earbuds case starts in a medium shot against near-black, a soft key light sweeping across its matte surface as it rotates on the concrete pedestal; the camera glides inward to a tight close-up on the hinge line as a low mechanical hum climbs half a tone under a constant faint metallic resonance, ending on a still, sharply lit close frame.
Verified against
Veo 3.1 3.1 · 2026-07-21
Changelog
- 2026-07-21 — Initial publish, verified against Veo 3.1 native-audio generation.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
