Verified against Midjourney · 2026-08-04
Animate a finished Midjourney still into a short native video clip
A brief for Midjourney's own native image-to-video feature that turns an already-generated still into a short motion clip — describing camera movement and subject motion the way you would for the still image itself, tuned to what a several-second extension of a single frame can plausibly do.
The prompt
Ready to copy — highlighted parts are example details you can swap.
SOURCE IMAGE a chosen, upscaled still of a woman in a red coat walking down a cobblestone street at dusk, mid-stride — the already-generated, already-chosen still this clip animates from; the video feature works from this specific frame, not from a fresh text prompt describing the whole scene again. MOTION TO ADD her continuing her walking stride naturally, coat hem swaying slightly with the movement — name one primary motion, not several competing ones. A still that shows a person mid-stride and a curtain in the background can plausibly animate the person continuing to walk and the curtain drifting in a breeze at the same time, since those are both small, physically consistent continuations of what the still already implies — but asking for a completely different camera angle, a new character entering, and a full scene change all within a several-second clip from one starting frame is asking the feature to do more than that short a clip generated from a single image can coherently deliver. CAMERA MOVEMENT, IF ANY locked camera, no movement, focus entirely on her walking motion — state explicitly if the camera should stay locked or move; a locked camera focusing purely on subject motion is generally the safer, more coherent choice for a first attempt from a still with complex detail, since a moving camera compounds however much interpretation the model already has to do to animate the subject itself. MOTION INTENSITY lower motion intensity, to keep her face and coat detail close to the original still — Midjourney's video controls typically offer a choice between a lower-motion, more subtle animation and a higher-motion, more dramatic one; lower motion keeps the result closer to the original still's exact detail and composition, higher motion allows more visible change but risks the subject drifting further from what the still originally showed. CLIP LENGTH start with the first short segment, extend only once if the walking motion looks natural — clips extend in short increments; treat the first generated segment as the base and only extend further if the motion established in that first segment is actually working, since extending a clip whose initial motion already looks wrong just compounds the same problem across a longer duration. OUTPUT A short video clip beginning from the exact source still, with the described motion applied, at the chosen motion intensity and camera behavior. IF THE MOTION LOOKS WRONG OR THE SUBJECT WARPS Lower the motion intensity before rewriting the motion description — a subject that warps or loses coherence during animation is very often a motion-intensity problem, where the model has been given more freedom to change the frame than the specific source image's detail level can support without breaking down, not a wording problem with how the motion was described.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Midjourney's native video feature animates from a specific already-generated frame rather than from a fresh text description of the whole scene, which means the actual creative decision that mattered most — composition, lighting, subject appearance — was already locked in when the still was generated and chosen; the video step's only real job is adding believable motion consistent with that frame, not re-deciding what the scene looks like. Framing the brief around one primary motion, rather than several competing changes, matters because a several-second clip generated from a single starting frame has a narrow amount of physically plausible change it can introduce before it stops looking like a continuation of that frame and starts looking like a different, disconnected scene stitched on afterward. Explicitly choosing whether the camera moves, rather than leaving it unstated, targets a real compounding-difficulty problem: animating subject motion from a still frame is already an interpretive task, since the model has to infer plausible continued movement from one instant; adding independent camera movement on top asks it to solve two overlapping problems in the same short clip, which is why a locked camera is the more coherent, more reliable choice for a first attempt, especially on a still with a lot of fine detail (hands, textured fabric, an expressive face) that is easiest to lose coherence in in exactly the way this kind of feature commonly does. The motion-intensity distinction between "lower, closer to the original" and "higher, more dramatic but more drift-prone" reflects the actual mechanical trade the setting makes: the model is not simply choosing how fast something moves, it is choosing how much freedom it has to reinterpret and regenerate detail across the clip's duration, so a higher setting on a detailed still is spending that freedom on can-be-imperceptible drift in exactly the areas (a face, a hand, fine fabric texture) that were the most carefully chosen parts of the original still. Building a clip up in short increments, and only extending further once the initial motion is confirmed to be working, rather than requesting a long clip in one go, exists because each extension segment continues from wherever the previous one left off — if the very first segment already shows the subject warping or the motion reading as physically implausible, extending it further only compounds that same defect across more seconds of footage rather than giving the model a chance to correct course, since there is no correction mechanism partway through an extension; the fix has to happen by regenerating from the source still with adjusted settings, not by continuing forward from a flawed first segment.
What you get back
A short clip of the woman continuing her walking stride down the cobblestone street, her coat swaying naturally with the motion, camera locked and steady, her face and coat detail staying close to the original chosen still rather than visibly warping over the clip's duration.
Verified against
Midjourney v7 · 2026-08-04
Changelog
- 2026-08-04 — Initial publish, verified against Midjourney v7 native image-to-video for still-to-motion animation.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
