Flux

Verified against Flux · 2026-07-29

Transfer an exact pose and layout onto a new subject with Flux.1 Depth

A structural-conditioning prompt for Flux.1 Depth that generates an entirely new subject, styling, and scene while locking the exact pose, camera angle, and spatial layout of a reference photo — useful for turning a rough reference or stock photo into on-brand creative without redrawing the composition from scratch.

Flux.1 Depth [pro]6 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

DEPTH REFERENCE IMAGE
stock_yoga_pose.jpg — a stock photo of a person in a seated forward-fold yoga pose, shot from a low three-quarter angle. This image's depth map — the exact spatial arrangement of near-to-far surfaces, the pose, and the camera angle — is the fixed structural skeleton for this generation. The reference's own colors, lighting, and subject identity are not being kept, only its geometry.

WHAT CARRIES OVER FROM THE REFERENCE
Only the structural layout: the exact seated forward-fold pose, the low three-quarter camera angle, and the floor-to-ceiling spatial proportions of the room — the pose, the relative position of foreground and background elements, and the camera's height and angle. Nothing about the reference's actual appearance — its lighting, its color palette, who or what is in it — should influence the new image beyond this geometry.

NEW SUBJECT AND STYLING
a middle-aged man in athletic wear, calm expression, mid-stretch, occupying exactly the spatial position and pose the depth map defines.

NEW SCENE AND ENVIRONMENT
a minimalist home studio with a single potted plant in the far background and morning light through a sheer curtain, built around the same near, mid, and far depth layers as the reference but with entirely new content in each layer.

LIGHTING AND MOOD FOR THE NEW IMAGE
soft, warm morning backlight creating a gentle rim on the subject's silhouette. This is generated fresh for the new subject and scene — do not carry over any lighting direction implied by the reference image's own shadows unless the new lighting note happens to call for something similar.

HOW CLOSELY TO FOLLOW THE STRUCTURE
match the pose exactly, including hand and foot placement — this is for a fitness-app tutorial frame where the pose itself is the instructional content. State whether the pose must match exactly — useful for a product held in a very specific way — or loosely, useful when only the overall composition matters and not finger-level accuracy.

WHAT THE RESULT SHOULD NOT RESEMBLE
Because there is no negative-prompt field, say this positively: the final image should look like an entirely original photograph of the new subject in the new environment, sharing only its spatial composition with the reference — not a filtered or restyled version of the reference photo itself, and not a visible depth-map artifact such as banding or flattened perspective anywhere in the output.

CONDITIONING STRENGTH NOTE
If the platform or interface exposes a depth-conditioning strength value separately from this text prompt, treat structural_adherence_note above as the description of what that numeric strength should achieve, not a replacement for setting it — a high strength value with a loosely-worded adherence note, or a low strength value with an "exact match" instruction, will pull against each other, so keep the language here consistent with whatever strength setting is actually in use for this generation.

CHECK BEFORE FINALIZING
Compare the new subject's key contact points against the reference — where a hand touches a surface, where a foot bears weight, where the body bends — since these are the specific points where a depth map's structural guidance is easiest to satisfy loosely without actually matching, and an otherwise-good generation can still get exactly these points subtly wrong.

OUTPUT
One fully rendered image with new subject, new environment, and new lighting, whose composition and pose match stock_yoga_pose.jpg — a stock photo of a person in a seated forward-fold yoga pose, shot from a low three-quarter angle's structural layout exactly.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

Depth conditioning works by extracting a depth map from the reference and feeding it as a structural constraint alongside the text prompt, which means the model is being given two genuinely separate inputs — geometry from the image, everything else from the words. The prompt's job is to be explicit about that split, which is exactly why WHAT CARRIES OVER FROM THE REFERENCE names the geometry alone and explicitly disclaims the reference's lighting, color, and subject identity; without that disclaimer, some of the reference's original color and lighting character can bleed through as an unwanted stylistic echo. Naming structural_adherence_note explicitly — exact versus loose — matters because depth-conditioning strength is not a fixed setting in practice: a fitness-pose tutorial needs finger- and joint-level accuracy preserved, while a use case that only needs "a similar layout of foreground subject and background environment" can tolerate loose adherence that gives the new subject's own proportions more natural room to render. Not stating which one is needed leaves the model to guess at a default strength that may be wrong for the specific job. The explicit "should not resemble a filtered version of the reference" line targets a real failure mode of structural-conditioning models generally: when the depth signal is strong and the new-content prompt is comparatively thin, the output can end up reading as a restyled version of the original photo rather than a genuinely new one. Stating the new subject, environment, and lighting in as much concrete detail as a from-scratch text-to-image brief would need is what prevents the reference from dominating more than its intended structural role. Warning against visible depth-map artifacts matters because a depth map is itself a lossy, low-fidelity representation of 3D space, and a poorly conditioned generation can visibly show that flattening — especially at object boundaries where the depth map's edges were imprecise. Naming it as an explicit thing to avoid gives the model a concrete failure to check its own output against, rather than leaving it as an invisible risk nobody flagged.

What you get back

A new photograph of a calm-looking man mid-stretch in a minimalist home studio, in the exact seated forward-fold pose and low three-quarter camera angle from the original yoga stock photo, with entirely new lighting, subject, and environment — none of the stock photo's original color or lighting carried over.

Verified against

Flux Flux.1 Depth [pro] · 2026-07-29

Changelog

  • 2026-07-29 Initial publish, verified against Flux.1 Depth [pro] for pose-locked, fully restyled generation from a stock-photo reference.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
Get a clean, ecommerce-ready product shot from Flux.2 with no negative promptsA structured product-photography brief for Flux.2 that produces a genuinely clean, seamless studio background purely through positive description — because Flux has no negative-prompt field to simply tell it "no clutter, no props, no shadows."Flux.2Flux.1 [pro]2026-07-22Brief Flux.2 for an editorial lifestyle photograph, steered by positive descriptionA camera-and-film-aware natural-language brief for Flux.2 that gets a genuine editorial-magazine look — because Flux has no negative-prompt field, unwanted elements have to be steered out through what you describe, not an exclusion list.Flux.22026-07-24Design a scroll-stopping social graphic on Flux.2 with real reserved caption spaceA platform-aware brief for a social campaign graphic on Flux.2 that reserves genuinely clean space for a text overlay added later in a design tool — achieved by describing the empty area explicitly, since Flux has no exclusion field to keep text or clutter out of it.Flux.22026-08-01Keep a character consistent across multiple scenes with Flux.1 KontextA sequential in-context editing workflow for Flux.1 Kontext that locks a character's face, outfit, and proportions across a series of new scenes by editing forward from one reference image each time, instead of re-describing the character from scratch and getting a visibly different person every generation.Flux.1 Kontext [pro]Flux.1 Kontext [dev]2026-07-26
All Flux prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY