Verified against Nano Banana / Gemini 3.1 Flash Image · 2026-07-22
Put a specific garment on a specific person without changing their identity
A virtual try-on edit that fits an uploaded garment onto an uploaded person photo, treating the person's face, body, and pose as fixed and the garment's own fabric physics — drape, wrinkle, fit — as the only thing that should adapt to their body.
The prompt
Ready to copy — highlighted parts are example details you can swap.
You are compositing a garment onto a person in an uploaded photo — a virtual try-on, not a new photoshoot. The person's identity and pose are fixed; the garment must adapt to fit their body, not the other way around. PERSON PHOTO a three-quarter-length shot of a woman standing, arms relaxed at her sides, facing slightly left, soft window light from the right. Their face, exact pose, body proportions, skin tone, and the photo's existing lighting must remain completely unchanged. GARMENT TO APPLY a cropped, oversized denim jacket in mid-wash blue with a soft, slightly stiff cotton-denim texture. Fit this garment onto the person as it would actually drape on a body their size and in their current pose — sleeves following their actual arm position, hemline falling according to real gravity and fabric weight, not floating independent of their posture. FIT AND POSE NOTES jacket should look slightly loose through the shoulders, sleeves pushed up to just below the elbow. Where the person's current pose would naturally create fabric folds, bunching, or stretch — an arm bent at the elbow, a seated pose compressing a hem — render those folds convincingly rather than showing the garment perfectly flat as if on a mannequin. BACKGROUND TREATMENT keep the original background exactly as it is — no replacement needed. Whatever you do with the background, it must not require altering the person's pose, crop, or the photo's original lighting angle to accommodate it. IDENTITY PRESERVATION her face, hairstyle, exact pose, and skin tone must be pixel-identical to the source photo. This is a hard rule: do not adjust facial features, body shape, skin tone, or apparent age to "match" the new garment stylistically. The only thing that should look different between the source photo and this output is the garment itself and, where genuinely necessary, the parts of the body it now covers or reveals differently than what they were wearing before. LIGHTING CONSISTENCY The garment must pick up highlights and shadows consistent with the original photo's existing light source and direction — if the source photo has a single soft key light from one side, the new garment's fabric should show believable highlight and shadow from that same side, not generic even studio lighting that doesn't match the rest of the image. OUTPUT One composited image. If the requested garment's cut or style genuinely conflicts with the person's current pose in a way that can't be resolved believably — for example, a garment that requires a standing pose applied to a photo where they're seated at a desk with only their upper body visible — say so explicitly and describe what pose would actually be needed instead of forcing an unconvincing result.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Nano Banana's multi-image input handling lets it condition a single output on more than one uploaded reference — a person photo and a separate garment reference — which is the specific capability this prompt is built around, but that same flexibility means the model has to be told explicitly which reference is the fixed anchor and which one is the adaptable element, or it will sometimes resolve conflicts between the two by quietly adjusting the person's pose or proportions to make the garment fit more easily instead of adapting the garment to the person, which is the opposite of what a real try-on needs to demonstrate. Second, the instruction to render fabric folds according to the person's actual pose — rather than a flat, mannequin-style drape — targets a specific and common quality failure in AI-generated try-on images: a garment rendered as if it were laid on a flat surface and then pasted onto a photo of a moving body reads as obviously synthetic the instant a viewer notices the fabric isn't responding to the arm's bend or the seated compression at the waist, which is precisely the kind of physical detail a model trained on real photographs has actually learned to render when it's explicitly told the pose matters to the outcome. Third, the lighting-consistency instruction is the detail that most separates a convincing composite from an obviously fake one in practice — a garment lit with generic, source-agnostic studio lighting sitting on a person photographed under a single soft directional light creates a visible mismatch a viewer registers as "something's off" even without being able to name why, and stating the original light's direction as a constraint the garment must also obey is what closes that specific, hard-to-articulate but easy-to-spot gap. Finally, giving the model explicit permission to refuse a physically implausible combination — a full-length garment applied to a seated, upper-body-only crop — rather than forcing a best-effort result matters because a silently-forced unconvincing render wastes the generation and gives no useful signal about what to change, whereas a stated refusal with a concrete alternative pose recommendation turns a dead end into an actionable next step.
Verified against
Nano Banana / Gemini 3.1 Flash Image Gemini 3.1 Flash Image · 2026-07-22
Changelog
- 2026-07-22 — Initial publish, verified against Nano Banana (Gemini 3.1 Flash Image) on a jacket try-on over a standing three-quarter photo.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
