DALL·E / GPT Image

Verified against ChatGPT · 2026-08-08

Generate a paid-social ad hero image with the headline already legible inside the frame

A hero-image brief for GPT Image that treats the ad's headline as text rendered directly into the pixels, not a caption bolted on after export, plus a conversational follow-up step for spinning off square, 4:5, and Story-ratio crops from the same approved image instead of re-rolling each one from scratch.

ChatGPT (GPT Image)7 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

PRODUCT OR OFFER
A 3-month specialty-coffee subscription box, first box 50% off

HERO VISUAL
an overhead shot of a steaming ceramic mug of pour-over coffee beside a torn-open kraft-paper shipping box with beans spilling out

HEADLINE TO RENDER ON THE IMAGE
Render this exact line of text directly into the image, not as a caption you describe afterward: "Your first box is half price". Set it in a bold, high-contrast sans-serif, large enough to stay readable at thumbnail size in a phone feed, and keep it to this one line — do not paraphrase it, shorten it, or add a second line the brief didn't ask for.

WHERE THE TEXT SITS
upper third of the frame, over the plain wood tabletop to the left of the mug, not over the beans or the box. Leave this area of the frame genuinely clear of busy detail before the text goes in — a headline dropped over a cluttered background is the single most common reason on-image ad text ends up unreadable, and no amount of bolding the type fixes a background that was never cleared for it.

BRAND COLOR AND MOOD
warm terracotta and cream tones, soft morning light, cozy rather than corporate

WHAT TO KEEP OUT OF FRAME
No second product, no competitor-style logo, no stock-photo watermark, no small print or fine-print text anywhere else in the image — one headline, rendered once, is the only text this image should contain.

FIRST PASS
Generate this as one image at 1:1, for an Instagram feed post.

FORMAT VARIANTS (do this after the first image is approved, in the same conversation)
Once the hero image above looks right, ask for that same approved image re-cropped and re-composed — not regenerated from a fresh description — into a 4:5 feed crop and a 9:16 Story/Reels crop, with the headline repositioned so neither platform's UI overlay covers it. Reference "the image above" explicitly rather than re-typing the brief as a new prompt; re-describing it as a fresh generation risks a different product angle, a different headline placement, or a shifted color grade landing in the second format even though nothing about the ad was supposed to change.

OUTPUT
One primary hero image with the headline legible inside the frame, followed by the additional format variants generated as in-conversation edits of that same approved image rather than independent re-generations.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

GPT Image generates text as part of the same native image-token stream as the rest of the picture rather than compositing it on afterward the way older diffusion-only pipelines effectively had to, which is the specific reason a short, exact headline string rendered directly into the frame comes back legible far more often than it did on prior-generation DALL·E models that reliably garbled anything beyond a couple of words — but that reliability drops fast once the surrounding composition is already busy where the text needs to sit, which is why the brief clears that area explicitly rather than trusting the model to find clean space on its own after the fact. The format-variant instruction leans on a specific mechanical difference between an in-conversation follow-up and a brand-new prompt: asking ChatGPT to re-crop "the image above" conditions the next generation on the actual approved pixels already in the thread, carrying forward the exact product angle, color grade, and headline styling that were just signed off, whereas typing a fresh description of the same ad into a new generation call has nothing to anchor it to that specific result and will re-synthesize the whole scene from language alone — close to the original, but rarely identical, and identical is the entire point once one version has already been approved for use. Naming everything that must stay out of the frame matters because GPT Image, trained on a huge corpus of real ad and product photography where competing badges, watermarks, and secondary products are common, will add one of those elements on its own initiative more often than a brief-writer expects unless it's told explicitly not to.

What you get back

A warm, morning-lit product shot with the exact line "Your first box is half price" rendered crisply over the clear tabletop area, followed by two additional crops of that same approved image in 4:5 and 9:16 with the headline nudged to stay clear of each platform's UI chrome.

Verified against

ChatGPT GPT Image (2026 release) · 2026-08-08

Changelog

  • 2026-08-08 Initial publish, verified against ChatGPT GPT Image (2026 release).

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All DALL·E / GPT Image prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY