Verified against Claude · 2026-07-27
Write a thumbnail concept brief a designer or image model can actually execute
Turns a video topic and title into a specific composition brief — subject, expression, contrast, text treatment, and what the title is left to say instead of the image — rather than a vague "make it eye-catching" note.
The prompt
Ready to copy — highlighted parts are example details you can swap.
You are writing a thumbnail concept brief for a YouTube video — a specification a designer, or an AI image tool, can execute directly, not a vague mood description. The brief must describe exactly what appears in the frame, because "make it pop" and "eye-catching" give the executor nothing to actually draw. VIDEO TOPIC AND TITLE A video testing whether a $30 knife set can outperform a $300 professional set on five real kitchen tasks FINAL OR WORKING TITLE I Replaced My $300 Knives With $30 Ones for a Week CHANNEL THUMBNAIL STYLE High-contrast close-ups on a plain dark background, bold white sans-serif text, red accent circle used as a recurring callout device KEY VISUAL ELEMENT AVAILABLE Close-up shot of the cheap knife slicing cleanly through a tomato, and a separate shot of the expensive knife visibly chipped on its edge DEVICE CONTEXT Mostly mobile home-feed and suggested-videos rail, surrounded mostly by warm-toned cooking-channel thumbnails BRIEF RULES Specify the single focal subject first — a face, an object, a before/after split — and justify why it, not something else in the video, is the strongest visual anchor; a thumbnail with two or three competing focal points loses to one with a single unmistakable subject at the small size thumbnails are actually viewed at. If a face is the focal subject, specify the exact expression in concrete terms — not "excited," but "eyebrows raised, mouth open mid-word, eyes wide, as if reacting to something just off-frame" — since a vague emotion word gives no direction and a genuinely legible expression at thumbnail size needs to be exaggerated well past how a person would actually look at that moment. Decide the text-on-thumbnail question deliberately: if the title text is already strong on its own, specify zero or minimal thumbnail text and say why, rather than defaulting to restating the title as an overlay — but if you do include text, cap it at three to five words in a typeface and color that stays legible at roughly 120 by 68 pixels, the actual size a thumbnail renders at on a mobile suggested-videos rail. Specify contrast and color deliberately relative to what a viewer will actually see it against — a thumbnail sitting in a feed of mostly blue-toned tech thumbnails should not also default to blue, and the brief should name the actual competing colors it needs to stand out from, not describe colors in isolation. State exactly what information the thumbnail is deliberately leaving for the title to carry, and what the title is leaving for the thumbnail to carry — the two should not duplicate the same fact. OUTPUT FORMAT 1. One-paragraph concept statement naming the focal subject and the single idea the thumbnail communicates in under one second. 2. A composition spec: subject placement in-frame, expression or object detail, background treatment, color and contrast direction relative to feed context, any text with exact wording, size, and placement. 3. One line on what the title is left to say that the image deliberately does not. 4. If an AI image tool will generate this, a ready-to-use generation prompt matching the spec exactly.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Specifying a single focal subject and rejecting multi-element compositions is grounded in the actual viewing conditions a thumbnail competes under: it is evaluated at roughly 120 by 68 pixels on a mobile suggested-videos rail, in a fraction of a second, alongside a dozen competitors doing the same thing — a composition with two or three points of interest that reads clearly on a full-size monitor becomes visual noise at that real render size, where only a single unmistakable shape survives the scroll-past glance. Demanding a concretely described expression rather than an emotion label addresses a specific gap between how a designer or image model interprets "excited" versus what is actually legible at thumbnail scale — a genuinely subtle, true-to-life expression of surprise disappears at that size, so the brief has to ask for the exaggerated, borderline-cartoonish version of the expression that reads instantly, which is a deliberate photographic choice, not an inaccuracy. Making the text-or-no-text decision explicit, with a real character cap and a legibility target tied to the actual render size, targets the common default failure of restating the title as a thumbnail overlay — when both elements say the same thing, the creator has spent two of the two available attention slots on one message instead of two, which is a wasted opportunity distinct from a bad thumbnail; a thumbnail with no text at all, letting a strong title carry the words, is frequently the higher-performing choice, but only if that's a deliberate call rather than an accidental default. Requiring contrast and color decisions relative to the actual feed context, not in isolation, matters because a thumbnail's job is fundamentally comparative — it needs to look different from whatever is next to it in a specific feed at a specific moment, and a well-designed thumbnail that happens to match the dominant color of everything around it will underperform an objectively less polished one that simply stands out, which is a context-dependent judgment a generic color-theory answer can't make without knowing what it's actually competing against. Splitting what the image says from what the title is left to say turns thumbnail and title into one coordinated two-part message instead of two independently optimized pieces that might duplicate or, worse, contradict each other.
Verified against
Claude Sonnet 4.6 · 2026-07-27
ChatGPT GPT-5.1 · 2026-08-03
Changelog
- 2026-07-27 — Initial publish, verified against Claude (Sonnet 4.6) and ChatGPT (GPT-5.1).
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
