Verified against Grok · 2026-08-02
Read a screenshot posted on X the way the thread actually means it
Analyzes an image attached to a specific X post together with the surrounding thread's own context and claims, so Grok's multimodal read catches what the screenshot actually shows rather than what the post's caption claims it shows.
The prompt
Ready to copy — highlighted parts are example details you can swap.
You are analyzing an image attached to a specific X post, using Grok's multimodal vision capability together with the actual surrounding thread context — the post's caption, the replies, and any competing interpretation already being argued about in the thread. Your job is to check what the image actually shows against what the post claims it shows, not to simply describe the image in isolation or accept the caption's framing as accurate by default. THE POST AND IMAGE A screenshot claiming to show an error message from a specific company's app, posted with the caption implying the company's service is currently broken for everyone THE CLAIM BEING MADE ABOUT IT The caption states this proves a widespread, ongoing outage affecting all users right now COMPETING INTERPRETATIONS IN THE THREAD Several replies argue the screenshot could be an old cached screenshot being reposted, since a similar error was reported on this app months ago; a couple of replies say the service is working fine for them right now WHAT WOULD SETTLE THE DISPUTE A visible current timestamp in the screenshot's own UI, or independent confirmation from other users experiencing the same error right now ANALYSIS RULES Describe only what is actually visible in the image first, independent of the caption's framing — timestamps, visible UI elements if it's a screenshot of an app or platform, any text legible in the image itself, signs of editing or cropping — before evaluating whether that description supports, contradicts, or is simply silent on the specific claim being made about it. Check for the specific, checkable signs of manipulation or selective framing that are common in screenshots circulating on social platforms: a visibly cropped edge that could be hiding contradicting context just outside the frame, inconsistent fonts or spacing suggesting an edited screenshot rather than a genuine capture, a timestamp or metadata detail that doesn't line up with the claimed date or context, or a caption that describes something the image itself does not actually show and requires the viewer to take on faith. State explicitly whether the image, on its own, is sufficient evidence for the claim being made, is consistent with the claim but doesn't independently prove it, or actually contradicts what's being claimed — these are three different findings, and treating consistent-with as equivalent to proves is the specific mistake that lets a genuinely ambiguous screenshot get treated as settled evidence. If the competing interpretations already circulating in the thread are visible to you, address each one directly against what's actually in the image, rather than only evaluating the original poster's framing and ignoring the pushback already happening underneath it. Name the specific additional detail — a fuller screenshot, the original unedited source, a second independent angle — that would actually settle the dispute if it existed, rather than declaring the matter closed based on what's currently visible if genuine ambiguity remains. OUTPUT FORMAT 1. Plain description of what's visible in the image, with no interpretation yet. 2. Verdict: sufficient evidence for the claim / consistent but not proof / contradicts the claim — pick one and defend it. 3. Response to each competing interpretation already circulating, addressed directly. 4. Any visible sign of cropping, editing, or inconsistency worth flagging. 5. What specific additional detail would actually settle this if the current image doesn't.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
The instruction to describe the image independent of the caption before evaluating the claim is the key structural safeguard here, because a caption primes interpretation — a vision model asked directly whether an image shows an outage is far more likely to read ambiguous visual evidence through the lens of the question it was just asked than one first asked to describe plainly what's visible and only then asked whether that description actually supports the claim, which is the same reason a careful human fact-checker separates description from interpretation as two distinct steps rather than one. The three-way verdict, sufficient evidence, consistent-but-not-proof, or contradicts, targets a specific and common sloppy move in screenshot-based claims: a single error screenshot is very often consistent with a wider outage without being remotely sufficient to prove one, since it's equally consistent with one user having one bad request, and collapsing that distinction into a binary supports-or-doesn't answer would force a genuinely ambiguous piece of evidence into a false confident bucket in either direction. Checking for specific, nameable manipulation signals, cropping, font inconsistency, a mismatched timestamp, rather than a vague sense of whether something looks edited, grounds the visual check in artifacts that are actually detectable from the image itself, which is the difference between a real forensic check and a model simply asserting confidence or suspicion with nothing underneath it that a skeptical reader could verify independently. Addressing competing interpretations already visible in the thread directly, rather than only evaluating the original poster's framing, matters because the replies underneath a disputed screenshot often already contain the actual counter-evidence, someone noting the error is months-old, someone confirming the service works fine for them right now, and an analysis that only engages with the caption while ignoring the pushback happening in the same thread is solving a narrower and less useful problem than the one actually in front of it. Naming the specific detail that would settle the dispute, rather than declaring a verdict as final when real ambiguity remains, keeps the analysis honest about the actual epistemic status of a single screenshot: a cropped or context-free image is frequently genuinely unresolvable on its own, and saying so plainly, with a concrete next check named, is more useful to someone trying to get to the truth than an overconfident verdict manufactured because a clean answer felt more satisfying to deliver than an honest not-settled-yet.
Verified against
Grok Grok 4.1 · 2026-08-02
Changelog
- 2026-08-02 — Initial publish, verified against Grok 4.1 multimodal vision with live thread context.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
