ChatGPT

Verified against ChatGPT · 2026-07-02

Get a structured, honestly-flagged read on an uploaded image

A vision prompt that forces ChatGPT to flag illegible or ambiguous parts of an image explicitly, instead of quietly guessing a plausible value for text or details it cannot actually confirm.

ChatGPT (Vision, GPT-5.1)

The prompt

Ready to copy — highlighted parts are example details you can swap.

Look at the image I just uploaded and answer using only what is actually visible in it.

WHAT I NEED FROM THIS IMAGE
The vendor name, invoice number, and every line-item amount

OUTPUT FORMAT
A table with columns: field, value, confidence

RULES
- If any text in the image is small, blurry, cut off, or otherwise hard to read with confidence, say so explicitly next to that item instead of guessing a plausible value.
- Do not infer information that is not visible just because it is typical or expected, for example, do not assume a total that is not shown just because the line items suggest one.
- If the image contains more than one distinct thing that could match The vendor name, invoice number, and every line-item amount, list all matches instead of silently picking the most likely one.
- End with a one-line confidence note: high, medium, or low, and what would resolve it if not high, for example a higher-resolution version or a different angle.
Customize the highlighted detailsoptional — the prompt above already works

Why this works

Vision models are documented to hallucinate plausible values for illegible or occluded regions of an image rather than reporting uncertainty by default, because the underlying generation objective favors a fluent, complete-looking answer over an honest gap, the same fluency-over-accuracy tendency text generation has, applied to pixels instead of facts. Instructing it to flag illegibility per item, rather than trusting a single end-of-response caveat, keeps the uncertainty attached to the specific value it belongs to instead of a blanket disclaimer that does not tell you which number to actually go double-check. The rule against inferring a typical value targets a specific vision failure mode on documents like receipts and invoices: the model has seen enough receipts with a printed subtotal that summing line items to a plausible total, when no total field is actually shown, is a completed pattern it is inclined to finish even though nothing asked it to compute anything. Asking for all matches rather than the single most likely one prevents silent disambiguation on an image with, say, two dates or two totals visible, where a confident wrong guess is worse than a surfaced ambiguity you can resolve yourself in two seconds.

What you get back

| field | value | confidence | |---|---|---| | vendor name | Riverside Office Supply | high | | invoice number | INV-08841 | high | | line item 3 amount | $4?.50, last digit obscured by a fold in the paper | low | Confidence note: medium overall, a straighter photo of the bottom third of the invoice would resolve the one unclear amount.

Verified against

ChatGPT GPT-5.1 (Vision) · 2026-07-02

Changelog

  • 2026-07-02 Initial publish, verified against ChatGPT GPT-5.1 Vision.

Building this for real?

This is a free starting point. If you'd rather have what Scult builds built and running for your business, that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All ChatGPT prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY