Research

Verified against ChatGPT · 2026-08-11

Build an evidence table that grades each source's reliability instead of just listing it

Produces a structured evidence table with a reliability grade and stated basis for that grade per source, so a reader can weigh conflicting claims instead of treating every row as equally trustworthy.

ChatGPT (GPT-5.1)3 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

Build an evidence table from the sources below. Every row needs a reliability grade with a stated reason — not just the claim itself — because a table that lists ten sources with no weighting implies they're all equally trustworthy, which is rarely true.

CLAIM OR QUESTION BEING EVALUATED
Whether a four-day work week measurably reduces employee burnout without reducing output.

SOURCES
A 2023 UK pilot study report, three company blog posts describing their own four-day-week experiments, and two news articles covering the pilot.

GRADING CRITERIA I CARE ABOUT
Sample size and whether the source has a financial incentive to report a positive result matter more to me than how recent it is.

For each source, extract: the specific claim it makes relevant to Whether a four-day work week measurably reduces employee burnout without reducing output., the type of source it is (peer-reviewed study, industry report, vendor material, journalism, forum/anecdote), and a reliability grade (High / Medium / Low) based on Sample size and whether the source has a financial incentive to report a positive result matter more to me than how recent it is.. State the one-sentence reason for each grade — sample size, disclosed methodology, conflict of interest, recency, or lack of any of those. Where two sources give conflicting claims, note the conflict directly in an adjacent row or a flagged note rather than letting it pass silently. Sort or group the table so the highest-reliability sources are easy to find rather than randomly interspersed with low-reliability ones.

WHAT NOT TO DO
Do not grade every source as "Medium" as a way of avoiding a real judgment call — if the grading criteria genuinely can't distinguish two sources, say so explicitly rather than defaulting to the middle grade as a non-answer.

OUTPUT FORMAT
A table with columns: Source | Claim | Source Type | Reliability Grade | Basis for Grade. Follow the table with a two-sentence summary of what the highest-reliability sources actually support, separate from what the full source list as a whole seems to suggest.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

An evidence table with no grading column implicitly treats every row as equally weighted evidence, and readers tend to take a table's structure at face value — ten rows of claims reads as ten independent confirmations, even if seven of them trace back to the same underlying press release or the same company's own marketing. Requiring an explicit reliability grade with a stated basis forces the model to actually interrogate each source's provenance rather than just extracting its headline claim, which is the step that catches a vendor blog post dressed up as independent evidence or a small pilot study being cited with the same confidence as a large peer-reviewed one. Explicitly forbidding a default-to-Medium pattern matters because "Medium" is the safest grade to assign when a model wants to avoid committing to a judgment — it sounds appropriately cautious without actually being useful, and a table where most rows land on Medium has quietly abdicated the one job the table exists to do. Separating the summary of what the highest-reliability sources support from what the full source list suggests as a whole directly addresses the averaging failure mode: if three low-quality sources and one high-quality source disagree, a naive synthesis leans toward the majority view by source count, when the correct read is usually to weight the single well-evidenced source more heavily than three weaker ones saying the opposite.

What you get back

Source: 2023 UK four-day-week pilot report | Claim: burnout scores dropped, output held steady | Type: independent pilot study | Grade: High | Basis: large multi-company sample, pre-registered methodology, no financial stake in outcome. Source: Company X blog post | Claim: "huge productivity gains" | Type: vendor/company self-report | Grade: Low | Basis: single company, no control group, incentive to report success publicly.

Verified against

ChatGPT GPT-5.1 · 2026-08-11

Changelog

  • 2026-08-11 Initial publish, verified against ChatGPT GPT-5.1.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All Research prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY