Verified against ChatGPT · 2026-08-08
Write a Deep Research brief that tells ChatGPT which sources outrank which before it starts pulling
Produces a Deep Research brief with an explicit source-priority order and named low-trust categories, so a long autonomous run doesn't quietly treat a vendor blog post and an independent benchmark as equally authoritative.
The prompt
Ready to copy — highlighted parts are example details you can swap.
Write a Deep Research brief for a multi-step autonomous research run. The brief needs to do more than state the topic — it needs to tell the research process which kinds of sources to trust more than others, because a long autonomous pass will pull from dozens of sources and has no way to know my priorities unless I state them upfront. RESEARCH TOPIC How reliable is on-device AI transcription accuracy for accented English speech compared to cloud-based transcription in 2026? TRUSTED SOURCE TYPES, IN ORDER 1) Independent benchmark studies with published methodology, 2) academic papers, 3) engineering blog posts from the model vendors themselves, 4) tech journalism. LOW-TRUST OR SELF-INTERESTED SOURCES TO FLAG Vendor press releases, affiliate-linked "best AI transcription tools" roundup articles, and any source that doesn't disclose its test methodology. RECENCY REQUIREMENT Published within the last 12 months — this space moves fast enough that older benchmarks are close to meaningless. CONTRADICTION HANDLING When sources disagree, do not average their claims into a middle-ground summary — that produces a number or conclusion no actual source supports. Instead, state each source's position separately, note which is higher-priority per my ordering, and flag the disagreement explicitly rather than resolving it silently. Any claim sourced only from a vendor's own marketing material or a paid placement must be labeled as such in-line, not folded into the narrative as if it were independent. RECENCY RULE Apply Published within the last 12 months — this space moves fast enough that older benchmarks are close to meaningless. strictly — a source published before that window should only be cited if no more recent source covers the same claim, and that fact should be stated. OUTPUT FORMAT Produce the final brief as: (1) the research question restated in one line, (2) the source priority order as a numbered list, (3) explicitly named low-trust categories, (4) the recency cutoff, (5) one line instructing the research process to log which tier each major claim in its final report came from.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Deep Research-style tools operate by issuing many search queries and synthesizing across whatever surfaces, and without an explicit priority ordering the synthesis step treats every retrieved passage as roughly interchangeable evidence — a benchmark with a disclosed methodology and a vendor's own marketing copy both just become "a source that says X," and the summarization step has a structural tendency to average disagreeing numbers into a plausible-sounding midpoint that no individual source actually reported, because a middle value reads as the most defensible synthesis even though it's evidence-free. Naming the low-trust categories upfront gives the model a concrete pattern to check retrieved content against — "is this a vendor's own claim about its own product" is a checkable question the model can apply source-by-source, whereas "be skeptical of biased sources" with no examples leaves the judgment call underspecified and inconsistently applied across the dozens of sources a long run touches. The recency window matters because search-based retrieval doesn't rank by publish date unless told to, so an older, more heavily-cited source can outrank a newer, more accurate one simply because it has more inbound links and mentions; stating the cutoff explicitly forces the model to check dates rather than default to citation volume as its proxy for authority. Requiring a per-claim source tier in the final output creates an audit trail that lets you catch it if the process quietly violated its own priority order somewhere in a long run.
What you get back
Research question: How does on-device transcription accuracy for accented English compare to cloud transcription in 2026? Findings — Tier 1 (independent benchmark, published methodology): a 2026 university study found on-device models trailed cloud by 4-7 WER points on non-native accents. Tier 3 (vendor blog, flagged as self-interested): Vendor X's own post claims near-parity, but discloses no test set. Disagreement noted: the two claims conflict; the independent benchmark is weighted higher per the stated priority order.
Verified against
ChatGPT GPT-5.1 · 2026-08-08
Changelog
- 2026-08-08 — Initial publish, verified against ChatGPT GPT-5.1 Deep Research.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
