Research

Verified against ChatGPT · 2026-08-10

Pressure-test a draft survey for leading questions and bad scales before it goes out to real respondents

Reviews a drafted survey question-by-question for leading phrasing, double-barreled questions, and scale mismatches, and rewrites only the flagged items — so you're not re-explaining your whole study to get one clean pass on wording.

ChatGPT (GPT-5.1)4 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

Review this draft survey instrument for wording problems before it goes out to respondents. I want a question-by-question audit, not a general commentary on survey design.

SURVEY GOAL
Deciding whether to keep or drop the optional onboarding webinar based on perceived usefulness.

DRAFT QUESTIONS (numbered, with response scale noted)
1. Did you find the onboarding webinar helpful and well-organized? (Yes/No) 2. How much did you enjoy using the product in your first week? (1-5 scale)

RESPONDENT POPULATION
New customers within their first 30 days, mostly small business owners with limited time to respond.

KNOWN SENSITIVE TOPIC
Whether they've considered canceling their subscription.

For every question, check it against four specific failure types and only flag the ones that actually apply — don't pad the audit with a pass/fail note on every category for every question:

1. LEADING PHRASING — wording that signals a preferred answer (e.g. "How much did you enjoy..." assumes enjoyment happened).
2. DOUBLE-BARRELED — asking two things in one question so a single answer can't be interpreted (e.g. "Was the onboarding fast and clear?" — fast and clear can diverge).
3. SCALE MISMATCH — a response scale that doesn't fit the question's actual range of possible answers, including scales missing a genuine "does not apply" option where one is needed.
4. SENSITIVE-TOPIC HANDLING — if Whether they've considered canceling their subscription. applies to a question, check whether the phrasing and answer options let a respondent answer honestly without visible judgment, and whether an opt-out is available.

For every flagged question, give the specific rewrite, not just the diagnosis — a respondent-facing survey needs the fixed wording ready to drop in, not a description of what's wrong with it. If a question has no issues, just list it as clean; do not manufacture a minor note to justify commentary on it.

After the question-by-question audit, do one pass across the whole instrument for question-order effects — whether an earlier question could prime how New customers within their first 30 days, mostly small business owners with limited time to respond. answers a later one — and flag only genuine ordering risks, not hypothetical ones.

OUTPUT FORMAT
A table: question number, issue type (or "clean"), and rewrite where applicable. Followed by a short list of any order-effect risks across the instrument as a whole.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

Naming four specific, checkable failure categories instead of asking for a general wording review matters because an open-ended "review this survey" prompt tends to produce GPT-5.1's default critique register — broad, hedged observations like "consider whether questions are neutral" — which sounds thorough but gives nothing a survey author can act on directly; a named failure type with a definition forces a verdict per question instead of vague commentary. The instruction to only flag categories that actually apply, rather than reporting pass/fail on every category for every question, prevents the output from ballooning into a padded audit where genuine issues get buried under routine "no issue here" notes — this is a common failure mode when a model is given a checklist and defaults to exhaustively confirming every box rather than surfacing only what matters. Requiring the actual rewrite rather than just the diagnosis closes the gap between "here's what's wrong" and something a researcher can paste directly into the field-ready instrument, which matters because leading and double-barreled phrasing is often genuinely hard to fix well, and a diagnosis without a fix just shifts the hard part back onto the person who asked for help. The sensitive-topic handling check exists because default survey phrasing frequently telegraphs an expected answer on topics respondents already feel judged about, and a model reviewing generically won't reliably catch this unless the specific topic is named as something to check for — naming it turns a general awareness into a targeted, checkable pass.

What you get back

Q1: double-barreled (helpful and well-organized can diverge) — rewrite as two separate questions, one per attribute, each on a 1-5 scale. Q2: leading phrasing (assumes enjoyment) — rewrite: 'How would you describe your experience using the product in your first week?' with a neutral 1-5 scale labeled from 'very negative' to 'very positive.' Order-effect risk: placing the cancellation-consideration question immediately after a satisfaction question may anchor respondents' answer — consider separating them with a neutral filler question.

Verified against

ChatGPT GPT-5.1 · 2026-08-10

Changelog

  • 2026-08-10 Initial publish, verified against ChatGPT GPT-5.1.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
All Research prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY