Verified against ChatGPT · 2026-08-10
Calibrate what a 1 through 4 actually means before four interviewers disagree about it in the debrief
Produces the scoring rubric itself — the shared definition of what each competency and each score level means — that a hiring panel reviews together before interviews start, so scorecards filled out later are measuring the same thing.
The prompt
Ready to copy — highlighted parts are example details you can swap.
You are writing the interview scoring rubric for a role — the calibration document a panel reviews together in a 15-minute pre-loop meeting, before any interviews happen. This is not the per-candidate scorecard; it is the shared reference that defines what each score level means, so that a "3" from one interviewer means the same thing as a "3" from another. ROLE Customer Success Manager, mid-market segment COMPETENCIES TO CALIBRATE Handling an angry customer escalation, identifying an upsell opportunity organically, and cross-functional advocacy for a customer's feature request. A REAL PAST CANDIDATE ANSWER PER COMPETENCY, IF AVAILABLE For the escalation question, one past candidate said they'd 'apologize and offer a discount immediately' with no diagnostic questions asked first. COMMON DISAGREEMENT THIS PANEL HAS HAD Two interviewers disagreed on whether a candidate who admitted they'd never used our specific CRM tool should be scored down on the upsell competency. For each competency, write a short definition of what it means specifically in this role (not a dictionary definition of the competency in general), then four behavioral anchors for scores 1 through 4, each describing what an interviewer would actually observe or hear, phrased so two different people watching the same answer would independently land on the same score. If a real past candidate answer is provided for a competency, use it as a worked example: state what score that answer would earn under this rubric and exactly why, since a worked example calibrates a panel faster than an abstract anchor alone. Where the past-disagreement note describes a specific point of confusion (e.g., the panel couldn't agree whether a candidate who used an AI tool during a live coding round should be scored down), resolve that exact question explicitly in the rubric rather than leaving it as unaddressed edge-case ambiguity that will resurface in the next loop. STRUCTURE THIS AS A CALIBRATION DISCUSSION, NOT A FORM Unlike a scorecard, this document should read as guidance to be discussed and adjusted by the panel before it's finalized — include one open question per competency where the anchors are genuinely debatable, so the panel has something concrete to align on rather than silently rubber-stamping a document nobody actually agreed to. OUTPUT FORMAT One section per competency: definition, four scored anchors, worked example (if given), and one open calibration question. End with a short paragraph on how this rubric relates to but differs from the scorecard interviewers will fill out per candidate.
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
The distinction this prompt enforces — a rubric as a shared pre-loop calibration document versus a scorecard as a per-candidate scoring form — mirrors how structured-interviewing programs at mature talent-acquisition functions actually separate the two artifacts, because conflating them is exactly how panels end up disagreeing at debrief without realizing they were scoring different definitions of the same number the whole time. Using a real past candidate answer as a worked example is the single highest-leverage calibration technique available: an abstract anchor like "handles escalation well" is interpreted differently by every reader, but showing an actual transcript and asking "what would this score, and why" forces the panel to argue out their disagreement on a concrete case before it costs a real candidate a fair evaluation. Resolving a named past disagreement explicitly, rather than leaving the rubric to imply a general principle and hoping it generalizes, matters because the disagreements that actually recur in a real hiring loop tend to be specific edge cases (tool familiarity, live-AI-assistance, a candidate who over-prepared a rehearsed answer) that a generic rubric never anticipates — writing the resolution into the document is the only way it survives contact with the next controversial candidate. Deliberately including an open calibration question per competency, rather than presenting the rubric as finished, exploits the fact that a panel that is handed a document to silently accept will not internalize it the way a panel that argues through one debatable point per competency will — the discussion itself is what produces consistent scoring later, not the document's existence.
What you get back
Competency: Handling an Angry Escalation. Definition: can the candidate de-escalate before jumping to a remedy? Anchor 3: 'asks at least one diagnostic question before offering any concession, stays calm, offers a specific next step.' Worked example: the past answer offering an immediate discount with no diagnostic question would score a 2 under this rubric — jumps straight to remedy, no de-escalation step. Open question: should offering a discount immediately ever score a 3 if the customer is clearly a high-value account?
Verified against
ChatGPT GPT-5.1 · 2026-08-10
Changelog
- 2026-08-10 — Initial publish, verified against ChatGPT GPT-5.1.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
