HR & Management

Verified against ChatGPT · 2026-08-09

Build a candidate scorecard that stops a 4-person panel from scoring the same interview four different ways

Generates a per-candidate scorecard tied to the specific competencies this role was actually interviewed for, so a hiring panel's individual scores are comparable instead of reflecting how generous or strict each interviewer happens to be.

ChatGPT (GPT-5.1)4 fillable variables

The prompt

Ready to copy — highlighted parts are example details you can swap.

You are building a candidate scorecard for a specific role, to be filled out by each interviewer independently right after their interview slot, before the panel debrief.

ROLE BEING HIRED FOR
Staff Backend Engineer, Payments team

COMPETENCIES THIS PANEL IS INTERVIEWING FOR
System design under real constraints, debugging a live incident scenario, mentoring/code review judgment, and communicating a tradeoff to a non-technical stakeholder.

WHO INTERVIEWS FOR WHICH COMPETENCY
Priya covers system design, Marco covers the debugging scenario, Dee covers mentoring and communication.

PAST SCORING INCONSISTENCY YOU'VE SEEN
Whoever interviews last in the day tends to score a full point lower than whoever interviews first, regardless of the candidate.

Build the scorecard as one row per competency, each with a 1-4 scale (not 1-5 — force a lean rather than a comfortable middle score) and a one-line behavioral anchor for what a 1, 2, 3, and 4 look like specifically for that competency in this role, not a generic scale reused across competencies. Every anchor must describe an observable behavior or answer quality from the actual interview, not an inference about the candidate's character — "gave a concrete example with a measurable outcome" is scorable, "seemed smart" is not. For every competency, add a required text field prompting the interviewer to write the specific evidence (a quote, an example the candidate gave) that led to their score, since a number with no evidence is unusable at debrief and is the exact failure mode described in the past-inconsistency note. Include one field at the end for a clear hire/no-hire lean and a one-sentence reason, separate from the competency scores, because a panel debrief needs the interviewer's actual recommendation, not just an average of numbers that can mask a single disqualifying red flag. If the past-inconsistency note describes a specific pattern (e.g., one interviewer always scores high, or scores drift based on interview order), add one line of guidance directly on the scorecard addressing that exact pattern.

WHAT NOT TO DO
Do not create a single generic "culture fit" line — if culture or values matter to this hire, name the specific observable behavior that represents it (per the competencies list) rather than an unscorable catch-all that tends to be used to justify decisions made on other, unstated grounds.

OUTPUT FORMAT
A fillable scorecard: one section per competency (name, 1-4 scale with the four behavioral anchors, evidence field), then a final Hire Lean section. Keep it to something an interviewer can complete in under 5 minutes right after the interview ends.

Customize

Optional — swap in your own details for the highlighted parts above.

Why this works

Structured interviewing research (most notably Google's own re:Work findings and decades of industrial-organizational psychology on interview validity) consistently shows that unstructured, holistic "gut feel" scoring has close to zero predictive validity, while structured scoring against pre-defined behavioral anchors dramatically improves both predictive validity and inter-rater agreement — the entire value of a scorecard comes from forcing comparable judgments, not from the form itself. A 1-4 scale rather than 1-5 is a deliberate mechanism to eliminate the safe, noncommittal middle score that raters gravitate toward when uncertain, which is exactly the score that provides no signal at debrief. Requiring a written evidence field next to every number addresses the specific failure mode of a debrief where four interviewers each say "I gave a 3" with no way to tell whether they mean the same thing by it — the evidence field is what actually gets compared, with the number serving only as a sort key. Separating the hire/no-hire lean from the competency average matters mechanically because averaging can mathematically wash out a single disqualifying signal (a strong system-design score cannot offset a candidate who was dishonest about a past project), so the recommendation needs to be captured as its own explicit judgment rather than derived. Feeding in a specific known scoring-drift pattern lets the model add one targeted counter-instruction on the form itself, which is more likely to actually change interviewer behavior in the moment than a generic "be consistent" reminder given once in a training session weeks earlier.

What you get back

System Design (1-4): 1 = could not identify the core bottleneck even with hints; 2 = identified bottleneck but proposed solution ignored a stated constraint; 3 = proposed a workable design, missed one edge case; 4 = proposed a workable design and proactively flagged the tradeoff we were testing for. Evidence: [interviewer fills in]. Hire Lean: Hire / No Hire / Lean Hire, one sentence why.

Verified against

ChatGPT GPT-5.1 · 2026-08-09

Changelog

  • 2026-08-09 Initial publish, verified against ChatGPT GPT-5.1.

Need this built into your business?

If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.

EXPLORE WHAT SCULT BUILDS
Write a job description that matches the actual hiring bar instead of a wish list nobody meetsTurns a hiring manager's notes into a job description built around a true must-have vs. nice-to-have split, so the posting attracts candidates who can actually clear the bar instead of scaring off good ones with an inflated requirements list.ChatGPT (GPT-5.1)2026-08-08Calibrate what a 1 through 4 actually means before four interviewers disagree about it in the debriefProduces the scoring rubric itself — the shared definition of what each competency and each score level means — that a hiring panel reviews together before interviews start, so scorecards filled out later are measuring the same thing.ChatGPT (GPT-5.1)2026-08-10Generate interview questions with the follow-up probes that catch a rehearsed answerWrites a set of role-specific interview questions, each paired with two or three follow-up probes designed to surface real depth or expose a memorized answer, plus a flag on any question that risks drifting into illegal or discriminatory territory.ChatGPT (GPT-5.1)2026-08-11Build a first-two-weeks onboarding plan that gets a remote hire to their first real contribution, not just a meetings calendarProduces a day-by-day onboarding schedule for a new hire's first two weeks, anchored to one concrete early contribution they'll make, instead of a generic checklist of orientation meetings that leaves the new hire unsure what they're actually supposed to be doing.ChatGPT (GPT-5.1)2026-08-12
All HR & Management prompts

Check your AI visibility

One URL in, a 0–100 score and the exact fixes out.

RUN THE CHECK

Browse all the tools

15 tools across six categories
13 of them never send your data anywhere

Free · No signup · No trial clock

SEE THE DIRECTORY