Verified against ChatGPT · 2026-08-12
Pick KPIs that are hard to game and actually track the thing your team is supposed to improve
Selects and stress-tests a small KPI set for a specific team goal — checking each candidate for how easily it can be gamed or how it might drive the wrong behavior before it goes on a scorecard.
The prompt
Ready to copy — highlighted parts are example details you can swap.
Help me choose the KPIs for a team scorecard. The real risk isn't picking a metric that's hard to calculate — it's picking one that's easy to game or that quietly incentivizes the wrong behavior once people start optimizing for it. TEAM AND GOAL Customer support team, goal is to resolve issues in a way that keeps customers from churning, not just to close tickets. CANDIDATE METRICS UNDER CONSIDERATION Average first response time, tickets closed per agent per day, customer satisfaction score (CSAT), ticket reopen rate. HOW THESE WILL BE USED Reviewed monthly by the support manager, and factored informally into individual agent performance conversations. AVAILABLE DATA Zendesk export with timestamps for first response, resolution, and reopen events, plus a post-resolution CSAT survey with roughly 30% response rate. For each candidate metric: 1. State in one sentence what behavior it would actually reward if someone optimized purely for this number and nothing else. 2. Identify the most plausible way someone could improve this metric without improving the underlying thing it's supposed to represent (a support team hitting a fast-response-time KPI by sending a low-quality canned reply immediately, for example) — be specific to this metric and this team, not generic. 3. Recommend keep, drop, or pair-with-a-counterbalancing-metric, and if pairing, name the specific counterbalancing metric that would catch the gaming behavior you identified. Then: 4. Recommend a final set of 3-5 KPIs total (not one per candidate — some will be dropped or merged), each with its exact definition (including edge-case handling, like how a cancelled or refunded transaction is treated) so two people computing it independently would get the same number. 5. Flag which of the final KPIs are leading indicators (predict future outcomes) versus lagging indicators (report what already happened), since a scorecard made entirely of lagging indicators tells a team what to feel bad about but not what to do differently. OUTPUT FORMAT - Per-candidate table: metric | what it rewards if gamed | specific gaming risk | recommendation - Final KPI set (3-5), each with a precise definition and edge-case handling - Leading vs. lagging label for each final KPI - One paragraph on what's deliberately NOT being measured and why that's an acceptable trade-off
Customize
Optional — swap in your own details for the highlighted parts above.
Why this works
Asking specifically what behavior a metric rewards if optimized in isolation, rather than asking whether it's a "good metric," reframes the task around Goodhart's-law-style failure, which is the actual mechanism by which KPIs go wrong in practice — a model asked generically to evaluate metrics will list textbook pros and cons (easy to measure, industry standard) without ever simulating what a person under pressure to hit the number would actually do, while asking it to state the rewarded behavior explicitly forces that simulation, surfacing gaming vectors like an agent closing tickets fast without actually resolving the issue, since "tickets closed per day" rewards speed of closure, not quality of resolution. Requiring the recommendation to be specifically keep/drop/pair-with-a-counterbalance, rather than an open-ended discussion, matters because pointing out a gaming risk without prescribing a structural fix just produces an anxious list of caveats attached to a scorecard that still ships unchanged — pairing forces a concrete answer (if ticket-closure speed is kept, it must ship next to a reopen-rate or CSAT metric that would catch the gaming) rather than a vague warning that gets ignored under deadline pressure. The leading-versus-lagging classification exists because a scorecard built entirely from lagging indicators (CSAT, reopen rate, all of which report on tickets already closed) tells a manager what already went wrong without pointing at anything actionable this week, and naming this split forces at least a conversation about whether a leading indicator (like first-response time, if it isn't dropped for gaming reasons) belongs in the set. The closing paragraph on what's deliberately not measured matters because every KPI set is an implicit statement of what doesn't count, and making that explicit prevents the team from assuming an unlisted dimension (like ticket complexity) was overlooked rather than deliberately excluded.
What you get back
Average first response time: rewards fast acknowledgment; gaming risk is an agent firing an instant canned non-answer to stop the clock without addressing the issue — recommend pairing with reopen rate to catch that. Tickets closed per agent per day: rewards raw closure volume; gaming risk is closing complex tickets prematurely — recommend dropping in favor of a weighted resolution-quality metric instead. Final set: CSAT (lagging), reopen rate within 7 days (lagging), first response time paired with reopen rate as its counterbalance (leading + lagging pair). Deliberately not measured: ticket complexity or issue category mix, since normalizing for that would require a taxonomy that doesn't exist yet — acceptable trade-off for a first version.
Verified against
ChatGPT GPT-5.1 · 2026-08-12
Changelog
- 2026-08-12 — Initial publish, verified against ChatGPT GPT-5.1.
Need this built into your business?
If a prompt isn't enough — what Scult builds, built and maintained for you — that's Scult's day job.
EXPLORE WHAT SCULT BUILDS
