How to build a candidate scorecard that reduces bias
Three interviewers, same candidate, same hour-long call. One writes "strong yes." Another writes "not sure, something felt off." The third didn't take notes and just says "I liked her" in the debrief. Here is the fix that's worked across every panel I've run.
I've sat in enough debriefs to know how they actually go. Everyone talks, the most confident voice in the room wins, and the decision gets written up afterward to sound more rigorous than it was. Nobody planned it that way. It just happens when there's no structure to push against.
A scorecard doesn't remove bias. I don't think anything fully does. What it does is slow the decision down enough that you have to say what you actually saw, not just how you felt.
Where the fake structure happens
I've reviewed a lot of "scorecards" that were really just a five-star rating with no explanation attached. They look rigorous in the ATS. They mean nothing.
The most common failure is scoring on vibes and calling it a competency. "Culture fit, 4/5" tells me the interviewer liked the person. It doesn't tell me why, and six weeks later when three candidates all scored a 4, nobody can explain the difference between them.
The second failure is scoring after the fact, from memory, an hour or two after the call ended. I did this constantly early on. You sit down to fill in the scorecard and what you remember isn't the interview, it's your impression of the interview, which has already been smoothed over into "good chat" or "wasn't quite it."
The third is copying the same generic rubric across every role. Communication, leadership, technical skill, problem solving, same four boxes whether you're hiring an SDR or a senior engineer. A scorecard that isn't built from that specific job's must-haves isn't testing for the job, it's testing for how likeable someone is in an hour.
A better test than a star rating
The question I actually ask myself for each requirement on the scorecard: could I defend this score to someone who wasn't in the room?
If the answer is "not really, I just felt it," that's not a score, that's a hunch wearing a number. Push past it with one follow-up: what did they say or do that led to this score? Not what impression you formed. What specific thing happened.
That single change, forcing a line of evidence next to every rating, is the whole trick. It's the difference between "communicates well, 4/4" and "walked through a database migration to a non-technical stakeholder using a moving-house comparison, no jargon." The second one is checkable. The first one is just you, agreeing with yourself.
What this looked like on a real scorecard
I had a candidate for an account manager role who came across brilliantly. Warm, articulate, funny. Two interviewers scored her a 4 on "client relationship management" almost on reflex.
When I asked for the evidence line, one interviewer wrote "great energy, would trust her in front of a client." The other actually went back through their notes and found she'd never been asked a single question about handling an angry client, renegotiating scope, or losing an account. The energy was real. The evidence for the actual competency wasn't there.
We rescored that line a 2, pending, and asked it directly in the next round. She answered it well and the 4 held up properly the second time. But if we'd trusted the first read, we'd have made the call on charm alone and gotten lucky or not.
The opposite happens too. I've seen quieter, less polished candidates score low on "communication" in the room, then the evidence line shows they gave the clearest, most structured answer of anyone that week. They just didn't perform it the way interviewers reward on instinct.
Why the evidence line does the actual work
A scorecard without evidence is a vote. A scorecard with evidence is a record you can compare across candidates weeks later, defend to a hiring manager who wasn't in the room, and actually learn from when a hire doesn't work out.
It also exposes weak job briefs fast. If you can't write a clean, testable requirement for the scorecard, that requirement was probably vague on the intake call too. Fixing the scorecard usually means going back and fixing the brief.
The other thing I'll say honestly: filling this in properly while also running the interview is hard. You're either present in the conversation or you're heads-down capturing exact quotes, rarely both well at the same time. That tension is real and it's the actual reason most scorecards end up reconstructed from memory instead of filled in live. It's part of why I started recording calls and pulling the evidence out afterward rather than trying to do both at once mid-conversation.
The short version
Score against the requirements from the job brief, not a generic rubric. Every rating needs one line of evidence you could defend to someone who wasn't there. If you can't write that line, you don't have a score, you have a feeling with a number attached. And if the requirement itself is too vague to test for, that's not a scorecard problem, that's an intake problem you'll want to catch before the search opens, not after the offer falls through.
Sonarnote records interviews locally, without a bot joining the call, and turns them into structured candidate briefs built from what was actually said. The evidence line is already there when you sit down to score. See how it works.
Was this worth your time?
Be the first to mark this helpful


