We Audit Every 'No Go' Batch By Hand. Here's Why That's Not Optional.
We Audit Every ‘No Go’ Batch By Hand. Here’s Why That’s Not Optional.
If your AI screen has a false-negative rate you can’t name, the fix isn’t a better model. It’s a habit. We’ve written about why that number stays invisible for most vendors; this is the specific practice we run to keep it from staying invisible for us.
The habit is simple to describe and easy to skip under deadline pressure: after every drive, someone pulls a sample from the No Go bucket, listens to the actual conversations, and checks whether the rejection matches the JD’s stated criteria. Not a dashboard review. Not a re-run of the score. A human, listening to the recording, asking one question — would I have advanced this person?
Why sampling, not a full re-review
Re-reviewing every rejected candidate defeats the purpose of screening at volume in the first place; a 3,000-candidate drive that requires a human to re-listen to most of the rejected calls hasn’t actually removed the bottleneck, it’s moved it. Sampling is the honest compromise: enough calls to catch systemic mistakes, few enough to fit inside the two-hour window the whole point of the screen was to create. In practice that’s a fixed percentage of the No Go bucket per drive, weighted toward calls where the score sat closest to the cutoff — a candidate rejected by a wide margin is far less likely to be a mistake than one rejected by a hair.
The borderline cases are where this earns its keep. A candidate who scored just under the line because of a rough opening stretch, background noise, or a question phrased in a way that didn’t land for them specifically — those are the calls where a human listener catches something a score alone won’t show. That’s also exactly the population most likely to include a real false negative, which is why the sample is weighted there instead of pulled at random.
What “catching a mistake” actually looks like
Most spot-checks confirm the rejection was reasonable — that’s the expected outcome, and it’s not nothing, because it’s the evidence that the classification is doing what it claims to do. Occasionally the listener disagrees: a candidate whose English was accented but perfectly clear got marked down for fluency instead of clarity, or a candidate answered a question technically correctly but in a way the transcript alone made look evasive. When that happens, the candidate gets reinstated and the specific session gets flagged, not as a one-off fix, but as a data point on where the classification needs tightening for that JD.
That flagging step is the part that makes this a QA loop rather than customer service. A single reinstated candidate doesn’t change much. A repeat pattern of reinstated candidates from the same JD, all on the same dimension, tells you something about the rubric, not the candidate — and that’s the signal a spot-check is actually built to surface.
How this differs from an audit trail
This isn’t the same exercise as the legal-compliance audit trail question — what your legal team will actually ask for is a different, narrower list: consent records, retention windows, model versioning. Those answer “can we defend this decision if challenged.” The No-Go spot-check answers a different question: “is the decision actually right, independent of whether anyone challenges it.” A screening system can have a pristine audit trail and still be quietly wrong about a slice of candidates nobody complained about, because most rejected candidates never complain — they just move on to the next application. Waiting for a complaint to trigger review means you only ever catch the mistakes someone was angry enough to flag.
The take
An AI screening system you never re-check isn’t accurate. It’s just unaudited, and the two get confused constantly because an unaudited system that hasn’t produced a visible complaint looks identical to an accurate one from the outside. The spot-check is cheap, a modest sample per drive, well under an hour of a TA lead’s time, and it’s the difference between trusting a classification because it hasn’t caused trouble yet and trusting it because someone actually checked.
If you’re running AI screening without a version of this in place, the honest first step isn’t switching vendors. It’s pulling a handful of calls from your last rejected batch this week and listening to them yourself. Talk to us if you want help setting up a sampling process against your own funnel.
See HireQwik in action
Book a 30-minute demo — bring a live JD and we'll screen your own candidates against it.