When Every Candidate Scores 100%, Check the Rubric
AI resume scores landing at 100 percent for nearly every candidate on a rubric almost always mean the rubric broke, not that the applicant pool got unusually strong. We saw this on one of our own screening runs this year: dozens of resumes came back with the exact same top score, all rated a perfect match. Nothing was wrong with the candidates. The rubric behind the score had one usable criterion, scored 1 to 3, and a rubric shaped like that can only ever say one of three things: 33 percent, 67 percent, or 100 percent. There is no number in between for it to reach for.
This sits next to three other places a screening run can go quietly wrong once it is already underway: stopping a campaign that is already running, pulling back interview invites already sent, or running into a slot that filled up. A scoring rubric with one row is the same family of problem. Something a recruiter configured, or did not configure, quietly took away the tool’s ability to tell candidates apart, long after the campaign was already live.
The Arithmetic Behind a Rubric That Can Only Say Three Things
Start with the simplest version of the problem: a rubric with exactly one scoring criterion, rated on a 1 to 3 scale. That criterion might be “years of relevant experience” or “communication clarity” or anything else a recruiter picked when the job description was set up. Whatever it measures, the math around it is fixed the moment there is only one row.
A candidate can score 1, 2, or 3 on that single criterion. Divide by the maximum of 3 and you get exactly three possible totals: 33 percent, 67 percent, or 100 percent. There is no path to 45 percent or 80 percent. The rubric is not being generous when a batch of candidates all land at 100 percent. It is doing the only thing it can do with one input.
This matters because a recruiter looking at a stack of 100 percent scores usually assumes the model is either very confident or badly miscalibrated. Neither read is right. The scoring engine is behaving exactly as designed. The design just has one moving part, and one moving part cannot produce a spread.
The Quieter Version: A Rubric With a 33 Percent Floor
A single-criterion rubric is the loud version of this problem. It clusters everyone at the top, so it is easy to notice. A ten-criterion rubric where every row went unanswered, an all-null rubric, has the same defect and it is much harder to see.
If a rubric has ten criteria and none of them ever get a real score, most scoring engines still assign each one a default baseline rather than a true zero. Run that baseline through the same math as above and the result is a structural floor: no candidate’s total can fall below roughly 33 percent, no matter how poor a fit they actually are. Set your auto-reject threshold at 30 percent, and it will never fire. Set it at 33 percent, and it sits exactly on a wall nobody can cross. The reject decision was never actually available.
This is a different failure than scores clustering in the 50s and 60s, which happens when a scorer counts total years of experience and degree level without ever weighing whether that background actually fits the role. That problem compresses scores toward the middle because irrelevant experience still counts for something. This one compresses scores toward the extremes, either a ceiling at 100 percent or a floor around 33 percent, because the rubric itself has too few working parts to produce a real spread. Different cause, different shape, same practical result: no threshold anywhere does its job.
Statisticians call the general pattern restriction of range: when the values feeding a measurement are artificially narrowed, the measurement can no longer separate a strong case from a weak one. The effect shows up the same way regardless of the field it appears in.
This Is a Scoring Defect, Not a Customer’s Bad Setup
It would be easy to describe a one-row rubric as a customer’s mistake: someone rushed the job description setup, skipped filling in criteria, and the tool just did what it was told. We do not think that is the honest way to describe it, and we are not going to write it that way.
The actual defect sits on our side. HireQwik’s scoring hierarchy checks for a job-specific rubric first, and if the rubric that comes back has one usable row or none at all, the engine should have flagged that spread and told a recruiter before any candidate was scored against it. It did not. A rubric that structurally cannot separate a strong candidate from a weak one should never have been allowed to run silently. Closing that gap is what we fixed, not a customer’s configuration.
Guidance on employment testing from the U.S. Equal Employment Opportunity Commission holds that a scored selection procedure should be validated, meaning it can actually distinguish candidates on a job-relevant basis, before it is used to include or exclude anyone. A rubric that can only output three numbers fails that bar before it ever reaches a legal question. That is one more reason to catch it inside the product, not leave it to be caught by a complaint.
The fix connects to a rule we already apply everywhere else: every role gets its own rubric, built from that role’s own job description, instead of one generic template stretched across every posting. We wrote about that principle in JD-aware AI screening: a customer success rubric and a site reliability rubric should never share the same axes, because the two roles do not share the same definition of a good answer. A one-row rubric is that same principle failing at the most basic level. Not the wrong criteria for the role, but not enough criteria to say anything at all.
The Seniority-Band Ceiling Compounds the Same Problem
A second HireQwik feature interacts with a thin rubric in a way worth naming directly: the seniority-band ceiling. A recruiter can configure an acceptable experience band for a role, say 2 to 4 years, and once that band is set, it decides whether a candidate reads as under, matched, over, or far over, rather than leaving that judgment to the model’s own guess.
Over-qualification is capped, not rewarded. A candidate above the configured band caps at 70. A candidate far above it, or far under it, caps at 25, which sits below most auto-reject thresholds. Without a configured band, nothing about this changes. With one configured, it applies across every scoring path we run, including the rubric path, though that took a second fix: the ceiling shipped on three of four scoring paths first, and the rubric path was the one where over-qualification kept sailing through uncapped until that gap closed.
Put a thin rubric and a configured band on the same role and the interaction gets confusing fast. A one-row rubric that would otherwise send everyone to 100 percent now has some candidates capped at 70 or 25 instead, because the band caught them even though the rubric could not tell them apart on its own. The scores start to look less uniform, but not because the rubric improved. The ceiling is doing the separating the rubric never could.
What a Real Fix Looks Like: 166 Candidates, 3 Scores to 101
Here is what fixing both the band and the rubric did on one real role, with the customer and role identity left out on purpose: before the fix, that role’s candidates spread across exactly 3 distinct scores. After the band was configured correctly and the rubric spread was fixed, the same 166 candidates spread across 101 distinct scores. The cluster at 100 percent broke apart into a real distribution, and 67 candidates landed below the auto-reject threshold where previously not one of them had.
That is not a small correction. Before the fix, a recruiter reviewing that role’s shortlist was looking at what appeared to be 166 nearly identical top candidates. Auto-reject never fired for anyone, which meant a human had to eyeball every profile by hand, exactly the work an AI resume screen exists to remove. After the fix, the candidates who genuinely did not fit landed where they belonged: in HireQwik’s Needs review band, or below the reject line entirely, on the HR dashboard where a recruiter actually reviews shortlists.
The number worth sitting with is 101 distinct scores from 3. That is not the model getting smarter about this batch of candidates. It is the same candidates, scored by the same underlying evaluation, finally being allowed to look different from each other.
How to Check Your Own Rubric for This Pattern
You do not need access to anyone’s scoring engine to check whether your own AI resume screen has this problem. Pull the full score distribution from your last completed campaign and look at two things.
First, count the number of distinct scores in the list. If a batch of 200 candidates produced only 3 or 4 distinct values, no matter what those values are, the rubric behind the scoring almost certainly has too few working criteria to separate anyone. A healthy rubric on a role with real variation in the applicant pool should produce dozens of distinct scores, not a handful.
Second, look at your lowest score in the batch. If your auto-reject threshold is 30 percent and your lowest score is 33 percent no matter how many candidates you run, you likely have a structural floor, not a lenient model. Try lowering the threshold in a test run and watch whether anything changes. If it does not, the floor is doing the work, not your settings.
Third, check whether your auto-decide bands are actually doing anything. If a role has auto-reject and auto-fanout thresholds configured but the Needs review band in the inbox is empty, holding every candidate at one extreme or the other, that is the same signal from a different angle. Auto-decide only works if the underlying scores have enough spread for a threshold to actually split the batch; a rubric stuck at three values gives auto-decide nothing to sort, so the bands sit there configured but functionally inert.
If any of these checks come back positive, the fix is the one we made ourselves: build the rubric with more than one working criterion per role, the way a per-JD screener build is supposed to work, and set the auto-reject threshold with the real floor in mind rather than an assumed one. Calibrating a rubric properly from the start, criterion by criterion, catches this before it ever reaches a live campaign.
Why This Took Longer to Notice Than It Should Have
A cluster of 100 percent scores does not look broken at first glance. It looks like good news: a strong shortlist, a job description that attracted exactly the right applicants. That is precisely why it survived as long as it did on the role where we first caught it. Nobody double-checks a shortlist that looks great.
The tell was not the top score. It was the auto-reject rate. A role running a healthy rubric against a real applicant pool almost always rejects a meaningful share of candidates outright, the ones who clearly do not fit. A role where auto-reject never fires, on a rubric that is supposedly doing its job, is the actual warning sign, not the celebration it can look like from a recruiter’s side of the dashboard. We now treat a silent auto-reject band as reason enough to pull up the score distribution and count, the same two checks described above, before assuming the shortlist is simply strong.
A rubric that can only say 33, 67, or 100 percent is not a sign your candidates are unusually strong or unusually weak. It is a sign the rubric has too few moving parts to say anything else. The fix is not a smarter model. It is a rubric built with enough real criteria to spread candidates across the full range, and a scoring engine that flags the problem before a recruiter ever sees a wall of identical scores.
Curious what your own score distribution would look like with the rubric and band fixed? Talk to us and we will walk through it on one of your live roles.
Frequently asked questions
Why do all my AI resume scores come back at 100 percent?
This usually means your rubric has only one real scoring criterion. A criterion rated 1 to 3 can only produce three totals: 33, 67, or 100 percent. If every candidate lands at the top, the rubric cannot tell them apart yet, not that your applicant pool is unusually strong.
Can a scoring rubric make an auto-reject threshold impossible to reach?
Yes. An all-null rubric with many unscored criteria still gets baseline values, which creates a structural floor around 33 percent. Any auto-reject threshold set at or below that floor will never trigger, no matter how poor a fit a candidate actually is.
Does an over-qualified candidate affect AI resume scoring?
Yes, once a recruiter sets an acceptable experience band for the role. An over-qualified candidate is capped at 70, and one far above or far below the band caps at 25, below most auto-reject thresholds. Without a configured band, nothing changes.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in