Would the Same Answer Get the Same Score Twice?
In July 2026 we switched HireQwik interviews to a newer evaluation model, and one pilot’s sales screening immediately came out roughly a point stricter across all its dimensions. Not on some candidates. On all of them. The new evaluator ranked people in almost the same order as the old one; it was just stricter across the board.
Nobody had done anything wrong, and no candidate had changed. That is what made it instructive. If an upgrade can move every score by a point overnight, then “is this score consistent?” is not a box you tick once. It is a question you keep asking. This post covers AI interview score consistency in HireQwik: what holds scores steady, where variation still creeps in, what we did about that July shift, and how you can test consistency on your own jobs.
What AI interview score consistency means
AI interview score consistency is the degree to which the same answer, judged against the same rubric, receives the same score each time, whichever candidate gave it and whenever it was scored. In HireQwik it rests on a fixed rubric per job, a written scale, an evidence-only rule and a low-randomness scoring setting.
Two kinds of consistency matter. The first is between candidates: two people who give equally strong answers should score the same, even if one was screened at 10 AM and the other at 9 PM. The second is over time: the same transcript scored today and next month should land in the same place unless the rubric changed on purpose.
Why consistency matters more at volume
A recruiter screening ten candidates can hold all ten in their head and compare them directly. A recruiter screening 3,000 for a campus drive cannot. At that scale, the score becomes the comparison. If it drifts between the morning batch and the evening batch, two equally good candidates can end up in different tiers for no reason except timing.
Human interviewers drift too, and usually more. Fatigue, mood, the order candidates appear in and the memory of the last strong answer all pull ratings around. That is why structured interviewing exists. The US Office of Personnel Management describes structured interviews as asking every candidate the same questions and evaluating every answer “using the same rating scale and standards for acceptable answers,” so that candidates are “assessed accurately and consistently.” An AI evaluator can apply that discipline without getting tired, but only if the inputs that drive the score are held steady.
Where HireQwik’s consistency comes from
Five things keep a HireQwik score stable:
- One rubric per job, applied to everyone. Every candidate on a job is scored against the same dimensions and the same rubric text. Nobody gets an easier version because they were screened later.
- A written scale. What a 7 means is fixed in advance, not left to the model. Every band on the 0 to 10 scale has its own fixed wording, and 0 is kept for dimensions the call never gave evidence on.
- An evidence-only rule. Scores must rest on what is in the transcript, and only concrete detail lifts an answer past mid-range. Tying scores to evidence removes much of the room for mood-like variation.
- A low-randomness setting. Language models can be run more or less deterministically. HireQwik scores with randomness turned low, which keeps identical input producing near-identical output instead of a spread of plausible answers.
- The full transcript, every time. The evaluator reads the whole conversation after the call, not a live impression formed halfway through. Every candidate is judged on their complete answers.
For the practical reading of each band, see what each score band claims; for the arithmetic, see the scoring pipeline end to end.
Where variation still creeps in
I would rather name the sources of variation than pretend there are none.
Answers between two levels. When an answer sits right on the line between two rubric levels, a re-score might land on either side. On a 3-point question that can mean half a point. Most answers are clearly one level or another; the borderline ones are where small movements happen.
The transcript itself. Speech-to-text is very good but not perfect. A mis-heard word, a dropped phrase on a weak connection or an accent the transcription handles less well can change what the evaluator reads. Two candidates who said the same thing can produce slightly different transcripts. On a noisy line this matters more, which is why the poor-audio guide treats bad recordings as a separate case.
Different conversations. No two calls are word-for-word identical. Candidates ask different questions, take different amounts of time and answer in different orders. The rubric is the same, but the evidence each call produces is not.
Changes to the evaluator or rubric. This is the largest source by far, and the one in our opening story. Any change to the scoring model, the rubric text or the scale can move every score at once.
What happened when we changed the evaluator
Back to July. When we moved to a newer evaluation model, we compared its scores with the old ones on interviews that had already been scored, and with decisions HR had already made on those candidates. For one pilot’s sales screening, the pattern was clear: the new evaluator was harsher by roughly a point across the board, yet its ranking of candidates barely moved.
That combination is good news and bad news. Good, because the ranking, which is what really decides who advances, was stable. Bad, because a uniform one-point drop would have pushed candidates across band lines: Go to On Hold, On Hold to No Go, and a recruiter looking at the dashboard would have had no way to tell why. A candidate’s tier would have changed because of our upgrade, not their answers.
So we lowered that screening’s bands by the same amount, which meant a Go still meant what it had meant the week before, and we verified the new bands against interviews HR had judged earlier. The lesson I took is that consistency over time is not automatic. Every change to the evaluator needs a before-and-after comparison on real, already-judged interviews, and band adjustments where the scale has shifted. How bands map to tiers covers the rest of that logic.
Re-scoring an interview: what the Regenerate button does
On a finished interview, HR can press Regenerate Evaluation to re-score the call. Before re-scoring, HireQwik archives the existing scores, so the previous result is kept and the two can be compared. The button is refused while a call is still in progress, and it needs a transcript to work from.
Regenerate exists for repair, not for shopping. The typical case is an evaluation that failed or never finished. If you re-score a normal interview, you should expect a very similar result. A large swing on a re-score is a signal worth investigating: it usually means the transcript or the job’s settings differ from the first run, not that the evaluator has changed its mind.
Each evaluation also records which evaluator version produced it, transcript-only or speech-augmented, and that label matters when comparing scores over time. Speech assessment is optional and off by default; where an account enables it, half of the communication score comes from audio, so the same call can legitimately get a different communication number under the two versions. Compare like with like.
How to test consistency on your own jobs
You do not need to take my word for any of this. Three simple tests will show you how consistent scoring is on your own roles:
- Compare similar answers. Find two candidates on the same job who gave similar answers to the same question. Their dimension scores should be close. If they are far apart, read both evaluation notes and see what the evaluator saw differently.
- Score alongside the AI. For ten interviews, write down your own dimension scores before reading the AI’s. Then compare. You are measuring how consistently a human and the evaluator apply the same rubric, which is often the most revealing test of all.
- Watch the distribution after any change. When a rubric is edited, compare the spread of scores in the next batch with the previous one. A sudden shift up or down across everyone points to a scale change, not a better or worse candidate pool.
If your test shows disagreement on one question, the usual culprit is fuzzy level wording, and the fix is to describe each level with a concrete example answer, as shown in the points rubric breakdown.
Changing a rubric mid-drive
Sometimes you will find a problem halfway through a campus drive: a level description that is too strict, or a question that confuses everyone. Fixing it is the right call. But be clear about what happens next.
Editing a job’s rubric does not quietly re-score the candidates who already finished. Their scores were produced against the old wording and stay that way unless someone regenerates them. So after a mid-drive change, your candidate list holds two groups scored against two versions of the rubric. A 7 in the first group and a 7 in the second do not necessarily mean the same thing.
Three habits keep this manageable:
- Write down the date and time of the change, and what you changed, in the job notes or your team’s tracker.
- Compare candidates within each group rather than across them, at least for the dimension you edited.
- Look again at borderline candidates from the first group, especially those just below a band line on the edited question. If the change would plausibly have moved them, regenerate those few evaluations, or give them a human review.
AI scoring did not invent this problem. A human panel that changes its marking guide halfway through a drive faces the same problem. The difference is that with an AI evaluator you can see exactly when the change happened and re-score cleanly if you need to.
Questions to ask any vendor about score consistency
If you are comparing AI screening tools, consistency is one of the easiest things to ask about and one of the most revealing. Here are the questions I would ask, including of us:
- If you re-score the same interview, how much does the score move? A vague answer like “it’s very accurate” is not an answer.
- What happens to old scores when you change the model? Look for a before-and-after comparison on real interviews and a plan for band adjustments.
- Are previous scores kept when something is re-scored? If old scores are overwritten, you lose the ability to audit.
- Is every candidate on a job scored against the same rubric? Adaptive or personalised scoring can sound appealing but makes comparison harder.
- Can I see which version of the evaluator scored each candidate? Without that, you cannot tell a real change in your pool from a change in the tool.
A vendor who answers these clearly is treating consistency as an engineering problem, which is what it is.
Consistent is not the same as correct
One caution to finish on. A consistent score is not automatically a fair or accurate one. A rubric that undervalues a skill will undervalue it consistently. A question with jargon that freshers have not met will cost every fresher the same points, every time. Consistency removes noise; it does not remove a bias that is built into the rubric.
That is why the other safeguards matter: rubrics written from the job description, a human decision that never overwrites the AI’s, and periodic checks of verdicts against who you actually hired, the loop described in recording what actually happened. Consistency is the foundation. It is not the whole building.
Try the side-by-side test this week
Pick a running job in your HireQwik account and choose ten completed interviews from it. Score one key dimension yourself before looking at the AI’s number, then compare. Note where you differed by more than two points and read those answers again. It takes about half an hour and tells you more about consistency on your roles than any claim from a vendor, including us. If you would like help reading the results, get in touch.
Frequently asked questions
If HireQwik re-scores the same interview, will the score change?
Usually very little. Scoring uses a fixed rubric, a written scale and low-randomness scoring, which means one transcript tends to produce the same scores. Small movements are possible on answers that sit between two levels. Re-scoring archives the previous scores first, so any change can be compared.
Can an update to the AI evaluator change candidate scores?
Yes. A change to the evaluating model or the rubric can shift scores across the board. After our July 2026 evaluator change, one pilot's sales screening came out roughly a point stricter across all dimensions while candidates stayed in nearly the same order, so we adjusted that screening's bands.
Is a consistent AI interview score the same as a fair one?
No. Consistency means the same answer is treated the same way each time. A rubric with a built-in bias would apply that bias consistently too. Fairness also needs a rubric tied to the job, human review of borderline cases and regular checks against real hiring outcomes.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in