Reading the Number: What Each Band of an AI Interview Score Means
Picture two recruiters on the same team looking at one candidate’s communication score of 6 and drawing opposite conclusions. One reads it as “fine, move on.” The other reads it as “barely average, worrying for a customer-facing role.” Neither is wrong about the candidate. They are working from different ideas of what a 6 means, because nobody ever wrote it down for them.
That kind of gap undermines any score, human or AI. So this post does the writing down. It covers AI interview score meaning in HireQwik band by band: what each number claims about an answer, why zero is not the worst score, why 5 is not a pass mark, what moves an answer from a 4 to a 7, and how to explain each band to a hiring manager in a sentence they can check.
What an AI interview score means in HireQwik
In HireQwik, an AI interview score is a judgement of how much evidence a candidate’s answers gave for one dimension of the job, on a written 0 to 10 scale. Each band has a fixed description the evaluator follows, from “very poor” at 1-2 to “exceptional” at 9-10, and 0 is reserved for “not enough data to assess.”
The important word is evidence. The score is not a rating of how pleasant, confident or polished the candidate sounded. It reflects what the candidate actually showed in their answers about the thing being measured.
The scale, band by band
Here is each band in short form, with my plain-English reading of what it tells a recruiter:
| Band | In short | What it tells you |
|---|---|---|
| 0 | Missing evidence | The dimension was never really tested in this call |
| 1-2 | Very poor | The candidate could not engage with the topic at all |
| 3-4 | Below average | They talked about it, but showed nothing concrete |
| 5 | Average | Adequate, with no detail that sets them apart |
| 6 | Above average | Real competence starting to show |
| 7-8 | Good to strong | Solid evidence you could repeat to a hiring manager |
| 9-10 | Exceptional | Unusual depth, worth a closer look whatever the role |
Read the table top to bottom and a pattern appears: the bands climb with specificity, not with enthusiasm. That is the core of how to read any number on the screen.
Why 0 is not the worst score
The easiest misreading is treating 0 as “terrible.” It is not. On HireQwik’s scale, 0 means the evaluator did not have enough evidence to judge the dimension. A call that dropped before the domain question was asked, or a short screen where a topic never came up, will show 0 for that dimension.
The evaluator is told to leave a gap visible rather than fill it with a guess. So a 0 should send you to the transcript to see what happened, not to the reject pile. Calls that end early are handled on their own track, flagged rather than scored.
A real 1 or 2 is different. It means the candidate was asked and could not engage at all. If you see a 1, the transcript will usually show an answer that wanders off topic or never forms a point.
Why a 5 is not a pass mark
School trains us to read 5 out of 10 as “half marks, scraped through.” The evaluator’s scale does not work that way. A 5 is described as “meets minimum expectations, nothing distinctive.” It is the answer of someone who understood the question and responded sensibly without showing anything you could point to.
That matters when you look at the overall total. On the default weighted bands, Go starts at 6.0 and On Hold at 4.5. A candidate who scored 5 on every dimension would land in On Hold, not Go. In other words, “adequate but undistinguished” goes to a reviewer’s desk, not straight through. I think that is right for a first screen: a person, not a formula, should decide whether “adequate” is enough for this particular role.
Points-based jobs work on a different total, with Go starting at 7.5 out of 10, but the same logic holds. Where all those points come from is explained in how a points rubric splits 10 points.
What moves an answer from a 4 to a 7
The evaluator’s calibration guidance boils down to one idea: claims without examples stay mid-scale, and concrete detail is what climbs it. Self-descriptions like “hard-working” or “good with people” with nothing behind them land in the 4 to 5 range. Real situations, numbers and depth push an answer to 6, 7 and beyond, and only genuinely standout answers reach 8 or more.
Take one invented question for a sales trainee role: “Describe a time you convinced someone to change their mind.”
- Around a 4: “I am very good at convincing people. In college my friends always listened to me when we planned trips.” The candidate claims the skill but gives no situation, no action and no result.
- Around a 6: “During our college fest I persuaded the principal to extend the event by one day. I showed him that ticket sales were higher than last year.” There is a real situation and one piece of evidence, but little about how they approached the conversation.
- Around a 7 or 8: “Our fest committee wanted to cancel the stalls because of low sign-ups. I called fifteen past vendors, found that most had missed our email, and brought eight back. I showed the committee the confirmed list and they kept the stalls.” Situation, personal actions, numbers and an outcome.
Notice that the third answer is not more polished or more confident than the first. It is more specific. That is the whole difference, and it is why I tell recruiters to read the evaluation note, not the number, when a score surprises them. The note will usually name the missing piece.
Why 9 and 10 are rare
It is natural to wonder why so few candidates get a 9. The scale marks 9 to 10 as exceptional and says such evidence is uncommon from entry-level candidates, and the guidance reserves 8 and above for genuinely impressive answers. High scores are kept scarce on purpose.
If every good answer earned a 9, the top of the scale would stop meaning anything and you could no longer tell a strong candidate from an exceptional one. Keeping 9 and 10 rare preserves room at the top for the answer that genuinely stands out, the one you would want a hiring manager to hear first. When you do see a 9, it is worth listening to.
Dimension scores vs the overall total
Everything above describes a single dimension. The overall total is a different kind of number: a weighted average of dimension scores, or on points jobs a sum of budgets. A total of 6.4 does not mean the candidate gave “6.4-quality” answers everywhere. It might mean an 8 on one dimension and a 5 on another.
So when a total surprises you, read the dimensions before the note, and the note before the transcript. That order usually finds the reason in under a minute. The arithmetic connecting dimensions to the total, with a worked example, is in from transcript to total, step by step.
Scores describe the conversation, not the person
A score is evidence about 15 to 20 minutes on a particular day, against a particular rubric. It says nothing final about someone’s ability, intelligence or future. A capable candidate who was nervous, distracted by a noisy room or caught off guard by an unfamiliar question can score lower than they would on another day.
That is not a reason to distrust scores. It is a reason to use them for what they are good at: sorting a large pool consistently so that human attention goes where it is most needed. The final call stays with a person and is saved in its own field, for reasons set out in why the human decision has its own record. If a nervous start seems to be costing good candidates, helping a nervous candidate through the screen has practical steps.
How to explain a score to a hiring manager
Whenever someone questions a number, translate the band into evidence. Keep each translation to one sentence, name the specific answer it rests on, and avoid adjectives the transcript cannot back up. A few templates that work:
| Score | A sentence you can say |
|---|---|
| 3-4 | “They talked about it in general terms but did not give a single concrete example.” |
| 5 | “A sensible answer that met the basics, without anything that stood out.” |
| 6 | “Some real detail, for example [the specific point from the notes], but not much depth.” |
| 7-8 | “A strong, specific answer: [the situation], what they did, and how it turned out.” |
| 9-10 | “One of the best answers we have heard on this; worth listening to the recording.” |
| 0 | “That topic never came up in the call, so there is no score for it.” |
Each sentence points back to the transcript. Anyone who doubts it can check, and a read-only scorecard link lets a manager do so without a login.
There is a broader principle under this. The US standards body NIST names “explainable and interpretable” among the traits of trustworthy AI in its AI RMF framework. In practice, for a screening score, explainable means that every number can be translated into a sentence about what the candidate said.
Four misreadings to avoid
Even with the scale written down, a few readings come up again and again. Each one leads to a worse decision than the number deserves.
“A high communication score means a good hire.” Communication is one dimension. A candidate can speak clearly and still give thin answers on the role’s most important questions. Read the content dimensions before you get excited about a 9 on communication.
“Two candidates with the same total are equal.” A 6.5 built from an 8 and a 5 is a different candidate from a 6.5 built from two 6.5s. The first has a clear strength and a clear gap; the second is steady but unremarkable. Which one you want depends on the role, and only the dimension scores show you the difference.
“A low score on one dimension is a red flag.” Sometimes it is. Often it means the question for that dimension was hard for this pool, or the topic barely came up. Check how other candidates on the same job scored on that dimension. If almost everyone is low, look at the question, not the people.
“The score already accounts for everything.” It accounts for what was said in this call. It does not know about a strong referral, a relevant portfolio or a candidate’s notice period unless those came up in the conversation. The resume score is separate, and the recruiter’s knowledge is separate again.
When the number and your instinct disagree
Every recruiter will sometimes read a transcript and think, “this person is better than their score.” That instinct is worth listening to, and worth testing. Write down what you think the score missed, in one sentence. Then find the evidence for it in the transcript. If you can point to it, record your own decision with that sentence as the reason. If you cannot, the score may have been right, and your instinct may have been reacting to something else, such as confidence, warmth or a familiar college name. Both outcomes are useful. Over a few weeks, your notes will show whether the scale, the rubric or your instinct needs adjusting.
Write your team’s score guide today
Copy the band table above into your team’s screening SOP, then sign in to HireQwik and choose three completed interviews: one high, one middle, one low. For each, read one dimension score and its note, and check that your team would describe the band the same way. Agreeing on what a 6 means takes ten minutes and saves a lot of Friday-evening debates. If you want help adapting the guide to a specific role, contact us.
Frequently asked questions
Is a 5 out of 10 a passing AI interview score?
Not on its own. On HireQwik's scale a 5 means the answer met minimum expectations with nothing distinctive. On the default weighted bands, Go starts at an overall 6.0, so a candidate averaging 5 would usually sit in On Hold for a person to review.
Why do so few candidates score 9 or 10 in an AI screening interview?
The evaluator's scale treats 9 to 10 as exceptional and unusual for entry-level candidates, and its guidance reserves 8 and above for genuinely standout answers. High scores are kept scarce on purpose so they still mean something when they appear.
What turns a 4 into a 7 on an AI interview answer?
Evidence. Generic statements without examples sit around 4 to 5 on HireQwik's scale. Answers with a specific situation, the candidate's own actions, real numbers or clear depth earn 6 to 7 and above. Fluency and confidence alone do not move the score.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in