AI screeningRecruitingHR techVoice AI

When a Good Candidate Gets a Low AI Interview Score

HireQwik October 2, 2026 10 min read

In April 2026 we ran two test candidates through a HireQwik screen on our staging system. Both spoke at an advanced, near-native level, at around 150 words a minute with almost no filler words. When we scored their communication from the transcript alone, they got a 5 and a 4 out of 10.

Nothing was broken in the usual sense. The evaluator did what it was told: it judged communication from text. And from text alone, short, plain answers from a fluent speaker can look ordinary. That result is exactly the gap our optional voice layer for communication was built to close, and it taught me something broader. When an AI interview score is too low for a candidate you rate, the cause is usually findable, and it is usually one of a handful of things. This post walks through six checks, in the order I would run them, and what to do once you find the answer.

Why a good candidate can get a low AI interview score

A good candidate can receive a low AI interview score when the call was cut short, when their answers were confident but short on specifics, when communication was judged from the transcript alone, when the audio was poor, or when the job’s dimensions or rubric text do not fit the role. Each cause leaves a visible trace on the scorecard or in the transcript.

The checks below follow those traces. Most take under a minute each, and together they usually tell you whether the score reflects the conversation or a problem with the setup.

Check 1: Was the call long enough to score properly?

Start with the simplest question. Look at the call length and the number of candidate responses in the transcript.

HireQwik does not score calls that are too short to judge. Anything under two minutes, or with only one or two lines of transcript, is recorded as incomplete rather than completed. Below six candidate responses, the evaluator is warned the call was cut short: dimensions it could not assess get a 0, the notes mention the early ending, and an “ended early” red flag is added.

So if you see several zeros and a note saying the interview ended early, the low total is a symptom of a short call, not a weak candidate. A 0 on HireQwik’s scale means “not enough data,” not “terrible.” The right response is usually a fresh slot, as covered in what happens to calls that stopped early, not a reject.

Check 2: Did the answers carry specifics, or only confidence?

This is the most common cause, and the hardest one for a human listener to notice, because confident speakers sound good.

The evaluator is instructed to keep unsupported claims mid-range and to reward only concrete evidence with higher numbers. A candidate who speaks warmly and fluently but never gives a concrete situation, action or result will score in the middle, however impressive they sound.

Read the evaluation note first. It usually names what was missing, for example “gave no example of handling a difficult customer.” Then read one answer in the transcript and ask: if I removed the tone of voice, what evidence is left? If the honest answer is “not much,” the score is probably fair, and the candidate may simply need a more probing second round. The band-by-band guide shows what a 5, a 7 and a 9 look like in real answers.

Check 3: Was communication judged from the transcript alone?

This is the check our April test pointed to. By default, HireQwik scores communication from the transcript, the same way it scores every other dimension. A transcript captures words, not delivery. It cannot hear that a candidate spoke smoothly, at a comfortable pace, with clear pronunciation.

Transcripts also lose things. Speech-to-text splits a real conversation into many short turns, often a few words each, so an answer that flowed naturally can look choppy on the page. Short spoken answers that would sound complete in person can read as thin in text.

Speech assessment is an optional feature, off unless an account has it enabled. Where it is on, audio-based delivery makes up half the communication score. Every evaluation carries a tag saying transcript-only or speech-augmented, so you can tell which kind you are looking at. If it is transcript-only and the communication score surprises you, listen to one minute of the recording. That minute will tell you more than the number. What to listen for in a communication score gives a short checklist.

Check 4: Was the audio poor?

Bad audio affects scores in two ways. Speech-to-text makes more mistakes on noisy or clipped recordings, so the transcript itself can contain wrong or missing words. And on accounts with speech assessment, a poor recording carries flags such as low signal-to-noise or mostly silence.

When those flags fire, the audio measurement is ignored, and the evaluator is instructed to go easy on pace, pauses and pronunciation, leaning towards a review tier rather than a rejection where communication is the only concern. But if the transcript itself is damaged, every dimension can suffer. If the recording sounds like a bad phone line in a busy room, treat the score as unreliable and consider a retry. What poor audio does to the score goes through the flags in detail.

Check 5: Do the job’s dimensions fit the role?

Open the job’s evaluation settings and look at what is being scored. A job set up through the screener builder has role-specific dimensions and rubric wording. Otherwise HireQwik uses dimensions derived from the job description, and failing that, six generic dimensions that run from communication through technical competency to an overall impression.

The default six are a reasonable starting point, but they do not suit every role. A strong candidate for a warehouse supervisor role may score modestly on “technical competency” simply because the questions never asked about anything technical, which drags the average down. If a job is running on the default six, that alone can explain a run of middling scores. The fix is to give the job proper settings, which is what a JD-aware rubric is for.

Check 6: Is the rubric written for this candidate pool?

On jobs with written rubric levels, read the descriptions for the question where the candidate lost the most marks. Two problems come up again and again.

The first is a level 3 that only an experienced professional could reach. If a campus rubric asks for “examples of managing a team through a quarterly target,” almost no final-year student will hit the top level, and every candidate loses the same points. The second is a domain question full of vocabulary your candidates have not met yet, which makes a knowledge gap look like a thinking gap. Both are rubric problems, not candidate problems. Calibrating a rubric for freshers has practical rewrites, and how rubric levels become points shows what each level is worth.

What to do once you know the cause

Finding the cause tells you what to do next:

What you foundWhat to do
A short or broken callGive the candidate a new slot for a complete interview
Confident but generic answersAccept the score, or advance to a probing human round if the resume is strong
Transcript-only communication that sounds fine on the recordingRecord your own communication score and decision with a note
Poor audioTreat the score as unreliable; offer a retry
Default dimensions on a specialised roleSet up proper job settings before the next batch
A rubric level out of reach for the poolRewrite the levels, then watch the next batch

When you disagree with a score, open the candidate’s interview page and use the review section to set your own decision, enter dimension scores of your own next to the AI’s and write a one-line reason. Nothing the AI produced is overwritten, so the disagreement itself becomes useful data. How overrides are kept on record explains why.

Finished interviews also have a Regenerate Evaluation option, which keeps an archived copy of the old scores before producing new ones. Use it to fix a processing fault, never to hunt for a nicer number.

What not to do

Two tempting reactions make things worse.

Do not lower the bands because of one candidate. A band change moves every future candidate on that job. If one score looks wrong, find out why first. If the cause is a rubric problem, fix the rubric; the bands may be fine.

Do not bulk-move a pile of low scores without reading them. If you find yourself overriding many low scores on one job, that is a strong signal of a setup problem in checks 5 or 6, and fixing it once is far better than overriding forever. Reading through the lowest scores every week, as in a regular No Go review, surfaces these patterns early.

A worked example: one surprising score, start to finish

Here is how the six checks might play out on an invented case. Priya, a final-year commerce student, applies for an inside sales trainee job. Her resume is strong: a sales internship, a college fest sponsorship role, good grades. Her interview total comes back at 4.2, a No Go on default bands. The recruiter is surprised and runs the checks.

Call length. The call ran 16 minutes with 14 candidate responses. Not a short call, so check 1 is clear.

Specifics. The evaluation note says her answers on objection handling “described a general approach without a specific example.” Reading the transcript, the recruiter agrees on one answer but notices that Priya did give a specific sponsorship story on the persuasion question, and it scored a 7. So the content is mixed, not uniformly thin. Partly explained, but not enough to drag a total that low.

Transcript-only communication. The evaluation is tagged transcript-only, and communication scored 5. The recruiter listens to a minute of audio: Priya sounds clear, calm and well paced. That gap is real, but on its own it moves the total by a fraction of a point.

Audio. No quality flags, and the transcript reads cleanly. Clear.

Dimensions. Here is the answer. The job was created quickly from a short job description, and it is running on the generic six dimensions. “Technical competency” scored 2, because no question in a sales screen asked anything technical, and “experience relevance” scored low because the call never asked about her internship. Two dimensions that barely applied to the role were pulling the average down.

Rubric. No written rubric on this job, so check 6 does not apply.

The fix has two parts. For Priya, the recruiter records an On Hold decision with the note “generic dimensions on a sales role; strong persuasion example and clear delivery,” and books her for a short human call. For the job, the recruiter sets up proper evaluation settings from the screener builder before the next batch, so that every later candidate is scored on what the role actually needs. One surprising score led to a setup fix that affects hundreds of future candidates. That is the real value of running the checks.

How often this happens, and why it is worth checking

Low scores for strong candidates are not rare in any automated screening, and some causes are unrelated to the candidate. A large field study of AI-run interviews with over 50,000 job seekers reported that around 7% of AI interviews ran into technical difficulties. Even a small share like that means a recruiter running a big campus drive will meet a few broken calls every week. The six checks exist so those cases get caught by a person instead of quietly filtered out.

Run the six checks on one surprising score today

Sign in to HireQwik and find the one candidate whose score surprised you most this week. Run the six checks in order and write down which one explained it. If none did, listen to the whole call and record your own decision with a reason. Either way you will learn something about your job’s setup. If a pattern keeps showing up across a role, share it with us and we will work through it together.

Frequently asked questions

Can a fluent English speaker get a low communication score in an AI interview?

Yes. When communication is judged from the transcript alone, delivery is invisible, and answers that are short or thin on detail can score modestly even from a clear, confident speaker. Reading the evaluation note and listening to a minute of audio usually shows whether that happened.

What should HR do after deciding an AI interview score was too low?

Record your own decision and a one-line reason in the HR Review panel, and enter your own dimension scores if useful. The AI's original scores stay on file beside yours. If the cause was a job setting, fix the setting for future candidates rather than editing past results.

Does a low AI interview score mean the candidate is automatically rejected?

Not necessarily. The score sets a verdict tier, and only the lowest band is No Go. A borderline total usually lands in On Hold for a person to review, and a recruiter can record a different decision on any candidate, whatever the AI's tier says.

See your own candidates screened

Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.