Reviewing an AI Interview Communication Score, Step by Step
Before you can review an AI interview communication score, you need to know what built it. Picture a scorecard with one line on it that you cannot stop looking at: communication 6.1. Is that good?
On its own, the number cannot tell you. A 6.1 could mean a clear speaker with muddled answers, or sharp answers from someone who froze up on delivery. Those are two different people, and they deserve two different decisions. A good review is mostly telling those two apart in under two minutes.
The Speech Analysis panel, one click into any screened candidate in the HR dashboard, exists so that you can. This is how I would read it, top to bottom.
Start With the Formula Line, Not the Final Number
The first line of the panel is the whole story in miniature. It reads something like: 0.50 × LLM 7.0 + 0.50 × Speech 5.2 = 6.1.
The LLM number is the transcript half. A language model read what the candidate said and judged structure, specificity and clarity of thinking. It was told explicitly to ignore delivery.
The Speech number is the voice half. It is calculated, not judged, from measurements of the candidate’s own audio.
The final number is the plain average of the two. So before you look at anything else, look at the gap between the halves. A small gap means both evaluators agree, and the final number is probably a fair summary. A big gap means they disagree, and the final number is hiding the interesting part. Why the two can disagree at all is the subject of why text and voice can tell different stories.
Just under the formula, a “Pre-blend” line shows the average and recommendation as they stood before the voice half was added. If that recommendation differs from the final one, delivery changed the outcome, and it is worth checking in which direction.
What Each Speech Chip Tells You
Under the formula sits a row of chips. Each one is a measurement, and each one points at something you can go and hear.
| Chip | What it measures | Where the delivery score starts to care |
|---|---|---|
| CEFR | Estimated spoken fluency level, A1 to C2 | Forms the base of the score with pronunciation |
| Pronunciation | Clarity of sounds, out of 100 | Forms the base of the score with CEFR |
| Pace | Words per minute, counted from syllable peaks in the audio | Outside 130 to 180 |
| Hesitation | Fillers, repeats and hedges per 100 words | From 3 per 100 words |
| Voicing | Share of speaking time with actual voiced sound | Below 75% |
| Pauses of 1s or more | Gaps between the candidate’s stretches of speech | Read together with fragmentation |
| Fragmentation | How chopped up the speech is | Moderate above 0.35, high above 0.5 |
The hesitation chip is worth reading in detail. It breaks down into fillers, repeats and hedges. Fillers are “um” and “uh.” Repeats are the same word twice in quick succession. Hedges are words like “like,” “maybe” and “basically,” and they count at half weight. We added hedges after learning the hard way that hesitant speakers often never say “um” at all. Fillers are the most frequent kind of disfluency in everyday speech, as a 2023 survey of filler research puts it, which is why a few of them cost nothing on this score.
If you want the background on what a CEFR level actually means for a hire, what CEFR fluency scoring measures goes into it. For why a few fillers are not a character flaw, see why talking fast shouldn’t sink a score.
Reading One Speech Panel, Chip by Chip
The chips are not decoration. The voice half is built from them, so you can check it yourself. Even the pace figure comes from a published, reproducible method: counting syllable peaks in the audio, as De Jong and Wempe described, rather than trusting a word count from the transcript. Here is an illustration with made-up readings for a very hesitant speaker, using a round starting number to keep the arithmetic easy.
- Fluency and pronunciation base: say 7.5. CEFR level and pronunciation are averaged into a starting point out of 10.
- Pace: 75 wpm. Far below the comfortable range, so the full two-point pace penalty applies.
- Pauses: 62 gaps of a second or more, averaging 2.4 seconds. Fragmentation read 0.71, well into high, which costs one point.
- Voicing: 65%. A third of the speaking time had no voiced sound in it, which costs half a point.
- Hesitation: 2.4 per 100 words. Just under the 3% line, so no hesitation penalty at all.
Start at 7.5, subtract 2 for pace, 1 for fragmentation and 0.5 for voicing, and the voice half comes to about 4.0. Put that next to a transcript half of 7.0 and communication lands at 5.5, a full point and a half below what the words alone earned.
The lesson for a reviewer is in that last chip. A hesitant call does not always show up as a high hesitation number. Some people go quiet rather than filling the silence with “um.” Read pauses, fragmentation and voicing together, and the picture is usually clear.
The Four Gaps and What Each Usually Means
Once you have the two halves, most candidates fall into one of four patterns. Here is how I read each one.
1. Transcript high, voice low. Good answers, shaky delivery. Look at which chip dragged the voice half down. If it is hesitation and pauses, you may have a nervous but capable candidate, and the content scores elsewhere on the card already protect them. If the role is customer-facing and the pace or fluency chip is the problem, the low score may be telling you something true about the job.
2. Transcript low, voice high. Clear, fluent speech, thin or scattered answers. Two common causes. The candidate may simply not have much to say. Or the transcription broke their answers into fragments that read badly. Read two answers in the transcript while the recording plays and the difference is obvious within a minute.
3. Transcript high, voice suspiciously smooth. Hesitation near zero, fast even pace, excellent answers. Check whether a “possibly scripted” chip is showing. If it is, catching read-aloud answers explains the ninety-second check.
4. Both low, both agreeing. The evaluators agree and the recording will almost always confirm it. This is the case where your time is least useful, so do not spend much of it here.
The review effort should go where the halves disagree, not spread evenly across every candidate.
What to Listen For in 90 Seconds of the Recording
When a gap sends you to the recording, you can skip most of the call. Pick one answer to an open question, ideally the longest one, and listen with a specific job in mind.
- If the voice half looks low: can you follow them easily? Would a customer or a teammate have trouble? Hesitation that sounds like careful thinking is different from hesitation that loses the thread.
- If the transcript half looks low: does the answer make sense when you hear it, even if the text looks broken? If so, the transcript was the weak link, not the candidate.
- If a long silence appears before answers: check the AI-assist panel on the same page. A slow-then-fluent pattern points to something different, covered in the slow, fluent chatbot pattern.
Then write one line in your notes about what you heard. It takes ten seconds, and when a hiring manager asks later why you advanced a candidate the score doubted, you have an answer.
When the Panel Says the Blend Was Skipped
Now and then an “Audio flags” row appears, with a note beneath it saying the blend did not run and communication rests on the text alone.
That means the audio was not reliable enough to measure. Most often the line was too noisy or the recording was mostly silence. In that case the communication score you see is the transcript’s alone, the voice half never entered it, and you should read it that way.
Do not read a skipped blend as a weak candidate. Read it as a missing measurement. Causes and fixes are in our post on noisy and broken recordings.
Where This Review Fits in Your Queue
You will not do this review on every candidate, and you should not. On roles where you have switched on auto-decide bands, candidates below your reject line and above your forward line never reach you. Only the uncertain middle lands in the Needs review queue, and that is where a communication score is most likely to be the deciding factor.
For screened candidates, the verdict filter in the review queue, which sorts screened candidates into the four verdict tiers, is the fastest way in. Start with On Hold. Those are the candidates the screen itself was unsure about, and a gap between the transcript half and the voice half is often the reason. Strong Go deserves a lighter pass, mainly to catch the smooth-delivery case in the third pattern above.
No Go is the last place to spend time, with one exception. A No Go that carries an audio flag, or a big gap between the halves, is worth ninety seconds before it closes.
Passing the Review On to a Hiring Manager
Often the person who has to agree with you never logs into a screening tool. The share button on a candidate’s page creates a read-only link with the verdict, each dimension’s score, the evaluator’s notes and the speech line, along with the transcript and a short-lived recording link. The link can be revoked at any time.
When you forward one, add a sentence about what you checked. “Communication is 6.1, but the transcript half is 8 and the delivery half is low only because of pauses. I listened, and the answers are clear.” That one sentence stops a hiring manager from reading the headline number and rejecting on it.
The Review Habits That Save the Most Time
Three habits make the difference between trusting your queue and re-listening to everything in it.
Sort by disagreement, not by score. A candidate at 6.0 with both halves at 6.0 needs no review. A candidate at 6.0 made of 8.0 and 4.0 needs ninety seconds. The same number, two different workloads.
Read the chips before the notes. The evaluator’s notes are written in words, and words can make a mediocre call sound fine. The chips are numbers you can check against what you hear.
Keep a record of every override. When you advance someone the score doubted, mark what happened to them later. Over a few months, that tells you whether your instinct or the score was right more often, which is the only honest way to calibrate either. The outcome record for exactly this is described in comparing screening verdicts with hiring outcomes.
What the Communication Score Is Not
It is worth closing on the limits, because a score looks more certain than it is.
The communication score is one line on the scorecard. Knowledge, experience and each job-specific question carry their own scores, built from what was said and nothing else. A low communication score does not mean the candidate did not know the answers.
It is not meant to be an accent score either. Pronunciation is one input among several in the voice half, and the voice half is only half of one dimension, which limits how far any accent effect can travel. We wrote about that risk directly in accent bias in AI voice screening.
And it is not a verdict on its own. It is evidence, laid out so you can check it. To see those chips filled in with your own voice, book yourself onto the demo interview and read the speech section of the scorecard it sends.
Frequently asked questions
What does the communication score in an AI voice interview actually measure?
In HireQwik it combines two things equally. Half is a language model's read of the transcript, judging structure, clarity and depth. Half is a delivery score measured from the candidate's audio, covering pace, hesitation, pauses, pronunciation and fluency level. The Speech Analysis panel shows both numbers and the formula that joins them.
What is a normal speaking pace in an AI interview?
The delivery score treats 130 to 180 words a minute as a comfortable range and applies no penalty inside it. Slightly outside, from 110 to 129 or 181 to 200, costs one point on the delivery half. Beyond that costs two, since very slow or very rushed speech is harder for a listener to follow.
Should a recruiter overrule a low AI communication score?
Sometimes. If the voice half is low because of hesitation but the transcript half is strong, listen to one answer before deciding. Nervous but capable candidates often look like this. If both halves are low and the recording confirms it, the score is probably right and your time is better spent elsewhere.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in