ai-screeningvoice-aihr-techrecruiting

AI Interview Transcript vs Voice: When the Two Disagree

HireQwik September 23, 2026 10 min read

The AI interview transcript vs voice problem showed up for us on a staging server in April 2026, before any customer saw it. Two test candidates spoke clear English at a C1 level, one at 154 words a minute and one at 142, with almost no fillers. Our transcript evaluator read their answers and gave them 5 and 4 out of 10 on communication.

Anyone who had listened to those calls would have said they communicated well. The transcript did not know that. It saw answers broken into short fragments by speech-to-text and scored the fragments.

That was the week I stopped treating a transcript as the interview. A transcript is a lossy copy of it. It keeps the words and throws away the way they were said, and sometimes the way they were said is the whole story. This post explains what HireQwik does when the two tell different stories about one candidate, and it covers both directions, because the disagreement runs both ways.

Why Does a Transcript Miss What the Voice Shows?

A transcript records what a candidate said, never the manner of saying it. Speed, hesitation, long gaps before an answer, whether a sentence came out whole or was assembled piece by piece: all of it is lost between the audio and the text. A language model reading that text is judging a written summary of a spoken event.

That gap cuts in two directions, and both hurt.

In the first, the transcript undersells someone. The two staging candidates above are the clean example. A fluent speaker whose answers were chopped up by the transcription looks scattered on paper and sounds fine on the call.

In the second, the transcript oversells someone. A candidate reading a prepared answer off a second screen produces a beautiful transcript. Every sentence lands. The words say “good answer.” The voice says it was read out, at an even speed, with no hesitations and not a single self-correction. We now flag that pattern separately, and the detail is in the audio tells of a candidate reading answers. A candidate getting live help from a chatbot leaves yet another trace, a long silence and then a suspiciously smooth answer, described in spotting chatbot help in a voice screen.

A human interviewer handles both cases without thinking about it. Strip the human out and you need something that listens, which is the reason the two-evaluator check exists at all.

What the Voice Layer Measures From the Audio

The delivery score does not come from a model’s impression. It comes from measurements you can check: speaking rate, the hesitation count, the number of one-second-plus gaps inside answers, the share of speaking time that carries voiced sound, pronunciation clarity, and an estimated CEFR band.

Two design choices matter more than the list.

First, the audio is the candidate’s alone. The call records the candidate’s microphone as a separate track from the interviewer’s voice, so the AI’s own speech can never be mistaken for the candidate’s. An early version measured silence across the whole call and counted the interviewer’s turns as the candidate’s pauses. One test candidate was charged with 32 long pauses that were really just the agent talking. We fixed it by analysing only the stretches where the candidate is speaking.

Second, it is audio only. No camera feed is scored, no facial expression, no “confidence” read. Article 5 of the EU AI Act prohibits emotion recognition in the workplace, and even setting the law aside, a face tells you far less about job performance than vendors like to claim.

The speaking-rate method we built on is published research, not a trade secret. It counts syllable peaks in the audio, following the approach De Jong and Wempe described in 2009, which means an auditor can reproduce the number from the recording.

How the 50/50 Communication Blend Settles a Disagreement

Here is the rule, in one sentence: the communication score is half the transcript evaluator’s score and half the delivery score, averaged.

The transcript evaluator is told to judge only structure, depth and clarity of thinking, and to leave delivery alone, because delivery is added downstream. Then the two halves meet.

CandidateTranscript readVoice readWhat happens to communication
Clear speaker, fragmented answersLowHighPulled up, often out of the reject zone
Polished answers, read aloudHighMiddling or flaggedPulled toward the middle, plus a “possibly scripted” flag
Strong on bothHighHighStays high
Weak on bothLowLowStays low

The second staging candidate from the opening is the case that saved the design. After we tightened the evaluator’s instructions, their transcript half scored communication 5.0. The delivery score was 7.79. Blended, communication came to 6.4, and the recommendation moved from Reject to Hold, which put a person back in the loop instead of an automatic rejection.

The part that surprises people is the other direction. Weak delivery can pull the communication score down. It is an average, so it moves both ways. Our own earlier posts have described the voice layer as one that can only help. That is true of the recommendation lift, which only ever moves a recommendation up. It is not true of the communication number itself, and a recruiter should know the difference.

The first staging candidate shows the quieter case, where the blend changes nothing that matters. Under the same revised instructions, their transcript half rose to 6.0 and the delivery half measured 7.91. Blended, communication came to 7.0. The recommendation was Hold before the blend and Hold after it. When both halves broadly agree, the blend just confirms what the transcript already said, and most calls in a normal drive look like this. The disagreements are the minority, which is exactly why they deserve attention when they happen.

One more detail matters here. The transcript evaluator is not allowed to sneak delivery back in. Its instructions say plainly that pace, accent and fillers are handled downstream, so a fluent answer delivered slowly should still earn a high transcript score for its content. Without that rule, delivery would get counted twice, once by the model’s impression and once by the measurements, and slow speakers would pay double.

What the Voice Layer Is Not Allowed to Touch

Delivery affects one dimension of the scorecard. It does not touch the rest.

Technical depth, relevant experience and every role-specific question get judged on the transcript alone. A candidate who explains a reconciliation process correctly, in halting English, gets full content credit for knowing the process. Their delivery only shows up in the communication line.

That limit is also what makes the size of the effect predictable. On our points rubric, communication is worth 2 points out of 10, with a scored minimum of 0.5. So even a very poor delivery reading can move the total by roughly one point. It can nudge a candidate across a band edge. It cannot bury someone whose answers were strong.

There are two more guards, and both exist because we got something wrong first:

  1. Empty interviews cannot be rescued by a nice voice. If the transcript evaluator returns zero on communication, meaning there was not enough in the interview to judge, the blend does not run. Clear speech on three one-word answers is still three one-word answers.
  2. Bad audio switches the blend off. When the recording is too noisy or too silent to measure, the delivery half is thrown out and the transcript score stands alone. What happens next is in scoring a call made on a bad line.

The Week the Voice Layer Was Wrong

It would be dishonest to write about the voice correcting the transcript without admitting the voice has been wrong too.

On April 23, 2026, one of us took a test interview and hesitated on purpose. Slow answers, “I don’t know” loops, repeated words, hedges. The delivery score came back at 7.72 out of 10, and it lifted a genuine No Go to On Hold. The layer built to catch weak delivery had rewarded it.

The cause was five separate bugs stacked on top of each other. The pace penalty never fired because it read the wrong field. A transcription setting quietly collapsed repeated words, so repetition was invisible. The filler list only knew “um” and “uh,” and hedging speakers do not say those, they say “like,” “maybe” and “basically.” Pauses between answers were not measured at all. And the share of time the candidate was actually voicing sound was computed and never used.

After the fix, the same interview scored 3.26 on delivery and stayed a reject. Hedges now count at half weight, repeats and fillers at full weight, and gaps between answers are measured.

There was a second, quieter miss. Until July 17, 2026, the blend and the communication floor did not reach verdicts on roles scored with a job-specific rubric at all. Those roles were being judged on the transcript alone without anyone noticing. A scoring layer that silently does not run looks exactly like one that runs and agrees. It is the best argument I know for printing both halves on every scorecard.

Where Both Halves Show Up for a Recruiter

Nothing about the blend is hidden. Every candidate record in the HR dashboard carries a speech panel that prints the sum itself, with the transcript number, the delivery number and the result, plus the recommendation as it stood before delivery was added. When a communication score looks odd, you can see which half moved it instead of guessing. Reading that panel quickly is its own skill, covered step by step in a recruiter’s guide to the communication score.

I no longer accept “the AI scored them a 6” from any tool, ours included. Six built from what? A number with its inputs attached can be argued with. A number without them can only be obeyed. We made that case at length in explainable AI in hiring.

Why the Voice Half Is Measured, Not Learned

There was a faster way to build this, and we turned it down. Modern speech models can turn a recording into a long list of numbers, and a classifier trained on those numbers would have produced a “communication” score in a week.

We did not do it, for three reasons.

An auditor has to be able to reproduce it. Bias audits of hiring tools, such as the ones New York City requires, ask how each group of candidates is scored. Words per minute, pause length and hesitation rate can be recomputed by anyone with the recording and a stopwatch. A learned score cannot be checked that way.

It has to make sense to the person being scored. “You spoke at 220 words a minute, and a comfortable range is 130 to 180” is feedback a person can act on. A position in a model’s hidden space is not.

Old verdicts should not quietly change. Interviews scored before the voice half existed kept their transcript-only verdicts and are tagged with the evaluator version that produced them. Nobody’s past result was rewritten by a new model after the fact.

The trade is that measurements are cruder than a trained model might be. I would rather defend a crude number I can explain than a precise one I cannot.

So Which One Wins?

Neither, and that is the point. The transcript wins on what a candidate knows. The voice gets half a say on how they communicate. No other score on the card is built from the audio.

If you run volume hiring for customer-facing roles, you already believe delivery matters, or you would not be screening by voice. The question is whether your tool lets delivery change the score silently, or shows you exactly how much it moved. Ask any vendor, us included, to show you the two halves on one real candidate.

Want both numbers from a real call? Book yourself the five-minute demo call. Five minutes, no account, and your own scorecard arrives by email with the delivery read printed on it.

Frequently asked questions

Can an AI interview score a candidate differently from what their transcript suggests?

Yes, but only on communication. In HireQwik the communication score is half the transcript evaluator's read and half a delivery score measured from the candidate's own audio. Technical and experience scores come from the transcript alone, so a strong answer keeps its content credit even when the delivery is weak.

Does the voice analysis look at facial expressions or emotions?

No. The analysis is audio only and measures physical things such as words per minute, hesitation rate, pauses and pronunciation. There is no video scoring, no emotion reading and no personality inference, partly because emotion recognition at work is prohibited under the EU AI Act.

Where can a recruiter see both the transcript score and the voice score?

On the candidate's Interview Details page, the Speech Analysis panel shows the formula with both numbers, the blended result, and the recommendation as it stood before delivery was added. A recruiter can see exactly how much of the communication score came from each half.

See your own candidates screened

Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.