Welcome back

Sign in to your screening dashboard

New to HireQwik? Book a demo

Book a demo

Tell us a little about your hiring. We'll reply within one business day.

Prefer email? interview@hireqwik.in
voice-aiai-screeningcampus-hiringhr-tech

Why Talking Fast Shouldn't Sink an AI Interview Score

HireQwik August 12, 2026 10 min read

A first-time campus candidate stumbles twice, restarts once, says “um” four times in ninety seconds, and still gives a technically correct, well-reasoned answer to a knockout question about debugging a production issue. That candidate is exactly who a badly built AI interview pace filler words scoring system is most likely to get wrong, not because the content was weak, but because nerves sound, on paper, a lot like incompetence to a system that doesn’t know the difference. This is genuinely one of the most common worries HR teams raise about voice AI, and it’s a legitimate one: if pace and hesitation feed directly into a reject decision at all, a screening tool ends up quietly filtering for confidence instead of competence, which is close to the opposite of what a first-round screen should be doing.

What pace and filler signals actually measure

HireQwik’s audio evaluator, running alongside the content-scoring evaluator on every interview, tracks four things straight from the raw audio rather than the transcript: pace in words per minute, how often the candidate hesitates or fills a gap with “um” and “like,” how clearly the words come across, and a CEFR-mapped fluency read. None of those four readings works as a flat penalty counter. A slower pace on a hard technical question usually means a candidate is actually thinking, not struggling, and the evaluator is built to weigh pace against the question being asked rather than judge it against one fixed threshold that would punish every deliberate speaker equally. Reading pace in context instead of as a raw number separates a system modeling real communication under pressure from one that simply rewards whoever talks fastest. A transcript-only scoring system has no access to any of this, since none of it survives the trip into text.

The research says fillers aren’t simply a defect

It’s worth pushing back on the instinct that “um” and hesitation are pure negative signal, because the research doesn’t actually support that as cleanly as most vendor marketing implies. Linguistic research on hesitation disfluencies describes “um” and “uh” as evidence of a language system doing exactly what it’s supposed to do: planning what to say next, repairing a false start, keeping a conversation moving rather than going silent while the speaker thinks. A separate study on auditory word recognition found that a filled pause like “um” can actually help a listener process what comes next, by signaling that more, and often more complex, speech is coming. None of this means fillers are irrelevant to how an answer lands. Heavy, constant filler use genuinely can make a response harder to follow. But it does mean a small number of “ums” from a nervous candidate is not automatically evidence of a weak answer, and a scoring system that treats it that way is measuring something closer to rehearsed confidence than actual competence.

This is an old bias, not a new one AI invented

None of this started with voice AI. Human interviewers have always tended to read fluent, confident delivery as a proxy for competence, and a meta-analysis of accent and fluency effects across employment interview studies found candidates rated as more fluent and standard-sounding consistently scored higher on hiring recommendations, independent of the actual content of what they said. A human recruiter running back-to-back phone screens on a tight schedule is prone to exactly the same shortcut a badly built AI system is: smooth talking gets a halo, hesitation gets a shadow, and neither one is actually measuring whether the candidate can do the job. Building a screening tool that’s careful about this isn’t correcting for a defect unique to AI. It’s an opportunity to do better than the baseline a live phone screen would have set anyway, and it’s one of the clearer cases where getting the design right actually matters more than getting the marketing right.

Where a naive system gets this wrong

Picture a symmetric scoring design, one where speech signal can push a score in either direction depending on what it finds. A candidate who hesitates twice and says “um” a handful of times while giving a strong, structured answer gets a delivery penalty stacked directly on top of an otherwise-solid content score, even though the hesitation had nothing to do with whether the answer was right. Multiply that across a large campus screening drive and a naive symmetric system doesn’t just misjudge one nervous fresher. It systematically down-ranks first-time interviewees, candidates from smaller colleges without formal interview prep, and anyone who takes hiring seriously enough to actually think before answering instead of reciting a rehearsed line. That’s a specific, measurable failure mode, not an abstract fairness concern, and it’s exactly the opposite of what a communication-first screen is supposed to protect against.

The asymmetric rule, applied to pace and fillers specifically

HireQwik’s asymmetric blend rule applies to pace and hesitation exactly the way it applies to accent: hesitation and pace can lift a borderline file into a Hold a human reviews, but neither has any lever to pull a strong content score back down. So the nervous fresher above, strong content, some hesitation, a handful of “ums”, keeps the outcome their content earned. The candidate whose content score genuinely was borderline, and who also hesitated and paused, is a different case. There, the hesitation is one data point a recruiter sees in the /inbox queue alongside the transcript, not a silent reason the file never reaches a human at all. Delivery signal narrows which candidates get a second look. It never quietly removes a candidate who answered well but sounded unsure while doing it.

A campus-hiring worked example

Run this across a real drive. Two candidates answer the same phase-0 knockout question early in the call. Candidate A answers instantly, evenly, no hesitation, technically adequate but generic, the kind of answer that could have come from anyone who memorized a standard response. Candidate B pauses, restarts once, says “um” twice, and then gives a specific, well-reasoned answer grounded in an actual project they worked on. A pace-and-fluency-only read might rate Candidate A higher purely on smoothness. HireQwik’s per-JD content rubric scores the substance of both answers independently first, and Candidate B’s specific, grounded example scores higher on content regardless of the hesitation around it. The delivery evaluator has no path to drag that content score down for taking a moment to get there. Across a large drive, that’s the difference between a shortlist selected for who sounds most rehearsed and one selected for who actually has something to say.

Interview coaching access is not evenly distributed

There’s a specific equity dimension to this worth naming directly. Candidates from metro campuses with strong placement cells often go through multiple mock interviews, structured coaching, and rehearsed answer banks before a real screen ever happens. Candidates from smaller colleges and first-generation job seekers frequently walk into their first-ever professional interview with none of that preparation, which shows up as exactly the pattern a pace-and-filler penalty would punish: more hesitation, more restarts, a rougher delivery around a genuinely capable answer. A scoring system that treats smoothness as a proxy for competence is, in practice, a system that quietly rewards access to interview coaching over the underlying ability the interview is supposed to be testing. That’s the opposite of what a screen serving 800 to 3,000-applicant campus drives pulling from a wide mix of institutes should be optimizing for, and it’s a large part of why the asymmetric rule isn’t a minor implementation detail. It’s load-bearing for whether the tool is actually fair at the volume Indian campus hiring runs at.

What to check when a candidate lands on Hold or No Go

If an HR team wants to verify a rejection wasn’t really about pace or fillers, the /inbox review breaks the verdict into content and delivery halves for every candidate, shown separately rather than folded into one number nobody can unpack. A recruiter spot-checking a No Go can see whether the verdict came from a content gap (a wrong answer, a missed knockout question, a JD mismatch) or was influenced by delivery, and because delivery can only ever have pushed a score up rather than down, a No Go verdict is never explainable by “the candidate sounded nervous.” That auditability matters just as much as the underlying rule itself, because a fairness rule nobody can independently verify from outside is really just a claim, and auditing a sample of rejected files by hand is exactly the habit that catches a rule quietly breaking in production before it costs real candidates a shot.

When pace and delivery honestly should matter more

It would be dishonest to claim delivery never matters. For a customer-facing or support role where staying composed and clear on a live call basically is the job, a per-JD rubric can weight the delivery read more heavily than it would for a backend engineering screen, because in that specific context, pace and clarity under mild pressure genuinely are part of what’s being tested, not just an artifact of nerves. Even there, the asymmetric rule still holds. Delivery can inform how a borderline answer gets prioritized for review, but a candidate who answers well doesn’t get demoted for a rough patch of hesitation on one question in a fifteen-minute call.

What this costs to get wrong at scale

Run the naive, symmetric version of this scoring across a large enough drive and the damage compounds in a specific, avoidable way. If even a modest share of genuinely strong candidates hesitate under first-interview nerves, and it’s a large share, especially among freshers with no prior professional interview experience, a symmetric delivery penalty doesn’t just cost a handful of edge cases. It systematically thins the top of the shortlist before a recruiter ever sees it. That’s a quieter failure than an obviously broken rubric, because nothing about the output looks wrong; the dashboard just shows a clean, plausible-looking shortlist that happens to skew toward whoever interviews smoothest rather than whoever answers best. The fix isn’t more sensitive fillers detection. It’s removing the path by which hesitation could ever cost a candidate an interview slot they earned on content, a guarantee the asymmetric rule is built to enforce structurally, not just hope holds true.

What to ask any vendor about this

Ask directly: if a strong-content candidate hesitates and says “um” a few times, does your system’s score go down, stay flat, or only ever move up? Most vendors haven’t thought about the question in those terms, because most communication scoring in this category is still built as a single blended confidence number with no visible internal rule at all. That’s a fair thing to press on in a demo, and it’s a sharper question than “do you measure communication,” which every vendor claims regardless of how the system is actually built.

A sharper follow-up worth asking in the same call: can they actually show a real production file where a hesitant candidate with strong content still made the shortlist, not just a slide claiming the system is fair in the abstract. A policy statement with no concrete production example to back it up is a meaningful gap between what gets marketed in a pitch deck and what has actually been verified against real interviews.

The takeaway

A nervous pause is simply not a competence signal, and treating it like one punishes exactly the candidates a fair screen is supposed to protect: first-generation applicants, smaller-college candidates, anyone without access to formal interview coaching. HireQwik’s pace and filler scoring exists specifically to catch genuine communication gaps, never to reward whoever simply sounds the most rehearsed, and the asymmetric rule is precisely what keeps those two things from getting confused on a live 1,099-interview pilot run or any drive at that scale. If you want to see exactly how a real transcript with real hesitation actually scores end to end on a live system, book a walkthrough and bring your single toughest edge case along to test it directly against.

Frequently asked questions

Do filler words like 'um' lower a candidate's AI interview score?

Not on HireQwik. Hesitation and filler signals can lift a borderline content score into a Hold that a recruiter reviews, but they have no lever to pull a strong content score down. Linguistic research treats 'um' as evidence of speech planning and repair, not weak content — though heavy, constant filler use can genuinely make an answer harder to follow.

Does penalizing hesitation in AI interviews favor coached candidates?

Yes. Interview coaching is unevenly distributed: metro campuses with strong placement cells run mock interviews and rehearsed answer banks, while smaller-college and first-generation candidates often face their first professional interview cold. Their capable answers arrive with more hesitation and restarts, so a smoothness penalty ends up rewarding access to coaching rather than the ability the interview is testing.

Should speaking pace and delivery count more for customer-facing roles?

For customer-facing and support roles, where staying composed and clear on a live call is basically the job, a per-JD rubric can weight the delivery read more heavily than for a backend engineering screen. Even there, delivery only informs how borderline answers get prioritized for review — a strong answer is never demoted for a rough patch of hesitation.

See your own candidates screened

Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.