What Poor Audio Quality Does to an AI Interview Score
Poor audio quality in an AI interview score starts as a measurement problem, and our own first microphone-clipping check proved it: on staging in April 2026, it flagged every single recording as clipped. Every one. A candidate speaking softly in a quiet room was marked the same as one shouting into a phone.
The cause was a units mistake. The check compared the loudness reading against a threshold written for a different scale, so normal speech always crossed it. We caught it before it reached production and fixed the threshold to the right scale, 85 dB. But it taught me something that shapes this whole post. A bad recording is not only the candidate’s problem, and a tool that cannot tell a bad line from a bad answer will punish people for their room.
Here is what HireQwik treats as unusable audio, how it changes the score, and the call a recruiter should make on a flagged interview.
Why Poor Audio Is a Screening Problem in India
For campus and volume hiring in India, clean audio is not a safe assumption. Candidates join from hostel rooms, shared flats, family homes and mobile data in a moving auto. Fans, traffic, a television two rooms away, a sibling on another call.
None of that tells you if the person could handle the role. A screen that scores the room instead of the person quietly removes candidates from exactly the colleges and towns that volume hiring is meant to reach.
Noise also hurts in two separate places. It degrades the delivery measurements, which is the part we control directly. And it degrades speech-to-text, which feeds the transcript. Research on speech recognition over real networks, such as a 2022 study of noise and network-distorted speech, shows that systems trained on clean audio lose a lot of accuracy when noise, jitter and packet loss arrive together. More recent work goes as far as building noise detection into the recogniser itself to cope with it.
So the design had to answer two questions: when to stop trusting the audio measurements, and what to do about a transcript that may also be damaged.
Why Cleaning the Audio First Is Not the Answer
The obvious fix sounds like noise removal. Run the recording through a filter, strip the fan and the traffic, then score the clean version.
It is less obvious than it sounds. A Bengaluru team studying medical speech recognition tested exactly this in late 2025 and found that adding a speech-enhancement step made recognition worse in all 40 configurations they tried, across four modern recognisers and nine noise conditions. The noisy originals were transcribed better every single time.
Delivery measurements have the same weakness. A filter that removes noise also reshapes pauses, loudness and voice quality, which are the very things being measured. A cleaned recording is not a clean one. It is a different recording with different numbers.
That is why detecting unusable audio, and changing how the call is scored when it happens, matters more to us than trying to scrub every call into shape.
The Audio Flags HireQwik Raises
After each call, the candidate’s audio is checked before any delivery measurement is used. These are the flags, in plain terms:
| Flag | What triggers it | What it means for the score |
|---|---|---|
| Low signal-to-noise | Voice barely stands out from background noise, below 5 dB | Voice half switched off |
| Mostly silence | More than 85% of the candidate’s audio is silent | Voice half switched off |
| Mic clipping | Loudness above 85 dB, typically a mic too close or too loud | Recorded as a warning on the panel |
| Noisy voice quality | Voice-clarity readings distorted by noise | Clarity reading dropped, pace and pronunciation kept |
| Under 30 seconds | Too little candidate speech to measure | No delivery analysis at all, transcript only |
Two details make these flags more trustworthy than they sound. First, the candidate’s mic is captured on a dedicated track, apart from the agent, so the interviewer’s speech can never be counted as noise on the candidate’s line. Second, the analysis keeps only the parts of the recording in which the candidate is talking, so a long gap while the interviewer asks a question is not counted as the candidate going silent.
What Happens to the Score When Audio Is Flagged
This is the core of it. When a serious audio flag is raised, the voice half of the communication score is thrown out, not averaged in.
On a normal call, communication averages a text-based read with an audio-based one, the design laid out in transcript versus voice. With a low signal-to-noise or mostly-silence flag, the delivery half is skipped completely. The communication score is the transcript half on its own.
Three more instructions go to the transcript evaluator at the same time:
- Do not penalise delivery. It is told that the speech numbers for this call are unreliable and must not be used to lower communication, and to judge only the clarity of what was said.
- Escalate instead of rejecting on communication. If the transcript alone would push the candidate to No Go on communication grounds, the evaluator is told to recommend On Hold instead, so a person can listen to the recording.
- Say so in the notes. The evaluation notes record that audio quality was flagged and that the communication score is based on the transcript only.
There is one gap here that I would rather you hear from us. Some roles carry a hard communication floor in their scoring setup, a minimum below which the verdict is No Go regardless of anything else. That floor is applied by the scoring code, and it reads the transcript-only score on a flagged call like any other. So on those roles, a flagged call whose transcript half lands under the floor can still come out as No Go. Until we close that, treat any flagged call sitting near a communication floor as one to open by hand.
The goal is that bad audio moves a candidate toward review, not toward rejection. The instructions above do most of that work. The floor gap is the part still on our list.
What Recruiters See on a Flagged Call
Open the candidate in the HR dashboard and the speech panel names the problem in an “Audio flags” row. Beneath that sits a line stating that the blend was skipped and only the language model’s score is in use.
Read that note as a missing measurement, not a weak candidate. The communication number on that card was built without any delivery information at all, so it can be higher or lower than the candidate’s real delivery.
The noisy-voice flag is gentler. It only tells the evaluator that the voice-clarity reading is unreliable. Pace, pauses and pronunciation still count, because those survive moderate noise far better than clarity measures do, so a call with only that flag still gets a full blend.
New to that panel? Our walkthrough of the communication score covers every chip on it.
When to Re-Invite Instead of Deciding
A flagged call is not always a call you can judge. Here is the rule of thumb I would use:
- Noise, but every answer is audible in the recording: judge the call. The transcript half is doing its job, and you can hear the delivery yourself.
- Answers are hard to follow even for you: the transcript may be damaged too, and any “possibly scripted” or AI-assist flag on the call deserves extra doubt, since noise corrupts the pauses and hesitations those flags read. The flags themselves are explained in candidates reading answers aloud and spotting live chatbot help. Do not reject on it. Re-invite the candidate with a short note about finding a quieter spot.
- Mostly silence, or under 30 seconds of speech: check whether the call dropped or the microphone failed. That is a technical failure, not an interview. Re-invite.
- Clipping warning, speech otherwise clear: usually harmless. The candidate was loud or close to the mic, and the words came through.
Re-inviting costs the candidate another slot, and it is the same instinct behind auditing No Go batches by hand: a wrong reject is the expensive mistake. Rejecting on a broken recording costs them the job, and costs you a candidate who may have been fine. When in doubt, re-invite.
Giving the Candidate a Second Slot
Re-inviting is only useful if it is easy, so here is how it works in practice.
A candidate can move their own booking three times at most, and only inside the 72 hours after they first booked. After that the slot picker locks, which stops endless rebooking from blocking seats. For a candidate whose first call was ruined by audio, you will often be past that point, or the call will already show as screened.
That is what the reset option on a scheduled or screened candidate is for. It emails the candidate a booking link to pick a new slot and clears their reschedule count. The broken call is not deleted. Its transcript and scores stay on record, so anyone auditing the decision later can see why a second interview happened. If the candidate is somehow still on a live call, the reset refuses rather than cutting them off.
Planning for Audio on a Campus Drive
On a campus drive the room problem is predictable, which means it can be planned for. Three things help.
Ask the placement cell for a quiet room. A handful of candidates will not have anywhere suitable to take the call. One classroom with a door and a few charging points, booked for the drive window, solves most of it.
Spread the slots out. When a whole batch joins from the same hostel at the same hour, rooms get shared and noise stacks up. The booking page shows the seats left in each slot and closes a slot when it fills, so candidates naturally spread across the day if you open enough of them.
Check the first ten calls. Open the first few completed calls in the drive and look for audio flags. If several share the same problem, it is usually the environment, not the candidates, and you can fix it for the next batch while the drive is still running.
How to Get Cleaner Audio Before the Call
Most audio problems are preventable, and the fix belongs in the invitation, not the scorecard.
Tell candidates three things before they join: find a quiet room with the door closed, use a headset with a microphone rather than the laptop’s built-in one, and join from a stable connection. The interview runs in Chrome, Edge or Safari, on desktop or mobile, with nothing to install, so the setup ask is small. We put together a fuller list in how to prepare candidates for an AI interview.
Candidates worry about this more than recruiters realise. Some assume that a noisy background will count against them no matter what, which makes them more nervous and quieter, which makes the audio worse. A single line in the invite saying that audio problems are flagged for a human to check, not scored against them, helps. We cover what else to say to anxious candidates in when a candidate is nervous about an AI interview.
What This Design Still Cannot Fix
Two limits are worth stating.
A damaged transcript is still damaged. Switching off the voice half removes one source of noise-driven error, but if the speech-to-text misheard half the answers, the transcript half carries that error. That is exactly why flagged calls go to a person instead of being rejected, and why listening beats reading on those calls.
Moderate noise below the flag thresholds still reaches the score. The flags catch audio that is clearly unusable. A slightly noisy call that stays above the thresholds is measured normally, and a little noise can nudge a delivery reading. The voice half only ever affects the communication line, which limits how far that can travel, but it is not zero.
Clean audio is a candidate’s job to attempt and a screen’s job to not punish. To hear how the call copes with your own connection and room, try the demo interview from wherever you are sitting right now, and check the flags on the scorecard it emails you.
Frequently asked questions
Does background noise lower a candidate's AI interview score?
Not through the voice measurements. When the audio is too noisy or too silent to measure reliably, HireQwik drops the delivery half and scores communication from the transcript only. Heavy noise can still make the transcript itself less accurate, which is why the evaluator is told to send a flagged call to a recruiter rather than reject it on communication.
What audio problems does an AI voice interview check for?
HireQwik checks for a poor signal-to-noise ratio, a recording that is mostly silence, a microphone that is clipping from being too loud, and voice readings distorted by noise. Under 30 seconds of candidate speech, no delivery analysis runs at all. Each problem shows up as a named flag on the candidate's speech panel.
What should a candidate do to get clean audio for an AI interview?
Use a quiet room, a wired or Bluetooth headset rather than a laptop's built-in microphone, and a stable connection. The interview runs in Chrome, Edge or Safari on desktop or mobile, with no install. If the call drops or the audio is unusable, recruiters can re-invite the candidate rather than judge a broken recording.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in