What If the AI Rejects Someone for Words They Never Said?
In early August 2026, a number on our internal board said roughly 15% of the quotes behind our knockout rejects were fabricated. If that was true, some candidates had been auto-rejected on words they never said. I remember reading it twice. For a company whose whole pitch is verdicts you can check, it was close to the worst finding possible.
It turned out to be mostly wrong, and the way it was wrong taught us more than a clean result would have. This post is about AI hiring hallucination in the one place it does the most harm, inside a reject: how we measured it, why the first measurement misled us, the check that now runs on every knockout quote, and what that check still cannot catch.
What an AI hiring hallucination looks like in a screening call
An AI hiring hallucination is output from a language model that reads as confident evidence about a candidate but is not supported by what the candidate actually said or submitted. In a voice screen, the risky form is a quote or reason attached to a decision that cannot be found anywhere in the transcript.
Researchers have measured how common unsupported output can be, even in specialised tools. A Stanford study of legal research assistants found they hallucinated on 1 in 6 or more benchmark queries. A 2026 field experiment on AI voice interviews lists keeping the agent free of hallucinations among the core challenges of automated interviewing. None of that is a reason to avoid AI screening. It is a reason to check the specific spots where an invented sentence would change someone’s outcome.
Why a knockout reject is where a fabricated quote hurts most
HireQwik’s knockout questions run at the start of the call and test hard requirements taken from the job description. A failed answer ends the conversation early and produces a No Go. The agent records which question failed and the candidate’s words as evidence, and the verdict engine looks at knockouts before anything else, as the verdict tier explainer lays out.
That design is efficient and fair when the quote is real. It is uniquely dangerous when it is not. There is no full interview to balance it, and no other dimension to lift the average. The candidate never reaches the remaining questions. And the recorded quote is precisely the sentence a recruiter would read out as the reason, which is why reading a No Go reason leans on it so heavily. A fabricated quote inside a knockout is a reject with a made-up justification stapled to it.
How we measured it, and why the first number was wrong
The 15% figure came from an earlier verifier. It compared the start of each quote with only the candidate’s last ten or so messages. That sounds reasonable until you see how speech-to-text chops up a real conversation. About 37% of candidate turns in our data were four words or fewer: “yes”, “okay, so”, “one second”. Ten turns like that cover roughly 20 seconds of speech. A candidate who answered the notice-period question early and then kept talking for a minute had their real answer fall outside the window.
So the verifier flagged 162 of the 166 knockout quotes in production as unverified. When we re-checked every one against the full transcript, 163 were there, word for word or close to it. One was genuinely missing: the agent had paraphrased, and its wording did not appear anywhere in that call. Two could not be checked, because the stored text was a bracketed description such as “no audible response” rather than a quote.
One real case out of 166 is small. It is not zero, and it was a real person who had been rejected on a sentence they never spoke. The lesson I took was not “our AI is fine.” It was that a checker can fail in a way that looks exactly like the problem it is meant to detect. Before acting on an alarming number, re-derive it from the source.
How HireQwik’s knockout quote check works now
Since 8 August 2026 there is one shared grounding check in our core code, so every place that verifies a quote applies the same rule:
- Normalise both texts. Lowercase, strip punctuation, collapse spaces, so “Ninety days.” and “ninety days” match.
- Short quotes need an exact match. A quote of six words or fewer must appear in full somewhere in the candidate’s transcript.
- Longer quotes need substantial overlap. The quote is split into runs of six consecutive words, and at least a third of those runs must appear in the transcript. That tolerates small transcription differences without letting an invented sentence through.
- Always the full transcript. Never a recent window. That was the original bug.
- Three answers, not two. Grounded, not grounded, or unknown.
A background task stamps the result on recent interviews about once a minute, and a one-off run stamped every historical knockout in production.
Here is how the rule plays out on an invented example. The job needs people for rotating night shifts, and the agent records this knockout evidence: “I cannot work night shifts because I take care of my mother at home every evening.” That is sixteen words, so it counts as a long quote. Cut into runs of six consecutive words, it gives eleven runs, and at least three must appear in the transcript for the quote to count as grounded.
Now suppose speech-to-text wrote the answer as “I can’t work night shifts because I take care of my mother at home every evening.” Even with punctuation stripped, “can’t” and “cannot” are different words, so the two runs that include it fail. The other nine match exactly, well above the three needed, and the quote passes. Small transcription differences cost a run or two, not the whole quote.
Suppose instead the candidate had said “Night shifts are fine for me, I take care of my mother in the mornings.” One run, “I take care of my mother”, does appear. One out of eleven is short of three, so the quote fails and the reject becomes an On Hold. That is the point of the threshold: a shared phrase is not the same as a shared answer. A short quote such as “ninety days notice” gets no such allowance; all three words must appear together.
Knockout rules themselves come from the job description, so the evidence check pairs naturally with the kind of question review described in what a candidate sees when a knockout fires.
The third answer is the one I care about most. If the check cannot run, because there is no transcript or the stored text is not a real quote, the verdict is left exactly as it was. Missing evidence never flips a result in either direction. A check that turned every broken recording into an On Hold would have flooded review queues and taught recruiters to ignore the warning.
What HR sees when a knockout quote fails the check
When a knockout reject’s quote is provably absent from the transcript, the verdict engine does not issue No Go. The candidate becomes On Hold, and the reason text explains that a knockout reject was recorded for a named question but the quoted answer was not found in the candidate’s transcript, then asks HR to review the call before acting. In the inbox and on the interviews list, a small warning chip marked “review” sits beside the verdict.
What to do with that row:
- Open the recording and jump to where the knockout question was asked. It will be near the start.
- Listen to the actual answer and decide whether it meets the requirement.
- Record your decision in the HR Review panel with a one-line note on what the candidate really said. Why the AI’s verdict and your call stay in separate fields is covered in the post on recruiter overrides.
- If the candidate did meet the requirement, remember the call ended early, so the rest of the screen never happened. Resetting their booking gives them a fresh slot for the complete screen.
Only one of 166 knockout quotes failed the check when we measured, so this should be a rare row, not a daily chore. If you start seeing it often on one job, the knockout question itself may be ambiguous enough that the agent and the candidate are hearing two different questions. Rewrite it.
Writing knockout questions that produce a clean quote
The best protection against a bad knockout quote starts before the call, with the question itself. A vague question invites a vague answer, and a vague answer is hard to quote accurately, for a model or for a person. A precise question tends to get a short, direct reply that is easy to find in any transcript.
Some before-and-after pairs from the kind of jobs our pilots run:
| Vague question | Precise question |
|---|---|
| “Are you okay with the shift timings?” | “This role works rotating night shifts from 10 PM to 7 AM. Can you work those hours?” |
| “Are you flexible on location?” | “The job is based in our Pune office five days a week. Can you work from there?” |
| “When can you join?” | “We need someone to start within 30 days. Can you join by then?” |
The precise versions share three habits. They state the requirement inside the question, so the candidate knows exactly what is being asked. They invite a yes or no, with room to explain. And they carry one condition each, never two. Notice period deserves particular care in Indian hiring, where 90-day notices and buyouts are routine, and asking about notice period on the first call covers how to word it so the answer is unambiguous. “Can you work nights and relocate?” produces an answer the agent may have to split, and splitting is where paraphrase creeps in.
There is a fairness benefit too. A candidate who hears the full requirement can answer honestly the first time, instead of guessing what “okay with the timings” means and being rejected for guessing wrong. If your recorded quotes often look like half-answers, such as “yes, mostly” or “depends”, the question is usually the problem, not the candidate and not the model.
What the quote check does not catch
I would rather list the gaps than have you find them.
It covers knockout quotes, the evidence behind the one verdict path with no human balance built in. The notes on other dimensions are evidence to confirm against the recording, not guaranteed quotations. A short quote that happens to match an unrelated phrase elsewhere in the transcript would pass. A longer quote that paraphrases heavily but accurately might fail and route to review, which is the safe direction to err. And the check says nothing about whether the knockout question was fair to ask. For calls that end early for reasons other than a knockout, the incomplete interview guide explains what is and is not scored.
Questions to ask any AI screening vendor about hallucinated evidence
Whether or not you use us, these five questions separate a vendor who has looked at this problem from one who has not:
- When your system rejects a candidate, is a quote or reason attached, and is it checked against the transcript?
- Is that check run against the full transcript or only a recent window?
- What happens to the verdict when the check fails, and when it cannot run at all?
- How many rejects failed the check the last time you measured, and how did you verify the checker itself?
- Can a recruiter jump straight to the moment in the recording that the reason refers to?
A vendor who answers the fourth question with an accuracy percentage and no method has not verified their checker.
Look for the review chip on your next drive
In your HireQwik account, filter a finished job’s Screened tab to On Hold and look for the warning chip beside the verdict. If you find one, listen to the first two minutes and decide. If you find none, that is the expected result, and now you know what it would look like if the check ever fired. Questions about how the check behaves on your roles are welcome: reach us here.
Frequently asked questions
Can an AI interviewer invent what a candidate said?
Language models can produce fluent text that is not supported by their input, so a recorded quote can in principle be wrong. HireQwik checks each knockout quote against the candidate's full transcript. If the words are provably absent, the reject is held as On Hold for a person to review.
What does the knockout review chip mean in the HireQwik inbox?
It means a knockout reject was recorded, but the quote behind it could not be found in what the candidate actually said. The candidate is placed On Hold instead of No Go, and the reason text asks HR to listen to the call before acting.
How many knockout quotes failed the grounding check?
When the check was first run across production in August 2026, 163 of 166 knockout quotes were found in the candidates' transcripts, one was not, and two could not be checked. The single ungrounded case now shows as On Hold for review.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in