Is Your Candidate Using ChatGPT During the AI Interview?
Will a candidate using ChatGPT during an AI interview get caught by the transcript? No. A copied answer from a chatbot reads well. That is the whole point of a chatbot.
The audio is a different story. On August 8, 2026, we shipped a set of AI-assist signals to production and ran them over our interview history. Out of 2,507 past interviews, two came back in the high band. One was a real candidate, passed to a person to review the recording. The other was an internal test call, made by a team member who, it turned out, really had been doing something else on screen while answering.
The signals told the truth about our own teammate. That was the moment I trusted them enough to write about them. This post covers the one audio pattern that gives live help away, the seven signals around it, and why none of them is allowed to reject anybody.
Why the Transcript Cannot Spot ChatGPT
A chatbot answer is designed to pass a text check. It is on-topic, structured and complete. If your screen judges only the words, a candidate reading from a chatbot scores like a strong candidate, because on paper they are one.
This is where the text and the audio part ways, the split we unpack in transcript versus voice in AI interviews. The words look earned. The timing and the delivery show how they were produced.
Recruiters have noticed. Technical assessments went first, as we noted when assessment cheating doubled in a year, and live interviews are next. Harvard Business Review asked last year whether hiring teams are interviewing a candidate or their AI, and interviewing.io showed in a controlled experiment how easy it was to cheat with ChatGPT in technical interviews without being caught.
The Slow-Then-Fluent Pattern in the Audio
Here is the pattern that gives live help away, and it runs against normal human behaviour.
When an honest candidate takes a long time to start an answer, it is usually because they are thinking. When they do start, they hesitate, restart and fill gaps, because they are still thinking. Slow starters hesitate more, not less.
A candidate using a chatbot does the opposite. They type or dictate the question, wait for the response, then read it out. The audio shows a long silence, then a clean, even, hesitation-free answer. Slow, then fluent. We call it the type-wait-read pattern.
Our signal fires on it only when both halves are extreme at once: a typical wait of 12 seconds or more before substantive answers, and 1.5 hesitations or fewer per 100 words once the candidate talks, over at least 150 words of speech. Honest candidates who are that slow are almost never that fluent.
This is different from someone reading a prepared script. A reader answers quickly and fluently, because the text is already in front of them. That pattern has its own flag, covered in how a candidate reading answers sounds. Slow-then-fluent is the signature of help arriving during the call.
The Seven AI-Assist Signals, One by One
Slow-then-fluent is the strongest signal, but it never acts alone. The composite checks seven things, all from data the call already records:
- Tab switches. The browser reports when the candidate leaves the interview tab. Only 3.4% of interviews show any at all.
- Long answer latency. A typical wait before substantive answers of 15 seconds or more, which is roughly the slowest tenth of honest candidates.
- Slow but fluent. The inversion described above, weighted most heavily.
- Read-speech acoustics. The existing possibly scripted flag, counted as one input.
- Essay structure. Long spoken answers that repeatedly open with written-style phrases such as “firstly,” “in conclusion” or “that’s a great question.”
- Zero self-repair. No self-corrections across 200 or more words. People revise live speech, readers do not.
- Camera off. Turned off at least twice. One accidental toggle happens to anyone, so this carries the smallest weight.
Each signal that fires adds weight, and the total lands in a band: low, elevated or high. Elevated starts at 35 points and high at 65. The heaviest single signal, repeated tab switching, tops out at 35, so one signal on its own cannot reach high. It takes several independent signals agreeing.
Two examples show how that plays out.
A candidate whose answers are slow-then-fluent (25 points) and who also trips the read-speech flag (20 points) lands at 45, elevated. That is worth a listen, but it is also consistent with a nervous person reading notes they wrote the night before.
A candidate with three tab switches (35 points), the slow-then-fluent pattern (25) and no self-corrections across a long call (10) lands at 70, high. Three different kinds of evidence, from the browser, the timing and the speech, all pointing the same way. That is the profile worth opening first.
What the list does not include matters as much. There is no gaze tracking, no eye-movement model, no screen recording and nothing for the candidate to install. Every signal comes from timing, audio and the browser session the interview already runs in.
How We Set the Thresholds From 1,447 Interviews
The easy way to build this would have been to guess. “More than eight seconds before an answer is suspicious” sounds reasonable. It would have been a disaster.
Before wiring anything in, we pulled 1,447 finished production interviews from the previous 60 days. The median wait before a substantive answer was 6.9 seconds. The slowest tenth waited 15.1 seconds or more. An eight-second rule would have flagged about 30% of honest candidates. Nervous people, careful people, people answering in their second language and people on a patchy line all take time. That last group has its own handling, set out in what poor audio does to scoring.
So every threshold sits at the tail of the honest distribution, not in the middle of it. Then we ran the whole engine over 2,507 past interviews before switching it on. The result was 97.4% low, 45 elevated, 2 high and 18 with too little data to judge. A signal that fired on a third of your candidates would train your team to ignore it within a week. One that fires on a handful gets looked at.
A threshold inside the honest crowd does not detect cheating. It manufactures suspicion. That rule shaped every number in this system.
Built From Data the Call Already Had
Every one of the seven signals is computed from something the interview was already recording: answer timestamps in the transcript, the speech measurements from the audio, and the tab and camera counts the browser session reports. Nothing new is asked of the candidate. There is no extra consent screen, no plugin, no second camera.
That choice had a practical payoff. Because the inputs already existed, the signals could be run backwards over every past interview on the day they shipped, which is how we got the 2,507-interview result above before a single new call came in.
It also had a bug worth admitting. The signals wait until a call’s speech analysis has finished, so they do not judge a call on half its data. During testing on staging we found that calls with no candidate audio at all, such as browser test calls, never finish speech analysis because it never starts. Those calls would have waited forever. We added a fallback that stamps them after two hours, and a test that fails if that fallback is ever removed. A check that silently never runs looks exactly like a check that found nothing.
When the Band Says Insufficient
Eighteen of the 2,507 backfilled interviews came back with a fourth result: insufficient. That is not a pass. It means the call carried no usable timing, speech or browser-session data at all, so there was nothing to judge either way.
Treat insufficient the way you would treat a missing reference. It tells you nothing about the candidate, good or bad. If that interview is otherwise borderline, listen to it yourself instead of reading the empty band as reassurance.
Why the Signals Never Change a Score
Nothing downstream reads the AI-assist band. Scores, verdicts and auto-reject bands are computed without it, and our test suite fails if scoring code ever starts looking at it.
The reason is the same asymmetry that shapes everything we build around delivery. Letting a cheater through means one extra interview that goes nowhere. Rejecting an honest person because a pattern looked odd loses you a good hire and takes away a fair chance they will never know they lost. The two harms are not the same size, so we do not treat them as if they were.
Missing data also fails in the safe direction. If a signal cannot be computed, because timestamps are missing or the transcription service was down, it simply does not fire. Partial data can only lower the band, never raise it.
On the candidate’s record, the panel shows the band, how many signals were checked, how many fired, and the evidence each one fired on. The copy under it reads “signals, not proof. Verify in the recording before acting.”
Handling an Elevated or High AI-Assist Band
Pull up the call in the HR dashboard and hold each fired signal up against what you actually hear.
- For slow-then-fluent: skip to a long pause. Listen to the silence, then the first ten seconds of the answer. Faint typing, a change in tone when the answer starts, or an answer that sounds read rather than spoken are what you are listening for.
- For tab switches: look at when they happened. One switch at the start, while the candidate found their headphones, means little. Switches clustered just before each answer mean more.
- For essay structure: check that the candidate can leave the script. A follow-up question that builds on their own answer is hard to paste into a chatbot fast.
Then decide what you would have decided if a human interviewer had raised the same concern. Usually that means inviting the candidate to a short live conversation, not rejecting them. Someone who used help will struggle there. Someone who was just slow and articulate will be fine, and you will have lost nothing.
Warn Candidates in the Invitation
The cheapest anti-cheating measure is not a signal at all. It is a sentence in the invitation.
Most candidates who reach for a chatbot mid-interview are not planning fraud. They are anxious, they assume everyone else is doing it, and they assume nobody will notice. A plain line telling them the call is a live conversation with follow-up questions, and that long pauses and leaving the interview tab are visible to the hiring team, removes the last two assumptions.
It also protects you. A candidate who was told up front and chose to switch tabs anyway is a much easier conversation than one who can reasonably say nobody mentioned it. Our notes on what goes in the pre-interview message covers the rest of what belongs in that message.
What These Signals Cannot Promise
It would be easy to oversell this, so here are the limits plainly.
A candidate using a second device with a fast reader could avoid long pauses. A candidate who reads a chatbot answer with fake hesitations could pass the fluency check. And a candidate who uses help on one question out of eight will barely move a typical wait time. These signals raise the cost of cheating. They do not make it impossible, and any vendor who claims a cheating detector catches everything is describing a demo, not a product.
What they do reliably is point your attention at the few calls worth listening to, without accusing the many that are not. Pair them with a screen that sounds like a conversation, one where voice beats a chat-based screen, and the easy routes to cheating get noticeably harder. If a flagged call sits in the review queue, reading the communication score panel covers the speech numbers sitting beside it.
To see a scored call from the candidate’s chair, run through the demo interview. Your scorecard follows by email a few minutes after you hang up.
Frequently asked questions
How do you know if a candidate is using ChatGPT during an AI voice interview?
The strongest audio clue is a long silence before an answer, followed by speech that is unusually fluent. Honest candidates who take a long time to answer tend to hesitate more once they start, not less. HireQwik treats slow-then-fluent as one of seven signals and asks a recruiter to verify in the recording.
Can the AI-assist signals reject a candidate automatically?
No. The signals are flag only. They never change a score, a verdict or an automatic reject rule, and tests in the codebase enforce that. The panel on the candidate's page says signals, not proof, and tells the recruiter to check the recording before acting on anything.
How often do AI-assist signals fire on honest candidates?
Rarely, by design. When the signals were run across 2,507 past interviews, 97.4% came back low, 45 elevated and 2 high. The thresholds were set from the real spread of 1,447 production interviews rather than guesses, because a guess like eight seconds is suspicious would have flagged about 30% of honest candidates.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in