AI screeningRecruitingHR techCompliance

The AI Recommends, the Recruiter Decides: Keep Both on Record

HireQwik September 30, 2026 10 min read

The sentence that shaped how HireQwik handles verdicts did not come from an engineer. It came from a pilot customer’s HR team in May 2026, soon after our recalibrated bands went live: the AI should be the filter, and HR should be the judge. On Hold, they said, was their real shortlist. What they needed from us was the obvious No Gos cleared away, not software pretending to make hire decisions.

We rebuilt our bands around that idea. It also settled an argument we had been having about the data model. When the AI says one thing and the recruiter decides another, which one counts as “the decision”? Our answer is both, stored apart. This post explains the AI verdict vs recruiter decision split in HireQwik, what the review panel records, why neither field ever overwrites the other, and a routine that keeps a busy team from quietly handing its judgement to the model.

AI verdict vs recruiter decision: two fields, on purpose

In HireQwik, the AI verdict is the recommendation computed after a screening call: Strong Go, Go, On Hold or No Go. The recruiter decision is a separate choice a person records using the same four options, saved with their name and the time. Setting the second never changes the first.

The verdict comes from a fixed rule order and a set of bands, laid out in how each verdict tier is decided. When the verdict is a reject, its recorded cause can be read and repeated, as shown in reading a No Go reason. This post picks up at the next moment, when a human opens the candidate.

What the HR Review panel on Interview Details records

Open any completed interview and scroll to the panel titled HR Review. It holds three things.

A decision row. Four buttons labelled No Go, On Hold, Go and Strong Go, the same vocabulary as the verdict, so nobody has to translate between two scales. Click one and it saves. Underneath, a line reads “Set by”, then the reviewer’s name, then the date and time.

Your scores next to the AI’s. Every dimension the AI scored is listed with its number and an empty box beside it. A recruiter can type their own score from 0 to 10 in half steps and save. None of this is required. It becomes useful the day a hiring manager asks “which part did you disagree with?”, because the answer is sitting there dimension by dimension.

Screening notes. A comment box for whatever neither number captures: “strong on process, could not give one customer example,” or “called back after the drop, much better second time.”

On the interviews list, the recruiter’s decision shows as its own pill beside the AI verdict, and completed interviews without an HR decision are counted as awaiting review. That count tells a team lead how much of a drive no person has looked at yet.

Why HireQwik never overwrites the AI verdict

One editable field would be simpler to build. We chose not to, for three reasons.

Disagreement is data. If verdict and decision share a field, every disagreement vanishes the moment it happens. Kept apart, a pattern can surface: say, recruiters keep moving On Hold candidates to Go for one role. That usually means the role’s bands or its communication floor are too strict, and you would never spot it with a single field.

Accountability needs a name. “Set by” plus a timestamp is a small thing. But when a candidate asks months later who made the call, “the system” is not an answer anyone should give. A named human decision next to an explained AI recommendation is.

The regulation points the same way. Article 14 of the EU AI Act asks that people overseeing high-risk AI, a category that includes hiring tools, stay aware of the pull to over-rely on its output and remain able to disregard or override it. You can only show an override happened if the original recommendation is still there to compare against.

Where the recruiter’s decision should differ from the AI

Here is the uncomfortable part. Two fields do not guarantee two opinions.

A University of Washington team ran a resume-screening experiment with 528 people and simulated AI recommendations that were deliberately biased. Without suggestions, participants picked candidates from different groups at equal rates. With a biased AI beside them, they followed its preferences up to about 90% of the time. The UW News summary is worth ten minutes of any TA lead’s time, and the full research paper is more sobering still: even people who rated the AI’s advice as low quality were swayed under some conditions.

I do not think the answer is hiding the AI verdict from recruiters. Reading every candidate blind at campus-drive volume is not realistic; that is why screening tools exist in the first place. The answer is to build moments into the routine where the recruiter forms a view before, or apart from, the verdict, and then to watch whether disagreement ever actually happens.

There is no correct disagreement rate, and I would distrust anyone who quotes one. But zero across a large drive is a signal worth checking. It usually means people are clicking the same button as the pill, often in bulk. HireQwik’s filters stop bulk actions from touching hidden rows, as described in the bulk-approve mistake, but no filter can turn a bulk click into real judgement.

A four-step review routine, one step per verdict tier

This is the routine I suggest to pilot teams. It takes about as long as clicking through, and it keeps the recruiter’s call their own.

  1. For On Hold, listen before you look. Open the recording or transcript before reading the scorecard notes. Form a Go or No Go view in your head. Then read the notes and see where you differ. On Hold is where the AI has said outright that it is unsure, so it is where human judgement adds the most.
  2. For Go, score one dimension yourself first. Pick the dimension that matters most for the role and enter your own score before glancing at the AI’s. If the gap is more than two points, read that answer again.
  3. For No Go, read the reason. Never close a reject on the pill alone. Read the recorded cause and decide whether you would say the same thing to the hiring manager.
  4. For Strong Go, spot-check. Take one Strong Go in ten and listen to the first question. On the points rubric a Strong Go is marked priority and tends to move fast, which is exactly why an occasional check matters.

Record the result in the decision row, and add one line of notes whenever your call differs from the verdict. That line is what makes the difference reviewable later. If On Hold rows are piling up untouched, set a deadline as described in giving every On Hold a decision date.

Three disagreements, and what each one teaches

Overrides are most useful when you read them as feedback on the setup, not just on the candidate. Here are three invented but typical cases from a customer support drive.

The AI says Go, the recruiter says No Go. Midway through the call, the candidate mentions they can only work from home. The job is office-based, but nobody made location a knockout, so the call carried on and scored well. The recruiter is right to reject. The lesson is not about the AI’s judgement at all: a hard requirement was missing from the knockout list. Add it before the next batch of invites, and the next candidate like this gets a clear, quoted reason at the start of the call instead of a Go that a person has to undo.

The AI says No Go, the recruiter says On Hold. The score sits at 4.3, under the 4.5 line. The recruiter listens and hears a nervous start and two solid answers near the end. Moving the candidate to On Hold for a second look is reasonable. If the same pattern keeps showing up on this role, with late-warming candidates just under the line, it is worth asking whether the first question is too hard as an opener.

The AI says Strong Go, the recruiter says Go. The candidate was fluent and confident, but the recruiter notices the answer to the process question was generic. Nothing is broken here. A Strong Go says the conversation passed every check by a wide margin, and a recruiter trimming it one tier is the review working as intended.

Notice that only one of the three says anything about the model. The other two point at the knockout list and the question order, both of which your team controls.

Writing an override note that still makes sense in six months

The note is the part of an override that people skip, and it is the part that matters most later. A decision without a reason tells an auditor that someone clicked. A decision with a reason tells them why.

Good notes share three traits. They name the evidence, they are specific to this candidate, and they are short. Compare:

Weak noteUseful note
“Not a fit”“Said at 04:10 they can only work remote; role is on-site in Chennai”
“Better than the score”“Last two answers strong; first answer rushed, likely nerves. Worth a panel round”
“Agree”No note needed when you agree; save the effort for disagreements

A timestamp from the recording is the single most helpful thing you can add. It lets the next reader, whether a hiring manager, a colleague covering your drive or a legal reviewer, jump to the exact moment you relied on. Keep notes about the answer, never about personal traits unrelated to the role.

How decisions and hiring outcomes close the loop

A third record comes after the verdict and the recruiter’s call: what actually happened. HireQwik can store the real outcome for each candidate, from hired to withdrew, and each job gets a card that sets the screener’s verdicts beside those outcomes, explained in recording what actually happened.

Put the three side by side and you can answer the question that matters: when a recruiter overruled the AI, who turned out to be right? We deliberately do not let that data change bands by itself. Moving a threshold stays a human choice, and the case against auto-tuning sets out why. But without two separate fields in the middle, there would be nothing to compare in the first place.

One practical note for teams with a weekly rhythm: the weekly screening recap now carries a hires line, which makes a monthly look at overrides against outcomes much easier to schedule.

What this split does not solve

Recruiters carry biases too. Separating the fields makes a biased human decision visible; it does not prevent one. The per-dimension HR scores are a working tool, not a formal calibration study. And “Set by” records who clicked, not how carefully they looked. The routine above exists because a data model alone is never enough.

Calls with a problem of their own need the same discipline. A knockout reject resting on words the transcript does not contain arrives as On Hold (see how quoted evidence is verified), and a call that dropped part-way is flagged rather than scored, as explained in the incomplete interview guide. In both cases the recruiter’s decision is the one that closes the file.

The design choice under all of this is simple, and I think it is right. The AI’s job is to make a clear, explained recommendation. A named person’s job is to decide. When the two disagree, both should still be on the page.

Run the routine on twenty interviews

Pick one running job in the HR dashboard. For the next twenty completed interviews, follow the four steps and record a decision on each. At the end, count how often your call differed from the verdict and reread the notes you left. That count, and the reasons behind it, says more about your screening than any vendor’s accuracy slide. If you want help reading the results, contact us.

Frequently asked questions

Does changing the HR decision in HireQwik change the AI verdict?

No. The recruiter's decision is saved in its own field with the reviewer's name and the time. The AI verdict stays as it was computed, so anyone looking at the candidate later can see the recommendation, the human call, and where the two differed.

How often should a recruiter disagree with the AI verdict?

There is no correct rate, and nobody has a trustworthy benchmark for it. What matters is that disagreement is possible and happens for stated reasons. If a team has not overruled a single verdict across a large drive, check whether people are reviewing or only confirming.

Can a recruiter score a candidate alongside the AI?

Yes. On Interview Details, the HR Review panel lists each scored dimension with the AI's number beside an empty box. The recruiter can enter their own score from 0 to 10 in half-point steps, save it, and add screening notes explaining the difference.

See your own candidates screened

Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.