How Every Screening Call Earns Its Verdict Tier
In July 2026 we deleted four numbers from a dropdown in our own product. The Decision filter on the interviews page had read “Strong Go (7.5+)”, “Go (6-7.4)”, “On Hold (4-5.9)” and “No Go (<4)” since April. By July every one of those ranges was wrong. Our thresholds had moved three times, and nobody had touched the labels. The filter still returned the right candidates, because it matched the verdict word rather than the number. But a recruiter who read those labels and explained a decision with them would have been quoting a rule that no longer existed.
That small embarrassment is the whole idea behind this post. An AI interview verdict is only worth something if you can say, out loud and correctly, the rule that produced it. Below is the exact order HireQwik applies after a 15-20 minute screening call, the bands that turn a score into one of four tiers, the two outcomes that are not tiers at all, and the places where I think a verdict stops being defensible.
What an AI interview verdict is in HireQwik
An AI interview verdict is HireQwik’s classification of one completed screening conversation into Strong Go, Go, On Hold or No Go. It is computed after the call from scored dimensions, knockout results and the job’s own settings. It is a recommendation to HR, stored apart from the decision a recruiter records.
It is also not the resume number. The pre-call match score has its own bands with similar names, a naming choice we explained when comparing the pre-call score with the post-call result. The verdict describes a conversation that actually happened, with a transcript and a recording behind it.
The four tiers are coarse on purpose. A recruiter working through a 3,000-candidate campus drive does not need a person ranked 7.3 against another at 7.1. They need to know which pile someone belongs in and what that pile asks of them. Four piles is about the most a busy team can act on without inventing its own sub-tiers in a spreadsheet.
The order HireQwik checks before assigning a tier
Most arguments about a verdict are really arguments about which rule applied. So the engine walks a fixed list, and the first rule that fits wins:
- Knockout reject. If a Phase-0 knockout question failed during the call, the verdict is No Go, and the failed question travels with a direct quote from the candidate. Such calls tend to be short, because the agent closes them as soon as the must-have fails.
- The job’s own evaluation settings. Jobs built with a structured screener carry dimensions, weights, thresholds and an optional communication floor, all drawn from the job description. When present, these decide.
- Built-in track schemas. A few screening tracks have fixed dimension sets and weights of their own.
- A scoring profile read from the job description. Weighted dimensions produced by our job-description analysis.
- A default. Six equal-weight dimensions, used only when nothing above exists.
Why write this down? Because “the AI said No Go” is not a reason anyone can check. “The candidate failed the knockout on night-shift availability and said, in their own words, that they cannot work nights” is. So is “their weighted average across this job’s five dimensions was 4.1, under the 4.5 line for On Hold.” Both are sentences a recruiter can check against the recording. The fixed order is what makes producing one possible.
The settings in step 2 come from the job description itself, which is the point of a per-JD screening rubric: a sales role and a support role should not be judged on identical dimensions just because both ran through the same tool.
How score bands turn a number into Strong Go, Go, On Hold or No Go
Once the rule path is settled, the tier is a comparison against bands. Two scoring styles run in production today:
| Scoring style | Strong Go | Go | On Hold | No Go |
|---|---|---|---|---|
| Weighted average, 0 to 10 (default) | 8.0 and above | 6.0 to under 8.0 | 4.5 to under 6.0 | under 4.5 |
| Points total, out of 10 | 9.0 and above | 7.5 to under 9.0 | 6.0 to under 7.5 | under 6.0 |
The points style is newer. Each core question has its own budget (the first two carry 3 points each, the third 2, and communication 2), and the total is compared with the bands. On that path a Strong Go is also marked as priority, so it rises to the top of the review list.
Both sets are defaults, not laws. A job can carry its own thresholds, and it can also switch on a communication floor. With the floor on, a communication score above zero but below the job’s threshold (4.0 out of 10 unless changed) turns the verdict into No Go whatever the average says, and the verdict carries a written line naming that threshold. Here is how that plays out on an invented inside-sales job with the floor switched on at the default. A candidate averages 6.8 across the job’s dimensions, comfortably inside Go. But her communication score is 3.5. Because 3.5 is above zero and under 4.0, the floor wins: the verdict becomes No Go, and the stored reason names the threshold of 4 she missed. A recruiter who sees a 6.8 next to a No Go does not have to guess; the sentence explaining the gap is already written. Jobs with no evaluation settings of their own also get a 4.0 floor on the default path, but only for non-engineering roles.
For a customer-facing role, most hiring managers would sign that rule. For a back-office data role they may not, which is why it is a per-job switch rather than a blanket default.
A band edge is still an edge. A 5.95 and a 6.05 sit in different tiers while being, for practical purposes, the same performance. I would rather you know that than believe the line means more than it does. It is one reason On Hold exists at all, and one reason the recruiter’s own call lives in a separate field, covered in who owns the final hiring call.
Why the verdict filter lost its numbers
Back to that dropdown. When we wrote those labels in April, they were correct. Then three things happened. Default bands moved as we calibrated against HR decisions. One pilot’s sales screening got a band set of its own. Jobs gained per-job bands. Each change was sensible. None of them touched a label that lived in a different file.
The recalibration that mattered most came when we changed the model that evaluates interviews. The new one scored about one point lower on every dimension for that sales screening. It did not reorder candidates; it ranked them in nearly the same order, just lower. Left alone, that shift would have quietly pushed good candidates from Go into On Hold and from On Hold into No Go, with nothing on screen to explain why. We moved that scheme’s bands down so each tier kept its meaning, and checked the result against interviews HR had already labelled before shipping.
The lesson we wrote in our engineering log is short: a hard-coded number that copies a value living somewhere else will rot. Because thresholds now differ by scoring style and by job, no single range printed on a filter can be right. So we removed the numbers. The candidate’s score still shows next to the verdict, and the rule that produced it sits on the job. If you run your own screening process, with or without us, the same rule applies to any slide deck or SOP that quotes a cut-off. Point to the live setting, or leave the number out.
Incomplete and Pending: two outcomes that are not tiers
Two labels can appear in the verdict column without being verdicts.
Pending is the state between hang-up and scoring. The inbox labels it “Awaiting eval”, and there is nothing to decide yet.
Incomplete means the call did not cover enough ground to judge. On the points path, if only one core question (or none) got a score, the result is Incomplete with an abandonment flag, and the inbox shows a chip explaining that the screen was flagged rather than scored low. The difference is not cosmetic: a candidate whose network died in minute four should not sit in the No Go pile beside someone who answered every question badly. What to do with those rows is the subject of the incomplete AI interview guide.
What each verdict tier asks HR to do next
A tier is useful only if it comes with an instruction. Here is how I would brief a TA team:
| Verdict | What it claims | What HR does next |
|---|---|---|
| Strong Go | The call cleared every rule with room to spare | Move to the next round quickly; see the Strong Go handoff checklist |
| Go | A clear pass with thinner margins | Advance after a short skim of the scorecard notes |
| On Hold | Not enough signal to advance or reject | A person listens and decides within a set deadline, as argued in putting a deadline on On Hold |
| No Go | A knockout failed, a floor was missed, or the score sat under the On Hold line | Read the stated reason before closing; how to read a No Go reason walks through each type |
On Hold deserves one more paragraph. In May 2026 a pilot HR team told us plainly what they wanted: the AI as a filter, HR as the judge. For them On Hold was the real shortlist, and the AI’s main job was to peel off the obvious No Gos. We redesigned our bands around that. So if your On Hold pile looks large, that may be the system doing what a careful team asked for, not failing to decide.
Where a verdict stops being defensible
Even with a fixed rule order, there are places I would not stand behind a tier without a person looking.
Evidence missing from the transcript. A knockout No Go that quotes words the candidate never said should not stand. Since August 2026 the engine routes those to On Hold for review. How we measured that problem, and how our first measurement was itself wrong, is in checking AI quotes against the transcript.
A threshold nobody has tested against reality. Bands are a hypothesis until you compare verdicts with who you actually hired. HireQwik can store what finally happened to each candidate, but that data deliberately does not move thresholds on its own; the reasoning is in why outcome data should not auto-tune thresholds.
A model or rubric change. Any change to the evaluator or the rubric can shift every score at once, as ours did. After such a change, re-read a sample from each tier before trusting the new distribution.
The US National Institute of Standards and Technology describes trustworthy AI through properties such as valid and reliable, accountable and transparent in its AI Risk Management Framework. I find that more useful than any vendor accuracy claim, including ones we could make. A defensible verdict is one where each of those properties maps to something you can open: the rule path, the band, the quote, the recording.
A recent field experiment on AI-led job interviews, covering more than 50,000 applicants, found AI interviewers followed a more consistent structure than human recruiters. Consistency is the precondition for a defensible verdict, not the proof of one. The proof is being able to show your working for any single candidate.
Set the rules before the first call
If you are setting up a new role, open it from the HR dashboard and read its evaluation settings before any invite goes out: which dimensions count, where the bands sit, whether the communication floor is on. Those are the sentences you will be quoting in three weeks when a hiring manager asks why someone was rejected.
Over more than 1,099 pilot interviews, the verdicts that caused trouble were almost never the ones where the model got it wrong. They were the ones where nobody could say which rule had applied. To have us review a job’s settings with you before a big drive, talk to us.
Frequently asked questions
In what order does HireQwik decide an AI interview verdict?
A recorded knockout reject is checked first. If there is none, the job's own evaluation settings decide, then built-in schemas for specific screening tracks, then a scoring profile read from the job description, and finally a default six-dimension average. The first rule that applies sets the tier.
Are Strong Go and No Go thresholds the same for every job?
No. The default weighted path marks 8.0 and above out of 10 as Strong Go, 6.0 as Go and 4.5 as On Hold. Jobs scored on the points rubric default to 9.0, 7.5 and 6.0 out of a 10-point total. Either set can be changed per job.
What should HR do with each AI verdict tier?
Move Strong Go candidates forward quickly, advance Go candidates after a short skim of the notes, give every On Hold candidate a human decision within a set deadline, and read the recorded reason before closing any No Go. Each tier is an instruction, not just a label.
See your own candidates screened
Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.
Existing customer? Sign in