AI screeningVoice AIHR techRecruiting

From Hang-Up to Number: How HireQwik Scores a Screening Call

HireQwik October 2, 2026 11 min read

The question I hear most in product demos is not “how accurate is it?” It is quieter than that. Someone points at a candidate’s row, sees 6.4, and asks: “Where exactly does that number come from?” Usually it is a TA lead who has been burned before by a tool that produced scores nobody could explain.

That question deserves a complete answer, not a reassuring one. An AI interview score is only useful if a recruiter can follow it backwards, from the total to the dimension scores, from each dimension score to the scale it was graded on, and from the scale to the words the candidate actually said. This post walks through every step HireQwik takes between the moment a screening call ends and the moment a number appears in your inbox, including the parts that are optional and the things the score deliberately leaves out.

How HireQwik defines an AI interview score

An AI interview score in HireQwik is a set of numbers produced after a screening call by an evaluator that reads the full transcript against the job’s scoring dimensions. Each dimension gets its own score on a written scale, with notes and flags explaining it, and those scores are combined into one total that decides the verdict tier.

Two things follow from that definition. The score describes one conversation, not a person’s career. And every step from transcript to total is a rule you can inspect, and the eight steps below show each one.

Step 1: The call ends, and nothing has been scored yet

During a screening call, the voice agent has one job: run a structured conversation well. It asks the job’s questions, answers a candidate’s questions about the company, and closes politely. It does not grade answers as it goes.

Separating the two is deliberate. An agent that tried to score in real time would be making judgements on half an answer, before a candidate who started nervously had a chance to recover. Scoring after the call means the evaluator always sees the whole conversation at once, the way a careful human reviewer would if they had time to read every transcript.

The one exception is a knockout question. If a must-have requirement fails, such as night-shift availability or a notice period the role cannot wait for, the call closes early and the result is a No Go with the candidate’s answer quoted. That path is about a stated requirement, not a score, and our transcript check on knockout quotes explains how we make sure the quoted words were really said.

Step 2: A substance check decides whether there is anything to score

Not every call contains enough to grade. Before any evaluation, HireQwik checks that the call has real substance. A call shorter than two minutes, or with fewer than three transcript turns, is marked incomplete rather than completed, and no score is produced from it.

There is a second, softer guard inside the evaluator itself. If the candidate gave fewer than six responses, the evaluator is told plainly that the interview was incomplete. It must score 0 for any dimension where the evidence is missing, state in the notes that the call ended early, and add an “ended early” red flag. The instruction is explicit that a 0 meaning “not assessed” is better than “a speculative middle score like 5 or 6 with no supporting evidence.”

I care about that sentence more than almost any other line in our scoring code. A system that fills gaps with 5s produces averages that look reasonable and mean nothing. What happens to a call that stops part-way is covered in the guide to incomplete interviews.

Step 3: The evaluator reads the whole transcript against the job’s dimensions

Once a call passes the substance check, the evaluator receives the full transcript and a list of dimensions to score. Which list it gets depends on how the job was set up, checked in this order:

  1. Settings written for this job. A job created from a structured screener brings its own dimensions, weights or point budgets, and rubric text for that role.
  2. A scoring profile built from the job description. If there are no explicit settings, our job-description analysis supplies weighted dimensions, plus role-specific red flags and green flags for the evaluator to watch for.
  3. A default set of six. Communication, technical competency, experience relevance, cultural fit, motivation clarity and overall impression, used only when nothing more specific exists.

The order matters because a sales role and a back-office role need different yardsticks. Reading dimensions from the job description is the idea behind matching the rubric to the role, and it is why I would always rather a customer spend twenty minutes on a job’s settings than rely on the default six.

Step 4: Each dimension gets a number on a written scale

The evaluator does not invent its own idea of what a 7 means. On the standard 0 to 10 scale, the meaning of each band is written into its instructions:

ScoreWhat it is told the score means
0Not enough data to assess
1-2Very poor: incoherent, no meaningful response, cannot perform the basics
3-4Below average: vague, generic answers without specifics or examples
5Average: meets minimum expectations, nothing distinctive
6Above average: shows competence with some relevant detail
7-8Good to strong: specific examples, clear depth, real capability shown
9-10Exceptional: outstanding evidence, rare at entry level

Alongside the scale sits calibration guidance. Scores must rest on evidence from the transcript. A vague answer such as “I’m a team player” or “I work hard”, with no example behind it, lands around 4 to 5. Specific project details, real numbers or clear depth earn 6 to 7 and above. Scores of 8 and up are reserved for genuinely standout answers.

That guidance is the single biggest reason two candidates with similar confidence can score very differently. Confidence is not evidence. A candidate who says “I handled an angry customer once, it was fine” and one who explains what the customer wanted, what they tried first and how it ended are giving the evaluator very different amounts to work with. What each band looks like in practice is the subject of what a 5, 7 or 9 actually tells you.

The scoring call itself uses a low-randomness setting, so the same transcript and the same rubric produce stable results rather than a different number each time. How stable, and where variation still creeps in, is a fair question that deserves its own post: AI interview score consistency.

Step 5: Notes, red flags and green flags travel with the numbers

A bare number is the least useful part of a score. Along with the dimension scores, the evaluator writes a two-to-three sentence evaluation note and two short lists: red flags (concerns worth checking) and green flags (strengths worth noticing). When a job’s description includes role-specific flags, such as “cannot explain a basic sales cycle” or “has handled field collections before”, the evaluator is told to look for them.

These are what turn “6.4” into something a recruiter can repeat. “Clear on why they want the role and gave a solid example of owning a delayed delivery, but could not explain how they would prioritise two urgent tickets” is a sentence a hiring manager can verify by playing the recording. The number summarises it; the note justifies it.

A word of caution: the notes are the evaluator’s summary of the conversation, not verbatim quotations. Treat them as pointers into the transcript or recording, and check them there before acting on a borderline score.

Step 6: Communication, and the optional voice layer

Communication is the one dimension where how someone speaks can matter as much as what they say. By default, HireQwik scores it the same way as every other dimension: from the transcript.

Some accounts also have speech assessment switched on. On those accounts, the candidate’s own audio track is analysed for pace, hesitation, pronunciation and fluency level, and that delivery score is averaged equally with the transcript score for the communication dimension only. Every other dimension stays transcript-only. Both halves are stored next to the blended number, so a reviewer can see which half moved the result, as explained in how to review a communication score.

If the recording is poor, with heavy noise or long stretches of silence, the voice half is dropped and communication falls back to the transcript score, with the notes saying so. Every evaluation is also tagged as transcript-only or speech-augmented, so you always know which kind of score you are looking at.

Step 7: From dimension scores to a single total

With every dimension scored, HireQwik combines them into one total. There are two ways this happens, and the job’s settings decide which one applies.

Weighted average (0 to 10). Each dimension carries a weight, and the total is the weighted average of the dimension scores. Consider an invented support role with four dimensions: problem solving weighted 40%, communication 30%, product understanding 20% and motivation 10%. A candidate who scores 7, 6, 5 and 8 on those gets 0.4 × 7 + 0.3 × 6 + 0.2 × 5 + 0.1 × 8, which is 2.8 + 1.8 + 1.0 + 0.8, or 6.4 overall.

Points total (out of 10). On points-based jobs, each part of the screen has its own budget: 3 points each for the first two core questions, 2 for the third, up to 2 for communication, and an optional domain bonus of up to 1, with the total capped at 10. The evaluator scores each part within its budget, and the parts are simply added up. The full layout, including why the budgets differ, is in where all 10 points go.

Either way, the arithmetic is plain. No hidden model blends the dimensions into the total; the total is a sum or an average you could work out on a calculator from the numbers on screen.

Step 8: The total meets the job’s bands

The last step turns the total into a verdict tier. On the default weighted path, 8.0 and above is Strong Go, 6.0 is Go and 4.5 is On Hold, with anything lower a No Go. Points-based jobs default to 9.0, 7.5 and 6.0. Any job may set its own bands, and it can switch on a communication floor that turns a low communication score into a No Go whatever the total says.

So the invented support candidate at 6.4 lands in Go on default bands. Had problem solving come in at 5 instead of 7, the total would drop to 5.6 and the result would be On Hold. One dimension, two points, a different tier. That is why the full rule order, laid out in the four tiers and the checks behind them, is worth knowing before you defend a decision to anyone.

What the AI interview score deliberately leaves out

A score is easier to trust when you know its boundaries. Four things are not part of it:

  • The resume. The resume match score is calculated before the call from the application. It is a separate number, and why the resume number stays separate explains the reasoning.
  • AI-assist signals. Indicators such as tab switches or answers that sound read from a script are shown to HR as flags to verify in the recording. They never change a score or a verdict.
  • The recruiter’s own judgement. A recruiter can record their own decision and their own dimension scores beside the AI’s. Those are stored separately and never overwrite the evaluator’s numbers.
  • Anything outside the conversation. On standard jobs the evaluator receives the transcript, the role and the job’s scoring settings, and is told to base every score only on evidence from the transcript. It is not given the resume, a photo or demographic details.

None of this replaces a recruiter. It changes where their time goes. A large 2026 field experiment on AI-led interviews, with more than 50,000 applicants, reported that the AI interviewer stuck to the planned structure more reliably than human recruiters did. Structure is what makes a written scale usable: the same scale, applied to every transcript at any volume, without the fatigue that sets in by a recruiter’s eightieth call of the day.

Trace one score end to end today

Open the HR dashboard, choose one completed interview and follow it backwards. Start with the total, then the dimension scores, then the evaluation note and flags, then the transcript itself. If every number leads you to something the candidate actually said, the score is doing its job. If one does not, that is the row to open the recording on.

We have run more than 1,099 pilot interviews, and the scores that worried customers were rarely the high or low ones. They were the ones nobody could trace. If you would like to walk through a real scorecard together before your next drive, get in touch.

Frequently asked questions

Does the HireQwik voice agent score candidates while the interview is happening?

No. The agent's job during the call is to run the conversation. Scoring happens afterwards, when a separate evaluation step reads the complete transcript against the job's dimensions and writes a score for each one, along with notes and red and green flags.

What does a score of 0 mean on a HireQwik interview dimension?

Zero means the evaluator did not have enough evidence to judge that dimension, for example because the call ended before the topic came up. It is a marker for missing evidence, and the evaluator is told to use it instead of guessing a middle score.

Is the resume match score part of the AI interview score?

No. The resume score is calculated before the call from the application, and the interview score is calculated after it from the conversation. HireQwik keeps them as two separate numbers so you can see when a strong resume and a weak call disagree, or the reverse.

See your own candidates screened

Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.