Welcome back

Sign in to your screening dashboard

New to HireQwik? Book a demo

Book a demo

Tell us a little about your hiring. We'll reply within one business day.

Prefer email? interview@hireqwik.in
ai-screeninghr-techrecruitingvoice-ai

What Your Own AI Demo Scorecard Actually Tells You

HireQwik August 31, 2026 10 min read

The first time I read one of our own scorecards as a buyer rather than as the person who built the thing, I got a Strong Go and learned almost nothing from it. The number was flattering and useless. What taught me something was the third item down, where the evaluation quoted a sentence I had rushed through and marked it as an unsupported claim. It was right. I had waved at an answer instead of giving one.

That is the whole skill of reading an AI interview scorecard, and it is not the skill most buyers bring to it. If you have taken a live vendor demo, or you are about to take the interview on our own demo page, the document that arrives afterwards is the most honest thing you will see in the entire sales process. Most people skim the verdict and stop.

Here is how to read the rest.

Start at the bottom, not the top

The verdict tier is at the top of every scorecard in this category, ours included: Strong Go, Go, Hold, No Go, or whatever four labels the vendor has chosen. It is also the least informative line on the page, because a tier is a compression of everything underneath it.

Read from the bottom item up. Find the lowest scored dimension and ask a single question: can I tell exactly what produced this number?

There are only two possible answers. Either the scorecard shows you the specific thing the candidate said or did not say, or it gives you an adjective. “Communication: 6/10, could be clearer” is an adjective. “Communication: 6/10, gave a two word answer to the scope question and did not elaborate when prompted” is evidence.

This distinction decides whether the document is usable six weeks later, when a rejected candidate emails asking why, or when your own hiring manager disagrees with a shortlist and wants to know the basis.

The three sections that carry actual information

Strip the branding off any scorecard in this market and you are looking for three things.

Per question evaluation with the candidate’s own words in it. Every score should point back at a specific moment in the conversation. On our scorecards, the per question evaluation quotes your answers directly, which means a recruiter reviewing a shortlist is reading the candidate’s sentence and the machine’s read of it side by side rather than trusting a number in isolation. When you look at a vendor’s scorecard, count how many scores have a quoted basis. If the answer is zero, you have a summary, not an evaluation.

A communication read that names what it measured. This is where most tools in the category are weaker than they appear. A great many read communication off the transcript and nothing else, so a candidate whose words looked tidy after speech to text is indistinguishable from one who actually spoke well. Our own layer listens to the recording instead, and what it measures there is speaking rate, how often the speaker hesitates or fills a gap, how clearly words are formed, and a CEFR band for overall fluency. Whether or not you buy from us, ask the vendor which of those come from audio and which from text. The answer is usually revealing and rarely volunteered. We went through the mechanics of that gap in the transcript only blind spot.

A verdict that maps to an action. Four bands are useful only if your team knows what each one triggers. Strong Go means schedule. No Go means send the rejection. The two middle bands are where the actual work is, and a scorecard that dumps 60% of a cohort into the middle has not saved anybody any time.

The test that tells you whether it is real

Here is the exercise I would run on any vendor’s scorecard, ours included, before signing anything.

Take the lowest scored item. Write, in two sentences, the message you would send to a candidate who was rejected on that basis and wrote back asking for a reason. Do it using only what is on the page in front of you.

If you can write those two sentences, the scorecard is a defensible artifact. If you find yourself inventing a reason, or falling back on machine language of the sort a person cannot argue with, that scorecard will fail the first time a real candidate pushes back. We have written the full 72 hour playbook for exactly that scenario in handling bias complaints from rejected candidates, and the single thing that makes that playbook workable is a scorecard with quoted evidence in it.

This test has a second benefit. It stops you from buying on the strength of a good looking interface. Interfaces are easy. An evidence trail that holds up under a complaint is not.

There is a regulatory version of the same test arriving from several directions at once. Under New York City’s Local Law 144, an automated employment decision tool already needs an annual independent bias audit, a published summary of the result, and notice to candidates ahead of use. Law firms tracking enforcement have flagged rising exposure for employers rather than a settled compliance regime. Academic work auditing compliance with that law found that a great many covered employers publish nothing verifiable at all. Why that bears on a scorecard sitting in your inbox today: nothing can be audited afterwards unless it recorded its basis at the time. A tool that produces adjectives cannot be audited later, whatever the vendor intends to do about it when the rules reach you.

Why your own result is the weakest data point on the page

You will probably score well. Take that seriously as a warning rather than as reassurance.

You chose the moment, the topic is your own daily work, nothing is riding on it, and no part of you is nervous. You are the easiest input the system will ever see. A fresher taking the same call at 9 PM after four rejections, on a patchy hostel connection, in their third language, is the input that matters, and nothing in your own transcript tells you how the system handles that person.

So do not read your score as a measure of accuracy. Read it as a measure of legibility. The question your own scorecard can genuinely answer is: when this system makes a judgment, can I follow the reasoning? That is worth knowing, and a single self administered interview cannot establish much else.

If you want a read on accuracy, you need something else entirely, which is a rejection audit on a real cohort. We described how to run one in the No Go audit.

Ask what it says about a candidate it could not evaluate

Here is a question almost nobody puts to a vendor, and it exposes more than any feature demo.

What does the scorecard look like when the interview did not really happen? The line dropped after two minutes. The candidate’s mic picked up a hostel corridor. They answered every question with three words and nothing the follow ups asked drew anything more out of them.

At campus volume this is not an edge case, it is a daily occurrence, and there are only two possible designs. Either the system scores what little it got and hands you a number, or it declines and marks the attempt as unevaluable so a human decides what happens next.

The first is worse than useless, because a low score from a broken call is indistinguishable in a shortlist from a low score from a real one. That candidate gets auto-rejected for their network connection. Ask to see a real example of each, not a description of the policy, and check whether the reason the evaluation is thin is stated anywhere on the page. If a vendor has never been asked this, you will hear it in the pause.

Four failure patterns in vendor scorecards

Across the scorecards we have seen from this category, four patterns come up repeatedly and each one is a reason to ask a harder question.

Scores with no denominators. A dimension scored 78 with no stated scale, no rubric, and no indication of what an 85 would have required. This is common and it is not a formatting issue. It usually means the rubric is a single generic template rather than one derived per job description, which is a distinction we unpacked in the per JD rubric.

A summary paragraph doing the work of the evidence. Fluent prose describing the candidate, with no quotes. This reads well and cannot be audited. If the paragraph could have been written from the resume alone, the interview added nothing.

Every candidate landing in the middle. If a vendor shows you sample scorecards and they all cluster in the same band, that is not caution, it is a scoring engine that cannot separate. Relevance blind scoring squeezes almost every candidate into one narrow middle band, which we took apart in why AI resume scores cluster.

A confidence or accuracy percentage attached to the model itself. Any vendor telling you their screening AI is “95% accurate” should be asked for the held out evaluation set, the labelling protocol, and the date. We do not publish an accuracy figure for our own scoring because calibration is ongoing and we do not have a defensible number yet. Neither, in all likelihood, does the vendor quoting you one.

The second scorecard is where it gets interesting

One scorecard tells you whether the reasoning is legible. Two tell you whether it is consistent, and consistency is the property that actually determines whether a shortlist is rankable.

Take the interview a second time, a day later, and answer roughly as well as you did the first time. Not identically, which is impossible and would test the wrong thing. Comparably. Then put the two scorecards side by side and look at three things.

Did the same dimension move a lot? If your communication score swung fifteen points across two similar performances, the engine has a variance problem, and that variance is being applied to every candidate in a cohort. On a drive of a thousand people, an engine that noisy will mis-sort a meaningful slice of them purely by chance. You will never see it, because you have nothing to compare each candidate against.

Did the evidence change or just the number? Compare the quoted reasons, not the scores. If the same weakness in your answers is named both times, the engine found something real. If two different explanations arrive for two similar performances, at least one of them was constructed after the fact to justify a number.

Did the verdict band hold? Scores drifting a few points inside a band is fine and expected. A verdict crossing from Go to Hold on two comparable conversations is not, because the band is the thing your team acts on. Anything that changes the action is a different class of problem from anything that only changes a number.

If you cannot take it twice, the cheaper version is to compare your scorecard against a colleague’s from the same demo. It is a weaker test, because you are now comparing two different people rather than isolating the engine, but a wild divergence between two similarly capable colleagues is still worth a question on the next call.

Almost nobody runs this test, which is why almost nobody discovers the variance until a hiring manager rejects a shortlist and cannot say why.

The take

A scorecard is not a report card. It is a claim, and every score on it is a claim that either carries its evidence or does not.

Read from the bottom. Find the weakest score. Try to write the rejection email from it. If you cannot, the number at the top was decoration, however good it looked. That test costs you four minutes and it is worth more than the entire demo call that preceded it.

If you have not taken one yet, the live interview on our demo page produces a real scorecard in your inbox, and you are welcome to run this test against it.

Frequently asked questions

What should an AI interview scorecard contain to be useful?

Three things: a per question score with the candidate's own words quoted as the reason, a communication read that names what was measured, and a verdict tier that maps to a decision. A score without a quoted basis is an opinion with a number attached to it.

Why does a high score on my own demo interview mean very little?

Because you are a hiring professional answering questions about hiring, on your own schedule, with no stakes. You are the easiest possible input. What the result tells you is whether the reasoning is legible, not whether the model is accurate.

Can an AI scorecard justify a rejection to a candidate who asks why?

Only if the evidence sits inside it. Test this by finding the lowest scored item and writing the two sentence explanation you would send. If you cannot write it from the scorecard alone, that scorecard will not survive a real complaint.

See your own candidates screened

Book a 30-minute demo. Bring a live JD and we'll screen against it, then start with a pilot on your own candidates before committing to anything.