AI recruiting agency · Updated

How to check a candidate shortlist before you interview anyone

Short answer

You can verify a candidate shortlist before your first interview using five fast checks: checking the rationale for every name, reviewing named open questions, running a blind evaluation on 20 profiles, measuring holdout accuracy, and testing the system on a previously filled role. Each check takes under an hour and works whether the list came from an agency, an AI tool, or your internal recruiter.

Why candidate shortlisting needs a new standard

The volume of applications is no longer a useful filter. According to Ashby's Talent Trends data, tech roles draw 610 inbound applications per hire in the Americas and 254 in EMEA. Gem's 2026 tech hiring guide puts the engineering application-to-hire rate at 0.2%, while sourced candidates in engineering and data science are hired 8 times as often. When inbound volume is overwhelming, the real differentiator is the quality of hire and the precision of the candidate shortlist, yet few teams measure it.

Five checks to evaluate any candidate shortlist

1. The rationale for every name (and every rejection)

Before agreeing to interviews, ask your provider a direct question: "Why is this person on the list and why were the others filtered out?"

A credible sourcing workflow keeps a clear audit trail, the same job a candidate scorecard does inside your own team. For instance, out of 1,191 processed profiles, 138 might be dropped for lacking specific credit risk experience, while 128 are filtered out because they are overqualified for the target role.

  • What to ask: "Why is this person on the list and why were the others removed?"
  • Good answer: Granular, objective rejection criteria aligned with your technical requirements.
  • Our benchmark: 1,191 profiles reviewed, with exact drop-off logs (138 removed for missing credit risk background, 128 set aside as overqualified). See both searches in detail.

2. Named open questions per candidate

A weak list presents every profile as a flawless match. A rigorous shortlist flags unknowns that need validation during the conversation.

  • What to ask: "What couldn't you verify about this candidate from public data?"
  • Good answer: Explicitly highlighted verification gaps (for example stack scale, leadership exposure, or a recent change of tools).
  • Our benchmark: Out of 40 profiles on a Java shortlist, 11 passed cleanly, while 29 came with a specific open question for the interview.

3. Blind evaluation on a sample subset

To test consistency, take 20 profiles, strip away all ranking scores and vendor notes, and have your internal engineers grade them blindly. Then compare the results against the vendor's ranking.

Be cautious: a raw agreement percentage is easy to inflate. A reviewer who approves almost everyone will match by chance. This is why a chance-corrected measure like Cohen's kappa is needed to separate real inter-rater reliability from luck (more in how we measure shortlist accuracy).

  • What to ask: Have your internal engineers grade 20 profiles blindly, then compare with the vendor's ranking.
  • Good answer: A high agreement rate, reported together with a chance-corrected measure.
  • Our benchmark: 84% agreement on a Java search, 72% on a data science search.

4. Holdout sample and kappa accuracy

Ask your provider how their matching performs on data it has not been tuned on. If it only works on the examples it was calibrated with, it will fail on your live open roles.

  • What to ask: "What is your accuracy on profiles your system wasn't trained or tuned on?"
  • Good answer: Holdout accuracy reported with a chance-corrected measure such as Cohen's kappa.
  • Our benchmark: 89% and 84% holdout accuracy; Cohen's kappa of 0.80 on the data science search (method).

5. Retrospective test on a filled vacancy

A test recruiters themselves recommend: give the vendor a role you have already closed and see whether the person you hired ranks near the top of their shortlist.

  • What to ask: Provide a role you have already closed and check whether your hired candidate lands at the top of their list.
  • Good answer: The successful hire appears in the top tier of the generated candidates.

What metrics prove quality vs. what is just noise

What a vendor shows you and what it proves
What a vendor shows youWhat it actually proves
Database size and total profile countsNothing about your specific bar or requirements
"93% match" or "100% success" claims without a methodologyNothing
Testimonials and customer reviewsThat another company was satisfied
Agreement with your internal team on a holdout sampleThat the ranking has learned your exact hiring bar

What these checks do not show

These evaluations measure the precision of the filtering and ranking process, but they cannot predict whether a candidate will accept an offer or how they will perform a year down the line. We do not track hiring outcomes after the shortlist is delivered, as too many internal variables influence offer acceptance and retention.

Calibration Audit

If you want to run these checks with us, try the Calibration Audit for $499. We evaluate one stalled vacancy: we build your grading matrix and deliver a test batch against your bar. The fee is fully credited toward a full search if you choose to proceed.

Ready to test your current pipeline? Check the pricing details on the main page or see how we work as an AI recruiting agency.

FAQ

How do you evaluate a sourcing agency before you sign?

Test the process before you interview anyone: read the rejection log, look for named open questions on each candidate, grade a sample of 20 profiles blind, and ask for accuracy on a holdout sample.

How many candidates should a shortlist have?

Ours has 40. A list of five shows you the winners and hides the ranking: you cannot see who was cut or why. Forty is enough for your hiring managers to compare candidates and still small enough to read in one sitting.

What is a good agreement rate between a recruiter and a hiring manager?

We have not found a published industry norm, so we publish our own: 84% on a Java search and 72% on a data science search. A raw rate means little alone, so check it with Cohen's kappa to rule out agreement by chance.

How do you evaluate an AI sourcing tool after the demo?

Look past the interface and the database size. Ask how the matching works, test it on a holdout sample, and run a blind comparison against your own team's grading of the same profiles.

How to evaluate candidates before the first interview?

Use a structured scorecard, check that every profile says what could not be verified from public data, and make sure your sourcing partner explains why each candidate was included or rejected.

Ready to see if we're calibrated to your bar?

Share the role details. We’ll review your stack and get back within one business day.

Is the hiring budget already approved?