Short answer
You can verify a candidate shortlist before your first interview using five fast checks: checking the rationale for every name, reviewing named open questions, running a blind evaluation on 20 profiles, measuring holdout accuracy, and testing the system on a previously filled role. Each check takes under an hour and works whether the list came from an agency, an AI tool, or your internal recruiter.
Why candidate shortlisting needs a new standard
The volume of applications is no longer a useful filter. According to Ashby's Talent Trends data, tech roles draw 610 inbound applications per hire in the Americas and 254 in EMEA. Gem's 2026 tech hiring guide puts the engineering application-to-hire rate at 0.2%, while sourced candidates in engineering and data science are hired 8 times as often. When inbound volume is overwhelming, the real differentiator is the quality of hire and the precision of the candidate shortlist, yet few teams measure it.
Five checks to evaluate any candidate shortlist
1. The rationale for every name (and every rejection)
Before agreeing to interviews, ask your provider a direct question: "Why is this person on the list and why were the others filtered out?"
A credible sourcing workflow keeps a clear audit trail, the same job a candidate scorecard does inside your own team. For instance, out of 1,191 processed profiles, 138 might be dropped for lacking specific credit risk experience, while 128 are filtered out because they are overqualified for the target role.
- What to ask: "Why is this person on the list and why were the others removed?"
- Good answer: Granular, objective rejection criteria aligned with your technical requirements.
- Our benchmark: 1,191 profiles reviewed, with exact drop-off logs (138 removed for missing credit risk background, 128 set aside as overqualified). See both searches in detail.
2. Named open questions per candidate
A weak list presents every profile as a flawless match. A rigorous shortlist flags unknowns that need validation during the conversation.
- What to ask: "What couldn't you verify about this candidate from public data?"
- Good answer: Explicitly highlighted verification gaps (for example stack scale, leadership exposure, or a recent change of tools).
- Our benchmark: Out of 40 profiles on a Java shortlist, 11 passed cleanly, while 29 came with a specific open question for the interview.
3. Blind evaluation on a sample subset
To test consistency, take 20 profiles, strip away all ranking scores and vendor notes, and have your internal engineers grade them blindly. Then compare the results against the vendor's ranking.
Be cautious: a raw agreement percentage is easy to inflate. A reviewer who approves almost everyone will match by chance. This is why a chance-corrected measure like Cohen's kappa is needed to separate real inter-rater reliability from luck (more in how we measure shortlist accuracy).
- What to ask: Have your internal engineers grade 20 profiles blindly, then compare with the vendor's ranking.
- Good answer: A high agreement rate, reported together with a chance-corrected measure.
- Our benchmark: 84% agreement on a Java search, 72% on a data science search.
4. Holdout sample and kappa accuracy
Ask your provider how their matching performs on data it has not been tuned on. If it only works on the examples it was calibrated with, it will fail on your live open roles.
- What to ask: "What is your accuracy on profiles your system wasn't trained or tuned on?"
- Good answer: Holdout accuracy reported with a chance-corrected measure such as Cohen's kappa.
- Our benchmark: 89% and 84% holdout accuracy; Cohen's kappa of 0.80 on the data science search (method).
5. Retrospective test on a filled vacancy
A test recruiters themselves recommend: give the vendor a role you have already closed and see whether the person you hired ranks near the top of their shortlist.
- What to ask: Provide a role you have already closed and check whether your hired candidate lands at the top of their list.
- Good answer: The successful hire appears in the top tier of the generated candidates.
What metrics prove quality vs. what is just noise
| What a vendor shows you | What it actually proves |
|---|---|
| Database size and total profile counts | Nothing about your specific bar or requirements |
| "93% match" or "100% success" claims without a methodology | Nothing |
| Testimonials and customer reviews | That another company was satisfied |
| Agreement with your internal team on a holdout sample | That the ranking has learned your exact hiring bar |
What these checks do not show
These evaluations measure the precision of the filtering and ranking process, but they cannot predict whether a candidate will accept an offer or how they will perform a year down the line. We do not track hiring outcomes after the shortlist is delivered, as too many internal variables influence offer acceptance and retention.
Calibration Audit
If you want to run these checks with us, try the Calibration Audit for $499. We evaluate one stalled vacancy: we build your grading matrix and deliver a test batch against your bar. The fee is fully credited toward a full search if you choose to proceed.
Ready to test your current pipeline? Check the pricing details on the main page or see how we work as an AI recruiting agency.