AI recruiting agency · Updated

How we measure shortlist accuracy: 84% and 72% agreement with client engineers, explained

Short answer

We compare our ranking with your engineers' own ratings of the same candidates. On live searches our ranking agreed with the client 84% of the time on a Java role and 72% on a data science role (89% and 84% on held-out samples). The data science search has a Cohen's kappa of 0.80. This page explains how each number is produced.

The method

Your engineering leaders label candidates one by one. Each verdict updates the ranking rules for your team. We then check our ranking against those verdicts on the same people: the share of candidates where both land on the same side of the line is the raw agreement. In one search the client labeled 71 profiles.

The published results

Accuracy on two live client searches
Java and Kotlin backend searchData science (credit risk) search
Raw agreement, live data84%72%
Raw agreement, held-out sample89%84%
Cohen's kappanot published yet0.80

What each number means

Raw agreement

The share of candidates where our ranking and your team's verdict land on the same side of the line. Simple, and easy to inflate, so we never report it alone.

Held-out samples

A slice of labeled profiles kept out of calibration. Scoring on those unseen profiles shows whether the ranking learned your bar or memorized examples.

Cohen's kappa

Agreement corrected for chance. Reviewers who approve most profiles agree often by luck, and kappa strips that out. 0.80 on the data science search is strong inter-rater reliability.

The funnel behind the numbers

Java and Kotlin search: 1,294 profiles gathered, 871 unique people, 110 cleared every hard requirement, 40 ranked and delivered. 11 cleared outright and 29 came with a single item to verify.

Credit-risk data science search: 1,191 profiles gathered, 812 unique people, 53 cleared every hard requirement (filters included 138 people without a credit-risk background and 128 overqualified candidates), 40 ranked and delivered.

In plain terms and what it does not show

Our ranking agrees with an experienced hiring manager about as well as two senior engineers agree with each other. The figures come from two client searches, so they show the method works, not that it will hit the same number on your role. That is what the $499 Calibration Audit is for: one stalled vacancy, scored against your bar before a full search. To run the same checks on any vendor, see how to check a candidate shortlist before you interview. See how the service works or what it costs against commission fees.

FAQ

What does 84% agreement mean?

On a live Java and Kotlin search, our ranking and the client's engineers put the same candidates on the same side of the line 84% of the time.

What is a held-out sample?

A slice of labeled profiles kept out of calibration. Scoring on those unseen profiles shows whether the ranking learned your bar or memorized examples.

What is Cohen's kappa in hiring?

Agreement between two raters corrected for chance. Ranked.I's data science search scored 0.80, which is strong inter-rater reliability.

Why publish accuracy?

Most sourcing services do not show how often their picks match a hiring team's judgment. Publishing the number lets you check the claim before you buy.

Ready to see if we're calibrated to your bar?

Share the role details. We’ll review your stack and get back within one business day.

Is the hiring budget already approved?