Short answer
A ranked shortlist starts as hundreds of raw profiles and ends as up to 40 named candidates, each with a reason. In the two searches below, between 299 and 1,505 profiles were gathered per role, 53 to 137 people cleared every hard requirement, and agreement with the client's own scoring was measured at 84% and 72%. Client names are withheld.
Beyond the hype of volume
Most engineering leaders think their hiring problem is a lack of candidates. Ashby's data says otherwise: tech roles draw 610 inbound applications per hire in the Americas and 254 in EMEA. The real bottleneck is not volume. The bottleneck is the noise floor.
When a team spends hours sorting through profiles that miss the mark on core requirements or seniority, the hiring pipeline stalls. A Ranked.I search collects hundreds of raw profiles, filters them against hard requirements, and delivers a ranked shortlist of up to 40 names with clear grading logic.
Below is a breakdown of two real searches, showing how the numbers move from raw data to delivered names.
Case 01: Consumer fintech, two remote roles
A consumer fintech company with two remote engineering roles, both requiring working-hours overlap with the team and two working languages.
Role A: Backend Engineer, Java and Kotlin
| Stage | Result |
|---|---|
| Profiles gathered from sources | 1,294 |
| Unique people after deduplication | 871 |
| Cleared every hard requirement | 110 |
| Delivered to client | 40 |
| Agreement with client scoring | 84% on live data, 89% on held-out profiles |
Out of 871 unique engineers, 761 were cut at the hard-requirement stage. 110 cleared every hard requirement, and from these we ranked the top 40 for the final shortlist.
Of the 40 delivered names, 11 cleared outright. Each of the remaining 29 came with a single named item to verify on the first screening call.
The client labeled 71 candidates one by one. Every verdict updated our ranking rules, so rejected candidates never resurface for this team.
Role B: Data Scientist, Credit Risk
| Stage | Result |
|---|---|
| Profiles gathered from sources | 1,191 |
| Unique people after deduplication | 812 |
| Cleared every hard requirement | 53 |
| Delivered to client | 40 |
| Agreement with client scoring | 72% on live data, 84% on held-out profiles |
| Cohen's kappa | 0.80 |
Credit risk modelling is an unforgiving domain. A general data science background is not enough.
138 candidates were filtered out because they lacked a credit-risk background. 128 were set aside as overqualified for the role. That left a tight pool of 53, from which 40 were ranked and delivered.
A Cohen's kappa of 0.80 shows that our ranking agreed with the client's engineers well beyond what chance would produce.
Case 02: AI company in the EU, three technical roles for one team
Three openings for one in-house team, from individual contributor to team lead, each shaped by a different constraint.
Role A: ML Team Lead, Speech and Language Models
| Stage | Result |
|---|---|
| Profiles gathered | 1,505 |
| Unique people | 853 |
| Cleared every hard requirement | 61 |
| Delivered | 40 |
The role came with a remote option, so location was ranked, not gated: a strong candidate in another country stayed on the list and simply scored lower on location. Among the candidates who cleared the bar, hands-on inference optimization broke the ties.
Role B: Senior ML Engineer, One National Language Market
| Stage | Result |
|---|---|
| Profiles gathered | 422 |
| Unique people | 416 |
| Cleared every hard requirement | 74 |
| Delivered | 40 |
Restricting the search to one national language market shrinks the talent pool sharply: 422 profiles against 1,505 for the team lead role. With a narrow market, the ranking shifted from breadth to depth of LLM work.
Role C: Senior DevOps and Infrastructure Engineer, Hybrid
| Stage | Result |
|---|---|
| Profiles gathered | 299 |
| Unique people | 291 |
| Cleared every hard requirement | 137 |
| Delivered | All 137, as a full market map |
This role shows why a fixed top-40 sometimes has to adapt to the market. The on-site requirement changed the job: the whole reachable local market that met the hard requirements was 137 people.
Instead of cutting to 40, we delivered all 137, so the team could see the local talent ceiling before setting salaries.
How to read our numbers
Raw agreement
The share of candidates where our ranking and the client's verdict land on the same side of the line. It is easy to inflate if requirements are loose, which is why we never rely on it alone.
Held-out profiles
Labeled profiles kept out of calibration. They show whether the ranking learned the client's bar or memorized examples.
Cohen's kappa
Agreement corrected for chance. A score of 0.80 indicates strong agreement between two independent raters. The full method is in how we measure shortlist accuracy.
What these numbers do not show
These metrics measure how closely our ranking matches an engineering team's own judgment. They do not track who accepted an offer or how a hire performed twelve months later. We focus on sourcing and grading precision; interviews and hiring stay in your hands. To run the same checks on any vendor, see how to check a candidate shortlist before you interview.
Ready to see how this applies to your stack?
Send us one stalled vacancy. We will build your custom grading matrix and deliver a test batch evaluated against your bar for $499. The fee is fully credited toward a full search if you decide to move forward. See pricing.