AI recruiting agency · Updated

What a ranked shortlist looks like: two real searches, inside the numbers

Short answer

A ranked shortlist starts as hundreds of raw profiles and ends as up to 40 named candidates, each with a reason. In the two searches below, between 299 and 1,505 profiles were gathered per role, 53 to 137 people cleared every hard requirement, and agreement with the client's own scoring was measured at 84% and 72%. Client names are withheld.

Beyond the hype of volume

Most engineering leaders think their hiring problem is a lack of candidates. Ashby's data says otherwise: tech roles draw 610 inbound applications per hire in the Americas and 254 in EMEA. The real bottleneck is not volume. The bottleneck is the noise floor.

When a team spends hours sorting through profiles that miss the mark on core requirements or seniority, the hiring pipeline stalls. A Ranked.I search collects hundreds of raw profiles, filters them against hard requirements, and delivers a ranked shortlist of up to 40 names with clear grading logic.

Below is a breakdown of two real searches, showing how the numbers move from raw data to delivered names.

Case 01: Consumer fintech, two remote roles

A consumer fintech company with two remote engineering roles, both requiring working-hours overlap with the team and two working languages.

Role A: Backend Engineer, Java and Kotlin

Role A: Backend Engineer, Java and Kotlin
StageResult
Profiles gathered from sources1,294
Unique people after deduplication871
Cleared every hard requirement110
Delivered to client40
Agreement with client scoring84% on live data, 89% on held-out profiles

Out of 871 unique engineers, 761 were cut at the hard-requirement stage. 110 cleared every hard requirement, and from these we ranked the top 40 for the final shortlist.

Of the 40 delivered names, 11 cleared outright. Each of the remaining 29 came with a single named item to verify on the first screening call.

The client labeled 71 candidates one by one. Every verdict updated our ranking rules, so rejected candidates never resurface for this team.

Role B: Data Scientist, Credit Risk

Role B: Data Scientist, Credit Risk
StageResult
Profiles gathered from sources1,191
Unique people after deduplication812
Cleared every hard requirement53
Delivered to client40
Agreement with client scoring72% on live data, 84% on held-out profiles
Cohen's kappa0.80

Credit risk modelling is an unforgiving domain. A general data science background is not enough.

138 candidates were filtered out because they lacked a credit-risk background. 128 were set aside as overqualified for the role. That left a tight pool of 53, from which 40 were ranked and delivered.

A Cohen's kappa of 0.80 shows that our ranking agreed with the client's engineers well beyond what chance would produce.

Case 02: AI company in the EU, three technical roles for one team

Three openings for one in-house team, from individual contributor to team lead, each shaped by a different constraint.

Role A: ML Team Lead, Speech and Language Models

Role A: ML Team Lead, Speech and Language Models
StageResult
Profiles gathered1,505
Unique people853
Cleared every hard requirement61
Delivered40

The role came with a remote option, so location was ranked, not gated: a strong candidate in another country stayed on the list and simply scored lower on location. Among the candidates who cleared the bar, hands-on inference optimization broke the ties.

Role B: Senior ML Engineer, One National Language Market

Role B: Senior ML Engineer, One National Language Market
StageResult
Profiles gathered422
Unique people416
Cleared every hard requirement74
Delivered40

Restricting the search to one national language market shrinks the talent pool sharply: 422 profiles against 1,505 for the team lead role. With a narrow market, the ranking shifted from breadth to depth of LLM work.

Role C: Senior DevOps and Infrastructure Engineer, Hybrid

Role C: Senior DevOps and Infrastructure Engineer, Hybrid
StageResult
Profiles gathered299
Unique people291
Cleared every hard requirement137
DeliveredAll 137, as a full market map

This role shows why a fixed top-40 sometimes has to adapt to the market. The on-site requirement changed the job: the whole reachable local market that met the hard requirements was 137 people.

Instead of cutting to 40, we delivered all 137, so the team could see the local talent ceiling before setting salaries.

How to read our numbers

Raw agreement

The share of candidates where our ranking and the client's verdict land on the same side of the line. It is easy to inflate if requirements are loose, which is why we never rely on it alone.

Held-out profiles

Labeled profiles kept out of calibration. They show whether the ranking learned the client's bar or memorized examples.

Cohen's kappa

Agreement corrected for chance. A score of 0.80 indicates strong agreement between two independent raters. The full method is in how we measure shortlist accuracy.

What these numbers do not show

These metrics measure how closely our ranking matches an engineering team's own judgment. They do not track who accepted an offer or how a hire performed twelve months later. We focus on sourcing and grading precision; interviews and hiring stay in your hands. To run the same checks on any vendor, see how to check a candidate shortlist before you interview.

Ready to see how this applies to your stack?

Send us one stalled vacancy. We will build your custom grading matrix and deliver a test batch evaluated against your bar for $499. The fee is fully credited toward a full search if you decide to move forward. See pricing.

FAQ

Why 40 names and not 5?

A ranked list of 40 lets your team see the actual shape of the market and the reason behind each candidate's position. Five names show you the winners and hide who was cut and why.

What does cleared every hard requirement mean?

The candidate passes every non-negotiable parameter you set: tech stack, seniority, location, time-zone overlap and language proficiency.

How fast is the turnaround?

5 to 7 days for a new role, 24 to 48 hours for a repeat search.

What is the pricing model?

A fixed fee: from $4,500 per new role ($6,000 for ML and LLM roles) and $1,500 to $2,000 for a repeat role. No commission on salary.

Can we test the methodology first?

Yes, through the Calibration Audit: one stalled vacancy for $499, fully credited toward a full search.

Ready to see if we're calibrated to your bar?

Share the role details. We’ll review your stack and get back within one business day.

Is the hiring budget already approved?