Waterfall Enrichment for Candidate Sourcing
Waterfall enrichment runs a list of people through several data providers in a fixed order, so each provider only works on the records the previous one missed. For candidate sourcing it matters more than for sales, because one provider typically returns about a third of a niche candidate market, and the contact detail you actually need is the personal one.
By Reyhan Khan, CEO & GTM Systems Architect at RecruiterGTM. Four years building recruitment systems, eight in GTM overall, and 650+ recruitment agencies audited in the last three years.
What the engine does with a job spec
You hand Claude the job spec or a CV of someone who would be perfect. It reads it, works your own database first, then goes outside for everything your database did not have, ranks what comes back against the actual role, and gives you a scored shortlist with the proof attached.
You give it
A job spec, or a CV of someone perfect for the role
Claude reads it
Turns the spec into the hard requirements a candidate must meet
Layer 0 · first, always
Your own ATS
Who you already hold, who you have already contacted, and who replied and was never followed up
then, only for what your database did not have
Outside sources, in order
Google X-ray + Exa, three independent nets
Reads what each person wrote about themselves
Prospeo, then Apollo, then SalesQL for the personal email
Each one only runs on what the one before it could not find.
Ranked
Scored against the actual role
Every candidate carries a word-for-word quote from their own profile proving why they qualify. No quote, no row.
You get
A scored shortlist, with the proof attached
You read the top of it and approve. That is the only part that still takes a person.
Key takeaways
- A waterfall is a fixed order: free providers run first, and paid ones only ever see the misses.
- Your own ATS is layer 0. On one real role, 63% of the qualified market had never been seen by the agency that owned the database.
- A candidate waterfall has a step in the middle that a sales waterfall does not: reading what each person wrote about themselves, before you pay for any contact details.
- Work emails are the wrong channel for passive candidates. The personal email is a separate layer and the reason the whole thing earns its place.
- Start with three layers if five feels like too much: ATS, then Prospeo, then SalesQL.
- Every row carries a verbatim quote proving why the person qualifies. No quote, no row.
Five words you will meet in this guide
- Enrichment
- Taking a name or a profile link and filling in the rest: current employer, job title, email address, phone number. A data provider does it for you, usually per record.
- Provider
- A company that sells that data. Apollo, Prospeo and SalesQL are all providers. Each holds a different slice of the market, and none of them tells you which slice it is missing.
- Scraping
- Reading what is publicly on a web page and saving it as data you can sort. Here it means reading the text people have written on their own public profiles.
- X-ray search
- A normal Google search aimed at one website only. Here it means using Google to read inside LinkedIn profiles, which surfaces people LinkedIn's own filters miss.
- Verified email
- An address checked against the receiving mail server to confirm somebody is actually there. Many providers guess addresses from a name and a company; a verified one has been tested.
What is waterfall enrichment?
Waterfall enrichment is the practice of passing each record through several data providers in sequence. Provider one runs on the full list, provider two runs only on what provider one failed to find, and so on down the chain. The result is higher coverage at lower cost, because the expensive sources only ever see what is left over.
The technique comes from B2B sales, where it is used to find buyers. Every major explainer on the subject, from Clay to Apollo to ZoomInfo, teaches it that way. Almost nobody writes about it for candidate sourcing, which is a problem, because the requirements are genuinely different and the differences are the parts that decide whether a shortlist is usable.
Why does one data source only find about a third of your candidates?
Because every provider builds its database from overlapping but incomplete sources, and none of them publishes where the gaps are. In a specialist niche, a single provider commonly returns coverage in the region of a third of the real market. You cannot tell from the inside, because the rows that come back look complete.
A list of 200 enriched candidates from one source looks like 200 candidates. It is actually 200 out of a market you never measured, and the 400 you did not see include the ones your competitors have not contacted either.
I tested this on a UK project controls agency. They had 18,195 candidate records on file, built over three and a half years. When I mapped the live market for the specific skill they needed to hire and checked it against their database, 14% of that market was already visible to them and 63% had never been seen by anyone. Their previous sourcing engine had contacted 5,056 people and surfaced none of the qualified ones, because it filtered on job titles and the thing they needed was a skill written in the body of a profile.
What makes a candidate waterfall different from a sales waterfall?
Three things. The personal email matters more than the work email. A qualifying step sits in the middle of the chain, where you read profiles before paying for contact data. And your own database is the first place you look. A sales waterfall optimises for a verified work address. A candidate waterfall that does the same thing produces a list you cannot safely use.
The work inbox is the wrong channel
When you approach a passive candidate you are asking them to consider leaving the employer whose inbox you would be writing to. That inbox is often monitored. Most enrichment providers return work addresses by default, which means a sales-shaped waterfall hands you the one channel least likely to get a reply and most likely to cause a problem for your client.
Qualification happens mid-waterfall, not at the end
In sales, the title is usually the qualifier. A VP of Finance is a VP of Finance. In recruitment, a large share of what you are asked to find is a skill, not a title: a certification, a security clearance, a named modelling tool, a specific methodology. No database holds those as a filter, because they live in free text that the person wrote about themselves.
So you read the profile text before you spend anything on contact data. On the role I ran, a title-led search returned roughly 60% people from a completely different profession who happened to share a word in their job title. A keyword-led search that read profile bodies returned a 44% genuine hit rate.
How do you build a candidate sourcing waterfall?
Six steps, in this order. Each one only runs on what the previous step could not resolve, and nothing paid runs until a candidate has passed the qualification screen. The order is the entire technique, so resist the urge to reshuffle it for convenience.
Step 1. Search your own ATS first (layer 0)
Before any external source runs, search the database you already own. It will rarely fill the role by itself, but it does two things no external source can. You learn your genuine coverage of your own market, and you stop approaching people twice, which is the fastest way to damage a candidate relationship.
It also surfaces the people who already replied positively and were never followed up. There were six of those sitting in the file on the role I ran. Six warm candidates get you to a placement faster than six hundred cold ones.
Here is the job spec I am working on: [paste it, or paste the LinkedIn profile of someone who would be perfect for the role]. Before you look anywhere outside, search my own ATS and tell me: 1. Who in there genuinely matches this spec, with the line from their record that proves it. 2. Who we have already contacted about anything, so we do not approach them twice. 3. Anyone who replied positively in the past and was never followed up. Give me the proof for each one. Do not include anybody you cannot evidence.
Step 2. Cast three independent discovery nets
Never one. Three, then remove the duplicates. The same person turns up in two or three of those searches and you want one row each. Match them on the profile address, meaning the LinkedIn URL, not the name. Names and job titles are written inconsistently across sources, so matching on those misses duplicates.
- An X-ray search through Apify. This is a normal Google search pointed at LinkedIn profiles only, written like
site:uk.linkedin.com/in "your skill term". It is the most accurate net for a skill-defined role, because Google reads the whole profile where LinkedIn's own filters do not. Use one rare keyword per query. Putting two exact phrases in quotes together usually returns nothing. And raise the number of pages you ask for, because Google only shows ten results per page. - Exa, a tool that searches by meaning instead of exact words, restricted to profiles. This catches the people who describe the work without ever using the acronym, and in my experience they are often the strongest candidates.
- Exa's find-similar feature. You point it at three or four people you have already confirmed are right, and it returns others who look like them.
The test nobody mentions: overlap between your nets should be low, around 5%. If all three keep returning the same people, you have one source wearing three hats, and your coverage is thinner than your row count suggests.
Step 3. Read what each person wrote about themselves
Now pull the text from every profile on your deduplicated list: headline, about section, every job description, skills and certifications. The tool I use is called harvestapi, which runs on Apify. You hand it a list of profile links and it returns the text from each one. Set it to return profile details without email addresses, which costs $4 per 1,000 profiles, four tenths of a penny each. You skip the email here on purpose, because the next three steps get contact details more cheaply and more accurately.
Then screen against a written list of checks and require a candidate to pass all of them. On the role I ran the five checks were: named software present on the profile, evidence of the technique described in the candidate's own words, a real programme behind the claim, current employer confirmed, and commutable to one of three cities. Roughly six in ten people with the right-sounding title failed the first two.
Here are the profiles you found: [attach the file]. Read the full text of each one, then score every person against this job spec. A candidate only passes if they meet ALL of these: 1. [your first hard requirement] 2. [your second] 3. [your third] For every person you keep, give me a word-for-word quote from their own profile that proves each requirement. No quote, no row. Tell me how many you rejected and the most common reason.
Step 4. Get work email and mobile, from the free source first
Prospeo runs first, on everyone who passed the screen. You give it a LinkedIn profile link and it returns a work email, a mobile number, and the person's full job history including how long they have been in each role. That history is useful on its own, because it tells you who is overdue a move. You will want their Pro tier at minimum, because the free allowance will not carry a real workload.
Step 5. Second pass on the misses only
Upload the leftovers to Apollo in one batch, strictly the records Prospeo could not find. Work email, sometimes a personal one, sometimes a phone number, plus current employment metadata you can use as a cross-check against what the profile claimed.
Step 6. Personal email and direct dial
SalesQL is the layer that makes a candidate waterfall different from a sales one. Steps 4 and 5 mostly return work addresses. This step returns the personal email and the direct number, which is the channel a passive candidate will actually answer.
On the test role, 33 of 50 candidates came back with a personal email and 32 of those verified. That number is step 6 doing its job, and without it the pack would have been work-addresses-only.
The five layers at a glance
| Layer | What it returns | Verdict |
|---|---|---|
| 0. Your ATS | Your true market coverage, duplicate removal, warm contacts never followed up | Never skip |
| 1. Apify + Exa | Discovery across three independent nets | Best for skill-defined roles |
| 2. Apify harvestapi | Full profile text for qualification | Best value at $0.004/profile |
| 3. Prospeo | Work email, mobile, job history with tenure | Run first, before anything paid |
| 4. Apollo | Second-pass contact data + employment cross-check | Best for closing gaps |
| 5. SalesQL | Personal email, direct dial | Best for passive candidates |
The three-layer version, if five is too many
You do not need the full chain to beat what you are doing now. Build this instead, and add the discovery and profile-reading layers once the habit is established.
| Layer | Why it is in the starter |
|---|---|
| Your ATS | Free, already paid for, and tells you what you actually hold |
| Prospeo | Returns tenure as well as contact data, and the Pro tier covers a real workload |
| SalesQL | The personal email, which is the whole point for candidate work |
What does a waterfall enrichment setup cost?
Less than most recruiters expect, because the structure is designed to keep spend off the records that will be rejected. The only fixed per-unit cost in the chain above is the profile scrape. Everything else runs on a plan you already hold, and only touches whatever is left over.
| Item | Cost | Note |
|---|---|---|
| Profile text scrape (Apify harvestapi) | $0.004 per profile | $4 per 1,000, with email addresses switched off |
| Prospeo | Pro tier or above | The free tier will not carry a real workload. Check their pricing page for the current figure |
| Apollo | Plan-dependent | Runs only on Prospeo's misses, so volume is a fraction of the list |
| SalesQL | Plan-dependent | Runs only on qualified candidates, never on rejects |
| Discovery (Exa, Google via Apify) | Pennies per run | You pay for processing time, not per person |
The single biggest cost control is the sequencing rule: never enrich a candidate you are about to reject. Qualification is step 3 and contact data starts at step 4 for exactly that reason. Agencies that enrich first and qualify second pay for the whole market instead of the shortlist.
What stops the list being wrong?
One rule and two review passes. The rule is that every candidate on the list carries a verbatim quote from their own profile proving why they qualify, and no quote means no row. The review passes are two independent checks that run before a human sees the pack, and the tool that built the list is never the one that approves it.
The quote rule does two things at once. Your consultant can verify any line against the public profile in seconds, which is what makes the list defensible in front of a client. And the automation cannot quietly invent a reason, because it has to cite the sentence.
Three failures worth borrowing from mine
- Evidence has to come from the current role to be claimed as current practice. On one run, 11 of 17 profiles I had recorded as current programme work were actually describing a previous employer. Still valid candidates. Wrong claim in the outreach.
- Constrain the outreach copy to quoted facts. Left unconstrained, 64% of my first-pass openers contained a factual error, including inventing a programme and describing a risk register as a modelling tool. A deliberately general opener does no damage. A confidently wrong one costs you the relationship.
- Never append rows to a finished list. I added 18 candidates after quality control had already passed. Those 18 were rejected at 44%, against 6% for everything that went through the full pipeline in order.
One concrete thing the second review pass catches: a shared team mailbox such as risk@theiremployer.com, sitting in the file marked verified. It passes every verification check precisely because it is a shared inbox that accepts anything sent to it. Sending to it means a poaching approach arrives in a team-wide inbox at the candidate's own employer, which costs you the client.
What did this return on a real role?
On one test role the waterfall returned 46 of 50 candidates contactable, 33 of them with a personal email, and surfaced 102 qualified candidates the agency had never seen. That was a UK project controls agency searching three cities for a skill no database holds as a filter. Here is the full output, with the method stated so the numbers can be checked.
| Measure | Result |
|---|---|
| Profiles delivered, each with a verbatim proof quote | 50 |
| Contactable on at least one channel | 46 of 50 |
| Personal email present / verified | 33 / 32 |
| Qualified candidates in their three cities they had never seen | 102 |
| Share of their own market previously visible to them | 14% |
Be clear about where the speed actually is, because this gets oversold. The longlist is the part that goes from days to minutes. Searching, reading profiles and filling in contact details across hundreds of people is a few minutes of processing. On a separate build we scraped 606 lawyers from one firm's directory, matched 89% of them to a LinkedIn profile, and verified email addresses for 604 of the 606, in under eight minutes. Profile matching and email discovery run as independent layers, which is why the email count is higher than the LinkedIn match count.
The shortlist still takes a person a few minutes to read and approve. There is no step where that disappears. The scoring gate runs automatically, then someone reads the top of the list, and what used to be a week of scrolling becomes a scored shortlist you sign off over a coffee, because the system did the scoring.
Where to go next
If you want the wider context around this, the candidate sourcing automation guide covers how the engine fits a full recruitment workflow, the sourcing automation playbook covers the campaign layer that sits on top of a waterfall, and the AI tools breakdown covers what else belongs in a recruitment stack in 2026.
Frequently asked questions
What is waterfall enrichment in recruitment?
It is running candidate records through several data providers in a fixed order, where each provider only processes the records the previous one failed to resolve. In recruitment it adds two things a sales waterfall does not have: your own ATS as the first source, and a personal-email layer at the end, because work addresses are the wrong channel for a passive candidate.
How many data sources do I actually need?
Three is a real waterfall and will beat a single provider: your ATS, then Prospeo, then SalesQL. Five is better for specialist roles, adding a discovery layer and a profile-reading layer. The order matters more than the count, because the order is what keeps paid providers off records that are going to be rejected.
Why not just use one tool that claims full coverage?
Because no provider publishes where its gaps are, and in a specialist niche a single source commonly returns around a third of the real market. The rows that come back look complete, so there is no internal signal that anything is missing. Sequencing several providers is the only way to measure your own coverage.
Is it worth enriching candidates before qualifying them?
No, and this is the main cost mistake. Qualification should sit in the middle of the waterfall, after discovery and profile reading but before any paid contact lookup. Enriching first means paying for the whole market rather than the shortlist. Never enrich a candidate you are about to reject.
How long does a waterfall take to run on a new role?
The longlist is minutes rather than days once it is built: searching, reading profiles and filling in contact details across hundreds of people takes a few minutes of processing. On one build, 606 lawyers were scraped from a single firm directory, 89% were matched to a LinkedIn profile, and 604 of the 606 had an email verified, all in under eight minutes. The shortlist then takes however long a person needs to read the top of a scored list.
