RecruiterGTMRecruiterGTM
    SourcingOS

    Waterfall Enrichment for Candidate Sourcing

    Waterfall enrichment runs a list of people through several data providers in a fixed order, so each provider only works on the records the previous one missed. For candidate sourcing it matters more than for sales, because one provider typically returns about a third of a niche candidate market, and the contact detail you actually need is the personal one.

    By Reyhan Khan, CEO & GTM Systems Architect at RecruiterGTM. Four years building recruitment systems, eight in GTM overall, and 650+ recruitment agencies audited in the last three years.

    What the engine does with a job spec

    You hand Claude the job spec or a CV of someone who would be perfect. It reads it, works your own database first, then goes outside for everything your database did not have, ranks what comes back against the actual role, and gives you a scored shortlist with the proof attached.

    You give it

    A job spec, or a CV of someone perfect for the role

    ▼

    Claude reads it

    Turns the spec into the hard requirements a candidate must meet

    ▼

    Layer 0 · first, always

    Your own ATS

    Who you already hold, who you have already contacted, and who replied and was never followed up

    ▼

    then, only for what your database did not have

    Outside sources, in order

    1Discovery

    Google X-ray + Exa, three independent nets

    2Profile text

    Reads what each person wrote about themselves

    3Contact data

    Prospeo, then Apollo, then SalesQL for the personal email

    Each one only runs on what the one before it could not find.

    ▼

    Ranked

    Scored against the actual role

    Every candidate carries a word-for-word quote from their own profile proving why they qualify. No quote, no row.

    ▼

    You get

    A scored shortlist, with the proof attached

    You read the top of it and approve. That is the only part that still takes a person.

    Key takeaways

    • A waterfall is a fixed order: free providers run first, and paid ones only ever see the misses.
    • Your own ATS is layer 0. On one real role, 63% of the qualified market had never been seen by the agency that owned the database.
    • A candidate waterfall has a step in the middle that a sales waterfall does not: reading what each person wrote about themselves, before you pay for any contact details.
    • Work emails are the wrong channel for passive candidates. The personal email is a separate layer and the reason the whole thing earns its place.
    • Start with three layers if five feels like too much: ATS, then Prospeo, then SalesQL.
    • Every row carries a verbatim quote proving why the person qualifies. No quote, no row.

    Five words you will meet in this guide

    Enrichment
    Taking a name or a profile link and filling in the rest: current employer, job title, email address, phone number. A data provider does it for you, usually per record.
    Provider
    A company that sells that data. Apollo, Prospeo and SalesQL are all providers. Each holds a different slice of the market, and none of them tells you which slice it is missing.
    Scraping
    Reading what is publicly on a web page and saving it as data you can sort. Here it means reading the text people have written on their own public profiles.
    X-ray search
    A normal Google search aimed at one website only. Here it means using Google to read inside LinkedIn profiles, which surfaces people LinkedIn's own filters miss.
    Verified email
    An address checked against the receiving mail server to confirm somebody is actually there. Many providers guess addresses from a name and a company; a verified one has been tested.

    What is waterfall enrichment?

    Waterfall enrichment is the practice of passing each record through several data providers in sequence. Provider one runs on the full list, provider two runs only on what provider one failed to find, and so on down the chain. The result is higher coverage at lower cost, because the expensive sources only ever see what is left over.

    The technique comes from B2B sales, where it is used to find buyers. Every major explainer on the subject, from Clay to Apollo to ZoomInfo, teaches it that way. Almost nobody writes about it for candidate sourcing, which is a problem, because the requirements are genuinely different and the differences are the parts that decide whether a shortlist is usable.

    What makes a candidate waterfall different from a sales waterfall?

    Three things. The personal email matters more than the work email. A qualifying step sits in the middle of the chain, where you read profiles before paying for contact data. And your own database is the first place you look. A sales waterfall optimises for a verified work address. A candidate waterfall that does the same thing produces a list you cannot safely use.

    The work inbox is the wrong channel

    When you approach a passive candidate you are asking them to consider leaving the employer whose inbox you would be writing to. That inbox is often monitored. Most enrichment providers return work addresses by default, which means a sales-shaped waterfall hands you the one channel least likely to get a reply and most likely to cause a problem for your client.

    Qualification happens mid-waterfall, not at the end

    In sales, the title is usually the qualifier. A VP of Finance is a VP of Finance. In recruitment, a large share of what you are asked to find is a skill, not a title: a certification, a security clearance, a named modelling tool, a specific methodology. No database holds those as a filter, because they live in free text that the person wrote about themselves.

    So you read the profile text before you spend anything on contact data. On the role I ran, a title-led search returned roughly 60% people from a completely different profession who happened to share a word in their job title. A keyword-led search that read profile bodies returned a 44% genuine hit rate.

    Classify the role before you touch a tool. Title-defined roles (Recruitment Consultant, P6 Planner, Quantity Surveyor) are fine for title and employer search. Skill-defined roles need keyword-led discovery and profile reading. Get this one decision wrong and every layer below it faithfully processes the wrong list.

    How do you build a candidate sourcing waterfall?

    Six steps, in this order. Each one only runs on what the previous step could not resolve, and nothing paid runs until a candidate has passed the qualification screen. The order is the entire technique, so resist the urge to reshuffle it for convenience.

    Step 1. Search your own ATS first (layer 0)

    Before any external source runs, search the database you already own. It will rarely fill the role by itself, but it does two things no external source can. You learn your genuine coverage of your own market, and you stop approaching people twice, which is the fastest way to damage a candidate relationship.

    It also surfaces the people who already replied positively and were never followed up. There were six of those sitting in the file on the role I ran. Six warm candidates get you to a placement faster than six hundred cold ones.

    Step 1 promptcopy and paste into Claude
    Here is the job spec I am working on: [paste it, or paste the LinkedIn
    profile of someone who would be perfect for the role].
    
    Before you look anywhere outside, search my own ATS and tell me:
    1. Who in there genuinely matches this spec, with the line from their record
       that proves it.
    2. Who we have already contacted about anything, so we do not approach them
       twice.
    3. Anyone who replied positively in the past and was never followed up.
    
    Give me the proof for each one. Do not include anybody you cannot evidence.

    Step 2. Cast three independent discovery nets

    Never one. Three, then remove the duplicates. The same person turns up in two or three of those searches and you want one row each. Match them on the profile address, meaning the LinkedIn URL, not the name. Names and job titles are written inconsistently across sources, so matching on those misses duplicates.

    • An X-ray search through Apify. This is a normal Google search pointed at LinkedIn profiles only, written like site:uk.linkedin.com/in "your skill term". It is the most accurate net for a skill-defined role, because Google reads the whole profile where LinkedIn's own filters do not. Use one rare keyword per query. Putting two exact phrases in quotes together usually returns nothing. And raise the number of pages you ask for, because Google only shows ten results per page.
    • Exa, a tool that searches by meaning instead of exact words, restricted to profiles. This catches the people who describe the work without ever using the acronym, and in my experience they are often the strongest candidates.
    • Exa's find-similar feature. You point it at three or four people you have already confirmed are right, and it returns others who look like them.

    The test nobody mentions: overlap between your nets should be low, around 5%. If all three keep returning the same people, you have one source wearing three hats, and your coverage is thinner than your row count suggests.

    Step 3. Read what each person wrote about themselves

    Now pull the text from every profile on your deduplicated list: headline, about section, every job description, skills and certifications. The tool I use is called harvestapi, which runs on Apify. You hand it a list of profile links and it returns the text from each one. Set it to return profile details without email addresses, which costs $4 per 1,000 profiles, four tenths of a penny each. You skip the email here on purpose, because the next three steps get contact details more cheaply and more accurately.

    Then screen against a written list of checks and require a candidate to pass all of them. On the role I ran the five checks were: named software present on the profile, evidence of the technique described in the candidate's own words, a real programme behind the claim, current employer confirmed, and commutable to one of three cities. Roughly six in ten people with the right-sounding title failed the first two.

    Step 3 promptcopy and paste into Claude
    Here are the profiles you found: [attach the file].
    
    Read the full text of each one, then score every person against this job
    spec. A candidate only passes if they meet ALL of these:
    1. [your first hard requirement]
    2. [your second]
    3. [your third]
    
    For every person you keep, give me a word-for-word quote from their own
    profile that proves each requirement. No quote, no row.
    Tell me how many you rejected and the most common reason.

    Step 4. Get work email and mobile, from the free source first

    Prospeo runs first, on everyone who passed the screen. You give it a LinkedIn profile link and it returns a work email, a mobile number, and the person's full job history including how long they have been in each role. That history is useful on its own, because it tells you who is overdue a move. You will want their Pro tier at minimum, because the free allowance will not carry a real workload.

    Step 5. Second pass on the misses only

    Upload the leftovers to Apollo in one batch, strictly the records Prospeo could not find. Work email, sometimes a personal one, sometimes a phone number, plus current employment metadata you can use as a cross-check against what the profile claimed.

    Step 6. Personal email and direct dial

    SalesQL is the layer that makes a candidate waterfall different from a sales one. Steps 4 and 5 mostly return work addresses. This step returns the personal email and the direct number, which is the channel a passive candidate will actually answer.

    On the test role, 33 of 50 candidates came back with a personal email and 32 of those verified. That number is step 6 doing its job, and without it the pack would have been work-addresses-only.

    The five layers at a glance

    LayerWhat it returnsVerdict
    0. Your ATSYour true market coverage, duplicate removal, warm contacts never followed upNever skip
    1. Apify + ExaDiscovery across three independent netsBest for skill-defined roles
    2. Apify harvestapiFull profile text for qualificationBest value at $0.004/profile
    3. ProspeoWork email, mobile, job history with tenureRun first, before anything paid
    4. ApolloSecond-pass contact data + employment cross-checkBest for closing gaps
    5. SalesQLPersonal email, direct dialBest for passive candidates

    The three-layer version, if five is too many

    You do not need the full chain to beat what you are doing now. Build this instead, and add the discovery and profile-reading layers once the habit is established.

    LayerWhy it is in the starter
    Your ATSFree, already paid for, and tells you what you actually hold
    ProspeoReturns tenure as well as contact data, and the Pro tier covers a real workload
    SalesQLThe personal email, which is the whole point for candidate work

    What does a waterfall enrichment setup cost?

    Less than most recruiters expect, because the structure is designed to keep spend off the records that will be rejected. The only fixed per-unit cost in the chain above is the profile scrape. Everything else runs on a plan you already hold, and only touches whatever is left over.

    ItemCostNote
    Profile text scrape (Apify harvestapi)$0.004 per profile$4 per 1,000, with email addresses switched off
    ProspeoPro tier or aboveThe free tier will not carry a real workload. Check their pricing page for the current figure
    ApolloPlan-dependentRuns only on Prospeo's misses, so volume is a fraction of the list
    SalesQLPlan-dependentRuns only on qualified candidates, never on rejects
    Discovery (Exa, Google via Apify)Pennies per runYou pay for processing time, not per person

    The single biggest cost control is the sequencing rule: never enrich a candidate you are about to reject. Qualification is step 3 and contact data starts at step 4 for exactly that reason. Agencies that enrich first and qualify second pay for the whole market instead of the shortlist.

    What stops the list being wrong?

    One rule and two review passes. The rule is that every candidate on the list carries a verbatim quote from their own profile proving why they qualify, and no quote means no row. The review passes are two independent checks that run before a human sees the pack, and the tool that built the list is never the one that approves it.

    The quote rule does two things at once. Your consultant can verify any line against the public profile in seconds, which is what makes the list defensible in front of a client. And the automation cannot quietly invent a reason, because it has to cite the sentence.

    Three failures worth borrowing from mine

    1. Evidence has to come from the current role to be claimed as current practice. On one run, 11 of 17 profiles I had recorded as current programme work were actually describing a previous employer. Still valid candidates. Wrong claim in the outreach.
    2. Constrain the outreach copy to quoted facts. Left unconstrained, 64% of my first-pass openers contained a factual error, including inventing a programme and describing a risk register as a modelling tool. A deliberately general opener does no damage. A confidently wrong one costs you the relationship.
    3. Never append rows to a finished list. I added 18 candidates after quality control had already passed. Those 18 were rejected at 44%, against 6% for everything that went through the full pipeline in order.

    One concrete thing the second review pass catches: a shared team mailbox such as risk@theiremployer.com, sitting in the file marked verified. It passes every verification check precisely because it is a shared inbox that accepts anything sent to it. Sending to it means a poaching approach arrives in a team-wide inbox at the candidate's own employer, which costs you the client.

    What did this return on a real role?

    On one test role the waterfall returned 46 of 50 candidates contactable, 33 of them with a personal email, and surfaced 102 qualified candidates the agency had never seen. That was a UK project controls agency searching three cities for a skill no database holds as a filter. Here is the full output, with the method stated so the numbers can be checked.

    MeasureResult
    Profiles delivered, each with a verbatim proof quote50
    Contactable on at least one channel46 of 50
    Personal email present / verified33 / 32
    Qualified candidates in their three cities they had never seen102
    Share of their own market previously visible to them14%

    Be clear about where the speed actually is, because this gets oversold. The longlist is the part that goes from days to minutes. Searching, reading profiles and filling in contact details across hundreds of people is a few minutes of processing. On a separate build we scraped 606 lawyers from one firm's directory, matched 89% of them to a LinkedIn profile, and verified email addresses for 604 of the 606, in under eight minutes. Profile matching and email discovery run as independent layers, which is why the email count is higher than the LinkedIn match count.

    The shortlist still takes a person a few minutes to read and approve. There is no step where that disappears. The scoring gate runs automatically, then someone reads the top of the list, and what used to be a week of scrolling becomes a scored shortlist you sign off over a coffee, because the system did the scoring.

    Where to go next

    If you want the wider context around this, the candidate sourcing automation guide covers how the engine fits a full recruitment workflow, the sourcing automation playbook covers the campaign layer that sits on top of a waterfall, and the AI tools breakdown covers what else belongs in a recruitment stack in 2026.

    Frequently asked questions

    What is waterfall enrichment in recruitment?

    It is running candidate records through several data providers in a fixed order, where each provider only processes the records the previous one failed to resolve. In recruitment it adds two things a sales waterfall does not have: your own ATS as the first source, and a personal-email layer at the end, because work addresses are the wrong channel for a passive candidate.

    How many data sources do I actually need?

    Three is a real waterfall and will beat a single provider: your ATS, then Prospeo, then SalesQL. Five is better for specialist roles, adding a discovery layer and a profile-reading layer. The order matters more than the count, because the order is what keeps paid providers off records that are going to be rejected.

    Why not just use one tool that claims full coverage?

    Because no provider publishes where its gaps are, and in a specialist niche a single source commonly returns around a third of the real market. The rows that come back look complete, so there is no internal signal that anything is missing. Sequencing several providers is the only way to measure your own coverage.

    Is it worth enriching candidates before qualifying them?

    No, and this is the main cost mistake. Qualification should sit in the middle of the waterfall, after discovery and profile reading but before any paid contact lookup. Enriching first means paying for the whole market rather than the shortlist. Never enrich a candidate you are about to reject.

    How long does a waterfall take to run on a new role?

    The longlist is minutes rather than days once it is built: searching, reading profiles and filling in contact details across hundreds of people takes a few minutes of processing. On one build, 606 lawyers were scraped from a single firm directory, 89% were matched to a LinkedIn profile, and 604 of the 606 had an email verified, all in under eight minutes. The shortlist then takes however long a person needs to read the top of a scored list.