Datum
Guide · Updated May 18, 2026

What is waterfall enrichment?

The short answer

Waterfall enrichment chains data providers in sequence: the first source fills what it can, the next takes the misses, then the next, until coverage is maxed. It exists because single-source data has a hard ceiling — contact accuracy plateaus near 78–84%, and one database has no fallback when a record isn't in its index. A well-configured waterfall regularly pushes match rates past 90% and lifts direct-dial coverage 20–40% above any single provider. The gain comes from sequencing, not from the number of providers.

How it works

Instead of accepting whatever one database returns, a waterfall sends the unmatched records to a second provider, then a third, and so on. Source A might fill 60% of a list; the remaining 40% flow to source B, then C, until the chain is exhausted.

The output is higher coverage than any single tool, because no one database covers everyone — but collectively they cover far more. The arithmetic is favourable in a way that's easy to underestimate: three providers that each cover 70% of a population, with imperfectly overlapping coverage, clear well past 90% between them, because each one's misses are not the same misses.

That's also the condition for it working. If your providers all assemble data the same way from the same public signals, their gaps overlap almost entirely and the chain adds cost without adding coverage. A waterfall of four similar databases is one database with four invoices.

The single-source ceiling

Major databases are effectively single-source for enrichment: if a record isn't in the index, there's no native fallback and you absorb the blank. That's why accuracy plateaus around 78–84% and drops sharply outside mainstream North American B2B.

The ceiling exists for structural reasons rather than because vendors aren't trying. Each provider builds from a particular set of inputs — public professional profiles, corporate sites, email and traffic signals, licensed feeds, contributed data from its own users. Those inputs determine which populations are visible. A buyer who doesn't maintain a public profile, works at a company with a minimal web presence, and uses a phone more than email is invisible to most of them at once.

Which is why upgrading a plan rarely fixes coverage. A higher tier buys more credits against the same index, not a different index. The fix is either a different collection method — another provider whose inputs differ from the first — or going to the source directly.

Sequencing is the whole design

The order of the chain determines both cost and result, and it's where most of the engineering judgment sits. The principle is straightforward: cheap and broad first, expensive and specialised last, so the premium provider only ever bills you for the records that everything else failed on.

Get the order backwards and you pay a specialist rate for records a commodity source would have filled. Since effective cost is list price divided by match rate divided by accuracy, and the specialist's premium is justified only on the hard tail, running it first can multiply the cost of a chain several times over for identical output.

Sequencing should also vary by segment rather than being one global order. A provider that's weak overall may be the best source in a specific geography or vertical, and a chain that routes by segment will beat a single fixed order. This is the main reason two teams using the same platform and the same providers get very different results — the tooling is identical and the routing isn't.

Then there's the stopping rule. Every chain has a point where the next provider's expected match on the remaining tail no longer justifies its price. Knowing where that point is, and stopping there rather than running every source on every record out of completeness, is the difference between a chain that saves money and one that just spends it in a specific order.

  • Broad and cheap first; specialised and expensive on the tail only
  • Route by segment — provider strength varies by geography and vertical
  • Set a stopping rule based on expected match against price
  • Choose providers whose collection methods differ, not just their logos

Verification, and why match rate isn't the goal

Match rate is the easiest number to move and the easiest to fool yourself with. A chain can hit 95% match by accepting low-confidence results from every source it touches, and deliver a dataset that's worse than one with 80% match and verified fields — because a wrong value is more expensive than an empty one. A blank prompts someone to check; a plausible wrong number gets dialled.

So verification belongs inside the chain, not after it. Where a provider returns a confidence level, use it: accept above a threshold, and let low-confidence results fall through to the next source rather than terminating the chain with a weak answer. Where two sources disagree on a field, that conflict should be recorded and resolved by a documented precedence rule, not silently won by whichever ran last.

The metric worth optimising is verified match rate — the share of records with a field you'd stake a rep's time on. It's a lower number than raw match, and it's the one that corresponds to anything real.

Where waterfalls still fall short

A waterfall is only as good as the sources in it. If none of your providers cover a market — common for niche verticals, the trades, regional operators, and data-poor segments — chaining them just stacks the same blanks. That's the point where custom sourcing, not more providers, is the fix.

It also does nothing about decay on its own. A chain enriches a record at a moment in time; roughly 30% of B2B data goes wrong within a year, mostly through job changes. Without scheduled re-enrichment, a beautifully executed waterfall produces a dataset that's excellent on the day it ran and quietly deteriorating from then on.

And a waterfall can't invent a field nobody collects. If your ICP turns on an operational attribute — a licence, a service line, a piece of equipment, a coverage area — no chain of contact databases will produce it, because none of them collect it. That gap is a sourcing problem, and recognising which of the two problems you have is most of the value of understanding how waterfalls work.

Should you build it or have it run for you?

Self-serve platforms have made chaining providers accessible to anyone willing to learn them, which is genuinely useful. What they hand you is the plumbing; the result still depends on provider selection, sequencing, segment routing, confidence thresholds, conflict resolution, and a stopping rule — none of which the tool decides for you.

Building it in-house makes sense if enrichment is a recurring core activity, someone owns it properly, and you'll keep tuning as provider quality shifts. It's a poor fit as a side responsibility, because a chain configured once and left alone degrades as providers change coverage and pricing, and nobody notices until the effective cost has quietly doubled.

Either way, the question to keep asking is the same: what is the verified match rate, at what effective cost per usable record, in each segment? A chain that can't answer that is being run on faith.

Common questions

  • A tool helps, but the result depends entirely on how well the chain is built and which sources are in it. Done poorly, a waterfall underperforms a single good source. Done well, it clears 90%.

  • Then a waterfall can't help — you have to source the data at its origin. That's the gap custom sourcing fills, and adding a fourth provider to a chain that's already missing the segment just stacks the same blanks.

  • Fewer than people assume, and chosen for difference rather than count. Two or three providers with genuinely different collection methods usually beat five similar databases, because similar sources miss the same records. Add a provider only when it measurably fills part of the tail the others leave.

  • No — and optimising for it is a common mistake. A chain can reach 95% match by accepting low-confidence results everywhere, producing a dataset worse than one at 80% with verified fields, because a plausible wrong number gets dialled while a blank gets checked. Track verified match rate instead.