TAM mapping & prospecting data
Define the ICP, map the whole market, and capture it as scored, ready-to-work data.
Your real total addressable market is larger than any database shows. Single-source tools plateau near 78–84% contact accuracy on mainstream B2B and fall away sharply outside it, so most teams plan against a fraction of their market. Datum maps and captures your entire TAM: we define the ICP from who actually closes, source the account population at its origin, run waterfall enrichment past 90% match, and score every account so reps work the ones most likely to close. You get a market you can defend, not a number a vendor happened to index.
Why your database undercounts your market
Every off-the-shelf database is a snapshot of what one vendor decided was worth cataloging. That bias is invisible until you sell to someone outside it. Coverage is deep where the vendor's customers already sell — venture-backed software, mid-market North American tech, roles with public LinkedIn footprints — and thin everywhere else: the trades, regional operators, franchise networks, family-owned manufacturers, licensed professionals, and any category that didn't exist when the index was built.
The damage compounds because the gap is silent. A filter that returns 3,000 accounts looks like an answer, not a shortfall. Nothing in the export tells you that 30,000 more companies match your ICP and simply aren't in the index. Teams then size the market, set quota, and plan territories against that number, and the error propagates into every downstream decision.
Freshness is the second half of the problem. B2B contact data decays roughly 30% a year, driven mostly by people moving: 65.8% of contacts change title or function within twelve months and 42.9% change phone number in the same window. Even a well-covered segment degrades between refreshes, so a market map is a process you run, not a file you buy once.
- Coverage skews to the segments vendors already sell into
- Missing accounts are invisible — an export never reports its own gaps
- Records decay ~30% a year, mostly through job changes
- Sizing built on one index inherits every one of its blind spots
Step one: an ICP defined from what actually closes
Most ICP documents describe the customer a company wants. We start from the customers it already has. We pull your closed-won and closed-lost history out of the CRM and look for the attributes that separate them — not just headcount and industry, but the operational tells that predict a fit: how many locations, which software they run, whether they hold a particular license or certification, how they buy, who signs.
Those distinguishing attributes matter more than the standard firmographic fields, because they are usually the ones no database sells as a column. They also tell us where to source. If your best customers are all licensed in a given trade, the state licensing registry is a better starting point than any contact database, and it is authoritative rather than derived.
We come out of this step with two things: a written ICP you can argue with, and a list of the source systems where that ICP is documented publicly. If those sources don't exist, we say so before you spend anything on capture.
Step two: sourcing the account population at the origin
Once we know where your market is documented, we capture it directly rather than through a reseller's index. Depending on the category that means public registries and licensing boards, industry association directories, marketplaces and review platforms, permit and filing records, franchise and dealer locators, or the trade bodies your buyers join.
Sourcing at the origin buys you three things a database can't. Completeness, because you get the whole population rather than the slice someone chose to index. Fields nobody sells, because the source carries the operational attributes your ICP actually turns on. And freshness on your schedule, because we re-run capture on a cadence you set instead of waiting on a vendor's refresh cycle.
It is not free. Custom sourcing is the right tool for maybe 10–20% of a typical data need — the genuine gaps — and it carries real upkeep, commonly cited at 20–30% of the original build cost each year as sites and defenses change. We scope it that way deliberately: buy the well-covered bulk, source only the gap, and own the maintenance so it doesn't quietly rot after handoff.
- Registries, licensing boards, and permit or filing records
- Association directories, marketplaces, and dealer locators
- Deduplication and entity resolution across every source
- Scheduled re-capture so the map stays current
Step three: waterfall enrichment past 90% match
Sourcing gives you accounts. Reaching them needs contacts, and no single provider has them all. A waterfall chains providers in sequence: the first source fills what it can, the unmatched records fall through to the second, then the third, until the chain is exhausted. Where a single source plateaus near 78–84% on mainstream B2B, a well-built waterfall regularly clears 90% match and lifts direct-dial coverage 20–40% above any one provider.
Order is the whole game. We sequence providers by cost and by where each one is genuinely strong, so the cheap high-coverage source runs first and the expensive specialist only ever sees the records everything else missed. That keeps effective cost down: list price divided by match rate divided by accuracy is the number that matters, and a cheap record that matches 40% of the time at middling accuracy is not cheap.
We also verify before anything reaches a rep, and we keep the chain honest about its own limits. If none of your providers cover a segment, stacking more of them just stacks the same blanks — that is the signal to go back to custom sourcing rather than add another vendor.
Step four: scoring, so a list becomes a work queue
A list ordered alphabetically is a list nobody works well. Scoring reorders it by probability, so the limited hours your reps have go to the accounts most likely to close. We build the model on your own closed-won and closed-lost outcomes rather than on invented point values, which is the difference between a score that reflects your business and a score that reflects whoever set up the rules two years ago.
In practice a model draws on three families of signal. Fit — the firmographic and operational attributes that separated your winners in step one. Intent and timing — hiring activity, expansion, technology changes, licensing or permit events, anything that suggests the problem you solve just got more urgent. And engagement — what the account has actually done with you, where you have that history.
We validate the way you would validate any model, not the way vendors demo one: hold out a slice of history, score it blind, and check whether the top decile really does convert better than the rest. If the lift isn't there, the model isn't ready, and we say so. Scores are written back into the CRM your reps already live in, with the reasons attached, because a number with no explanation gets ignored.
Scoring models drift. Your market moves, your product changes, the signals that predicted a win last year stop predicting one. We re-fit on a schedule and watch for the common failure modes: training on MQLs instead of revenue, leaking a post-sale field into the features, or scoring so aggressively that reps never see anything outside a narrow band and the model never learns anything new.
- Trained on your closed-won and closed-lost outcomes
- Fit, timing, and engagement signals combined
- Validated on held-out history, reported as lift by decile
- Written back into the CRM with the reasons attached
What the engagement actually looks like
The first couple of weeks are scoping and feasibility. We work through your closed-won history, agree the ICP, and identify the sources where that ICP is documented. You get an honest read on how much of your market is reachable and what it would cost before any build starts.
From there we stand up capture on a first segment rather than the whole market, because a narrow slice proves the sourcing and the fields quickly and cheaply. Enrichment and verification come next, then the first scored delivery into your CRM or warehouse — usually within the first month or so, depending on how many sources the ICP spans.
After that it becomes a loop. Capture re-runs on a schedule, enrichment refreshes against decay, the scoring model re-fits as outcomes accumulate, and we sit with your reps to hear which accounts were actually good and which weren't. That feedback is the most valuable input the system gets, and it's the reason we stay embedded rather than delivering a file and leaving.
When this isn't what you need
If you sell to mainstream North American software companies and your reps are working a list they trust, you probably don't have a sourcing problem. Buy the data, spend the money on the motion instead, and come back when coverage starts to cap out.
If your list is small enough for a person to build in a week, custom capture is over-engineering. If you need a contractual accuracy guarantee or an SLA on the data itself, a licensed vendor is the right answer, because we can report provenance and match rates but we can't sell you someone else's warranty. And if nobody documents your buyers publicly at all, there is nothing to source — we'd rather tell you that on the first call than take the engagement.
Buy tools, hire a team, or embed us
The comparison that decides most engagements: do it with more tools, a new hire, or a senior team embedded with yours.
| Dimension | Buy more toolsZoomInfo · Clay · Outreach | Hire in-houseA RevOps / GTM engineer | Embed DatumGTM engineering, executed |
|---|---|---|---|
| Your real TAM | Only what's already in the index | As far as one person can map it | We map and capture all of it |
| Time to first pipeline | However long you take to build it | 3–6 months to hire and ramp | We execute from week one |
| Who runs it day to day | Your reps, off the side of their desk | One seat, one point of failure | Embedded with your reps, in your Slack |
| Reporting | Dashboards you wire up yourself | If they get to it | RevOps reporting, actually analyzed |
| When it breaks | Your problem | Their problem, then yours | We own it and iterate |
Common questions
We estimate from the structure of the source records themselves — how many entries a registry or directory holds, what share pass your ICP filter, and how much overlap there is between sources. It's a bounded estimate with the method shown, and it firms up once capture runs on the first segment. You see how the number was built, not just the number.
We source publicly available information and respect each source's terms and rate limits. We don't bypass authentication, and we don't resell another vendor's licensed database. Scraping data that's public and logged-out sits on much firmer ground than data behind a login, and privacy law still applies to personal data however public it was — so scope and compliance are part of the discovery call, not an afterthought. This isn't legal advice.
Everything is built inside your stack — your CRM, your warehouse, your accounts — so the data lands in systems you already control. How each build is owned and handed off is scoped per engagement rather than promised as a blanket policy, and we document provenance for every record either way.
Those are tools you operate. A database sells you its index and stops where the index stops; a waterfall platform gives you the plumbing but the result depends entirely on how well you configure it. We do the work — including going and getting the data that isn't in anyone's index — and keep it running as sources change.
No. Everything we do is pre-outreach: sourcing, enrichment, scoring, and the systems that feed your reps. Your team owns every message that goes out.