Datum
Use case

Capturing the total addressable market databases miss

The short answer

If your model says 40,000 accounts but your database surfaces 4,000, the gap isn't your market — it's the index's blind spot. Mainstream tools plateau near 78–84% coverage on well-documented segments and fall away sharply elsewhere, so any sizing built on one of them undercounts. Datum maps and captures your entire real TAM from primary sources — registries, directories, marketplaces, filings — dedupes it into a defensible account universe, and scores it. You get a market you can plan territories against and show a board, not a number that quietly excludes most of your buyers.

Market sizing pulled from a single database inherits that database's gaps. For mainstream software that's tolerable. For most other markets it means your "TAM" is really "the slice one vendor happened to catalog" — and nothing in the export tells you which is which.

When that number drives territory planning, quota, and fundraising, the undercount has consequences well beyond the sales team. Capturing the full population at the source fixes the input. Scoring it tells you which part of the population to work first.

Why the number in your model and the number in your tool disagree

The two numbers are built in opposite directions, which is why they rarely meet. A top-down model starts from a market and reasons inward: how many businesses of this type exist, what share fit the profile, what each is worth. A database export starts from an index and reasons outward: here is what we have that matches your filters. One is an estimate of reality; the other is an inventory of a vendor's coverage.

Coverage isn't randomly distributed either, which makes the gap worse than a simple shortfall. Databases are deepest where their own customers sell — venture-backed software, mid-market North American technology, roles with an active public professional profile. They thin out for owner-operated businesses, regional and franchise networks, licensed trades, manufacturers, and anything that files with a regulator rather than announcing itself online.

So the export isn't a random 10% sample of your market that you could scale up. It's a biased slice, systematically missing the segments where the buyer is a person who has never updated a professional profile in their life. Multiplying it by a factor doesn't recover the missing accounts; it just produces a bigger wrong number.

Signs your market map is wrong

  • TAM vs. reality gapThe market is obviously bigger than what your tools return, and nobody can explain the difference.
  • Territory planning stallsReps run out of named accounts that clearly exist in the wild — they've driven past them.
  • The number won't defendYour addressable market can't be justified past the database it came from, which makes board conversations uncomfortable.
  • Segments look emptyA vertical you know is real returns a handful of rows, and the handful are the largest players only.

How we rebuild the account universe

  1. 01We identify the authoritative sources for your category — regulators, licensing boards, industry associations, marketplaces, filing systems, franchise and dealer networks — and capture the account population directly from them rather than through a reseller's index.
  2. 02We resolve entities across those sources, which is most of the real work: the same company appears with different legal names, trading names, addresses, and identifiers in every system, and a naive merge produces either duplicates or false matches.
  3. 03We enrich the resolved population, verify before delivery, and score it on the attributes that separated your closed-won accounts — then land it in your CRM or warehouse as a working account universe with firmographics, segments, and a fit score attached.
  4. 04We re-run capture on a schedule, because the population itself moves: businesses open, close, merge, and change status, and a market map is only true on the day it was built.

Entity resolution is the part everyone underestimates

Pulling records from five sources is a week's work. Deciding which of those records are the same company is the part that determines whether the output is usable. A single business can appear as a legal entity in a state registry, a trading name in a directory, a location record in a marketplace, and a parent company in a filing — with four different addresses, none of them wrong.

Get it wrong in one direction and you deliver a TAM inflated by duplicates, which is worse than the undercount you started with because it's confidently wrong. Get it wrong in the other and you collapse genuinely separate locations into one account, which breaks territory assignment for exactly the multi-site businesses that are often the best customers.

We resolve on a combination of identifiers rather than name similarity alone — registration numbers, licence numbers, normalised addresses, domains where they exist — and we keep the source records attached rather than discarding them into a merged row. When a rep asks why two branches are separate accounts and a third isn't, there's an answer.

A TAM number you can actually defend

The output isn't a headline figure. It's a population you can interrogate: how many accounts, from which sources, passing which ICP criteria, with what confidence, and what share of them have a reachable contact. Anyone who wants to argue with the number can argue with a specific step rather than with the total.

That matters most in the rooms where the number gets used. A board asking how you arrived at your addressable market wants the method, not the figure. A territory plan that assigns named accounts needs the accounts to actually exist and be assignable. A pricing or packaging decision that assumes a segment is large needs to know whether that segment is large or merely under-indexed.

It also changes what you do next, which is the point. A market that turns out to be four times larger than the export suggested is a hiring and coverage question. One that turns out to be genuinely small is a positioning question. Both are better problems than not knowing.

Common questions

  • Because they come from systems that exist for a reason other than selling you data — a licensing board records a licence because the law requires it. We keep the source record attached to every account, so any row can be traced back to where it came from and when it was captured, and you can spot-check as many as you like.

  • They will, constantly — different addresses, different names, different status for the same business. We keep every source record rather than flattening them, apply a documented precedence rule per field, and flag genuine conflicts instead of silently picking a winner.

  • Then you've learned something worth more than the engagement cost. A market that's genuinely small is a positioning and pricing problem, not a coverage one, and it's far cheaper to find that out from a defensible count than from two years of missed quota.

  • The account population moves more slowly than contact data, but it does move — businesses open, close, merge, and change licensing status. Most markets warrant a re-capture quarterly, with faster cycles where the category is churning.