Does lead scoring actually work?
Yes, when it's fit on your own closed-won outcomes rather than on guessed point values — and when the data underneath it is complete enough to rank. That second condition is where most scoring quietly fails: B2B records decay about 30% a year and single-source enrichment plateaus near 78–84% match, so a model reading half-empty, stale records is ranking noise with impressive-looking confidence. Build it on real outcomes, validate it on held-out history by reporting lift by decile, and write the reasons back alongside the score. Skip any of those and you get a number reps learn to ignore.
What a score is actually for
A score is not a truth claim about an account. It's a ranking device: an estimate of relative probability, used to decide what order a finite amount of rep attention gets spent in. That framing matters because it sets the bar for success. A model doesn't need to be right about any individual account. It needs the top of the list to convert measurably better than the middle.
Teams get into trouble when they treat the score as a verdict. Thresholds get set, leads below the line stop being worked at all, and the model stops receiving outcomes from anything it scored low — so it can never learn that it was wrong about them. Within a couple of quarters the score has become self-fulfilling, and its apparent accuracy is an artefact of what reps were allowed to touch.
Used as a ranking device, the same model behaves much better. The list is worked top-down rather than filtered, low-scored accounts still get worked when capacity allows, and the disagreements between the model and the reps stay visible — which is where nearly all of the improvement comes from.
Why hand-built point systems drift
The classic scoring model is a spreadsheet of point values: plus ten for a director title, plus five for a whitepaper download, minus fifteen for a free email domain. Someone reasonable assigned those numbers once, based on what seemed true at the time, and they were probably not unreasonable.
The problem is that nothing updates them. The market moves, the product changes, a channel that used to attract buyers starts attracting students, and the point values stay exactly where they were. Nobody notices, because a point system produces a confident number regardless of whether its weights still correspond to anything — there's no mechanism by which being wrong feels different from being right.
They also encode a specific and usually false assumption: that signals combine additively and independently. In reality the interactions carry most of the information. A particular title matters enormously at one company size and not at all at another. Two weak signals occurring together within a fortnight may be far more predictive than either alone. A point system can't express any of that, so it flattens the structure that would have been useful.
None of which means points are worthless. A point system built and maintained by someone close to the market beats no prioritisation at all, and it has the real advantage of being legible to everyone. It's a reasonable starting place — it just shouldn't be the ending place, and it needs someone whose job it is to revisit the weights.
What a model built on outcomes uses
Fitting on outcomes means training on your own closed-won and closed-lost history and letting the weights come from what actually differentiated them, rather than from what anyone believed. In practice that draws on three families of signal, and the useful models combine all three.
Fit: the firmographic and operational attributes that separate your winners. Standard fields do some work here, but the discriminating attributes are usually the ones no database sells as a column — how many locations, which software they run, whether they hold a specific licence or certification, how they're structured, whether they've expanded recently.
Timing: events suggesting the problem you solve just became urgent. Hiring in a relevant function, opening a location, a permit or licensing event, a technology change, a funding or acquisition event. Timing signals decay fast, which is the argument for scoring continuously rather than at import.
Engagement: what the account has actually done with you, where that history exists. Powerful when present, and the reason it can't be the whole model is that it only exists for accounts that already found you — score on engagement alone and you systematically deprioritise every account you should be reaching out to first.
- Fit — the operational attributes that separated your closed-won accounts
- Timing — hiring, expansion, permits, licensing, technology changes
- Engagement — real interaction history, where it exists
- Recency and decay weighting, so stale signals stop counting
How to validate it honestly
The validation that matters is simple and rarely done: hold out a slice of history the model never saw, score it blind, and check whether the top decile converted better than the rest. Report it as lift by decile — top 10% converted at this rate, next 10% at this rate, and so on. A model whose top decile doesn't clearly outperform the bottom isn't ready, whatever its headline accuracy figure says.
Be suspicious of accuracy as a metric here at all. In a dataset where most accounts don't convert, a model that predicts nothing ever converts scores impressively on accuracy and is worth nothing. Lift, precision at the volume your reps can actually work, and the shape of the decile curve tell you something useful; a single accuracy percentage mostly doesn't.
Then watch for leakage, which is the most common way a model looks brilliant in testing and fails in production. If a feature is only populated after a deal progresses — an opportunity field, a post-qualification note, anything a human filled in because the deal was going well — the model has learned to read the future. It'll show extraordinary validation numbers and rank new accounts no better than chance.
The other failure worth naming: training on MQLs or on qualification decisions rather than on revenue. That teaches the model to predict what your qualification process already does, including its mistakes. If the goal is more closed business, the label has to be closed business.
The data underneath decides the ceiling
A scoring model can only rank on what it can see. If a third of your records are missing the fields that carry the signal, the model isn't ranking those accounts badly — it's ranking them on absence, which usually pushes them down regardless of how good they are.
Two forces cause that. Coverage: single-source enrichment plateaus near 78–84% match on mainstream B2B and drops further outside it, so a meaningful share of records arrive incomplete. And decay: B2B data degrades about 30% a year, so fields that were populated correctly become quietly wrong — which is worse than blank, because the model treats them as evidence.
This is why serious scoring work starts with the enrichment layer rather than the model. Chaining providers so unmatched records fall through to the next source lifts match rates past 90%, and re-verifying on a schedule keeps populated fields from silently rotting. A modest model on complete, current data will beat a sophisticated one on sparse, stale data every time — and the gap isn't close.
Making reps trust it
A score with no explanation gets ignored, and reasonably so. Nobody reorders their week because a number in a column went up. Every score should carry the two or three factors that drove it, in plain language — recently expanded, runs the software our best customers run, holds the relevant licence — because that's also what the rep needs for the first thirty seconds of the conversation.
Put it where they work. A score that lives in a dashboard nobody opens changes nothing; a score that orders the queue in the CRM changes the day. And make disagreement cheap: a rep who thinks the model is wrong about an account should be able to say so in one click, because that objection is the highest-value data the system receives and it's usually thrown away.
Then refresh. Models drift because markets drift, so re-fit on a schedule as outcomes accumulate, and re-check the decile curve each time. The moment the top decile stops outperforming, something upstream has changed — and finding out what is more valuable than the score itself.
Common questions
Less than people expect. Even a few hundred closed-won and closed-lost outcomes beat a hand-built point system, and the model improves as outcomes accumulate. Below that, a documented point system maintained by someone close to the market is the honest interim answer.
Lift by decile on held-out history the model never saw: does the top 10% convert measurably better than the middle? That's the only test that matters. Be sceptical of a headline accuracy figure — in a dataset where most accounts don't convert, predicting 'no' every time scores well on accuracy and is worth nothing.
No, and doing that breaks the model. If low-scored accounts are never worked, the model never learns it was wrong about them, and its apparent accuracy becomes self-fulfilling within a couple of quarters. Work the list top-down rather than filtering it, and keep sampling below the line.
Quarterly is a reasonable default, more often in a fast-moving category. The trigger to watch isn't the calendar though — it's the decile curve flattening. When the top decile stops outperforming the middle, something upstream changed and the model needs attention.