The blog · Pipeline Strategy · From The Demand Compass · 6 min

Should you use AI or an LLM to score leads?

Two 0-to-10 scores built from fixed weights and decay, never a model's guess. Where AI reads a signal inside the engine, and the one line it never crosses.

A lighthouse with its lamp room open like a clockwork engine, orange gears turning inside, one steady beam over a dark sea. Artwork from The Demand Compass.

The Demand Compass scores every account with deterministic arithmetic, never with an AI model or a large language model. Each account gets two numbers, one per axis, built by summing weighted sub-scores and applying decay, so the same signals always produce the same score. AI does have a job inside the engine: reading messy, unstructured evidence, a job post, a filing, a news item, and turning a judgment call into one tagged input. What it never does is set the score itself, because a model that just says "this looks like an 8" cannot show its work, and a sales team that has already learned not to trust one black box will not start trusting a new one.

Why does a lead score need to be explainable, not just accurate?

A VP of Sales whose reps kept refusing marketing's leads is the failure mode the Demand Compass engine was built against. The leads were not all bad; the reps simply had no reason to believe the score, so a score everyone ignores is worth zero however accurate it is underneath. The engine's first job is not to be right. It is to be explainable, because a number nobody trusts never gets worked.

How does the Demand Compass calculate each account's score?

The Demand Compass gives every account two scores, one for brand awareness and one for readiness to buy, each on a zero-to-ten scale and each a sum of sub-scores capped at the top of the range. Readiness adds up the seniority and function behind a relevant job post, web intent on the pages that matter, developer or product activity, sharp company- and industry-level events, and one research signal that needs judgment rather than arithmetic. Awareness adds up connections and follows into the team, reactions to what gets published, recency-weighted newsletter engagement, event and badge scans, and ad engagement. Add the sub-scores, apply the decay that makes old evidence count for less, cap the total, and that is the axis.

Alongside the two numbers, the engine writes tags: up to a couple dozen typed signal tags per account, each a plain-language record of a specific thing that fired. "Senior security role posted." "Pricing page, three visits." "Spoke at a conference." "Newsletter, high recent engagement."

One honest note on the weights themselves: the structure transfers to any business, the exact weights do not. How much a pricing visit is worth against a job post against a funding round is not a universal truth, it is a hypothesis calibrated against your own outcomes. Copying another company's weights gives you someone else's mistakes.

Where does AI fit inside the scoring engine, and where is it never allowed?

AI reads at the edge of the Demand Compass engine and never touches the arithmetic at its core. The core is deterministic: the same signals, run through the same weights with the same decay, always produce the same score, and that repeatability is the source of the engine's credibility, not a limitation of it. There is exactly one place determinism gives way to judgment. Is this account actually building with AI agents, or does it just use the word on its site? Was there a real security incident here recently? Counting page visits cannot answer either question, so an AI agent reads the open web and answers instead, producing one tagged input that gets scored like any other signal.

Inside the engineAI does thisAI never does this
Reading a job post, filing or news itemYes, to judge what it meansNo
Turning a judgment call into a tagged sub-scoreYesNo
Setting the final readiness or awareness numberNoNever
Deciding whether an account is In-Market or Sales-ReadyNoNever
Replacing the audit trail with an opinionNoNever

An LLM that emits "this looks like an 8" cannot show its work, and will answer differently on Tuesday than on Monday.

That line is the one the engine will not cross. Magic in a demo is poison in production, and a black box is exactly the thing a sales team already learned not to trust once, back when the MQL score turned out to hide four different questions inside one number.

Why does auditability matter beyond winning over your own reps?

Auditability is what lets a rep open a routed account and believe it in four seconds, not what it costs to build. Signals are immutable events and scores are derived from them, so any score can be recomputed and explained rather than taken on faith. The routing screen does not say "this account is hot." It says why: new VP of Sales twelve days ago, pricing page visited three times this week, docs read twice, readiness eight, awareness four, In-Market. That is a case a human can read and check, and it is what turns the rejection that killed the MQL into action instead.

There is a second audience, and for a company selling into security buyers it is not optional. A security buyer runs a vendor assessment on everyone it evaluates, and a black-box AI that cannot explain itself fails that review on sight, while an engine where every score decomposes into named, dated signals passes it. Auditability is not only how you win over your own reps; it is how you survive your customers' diligence. That discipline paid for itself the first time a client asked, in a renewal meeting, why one specific account had been routed to their best rep, and got a complete answer from the underlying signals in thirty seconds.

What does an auditable score look like on a real account?

Following one account inside a cybersecurity deployment shows what the engine produces when it works. The two axes split clean: brand awareness was low, nobody there followed the company, no event scans, almost no LinkedIn footprint with the team, while readiness to buy was high, a sector breach, a new CISO, compliance and pricing visits, a stack-naming job post, all still inside the decay window. Low awareness, high readiness. In-Market.

The score did not arrive as a vague "hot." It arrived as a line a rep could audit: In-Market because of a sector breach, a new CISO hired, six compliance and pricing visits, and a security-engineer job post naming the stack, all in the last thirty days. Because every clause pointed at a dated event, the team trusted it.

0 to 10each axis's score range
~24plain-language signal tags an account can carry
30 secondsto explain a routing decision from the raw signals

You can argue with a hunch. You cannot argue with the log.

Can an LLM be trained to output a lead score directly?

It can, but the Demand Compass engine does not let it. A model can be prompted for a number, but it cannot show its work the way a summed, weighted, decayed score can, and it will drift, answering slightly differently on the same account from one day to the next.

What does AI actually read inside the engine?

Unstructured evidence a fixed rule cannot parse cleanly: a job post that may or may not describe building with AI agents, a filing, a news item about a possible security incident. Each becomes one tagged, scored input, never the final number.

Why not just trust a good AI model to judge the whole account?

Because trust has to survive a rep asking "why," a renewal meeting, and a security buyer's vendor assessment. A black-box judgment fails all three.

Do the scoring weights work the same for every company?

No. The structure, two axes, sub-scores, decay, a cap, transfers to any business. The exact weights are a hypothesis calibrated against your own outcomes; copying another company's weights imports their mistakes too.

What is a signal tag, and why does the engine bother writing them?

A plain-language record of one specific thing that fired, like "senior security role posted" or "pricing page, three visits." It is what a rep reads instead of trusting a bare number.