Optimized Sales Optimized Marketing Target Accounts For CROs For CFOs For CMOs Blog News Glossary Compare Tools About Schedule a Demo
Retention & Growth

How to Build a Customer Health Score That Predicts Renewals

Pete Furseth 7 min read
customer healthchurnretention
How to Build a Customer Health Score That Predicts Renewals
Home/ Blog/ How to Build a Customer Health Score That Predicts Renewals

What should a health score actually predict?

One thing: whether the account renews at or above its current ARR. Narrowing the claim is what makes the score useful, because a specific prediction can be checked against what happened. Scores designed to capture general account wellness cannot be tested, so nobody ever finds out they are wrong and nobody changes their behavior because of them.

Write the definition down before you pick a single input. Renewal at or above prior ARR, measured at the contract date, per account. That definition determines your training data, your validation method, and the argument you can make to a skeptical account executive who disagrees with a red score.

Put this to work on your numbers
Run your own numbers with the free Pipeline Velocity Calculator, then see how ORM builds it into a custom model.

Which inputs earn a place in the model?

Inputs the customer has to spend effort to produce. Cheap signals move for reasons that have nothing to do with renewal intent. A dashboard that auto-refreshes generates logins. An email opened by a spam filter counts as engagement. Neither predicts anything.
InputWhat it measuresCommon trap
Support case patternWhether the product is in real useTreating fewer cases as healthier
Seat or license deploymentValue the customer can defend at renewalCounting provisioned seats instead of active ones
Executive meeting attendanceWhether a sponsor still existsAccepting a delegate as a sponsor
Response latency to outreachWhether anyone is engagedCounting your own activity instead of theirs
Stakeholder continuityWhether your champion is still thereMissing role changes that never reach the CRM
Expansion conversation movementForward intentReading a stalled upsell as neutral
Support volume is the input most teams model backward. ORM data shows a curve rather than a line: zero cases in a year marks an account nobody is using, seven or more marks an account fighting the product, and three to five ordinary tier 2 or tier 3 cases marks the healthiest group in the base. Score it as bands.

How do you weight the inputs?

From outcomes, using accounts that already renewed or churned. Pull two years of completed renewals, split them into two groups, and compare how each candidate input behaved in the six months before the decision. An input that looks the same in both groups has no predictive value, whatever intuition says about it.

Machine grouping does this better than a spreadsheet once you have enough history. In our forecasting work, ORM groups opportunities with a machine learning model and predicts a separate close curve for each group, and the same principle applies to accounts. Similar accounts behave similarly, and the groups rarely match the segments a human would draw. Absent that, a simple weighted model built from outcome comparison beats a workshop consensus every time.

Two rules keep the model honest. Cap any single input at a share of the total that cannot swing an account between tiers on its own, and require at least one input from stakeholder continuity so the score cannot stay green after a champion leaves.

Why do most health scores fail?

They are never backtested, so nobody knows whether green means anything. The build usually goes: gather available fields, assign weights in a meeting, launch the score, and never check it against renewals. Twelve months later the score is decorative.

Run the validation before launch. Score every account as of a date twelve months in the past using only data available then, then compare against what actually happened. Two numbers come out. The first is separation, meaning the difference in churn rate between red and green accounts. The second is coverage, meaning the share of churned ARR that carried a red or yellow score more than 90 days before the loss. A score where red accounts churn at close to the same rate as green ones is worse than no score, because it directs effort at random while carrying an air of authority.

Watch for the retention version of a slipping close date as well. An expansion conversation that keeps sliding predicts trouble the same way deal slippage does on new business, and it moves earlier than usage decline.

Should the score be one number?

One number for routing, three components for diagnosis. A single score answers who to work first. It cannot answer what to do, and an owner handed a red account with no reason opens a support ticket and asks the customer if everything is fine.

Publish the score alongside its worst-performing components. Red on deployment sends the owner into a rollout conversation. Red on stakeholder continuity sends them into a reintroduction with a new economic buyer. Same tier, different play, and the component tells them which.

Keep account value out of the score entirely. Risk and ARR are separate dimensions, and blending them produces a number where a large healthy account and a small dying one land in the same bucket. Multiply them at the routing stage instead.

How do you keep the model honest?

Recalibrate every two quarters against completed renewals and log every override. Owners will disagree with scores, and they are sometimes right. Require them to record the override with a reason. Those reasons are training data: when the same override reason appears twenty times and the accounts renew, you have found an input the model is missing.

Publish the model's own scorecard next to it. Separation between red and green churn rates, coverage of churned ARR flagged 90 days out, and the count of overrides that turned out correct. A score with a visible track record survives disagreement, because the argument moves from opinion to evidence. A score with no track record loses every argument it has with a confident account executive, which is how these systems quietly stop being used.

Rebuild rather than tune when the business changes. New pricing, a new segment, or a materially different product means the base no longer resembles the accounts the model learned from. Hold the whole system accountable to gross revenue retention rather than to net revenue retention, since expansion can lift NRR while the loss rate underneath it accelerates. The same measurement discipline that keeps a forecast credible applies here, and the forecasting practices worth copying start with checking predictions against outcomes on a schedule.

Frequently Asked Questions

What should a customer health score actually measure?

The probability that an account renews at or above its current ARR. That is a testable claim, which means the score can be validated against outcomes. Scores built to measure vague account wellness cannot be proven right or wrong, so they never improve and nobody trusts them enough to reallocate work.

Which inputs belong in a health score?

Inputs that cost the customer effort to produce. Executive meeting attendance, support case patterns, seat deployment, and responsiveness to outreach all require someone at the account to do something. Login counts and email opens are cheap signals that move for reasons unrelated to renewal intent.

How should health score inputs be weighted?

From your own renewal outcomes, not from a workshop. Pull two years of renewals, split them into renewed and churned, and compare how each candidate input behaved in the six months before the decision. Inputs that fail to separate the two groups get dropped regardless of how sensible they seem.

Is more support activity a bad health signal?

Not in a straight line. ORM data shows a curve: accounts filing zero support cases carry elevated churn risk, accounts filing seven or more in a year also carry elevated risk, and accounts filing three to five ordinary tier 2 or tier 3 cases are the safest group. Score support volume as bands rather than treating fewer tickets as healthier.

How often should a health score model be rebuilt?

Recalibrate every two quarters against completed renewals and rebuild whenever the product, pricing, or ideal customer profile changes materially. A score trained on a base that no longer resembles your current customers will keep firing on patterns that stopped mattering.

PF
Pete Furseth
ORM Technologies
Pete has built custom revenue forecast models for B2B SaaS companies for over a decade.

See how ORM turns these insights into action

ORM builds custom revenue forecast models for B2B SaaS companies. Not dashboards. Prescriptive analytics that tell you what to do next.

Schedule a Demo