Optimized Sales Optimized Marketing Target Accounts For CROs For CFOs For CMOs Blog News Glossary Compare Tools About Schedule a Demo
Retention & Growth

Why Health Scores Fail To Predict Churn

ORM Technologies
Home/ Glossary/ Why Health Scores Fail To Predict Churn
Definition Most customer health scores fail because they are built from opinions about what should matter rather than from the signals that actually separated churned accounts from renewed ones. The result is a score that tracks engagement and misses cancellations.
Health scores fail for a small number of repeatable reasons, and every one of them is a design choice rather than a data problem. The common thread is that the score was assembled from what felt important in a workshop and then never tested against what actually happened at renewal.

The weights were set in a room

Most scores get their weights from a discussion. Usage feels important, so it gets 40 percent. NPS feels important, so it gets 20. Nobody checks whether either number separated churned accounts from renewed ones in the historical data.

Fix it by backtesting. Recompute each candidate signal as it stood 90 days before renewal across two years of outcomes, then keep the signals that produced real separation and cut the rest, however obvious they seemed. Weights derived this way survive scrutiny because someone can point to the accounts they came from.

Nonlinear signals get read as straight lines

Scoring assumes more is better and less is worse. Several of the strongest retention signals do not behave that way.

Support volume is the clearest case. ORM's read across its customer base is that an account with no support cases is at risk of churn, an account with seven or more cases in the last year is at risk, and accounts with three to five routine tier 2 or tier 3 cases are the least likely to leave, because those customers are engaged and getting help. A linear score that rewards fewer tickets gives the highest marks to the accounts most likely to disappear.

Any signal with a healthy middle needs a curve, not a slope.

Silence scores as health

Quiet accounts generate no alerts, so they sit at the bottom of every queue while noisy accounts absorb the attention. ORM applies the same reading to deal risk, where the earliest signal is the absence of a signal, meaning no activity, no data changing, and no notes. Retention works identically. An account that stopped filing tickets, stopped attending reviews, and stopped adding users is not stable, it is unmeasured.

Score absence explicitly. Days since last meaningful event should carry weight in its own right rather than being inferred from the other inputs.

The model never got updated

A score calibrated at launch describes the customer base as it existed at launch. The failure is the same one that breaks revenue forecasts, where a model built on old assumptions misses because something in the business or the market changed and the model did not respond. New segments, a pricing change, or a competitor creating price pressure all shift which signals matter.

Rerun the calibration every two quarters and any time a scoring input changes. Then check the score the way a forecast gets checked, against outcomes, the same discipline applied to forecast accuracy and to the retention numbers it drives, including net revenue retention. A score nobody has backtested is a shared opinion with a number printed on it.

Frequently Asked Questions

How do you test whether a health score predicts churn?

Backtest it. Recompute the score as it stood 90 days before each renewal in the last two years, then compare the score distribution of churned accounts against renewed ones. If the two distributions overlap, the score is measuring engagement rather than risk. Scoring current accounts tells you nothing, because you have no outcome to check the score against.

Why do green accounts still churn?

Usually because the score rewards logins and rewards silence. An account with steady logins from a shrinking group of users, no support cases, and a departed executive sponsor scores well on most models and cancels anyway. The inputs that would have caught it are absent from the formula.

Are more inputs better?

No. Adding weakly predictive signals dilutes the strong ones and makes the score harder to explain. Four inputs that separated churn in a backtest beat twelve inputs picked because the data was available.

Should health scores be replaced by a churn model?

They should be validated the same way a churn model is validated. A health score that gets backtested, weighted from history, and recalibrated on a schedule is a churn model with a friendlier interface. One that never gets checked against outcomes is a dashboard.

Put these metrics to work

ORM builds custom revenue forecast models that turn concepts like why health scores fail to predict churn into prescriptive action for your team.

Schedule a Demo