The weights were set in a room
Most scores get their weights from a discussion. Usage feels important, so it gets 40 percent. NPS feels important, so it gets 20. Nobody checks whether either number separated churned accounts from renewed ones in the historical data.
Fix it by backtesting. Recompute each candidate signal as it stood 90 days before renewal across two years of outcomes, then keep the signals that produced real separation and cut the rest, however obvious they seemed. Weights derived this way survive scrutiny because someone can point to the accounts they came from.
Nonlinear signals get read as straight lines
Scoring assumes more is better and less is worse. Several of the strongest retention signals do not behave that way.
Support volume is the clearest case. ORM's read across its customer base is that an account with no support cases is at risk of churn, an account with seven or more cases in the last year is at risk, and accounts with three to five routine tier 2 or tier 3 cases are the least likely to leave, because those customers are engaged and getting help. A linear score that rewards fewer tickets gives the highest marks to the accounts most likely to disappear.
Any signal with a healthy middle needs a curve, not a slope.
Silence scores as health
Quiet accounts generate no alerts, so they sit at the bottom of every queue while noisy accounts absorb the attention. ORM applies the same reading to deal risk, where the earliest signal is the absence of a signal, meaning no activity, no data changing, and no notes. Retention works identically. An account that stopped filing tickets, stopped attending reviews, and stopped adding users is not stable, it is unmeasured.
Score absence explicitly. Days since last meaningful event should carry weight in its own right rather than being inferred from the other inputs.
The model never got updated
A score calibrated at launch describes the customer base as it existed at launch. The failure is the same one that breaks revenue forecasts, where a model built on old assumptions misses because something in the business or the market changed and the model did not respond. New segments, a pricing change, or a competitor creating price pressure all shift which signals matter.
Rerun the calibration every two quarters and any time a scoring input changes. Then check the score the way a forecast gets checked, against outcomes, the same discipline applied to forecast accuracy and to the retention numbers it drives, including net revenue retention. A score nobody has backtested is a shared opinion with a number printed on it.
Frequently Asked Questions
How do you test whether a health score predicts churn?
Backtest it. Recompute the score as it stood 90 days before each renewal in the last two years, then compare the score distribution of churned accounts against renewed ones. If the two distributions overlap, the score is measuring engagement rather than risk. Scoring current accounts tells you nothing, because you have no outcome to check the score against.
Why do green accounts still churn?
Usually because the score rewards logins and rewards silence. An account with steady logins from a shrinking group of users, no support cases, and a departed executive sponsor scores well on most models and cancels anyway. The inputs that would have caught it are absent from the formula.
Are more inputs better?
No. Adding weakly predictive signals dilutes the strong ones and makes the score harder to explain. Four inputs that separated churn in a backtest beat twelve inputs picked because the data was available.
Should health scores be replaced by a churn model?
They should be validated the same way a churn model is validated. A health score that gets backtested, weighted from history, and recalibrated on a schedule is a churn model with a friendlier interface. One that never gets checked against outcomes is a dashboard.
Put these metrics to work
ORM builds custom revenue forecast models that turn concepts like why health scores fail to predict churn into prescriptive action for your team.
Schedule a Demo