Calibration asks a narrow question. When a model says 70 percent, does 70 percent happen? Score 200 deals at that level and roughly 140 should close. If 90 close, the model is miscalibrated, and every dollar figure built on its probabilities runs hot.
Ranking Is Not Calibration
A model can sort deals perfectly from most to least likely and still attach numbers that are systematically too high. Ranking answers which deals deserve attention this week. Calibration answers what number goes in the forecast. Weighted pipeline math depends entirely on the second, which is why teams that adopted probability scoring for coaching often get worse forecasts out of it.
How to Check Calibration
Pull resolved opportunities from the last several quarters. Bucket them by the probability they carried at a fixed reference point, such as 30 days before their predicted close date. Compare predicted probability against observed close rate inside each bucket.
Run the same check by segment and by rep. Aggregate calibration hides compensating errors, where an optimistic enterprise model and a pessimistic mid-market model cancel out at the roll-up and mislead every decision underneath it.
Why Stage Percentages Fail the Test
Static stage weights are the most common probability source in B2B SaaS and the least defensible. They get set once during a CRM implementation, applied uniformly to every deal in the stage, and almost never retested against what actually closed. A deal that entered stage four last week and one that has sat there for five months carry the same number.
Amount inflation compounds the problem. ORM's point is that pipelines routinely carry deal amounts well above what those deals actually close for, and probabilities can be perfectly calibrated while the dollar forecast still runs high. Calibrate expected value alongside likelihood.
When to Recalibrate
The true rate moves when conditions move. ORM's account of forecast misses is that something changed in the business or the market while the model kept running on old assumptions. A new competitor creates pricing pressure and average deal size falls. Capital tightens and buyers cut spending, so win rates drop. Uncertainty stretches cycles from qualified to closed.
Seasonality shifts the baseline too. ORM sees Q2 and Q4 running stronger than Q1 and Q3, with the third month of a quarter outperforming the first two. A model calibrated on blended history will read a normal Q1 as a collapse.
Track calibration as a standing measure alongside forecast accuracy and win rate, and reconcile it against your weighted pipeline math every quarter.
Frequently Asked Questions
What does a calibrated forecast model mean?
It means the stated probability holds up in practice. Score 200 deals at 70 percent and roughly 140 should close. If 90 close, the model is miscalibrated regardless of how well it ranked them.
How do you test calibration?
Take resolved deals from recent quarters, bucket them by the probability they carried at a fixed point such as 30 days before predicted close, then compare each bucket's predicted probability to its actual close rate.
Are CRM stage percentages calibrated?
Almost never. Stage weights are set once, applied to every deal in the stage, and rarely retested against outcomes. They also ignore deal age, source, and segment, which all shift the true rate.
Can a calibrated model still produce a wrong dollar forecast?
Yes, when the amounts are inflated. ORM illustrates the gap with a pipeline averaging 80000 dollars per open deal against closed-won deals averaging 40000 dollars. Correct probabilities on wrong amounts still double the forecast.
Put these metrics to work
ORM builds custom revenue forecast models that turn concepts like forecast model calibration into prescriptive action for your team.
Schedule a Demo