Optimized Sales Optimized Marketing Target Accounts For CROs For CFOs For CMOs Blog News Glossary Compare Tools About Schedule a Demo
Revenue Operations

How to Backtest a Sales Forecast Model Before You Trust It

Pete Furseth 6 min read
predictive sales analyticsforecast accuracymodel validationRevOpspredictive analytics
How to Backtest a Sales Forecast Model Before You Trust It
Home/ Blog/ How to Backtest a Sales Forecast Model Before You Trust It

A forecasting model that has never been scored against completed quarters is a hypothesis. Backtesting turns it into evidence. The procedure is not complicated, but three details decide whether the result means anything: which snapshots you score, whether future information leaked into the test, and whether you kept the quarters that went badly.

What is backtesting a sales forecast model?

A backtest reruns the model across completed quarters using only the data that existed at each point in time, then compares the prediction to what actually closed. The model does not get to know the answer.

The distinction between a backtest and a demo matters here. A demo shows what the model says about your current open pipeline, which nobody can verify for another eleven weeks. A backtest shows what it would have said last January, against a number you already have in the bank.

Put this to work on your numbers
Run your own numbers with the free Forecast Accuracy Scorecard, then see how ORM builds it into a custom model.

How much history does a valid backtest need?

Eight completed quarters. Four is the floor and it comes with caveats. Eight periods is enough to separate a repeatable pattern from one strange quarter, and it usually captures at least one full seasonal cycle.

Seasonality is a specific reason for the depth. Most B2B SaaS teams see Q2 and Q4 outperform Q1 and Q3, and the third month of any quarter outperforms the first two. A four-quarter test that happens to start in a strong period will read as a model that runs hot. Two years of periods lets the seasonal shape cancel out.

What do you measure at each snapshot?

Four metrics at four points in the quarter, scored identically every time. Fix the snapshots before you run anything, because the temptation to pick a flattering measurement point arrives the moment the first result appears.
SnapshotWhat it testsMetric to record
Day one of the quarterWhether the model can shape the quarter in advanceSigned error and absolute percentage error
End of month oneWhether early in-quarter creation is modeledAbsolute percentage error
End of month twoWhether the model reacts to slippageAbsolute percentage error and direction change
Friday before closeWhether the model convergesAbsolute percentage error
Record signed error at day one, not only the absolute value. The sign tells you whether the model runs optimistic or conservative, and a model with consistent direction is correctable through calibration. A model with random direction is not. See forecast accuracy for how each metric is defined.

Why does day-one error matter more than week-twelve error?

Because a forecast that is right at the end of the quarter has no operational value. By the final week the quarter has already happened, and nothing in the number can change the outcome.

The value of a forecast is knowing the likely shape of the quarter on day one, early enough to act. That is also the hardest prediction to make, which is why day-one error is the metric that separates products. Around 90 percent accuracy on new and expansion revenue is typical across the market, though most teams reach it through manual effort and lose it whenever conditions shift. ORM targets 95 percent without manual adjustment and holds it from day one through day ninety, updating as the quarter progresses.

How do you avoid leaking future information?

Reconstruct every field as of the snapshot date, including stage, amount, close date, and owner. Scoring a day-one prediction against today's field values is the most common way a backtest lies.

The leak is easy to miss because CRM reports show current values by default. A deal that was in stage two on January 2 and closed in March now reads as closed won in stage six. Feed that record into a day-one backtest and the model appears to have known the outcome. Rebuild the snapshot from field history, and if field history is not available for the period, drop that period rather than approximating it.

A second leak comes from opportunities created after the snapshot date. In-quarter creation is a legitimate source of revenue, and a good model predicts it as a category. What it cannot do is include the specific deals that were created later. Keep the two separate: predicted in-quarter creation is an output, and the actual records are the answer key.

How do you separate model error from structural error?

Sort the misses by what they correlate with. If error tracks a rep, it is judgment. If error tracks stage, deal size, segment, or age, the structure of your pipeline is producing it and no model will fix it for you.

Two structural patterns show up in nearly every backtest. The first is aged inventory. More than 10 percent of a typical pipeline has gone twelve months without a change in stage, close date, or amount, and it inflates the denominator of every ratio you compute. The second is a deal size gap. A pipeline averaging $80,000 that produces closed deals averaging $40,000 will read as over-forecasting quarter after quarter, which is arithmetic rather than a model defect. Strip aged records and check both averages before you judge any prediction. The aging patterns are covered in deal slippage.

What results should make you reject a model?

Three findings are disqualifying. First, day-one error no better than the process you already run. Second, error concentrated in one segment while the roll-up looks fine, which means offsetting mistakes are hiding two broken forecasts. Third, strong performance in stable quarters and a collapse in the quarter something changed.

The third is the important one. Forecasts miss because something in the business or the market moved and the model still runs on old assumptions. A competitor arrives and pricing pressure cuts average deal size. Capital tightens and win rates fall. Uncertainty stretches cycles from qualified to closed. Territories get redrawn and execution slips while coverage still looks healthy. Any model can predict a quarter that resembles the last one. You are buying the quarters that do not.

If you have a quarter in your history where the business visibly changed, score it separately and weight it heavily. A vendor asking you to exclude it is answering the question for you.

How often should you re-run the backtest?

Every quarter, adding the completed period and dropping nothing. A backtest is not a purchase gate you clear once. It is the running record of whether the model still describes your business.

Recalibrations, territory changes, and stage redefinitions all reset comparability. Batch those changes into one annual event, mark the date on the chart, and never compare across the line without labeling it. The rest of the operating discipline around this sits in our sales forecasting best practices guide, and the shared vocabulary lives in sales forecasting.

Frequently Asked Questions

What is backtesting a sales forecast model?

Backtesting reruns a forecasting model against completed quarters using only the information that existed at the time, then scores what the model would have predicted against what actually closed. It answers whether the model would have helped you, using periods where the outcome is already known.

How many quarters should a forecast backtest cover?

Eight quarters separates a pattern from a fluke. Four shows direction but will not survive a territory change or a comp plan change inside the window. If you only have four, start there and keep extending as quarters complete.

What is data leakage in a forecast backtest?

Leakage happens when the model sees information that did not exist at the moment it is supposed to be predicting. The most common form is scoring a day-one forecast using current field values instead of the values as of day one. It makes a mediocre model look excellent and is the single most common reason a backtest overstates accuracy.

Should a backtest be scored at quarter end or at the start of the quarter?

Start of the quarter is the score that matters. Every model looks accurate in the final week because the deals have already resolved. A model that is excellent at week twelve and poor at day one is a reporting system, not a forecasting system.

What backtest result should make you reject a model?

Reject it if day-one error is materially worse than your current process, if error is concentrated in one segment rather than spread, or if the model performed well in stable quarters and collapsed in the quarter your business changed. The last one is disqualifying, because change is what you are buying protection against.

PF
Pete Furseth
ORM Technologies
Pete has built custom revenue forecast models for B2B SaaS companies for over a decade.

See how ORM turns these insights into action

ORM builds custom revenue forecast models for B2B SaaS companies. Not dashboards. Prescriptive analytics that tell you what to do next.

Schedule a Demo