Optimized Sales Optimized Marketing Target Accounts For CROs For CFOs For CMOs Blog News Glossary Compare Tools About Schedule a Demo
Sales Forecasting

How to Build a Forecast Accuracy Scorecard for a B2B SaaS Team

Pete Furseth 6 min read
forecast accuracyforecast trackingRevOps reportingsales forecasting
How to Build a Forecast Accuracy Scorecard for a B2B SaaS Team
Home/ Blog/ How to Build a Forecast Accuracy Scorecard for a B2B SaaS Team

Most revenue teams review the forecast every week and review forecast accuracy almost never. A number gets submitted, the quarter closes, and the gap between the two disappears into a slide nobody opens again. A scorecard fixes that by turning error into a tracked metric with an owner, a cadence, and a history you can argue with.

What is a forecast accuracy scorecard?

A forecast accuracy scorecard is a fixed set of error metrics, captured at fixed points in the quarter, tracked at rep, segment, and roll-up level, and reviewed on a published schedule. The words "fixed" and "published" carry the weight. Teams that recalculate accuracy after the fact, using whichever snapshot flatters the quarter, produce a number that cannot be compared across periods. The scorecard exists so that a Q3 miss and a Q1 miss mean the same thing.

The scorecard is a measurement instrument, not a management tool. It tells you the size and direction of your error. It does not tell you why. Diagnosis comes after, and it comes from different data.

Put this to work on your numbers
Run your own numbers with the free Forecast Accuracy Scorecard, then see how ORM builds it into a custom model.

Which four metrics belong on the scorecard?

Absolute error, signed bias, hit rate, and category conversion. Everything else is a derivative of those four.
MetricQuestion it answersHow to calculateReview level
Absolute percentage errorHow big are our misses?Average of \actual minus forecast\divided by actualRoll-up and segment
Signed biasWhich way do we lean?Average of (forecast minus actual) divided by actual, keeping the signRep and roll-up
Hit rateHow often do we land in the band?Percent of periods inside your tolerance bandRoll-up
Category conversionDo our categories mean anything?Closed-won dollars divided by dollars in each category at snapshotRep and segment
Category conversion is the one most teams skip, and it is the one that exposes broken forecast language fastest. If commit converts at 78 percent for one rep and 96 percent for another, the word "commit" has two meanings on the same team. You can read more on how the categories should be defined in our guide to sales forecasting.

When in the quarter should you take the snapshots?

Take four snapshots: day one, end of month one, end of month two, and the Friday before close. Comparing a day-one forecast against a week-twelve forecast without labeling which is which is the most common way accuracy tracking gets corrupted.

Day one matters more than the others. Getting the forecast right in the final week of the quarter does not help anybody, because by then the quarter has already happened. The value of a forecast is knowing the shape of the quarter early enough to change it. A team whose week-twelve accuracy is excellent and whose day-one accuracy is poor does not have a forecasting capability. It has a reporting capability.

On new and expansion business, accuracy around 90 percent is a typical result, but it usually takes heavy manual effort to produce and it stops being true the moment conditions shift. ORM targets 95 percent without manual adjustment and holds it from day one through day ninety, updating as the quarter progresses. The model behind that number trains on your own historical sales performance and takes four to six weeks to train fully.

How do you score reps without punishing honesty?

Score direction and consistency across quarters, never the size of a single quarter's miss. A rep who misses by 15 percent in one direction and 15 percent in the other across two quarters has noise. A rep who misses by 8 percent in the same direction for six straight quarters has bias, and bias is correctable.

Publish rep-level signed bias. Keep rep-level absolute error internal to the manager. Absolute error at the rep level is dominated by deal size lumpiness, and ranking reps on it rewards whoever happened to own smaller deals. Signed bias is the number that reflects judgment.

One rule protects the whole system. When a rep raises their commit and misses, the coaching conversation is about the deal. When a rep lowers their commit and beats it, the coaching conversation is about the pattern. Reversing that teaches everyone to sandbag by the second quarter of tracking.

How do you tell a data problem from a judgment problem?

Sort the errors by what they correlate with. If error correlates with the individual, it is judgment. If error correlates with stage, deal size, segment, or age, it is structure.

Structural error has tells. Deals sitting in pipeline with close dates inside the quarter close far less often than people assume, and stale opportunities inflate the denominator on every coverage ratio you calculate. Pipeline that has not been touched in twelve months is dead weight, and a meaningful touch means a change in stage, close date, or amount, not a logged email. Strip the untouched inventory out before you blame a rep for a bad call.

Also check average deal size on both sides of the line. A pipeline carrying an average deal size of $80,000 that produces closed-won deals averaging $40,000 will miss every quarter regardless of who owns the deals. That is arithmetic, not judgment. See deal slippage for the aging patterns that drive it.

How often should the scorecard be reviewed?

Monthly with sales leadership, quarterly with the full team, and never in the same meeting as the forecast call. Mixing the two is how accuracy review becomes a negotiation about this week's number.

The monthly review has one job: decide whether last month's error was structural or behavioral, and assign the fix. The quarterly review has a different job: check whether the fixes moved the metric. Both are short. A scorecard review that runs longer than thirty minutes has become a deal review in disguise.

What breaks a forecast accuracy scorecard?

Changing the definitions mid-year. Every recalibration of a stage, every reassignment of a territory, every change to what counts as commit resets your history. Batch definition changes into one annual event, document the date, and mark it on the chart so nobody compares across the line without knowing.

The second failure mode is measuring only the roll-up. Company-level accuracy hides offsetting errors. One segment over-forecasting by 20 percent and another under-forecasting by 20 percent produces a perfect roll-up and two broken forecasts. Track segments separately or you will keep congratulating yourself on luck.

Start with two quarters of history, four snapshots, and the four metrics above. Add the forecast accuracy definitions to your operating docs so the terms hold still, and revisit the pipeline coverage inputs once the scorecard shows you where the error actually lives.

Frequently Asked Questions

What goes on a forecast accuracy scorecard?

Four metrics cover it. Absolute percentage error tells you how large your misses are. Signed bias tells you which direction they lean. Hit rate tells you how often you land inside your tolerance band. Category conversion tells you what share of commit, best case, and pipeline actually closed. Every metric is tracked at rep, segment, and company roll-up.

How far back should a forecast accuracy scorecard go?

Eight quarters is enough to separate a pattern from a bad quarter. Four quarters will show you direction but will not survive a territory change or a comp plan change in the middle of the window. If you only have four, start there and keep adding.

Should forecast accuracy be tied to compensation?

No, at least not in the first year of tracking. Paying on accuracy teaches reps to submit numbers they can hit rather than numbers they believe. Track it, publish it, coach against it, and leave it out of the comp plan until the underlying data and the stage definitions are stable.

What is a reasonable tolerance band for hit rate?

Set the band against your own historical error distribution rather than a borrowed figure. Rep-level bands need to be wider because rep-level revenue is lumpier. Set it once, write it down, and stop renegotiating it mid-quarter.

Who owns the forecast accuracy scorecard?

RevOps builds and publishes it. Sales leadership acts on it. If the sales team owns both the production and the interpretation of its own accuracy scores, the definitions drift toward whatever makes the current quarter look reasonable.

PF
Pete Furseth
ORM Technologies
Pete has built custom revenue forecast models for B2B SaaS companies for over a decade.

See how ORM turns these insights into action

ORM builds custom revenue forecast models for B2B SaaS companies. Not dashboards. Prescriptive analytics that tell you what to do next.

Schedule a Demo