The question comes up in the first vendor call and gets answered badly in both directions. Some teams delay a forecasting project for a year because they think their CRM is too messy. Others connect a system with 60 closed deals and wonder why the predictions swing. The real requirement has three parts: depth of history, volume of closed outcomes, and consistency of the fields you already capture.
How much sales history does a forecast model need?
Enough completed sales cycles that the model sees a large number of resolved outcomes. A model learns by watching deals move from creation to a resolved outcome. What it needs is enough of those complete journeys, so the calendar requirement scales with your cycle length.A team closing 900 deals a year with a 45 day cycle has more usable signal after twelve months than an enterprise team closing 40 deals a year has after three. Count outcomes, not birthdays. The second reason to want depth is seasonality. Most teams see stronger Q2 and Q4 than Q1 and Q3, and the third month of a quarter outperforms the first and second. A model needs to see that pattern repeat at least twice before it can separate a seasonal dip from a real decline.
How many closed deals do you need per segment?
Enough closed outcomes per segment that one deal cannot move the rate. The mistake is not the total. It is the split.A company with 4,000 closed deals looks well supplied until you slice by segment, region, and product and discover the enterprise motion contributes 90 of them. Segments too thin to hold a stable rate should roll up into a parent group for modeling and be reported separately for management. Forecast at the level where the data supports it, not at the level of the org chart.
Does messy CRM data disqualify you?
No. Inconsistent data disqualifies you. Messy data does not. Everybody thinks their data is uniquely bad, and that belief stalls more forecasting projects than any technical constraint. It is not true, and it does not matter as much as people assume. As long as the data is wrong in a consistent way, a model can learn the pattern and predict accurately around it.The distinction matters in practice. If reps routinely enter amounts 40 percent higher than what closes, the model learns the ratio. If half the team started inflating amounts in March because the comp plan changed, the model learns nothing useful from either period. Consistency across time is the requirement. Cleanliness is a preference.
One consistent pattern worth checking before you start: compare average deal size in pipeline against average deal size on closed won. A pipeline averaging $80,000 that produces closed deals averaging $40,000 is not a data emergency. It is a known ratio, and a model handles it. What it cannot handle is that ratio changing every quarter for reasons nobody records.
Which fields carry the most weight?
Stage, amount, close date, create date, and owner, plus the timestamped history of every change to them. The history matters more than the current value.| Field | What it contributes | What breaks it |
|---|---|---|
| Stage | Position in the funnel and conversion base rates | Stage definitions that get renamed or resequenced |
| Amount | Deal size distribution and revenue weighting | Amounts entered once and never revised |
| Close date | Timing prediction and slippage signal | Bulk pushes at quarter end |
| Create date | Age, aging curves, and cycle length | Records backdated during migrations |
| Owner | Rep-level bias correction | Territory changes with no history preserved |
| Change history | The strongest single predictor set | Systems that overwrite instead of logging |
What counts as activity worth recording?
A change in stage, close date, or amount. Logged emails and calendar invites are weaker signals. Meaningful activity means something moved on the opportunity record.This definition also sets the rule for aging. Pipeline that has gone twelve months without a meaningful change is dead inventory, and more than 10 percent of a typical pipeline sits in that state. It is not a data gap the model needs to fill. It is a set of records that should be excluded from coverage math, which is one of several reasons a raw coverage ratio misleads. See pipeline coverage for how to calculate it once the dead inventory is out.
How long until the model is usable?
Four to six weeks from connection to a fully trained model based on your company's historical sales performance. Plan the rollout around the training window rather than the contract date. Run the model in parallel with your existing process through one full quarter, compare day-one predictions against actuals, and only then move the model into the forecast call.What if you have almost no history?
Run a bottom-up build as the primary forecast and let the model earn weight as outcomes accumulate. A model with thin history can still do useful work. It can group similar opportunities and predict how long each group takes to close, and those timing curves stabilize faster than revenue predictions do.What it cannot do yet is give you reliable segment-level revenue. Treat the early output as a second opinion on timing and risk, not as the number you commit to the board. Our guides to creating a sales forecast and forecasting revenue cover the manual build that carries you through the first few quarters.
The short version: enough completed cycles to produce a large number of resolved outcomes, enough closed outcomes per forecasted segment that one deal cannot move the rate, field history switched on, and consistency you can point to. Nobody needs perfect data. They need data that means the same thing in January as it does in October.
Frequently Asked Questions
How much sales history does a machine learning forecast need?
Enough closed deals to cover several full sales cycles. The requirement is measured in completed cycles and closed outcomes, not in calendar time, so a high-volume team clears the bar faster than an enterprise team with 40 deals a year.
Does our CRM data need to be clean before we start?
No. It needs to be consistent. Everybody believes their data is uniquely bad and that this is why they cannot run the business the way they want. It is not the blocker. Garbage in does not have to equal garbage out, because a model can learn from data that is wrong in the same direction every quarter.
How long does it take to train the model once data is connected?
Four to six weeks for a fully trained model built on your company's historical sales performance.
Which CRM fields matter most to a forecast model?
Stage, amount, close date, create date, owner, and the timestamped history of changes to each. The change history carries more predictive weight than the current values, because a close date that has moved three times tells you more than the date itself.
Can you forecast with machine learning if the company is under two years old?
Partially. With thin history a model can still group opportunities and predict close timing, but confidence intervals stay wide and segment-level predictions are unreliable. Run the model alongside a bottom-up build and give the model more weight each quarter as outcomes accumulate.
See how ORM turns these insights into action
ORM builds custom revenue forecast models for B2B SaaS companies. Not dashboards. Prescriptive analytics that tell you what to do next.
Schedule a Demo