An incrementality holdout test measures what a marketing program actually caused. You randomly assign accounts or regions into a treated group that receives the program and a control group that does not, run both for a defined period, and compare the outcomes. The difference is incremental lift. Everything the control group produced without the program was going to happen anyway.
Why Attribution Cannot Answer This
Attribution models divide credit among the touches that appear on converted records. They have no view of the deals that would have converted with no marketing touch at all. Retargeting is the classic case: it reaches people who already visited your site, so it collects credit for buyers who were on their way regardless. A holdout is the only method that separates influence from coincidence.
Designing the Test for B2B
Three design choices decide whether the result means anything.
- Randomize at the account or geography level. B2B buying committees talk to each other. Splitting individual contacts inside the same account leaks treatment into the control. - Balance the groups on the variables that predict outcomes. Match on segment, employee count, and existing engagement before you split, so the two groups are comparable on day one. - Pick a measurable outcome inside the test window. Closed revenue is the outcome you care about, but a long cycle means you may need to measure opportunity creation or win rate on created opportunities instead.
Sizing and Duration
The test needs enough outcome events in the control group to detect the size of effect you care about. A holdout that produces four opportunities in control and six in treatment tells you nothing. Before launching, estimate your baseline conversion count for the control population over the planned window, and decide what lift would change your budget decision. If the population cannot produce enough events, expand the geography, extend the window, or accept that this program has to be judged another way.
Duration has to cover one full sales cycle plus reporting lag. ORM's opportunity data shows close-time curves running from 1 to 80 weeks with most expected volume before week 12, which means a program judged on closed revenue after a single quarter is being judged on a fraction of its output.
Reading the Result
Report lift as the difference in outcome rate between groups, with the absolute event counts shown next to it. A 29% lift on nine opportunities versus seven is not a finding. Feed confirmed lift back into budget allocation, and feed it into pipeline coverage planning only after the effect repeats in a second test. One clean experiment is evidence. One clean experiment repeated is a planning input, and it belongs alongside the rest of the demand assumptions in how to forecast revenue.
Frequently Asked Questions
How is a holdout test different from attribution?
Attribution divides credit for outcomes that already happened. A holdout test creates a counterfactual by withholding the program, which measures what would have happened without it. Only the second approach can prove a channel caused anything.
How long should a B2B holdout test run?
At least one full sales cycle plus the reporting lag, and long enough to accumulate enough opportunities in the control group to detect a difference. For teams with multi-quarter cycles, that usually means measuring an intermediate outcome such as opportunity creation.
What do you split on in a B2B holdout test?
Accounts or geographies, not individual contacts. Multiple buyers from the same account will see each other's exposure, which contaminates a contact-level split.
When is a holdout test the wrong tool?
When the program is small enough that any real effect would be smaller than normal quarter-to-quarter noise, or when withholding it damages a committed pipeline number you cannot afford to miss.
Put these metrics to work
ORM builds custom revenue forecast models that turn concepts like incrementality holdout test into prescriptive action for your team.
Schedule a Demo