Holdout Testing
Holdout Testing is an experiment that withholds marketing from a randomly chosen group so its results can be compared against an exposed group to measure true incremental impact.
Also known as: control group testing, suppression test, PSA holdout
Holdout Testing deliberately excludes a randomized portion of your audience from a campaign or program. The withheld group serves as a control, revealing what would have happened with no marketing at all. The difference between exposed and withheld outcomes is the genuine incremental effect of the activity, making it one of the most credible ways to prove causal impact.
What Holdout Testing Means
Holdout Testing is a controlled experiment in which a random subset of an audience is suppressed from receiving a marketing intervention. The withheld group serves as a control whose outcomes show what would have happened without marketing. Because the random assignment controls for everything except exposure, the difference in outcomes is the genuine incremental effect of the activity. This is the gold standard for proving incremental impact and the foundation for defensible ROAS claims that survive a CFO review.
How Holdout Testing Works
The audience is randomly split, one group exposed and the other suppressed, then outcomes are compared after the conversion cycle completes. Sample size depends on conversion rates and the size of the lift expected to be detectable; pre-test power calculations should set it, not a default percentage. Holdout sizes of 5 to 20 percent of the audience are common. Test duration should cover a full conversion cycle plus a statistical confidence buffer, often 4 to 12 weeks for direct response and 6 to 12 months for brand or always-on programs with delayed effects.
Common Pitfalls and Misconceptions
The tension is that holdouts require deliberately not marketing to some prospects, which feels like leaving money on the table. In practice the cost of a small holdout is modest, and the clarity it provides about real incremental value usually outweighs the foregone short-term revenue. Teams that resist holdouts on principle often end up unable to defend their largest spend lines when finance asks how they know the spend works. The second pitfall is running too short a test, which catches only immediate response and misses the lift the program was designed to produce over the conversion cycle.
Holdout Testing in Practice
The practitioner sophistication in B2B is matched-market holdout. Pure random holdouts at the account level can be impractical when account universes are small, so many B2B teams hold out by geography (suppressing campaigns in one region matched against another similar region) or by account cohort (holding out half of a segmented tier). The methodology is similar but less statistically clean than pure randomization, which is why these tests need larger sample sizes and longer durations to draw confident conclusions. They are still better than no holdout, and the cleanest programs run at least one holdout per major spend line per year on a documented calendar.
Frequently asked questions
-
How large should a holdout group be?
Large enough to detect a meaningful effect with statistical confidence, often 5 to 20 percent of the audience. The exact size depends on your conversion rates and the size of the lift you expect to see. Pre-test power calculations should set the size, not a default percentage.
-
How is holdout testing different from A/B testing?
A/B testing compares two versions of marketing against each other to find which performs better. Holdout testing compares marketing against no marketing, so it measures whether the activity creates incremental value at all, not just which variant performs better. Both are useful, but they answer different questions.
-
What can a holdout test prove?
It establishes causation, not just correlation. Because the only difference between exposed and withheld groups is exposure to the campaign, any outcome gap can be confidently attributed to the marketing itself. This is the gold standard for proving incremental impact and the foundation for defensible ROAS claims.
-
When should I use a holdout?
Use holdouts when you need to validate the true incremental contribution of a channel or program, such as retargeting, brand campaigns, or always-on nurture, where attribution models tend to overstate impact. They are especially valuable before a major budget decision or platform renewal.
-
Are holdout tests practical in B2B?
They can be, though small account universes make randomization harder. Many B2B teams run holdouts at the contact level within large nurture programs or at the account level for sizable target lists. Geo-holdouts (matched-market) are a common alternative when account counts are too low for pure random assignment.
-
What is a matched-market holdout?
A matched-market holdout suppresses marketing in one geography and measures lift against a similar control geography. It is the standard alternative when pure account-level randomization is impractical, common for TV, OOH, and broad-reach brand campaigns. The match quality between markets is critical to the validity of the result.
-
How long should a holdout test run?
Long enough to cover a full conversion cycle plus statistical confidence buffer, often 4 to 12 weeks for direct response and 6 to 12 months for brand or always-on programs with delayed effects. Running shorter than the conversion cycle catches only immediate response and misses the lift the program is designed to produce.