Holdout size for journey tests: Base holdout on detectable effect, eligible audience and time available; Use a sample calculation—no fixed percentage works for all tests; Smaller holdouts reduce customer impact but need enough data to compare accurately
Image: Lifecycle Marketing Lab

Lifecycle Experiments

Part of Lifecycle experiments and holdouts

Choosing a holdout size before launching a journey test

Choose a journey holdout using eligible volume, baseline outcomes, the smallest useful effect and the time available for a fair comparison.

Base the holdout size on the outcome you need to detect, the eligible audience you expect and the time available. A convenient percentage cannot replace a prospective sample calculation. A smaller holdout means fewer customers miss the proposed journey, but it usually gives a less precise comparison.

Define the test before the split

Define what the holdout misses: one promotional email or the whole marketing journey. Name the primary outcome and whether it belongs to a person or an account. For an account purchase, assign and count accounts, not every contact as an independent purchase opportunity.

Estimate the baseline outcome rate from a population resembling those who will qualify. For a numerical outcome, use its mean and variation. Whole-site traffic can misstate both the volume and outcome behaviour of a narrow journey audience. Estimate how many new eligible units will arrive while the journey, offer and outcome definition can remain stable.

Set the smallest effect that would change the decision. If a tiny gain would not justify running the journey, do not size the test to detect that tiny gain. Agree the analysis method, significance level and power before launch. These are planning choices, not numbers to choose after seeing the dashboard.

Compare feasible allocations

Calculate the sample requirement for the proposed split. For a fixed total audience, an even split generally estimates a difference more precisely than a very small holdout. A smaller holdout may be justified by the cost to customers of withholding the journey, but it still needs enough comparison observations to answer the question.

A planning worksheet should record:

Input / What to specify

Eligible units
Entry rule, counting unit and expected arrivals
Baseline
Outcome rate, or mean and variation, in a comparable population
Worthwhile effect
Smallest change that would alter the decision
Allocation
Journey and holdout shares
Analysis
Test method, significance level and power
Follow-up
Time each assigned unit needs before its outcome is final

Calculate the required units in each group and the likely calendar duration. Recruitment and outcome follow-up both take time: the last entrant may still need observation after the last journey message is sent.

Some platforms recommend small, long-running holdouts to measure the combined effect of many feature launches. That programme use case does not establish a suitable percentage for one marketing journey.

Key Inputs for Holdout Size Calculation

Eligible Units
Entry rule, counting unit (e.g., accounts), expected arrivals
Minimum Detectable Effect
Smallest meaningful change (e.g., +2% conversion)
Required Sample Size
Calculated based on power, significance, and effect size

Decide what happens if the test cannot finish

If the required sample cannot arrive before the offer, product or audience changes materially, record that limit before launch. Options include a larger still-relevant population, a more common outcome that genuinely supports the decision, or a larger minimum effect worth detecting. If none makes the test feasible, use other evidence for a narrower decision. Do not declare an underpowered result decisive.

Keep assignment and outcome definitions fixed during the test. Investigate a statistically meaningful, persistent mismatch between intended and observed allocation. Record opt-outs, skipped sends and competing campaigns so a weak treatment contrast is visible at readout.

There is no universal holdout percentage. Choose a split that respects the cost of withholding the journey and can resolve the stated decision within the available time.

More from Lifecycle Experiments

Lifecycle Experiments

Testing journey timing without changing the offer

Compare two journey schedules fairly while keeping the offer fixed, using one assignment clock and a shared customer-outcome window.