
Lifecycle Experiments
Part of Lifecycle experiments and holdouts
Choosing a holdout size before launching a journey test
Choose a journey holdout using eligible volume, baseline outcomes, the smallest useful effect and the time available for a fair comparison.
Base the holdout size on the outcome you need to detect, the eligible audience you expect and the time available. A convenient percentage cannot replace a prospective sample calculation. A smaller holdout means fewer customers miss the proposed journey, but it usually gives a less precise comparison.
Define the test before the split
Define what the holdout misses: one promotional email or the whole marketing journey. Name the primary outcome and whether it belongs to a person or an account. For an account purchase, assign and count accounts, not every contact as an independent purchase opportunity.
Estimate the baseline outcome rate from a population resembling those who will qualify. For a numerical outcome, use its mean and variation. Whole-site traffic can misstate both the volume and outcome behaviour of a narrow journey audience. Estimate how many new eligible units will arrive while the journey, offer and outcome definition can remain stable.
Set the smallest effect that would change the decision. If a tiny gain would not justify running the journey, do not size the test to detect that tiny gain. Agree the analysis method, significance level and power before launch. These are planning choices, not numbers to choose after seeing the dashboard.
Compare feasible allocations
Calculate the sample requirement for the proposed split. For a fixed total audience, an even split generally estimates a difference more precisely than a very small holdout. A smaller holdout may be justified by the cost to customers of withholding the journey, but it still needs enough comparison observations to answer the question.
A planning worksheet should record:
Input / What to specify
- Eligible units
- Entry rule, counting unit and expected arrivals
- Baseline
- Outcome rate, or mean and variation, in a comparable population
- Worthwhile effect
- Smallest change that would alter the decision
- Allocation
- Journey and holdout shares
- Analysis
- Test method, significance level and power
- Follow-up
- Time each assigned unit needs before its outcome is final
Calculate the required units in each group and the likely calendar duration. Recruitment and outcome follow-up both take time: the last entrant may still need observation after the last journey message is sent.
Some platforms recommend small, long-running holdouts to measure the combined effect of many feature launches. That programme use case does not establish a suitable percentage for one marketing journey.
Key Inputs for Holdout Size Calculation
- Eligible Units
- Entry rule, counting unit (e.g., accounts), expected arrivals
- Minimum Detectable Effect
- Smallest meaningful change (e.g., +2% conversion)
- Required Sample Size
- Calculated based on power, significance, and effect size
Decide what happens if the test cannot finish
If the required sample cannot arrive before the offer, product or audience changes materially, record that limit before launch. Options include a larger still-relevant population, a more common outcome that genuinely supports the decision, or a larger minimum effect worth detecting. If none makes the test feasible, use other evidence for a narrower decision. Do not declare an underpowered result decisive.
Keep assignment and outcome definitions fixed during the test. Investigate a statistically meaningful, persistent mismatch between intended and observed allocation. Record opt-outs, skipped sends and competing campaigns so a weak treatment contrast is visible at readout.
There is no universal holdout percentage. Choose a split that respects the cost of withholding the journey and can resolve the stated decision within the available time.


