
Lifecycle Experiments
Lifecycle experiments and holdouts
Plan a lifecycle test with a clear customer outcome, fair assignment, a defined holdout and an honest readout of what changed.
A lifecycle experiment tests whether a journey change improves a customer outcome. Define eligibility, assign units to experiences before the change can affect them, and compare the same outcome over the same follow-up period.
A holdout receives the agreed comparison experience. It may miss a marketing message but still receives necessary account and service communication.
Start with the decision
Ask one question the result can answer. “Does this setup email help new accounts complete their first useful task?” needs an email versus no-email comparison. “Should it arrive sooner?” needs two schedules using the same message.
If the offer, channel and timing all change, the test compares two packages and cannot identify which change made the difference.
Choose a primary outcome tied to the customer’s task. Opens and clicks can diagnose delivery or engagement but do not show that the task was completed.
Record the event source, counting unit and follow-up window. If the outcome is an account purchase, count accounts even when several contacts receive messages.
Design choice / Record before launch
- Population
- Entry event, eligibility and exclusions
- Assignment
- Person, account or task; allocation and stable group identifier
- Comparison
- What each group may receive
- Outcome
- Verified event, counting unit and follow-up window
- Decision
- Improvement worth acting on, analysis rule and guardrails
Protect the comparison
Assign eligible units at a defined point before the journey can influence them, and keep each unit in its assigned group. Where contacts share an account and can influence one account outcome, assigning them to different groups can mix the experiences. Account-level assignment may be appropriate if account relationships are reliable.
People who did not open the email are not a holdout: opening may reflect motivation that also predicts the outcome. Keep customers who opt out, encounter a delivery failure or complete a task before a delayed send in their original groups for the primary comparison. Report sends, deliveries and skips separately.
Specify what is withheld. A one-email holdout tests that email within the surrounding programme; a journey-level holdout tests the defined marketing steps.
Check whether another campaign delivers the same prompt or offer to the holdout. A long-running holdout across several initiatives answers a broader question and needs a separate design.
Keep any compliance review separate from the experiment design. Keep necessary service communication on its appropriate path for both groups.
Implement the intended absence
A holdout must be absent from the tested marketing message, not merely labelled as one in a report. Generate a delivery for the holdout and route it to an internal message trap, so conversions can be compared with those of people who receive the message.
For a workflow test, a random cohort branch can route units into the message and holdout paths; additional paths can support more than two branches. In Customer.io, the holdout variation must use the Queue Draft or Send Automatically sending behaviour. The Don’t Send setting does not generate the holdout delivery that test needs.
Plan for an answer the audience can support
Estimate eligible volume and the baseline outcome from a comparable population. Choose the smallest effect that would change the decision, then calculate the sample and duration for the planned allocation and analysis.
A smaller holdout means fewer customers miss the proposed journey but usually gives a less precise comparison. If the required duration extends beyond a stable offer or product state, narrow the question or use other evidence; do not call an inconclusive test decisive.
Start each customer’s follow-up clock at the same kind of event and allow the full outcome window before including them in a final readout. Set the analysis rule in advance. Repeatedly stopping a conventional fixed-horizon test when an interim result looks favourable increases false-positive risk; a planned sequential method requires its own rules.
Make the planning trade-offs explicit
A power analysis connects the historical mean and variance of the outcome with eligible traffic, the minimum detectable effect, experiment duration and allocation. The minimum detectable effect is the smallest change the design can reliably detect; a smaller true effect is less likely to produce a statistically significant result.
More time generally brings more observations and tighter confidence intervals, while allocating more eligible traffic can reduce the minimum detectable effect. If a negative impact is a concern or experiments need mutually exclusive audiences, allocation may instead be constrained. Base the estimate on a population relevant to the planned experiment, such as a similar user base or product area from a past experiment.
Read the result with its limits
Check identifiers, entry rules and outcome feeds. Then inspect what each group actually received, including skips, other campaigns and sales or support contact. Sound assignment cannot rescue a comparison in which the intended experience difference disappeared.
Report the primary outcome count and rate for each assigned group, their absolute difference, uncertainty and follow-up window. Put guardrails such as complaints or wrong sends beside the result. If the estimate remains too uncertain to resolve the decision, say which useful gains or harms remain plausible.
Record the decision and its reason: keep the journey, test a specific revision, investigate a data problem or stop. Preserve the original question and preselected outcomes even when the result is unfavourable.
Check assignment health before interpreting lift
A sample ratio mismatch occurs when observed group sizes differ from the configured allocation by more than chance would explain. Statsig uses a chi-squared test on exposures to check whether observed assignment frequencies match the expected split. A persistent mismatch is a data-quality problem: the groups may no longer be comparable, so do not trust apparent metric lift until the cause is addressed.
Investigate whether exposure logging is losing units, such as when a client crashes before logging an exposure, or whether a conditional dependency exposes different kinds of units across groups. Also check for bulk exposure logging that is truncated before every group is recorded. These failures can skew the measured groups even when the configured allocation was correct.
In this guide
- Choosing a holdout size before launching a journey testChoose a journey holdout using eligible volume, baseline outcomes, the smallest useful effect and the time available for a fair comparison.
- Testing journey timing without changing the offerCompare two journey schedules fairly while keeping the offer fixed, using one assignment clock and a shared customer-outcome window.
- Distinguishing a message effect from natural customer progressionUse an eligible holdout and a shared outcome window to separate what a lifecycle message may add from progress customers make anyway.
- Recording a negative result from a lifecycle experimentDocument an unfavourable or inconclusive journey test with its original decision, uncertainty, integrity checks and next action.



