
Lifecycle Experiments
Part of Lifecycle experiments and holdouts
Recording a negative result from a lifecycle experiment
Document an unfavourable or inconclusive journey test with its original decision, uncertainty, integrity checks and next action.
Record an unfavourable or inconclusive result with the original question, actual test conditions, outcome counts, estimated difference and uncertainty. “No lift” can mean an observed harm, a result too small to justify the journey, or too little information to decide. A useful record shows which conclusion the evidence supports.
Classify the result before naming it
Use the primary outcome and decision rule chosen before launch. A treatment rate below the comparison rate is an observed negative difference. Whether it supports a conclusion of harm depends on uncertainty and test quality.
If the interval includes both a worthwhile gain and harm, the result is inconclusive. If its upper bound is below the improvement needed to justify the journey, the test can rule out that worthwhile gain under the tested conditions, assuming the planned analysis and implementation were sound.
A non-significant result does not prove that two experiences are identical. A favourable secondary metric also does not replace a disappointing primary outcome after the fact. Report secondary findings as context and label unplanned comparisons exploratory.
Record field / What to keep
- Original decision
- Hypothesis, primary outcome and improvement worth acting on
- Population
- Eligibility, assignment unit and entry dates
- Experiences
- Intended and actual delivery in each group
- Result
- Assigned units, verified outcomes, rates, difference and uncertainty
- Integrity
- Allocation, missing data, cross-group exposure and material changes
- Decision
- Stop, keep, repair or retest, with an owner and reason
Critical experiment record fields to preserve
- Original decision
- Hypothesis, primary outcome and minimum worthwhile improvement
- Population
- Eligibility criteria, assignment unit and entry dates
- Experiences
- Intended and actual delivery in each group
- Result
- Assigned units, verified outcomes, rates, difference and uncertainty
- Integrity checks
- Allocation consistency, missing data, cross-group exposure, material changes
- Decision
- Stop, keep, repair or retest — with owner and reason
Separate the idea from the test’s validity
Check that the intended comparison occurred before interpreting customer behaviour. A persistent, statistically meaningful allocation mismatch may signal an assignment or logging problem.
A broken outcome feed, an offer that expired in one branch, or another campaign that reached the holdout can also undermine the result.
Mark the conclusion unresolved, preserve the observed record and identify the repair needed before drawing a causal conclusion.
If the test ran as designed, state its conclusion for the tested population, intervention and window.
“The delayed reminder did not show an improvement large enough to justify it for eligible new accounts within the planned follow-up” is more informative than “reminders do not work”. A different message, task or audience remains a separate question.
Make the next decision traceable
Record the planned sample and readout rule beside the achieved sample and actual stopping reason. If recruitment ended early, say why.
For a conventional fixed-horizon analysis, repeatedly stopping at a favourable interim result increases false-positive risk. Do not retrofit a threshold after seeing the data.
Include operational outcomes that matter to the decision, such as wrong sends, complaints or extra service work, without claiming the test established their causal effect unless those outcomes were measured and analysed for that purpose. An inconclusive primary result may still expose a fixable implementation fault.
End with one action and its reason: stop when the credible range of benefit cannot justify the cost or observed harm under the organisation’s criteria.
Retest a specific revision when the original comparison was sound but the estimate remains too uncertain.
Repair instrumentation and rerun when the test could not answer its question.
Keep the record discoverable so the same experiment is not repeated without its earlier result.


