Background
Product metrics change for many reasons at once: seasonality, marketing, customer mix, incidents, other releases, and chance. An outcome review needs a way to distinguish a coincident trend from the effect of the change being reviewed.
Receives the new experience.
Outcome: 42%Receives the prior experience under the same conditions.
Outcome: 39%The difference is not automatically proof; assignment, exposure, telemetry, statistical uncertainty, and guardrails still matter. But random assignment makes the two groups much more comparable than a before/after chart.
External evidence
Research on online controlled experiments describes randomized control and treatment groups as a strong design for identifying a causal relationship between a change and user-observable behavior. It also documents that plausible, expert-backed changes often produce surprising or opposite results. Work on Microsoft’s experimentation platform emphasizes that trust depends on the entire chain: experiment execution, logging, and analysis—not only the final number.
A holdout is organizational memory of the road not taken. It keeps a small, live counterfactual available long enough to stop a post-launch narrative from hardening into “proof.” The practical review artifact is a claim bundle: outcome, comparison, guardrail, and a trust check.
Use it
| Part | Ask before release |
|---|---|
| Outcome | What customer or business result should change? |
| Counterfactual | Can a randomized holdout or staged comparison be retained? |
| Guardrail | What harm would make an apparent win unacceptable? |
| Trust | How will assignment, exposure, and telemetry be checked? |
For a change that cannot have a holdout—such as a security fix or an irreversible migration—record that limitation. Use the closest defensible comparison and call the conclusion evidence-informed, not causally proven.