Decide Metrics Before Testing
MinutesBefore an A/B test or pilot starts, write down the primary metric, the harm metrics that must not get worse, the sample size or stopping rule, and the planned analysis, so that the result can tell you something you did not already want to hear.
How to do it
- 1Write a one-page plan before any data arrive: the hypothesis, the single primary metric, the minimum effect that would matter, the sample size or duration, and the rule for stopping. Date it and store it where it cannot be quietly edited.
- 2Choose a primary metric that reflects real value to the person on the other end, not only the next click. Completed sign-ups are easy to raise by hiding the price; retained, satisfied customers are not.
- 3Name guardrail metrics in advance: refunds, cancellations, complaints, unsubscribe and spam reports, support contacts, chargebacks, and outcomes for vulnerable segments. State the level of deterioration that stops the variant, however well the primary metric does.
- 4Fix the stopping rule. Checking results daily and stopping when they look good inflates false positives dramatically. If you need to peek, use a sequential method designed for it.
- 5List planned subgroup analyses before you look. Treat anything found afterwards as a hypothesis for the next test, not as a finding.
- 6Report everything you ran: all variants, all metrics, and the tests that showed nothing. A team that remembers only its wins will steadily overestimate what works.
- 7Run a longer holdout for changes that might trade long-term trust for short-term conversion. Many manipulative patterns win two-week tests and lose the year.
When to use it
- •Any A/B test, pilot, message test, or campaign evaluation whose result will be used to make a decision or a claim.
- •An optimization program is producing a steady stream of "wins" that do not show up in overall results.
- •A variant that improves conversion relies on urgency, friction, defaults, or obscured information.
- •You are reading someone else's case study or vendor claim and want to know what questions to ask.
Counters
Evidence and how strong it is
Simmons, Nelson and Simonsohn (2011) showed by simulation and by demonstration that a few common, individually defensible choices made after seeing data (adding observations, choosing among outcome measures, including covariates, dropping conditions) can raise the false-positive rate from the nominal five percent to more than sixty percent, and recommended deciding and disclosing these choices in advance. Preregistration is now standard in clinical trials and increasingly so in psychology (Nosek et al. 2018); in medicine, comparisons of registered protocols with published reports have repeatedly documented outcome switching. The same problems occur in commercial experimentation: Kohavi, Tang and Xu (2020), drawing on large-scale practice at technology companies, describe the inflation caused by peeking and multiple comparisons, recommend a predefined overall evaluation criterion, and stress guardrail metrics because short-term gains frequently come at long-term cost. Campbell's and Goodhart's laws describe the broader danger that an indicator used for decisions becomes corrupted by the pressure to move it. Evidence strength: strong statistical and meta-scientific evidence for the problem and for pre-specification as the remedy; evidence that pre-specifying guardrails reduces the shipping of manipulative designs is practitioner experience only.
- A plan is not a straitjacket. You can deviate, and you can explore, provided you say so and label exploratory results as exploratory. Hidden flexibility is the problem, not flexibility.
- A pre-specified metric can still be the wrong metric. Deciding in advance to maximize a number that rewards deception makes the deception more rigorous, not more ethical; the choice of metric is where the ethics enters.
- Statistical significance is not importance. A reliable effect that is too small to matter is a reason to stop testing that idea.
- Experiments on people raise their own ethical questions. Variants that could cause real harm, deceive participants about material facts, or target vulnerable groups need review before they run, not a guardrail that stops them afterwards.
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366The demonstration that undisclosed analytic flexibility can raise false-positive rates above sixty percent, and the recommendation to fix and disclose stopping rules and measures in advance.
- Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600-2606The case for preregistration as the way to keep the distinction between prediction and after-the-fact explanation.
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University PressPractitioner guidance on overall evaluation criteria, guardrail metrics, peeking, multiple comparisons, and long-term holdouts in commercial experimentation.
- Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67-90The statement of Campbell's law: the more a quantitative indicator is used for decision-making, the more subject it is to corruption pressures.