Decide Metrics Before Testing

Minutes

Before an A/B test or pilot starts, write down the primary metric, the harm metrics that must not get worse, the sample size or stopping rule, and the planned analysis, so that the result can tell you something you did not already want to hear.

How to do it

  1. 1Write a one-page plan before any data arrive: the hypothesis, the single primary metric, the minimum effect that would matter, the sample size or duration, and the rule for stopping. Date it and store it where it cannot be quietly edited.
  2. 2Choose a primary metric that reflects real value to the person on the other end, not only the next click. Completed sign-ups are easy to raise by hiding the price; retained, satisfied customers are not.
  3. 3Name guardrail metrics in advance: refunds, cancellations, complaints, unsubscribe and spam reports, support contacts, chargebacks, and outcomes for vulnerable segments. State the level of deterioration that stops the variant, however well the primary metric does.
  4. 4Fix the stopping rule. Checking results daily and stopping when they look good inflates false positives dramatically. If you need to peek, use a sequential method designed for it.
  5. 5List planned subgroup analyses before you look. Treat anything found afterwards as a hypothesis for the next test, not as a finding.
  6. 6Report everything you ran: all variants, all metrics, and the tests that showed nothing. A team that remembers only its wins will steadily overestimate what works.
  7. 7Run a longer holdout for changes that might trade long-term trust for short-term conversion. Many manipulative patterns win two-week tests and lose the year.

When to use it

  • Any A/B test, pilot, message test, or campaign evaluation whose result will be used to make a decision or a claim.
  • An optimization program is producing a steady stream of "wins" that do not show up in overall results.
  • A variant that improves conversion relies on urgency, friction, defaults, or obscured information.
  • You are reading someone else's case study or vendor claim and want to know what questions to ask.

Counters

Evidence and how strong it is

Simmons, Nelson and Simonsohn (2011) showed by simulation and by demonstration that a few common, individually defensible choices made after seeing data (adding observations, choosing among outcome measures, including covariates, dropping conditions) can raise the false-positive rate from the nominal five percent to more than sixty percent, and recommended deciding and disclosing these choices in advance. Preregistration is now standard in clinical trials and increasingly so in psychology (Nosek et al. 2018); in medicine, comparisons of registered protocols with published reports have repeatedly documented outcome switching. The same problems occur in commercial experimentation: Kohavi, Tang and Xu (2020), drawing on large-scale practice at technology companies, describe the inflation caused by peeking and multiple comparisons, recommend a predefined overall evaluation criterion, and stress guardrail metrics because short-term gains frequently come at long-term cost. Campbell's and Goodhart's laws describe the broader danger that an indicator used for decisions becomes corrupted by the pressure to move it. Evidence strength: strong statistical and meta-scientific evidence for the problem and for pre-specification as the remedy; evidence that pre-specifying guardrails reduces the shipping of manipulative designs is practitioner experience only.

Cautions
  • A plan is not a straitjacket. You can deviate, and you can explore, provided you say so and label exploratory results as exploratory. Hidden flexibility is the problem, not flexibility.
  • A pre-specified metric can still be the wrong metric. Deciding in advance to maximize a number that rewards deception makes the deception more rigorous, not more ethical; the choice of metric is where the ethics enters.
  • Statistical significance is not importance. A reliable effect that is too small to matter is a reason to stop testing that idea.
  • Experiments on people raise their own ethical questions. Variants that could cause real harm, deceive participants about material facts, or target vulnerable groups need review before they run, not a guardrail that stops them afterwards.
  1. Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366
    The demonstration that undisclosed analytic flexibility can raise false-positive rates above sixty percent, and the recommendation to fix and disclose stopping rules and measures in advance.
  2. Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600-2606
    The case for preregistration as the way to keep the distinction between prediction and after-the-fact explanation.
  3. Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press
    Practitioner guidance on overall evaluation criteria, guardrail metrics, peeking, multiple comparisons, and long-term holdouts in commercial experimentation.
  4. Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67-90
    The statement of Campbell's law: the more a quantitative indicator is used for decision-making, the more subject it is to corruption pressures.
Last reviewed

More in Campaign integrity (for teams that persuade)

TARES Self-Audit

Before a campaign, advertisement, fundraising appeal, or pitch goes out, test it against the five TARES duties: Truthfulness of the message, Authenticity of the persuader, Respect for the audience, Equity of the appeal, and Social responsibility for the common good.

The Front-Page Test

Ask whether you would be comfortable seeing the tactic, including how it works and why you chose it, described accurately on the front page of a newspaper read by your audience; if the tactic only works when the audience does not know about it, treat that as a finding.

Harm-Class Inventory

List every persuasive element in a campaign, tag each one as neutral craft, dual-use, or manipulative by design, and for each dual-use element write down where the line is and which side of it you are on.

Disclosure Checklist

Before publishing sponsored, affiliated, incentivized, or endorsed content, check that every material connection between the speaker and the brand or cause is disclosed clearly, conspicuously, in plain language, and in the same place and format as the claim it qualifies.

Vulnerable-Audience Screen

Before launch, ask who will actually receive the message, which of them are least able to evaluate or resist it, and what it does to them; then change the targeting, the tactic, or the safeguards so that the campaign does not get its results from the people least able to say no.

Red-Team Your Campaign

Before launch, give people who did not build the campaign the explicit job of attacking it as a skeptical journalist, a regulator, a competitor, a harmed customer, and a bad actor would, and fix what they find while it is still cheap to fix.