LogicalMANIPULATIVE

P-Hacking

What it is

Trying many analytic choices — outcomes, subgroups, covariates, exclusions, stopping rules — and reporting the one that crossed the significance threshold as though it were the only test planned.

How it works

A p-value below 0.05 means that a result this extreme would arise by chance about one time in twenty if nothing were going on — for one test, specified in advance. Flexibility destroys the guarantee. Simmons, Nelson and Simonsohn showed in 2011 that four ordinary choices — measuring two outcomes and reporting either, adding ten participants if the first batch fell short, including or dropping a covariate, and dropping one experimental condition — raise the false-positive rate from 5 percent to about 61 percent when combined; with them, they published evidence that listening to “When I'm Sixty-Four” made undergraduates a year and a half younger. Gelman and Loken added that no one needs to run many tests: it is enough that the analysis was chosen after seeing the data, because a different dataset would have prompted a different analysis — the garden of forking paths. The result is not fraud in the ordinary sense; each choice is defensible on its own. The tell is a confirmatory frame on an exploratory search: an unexpected subgroup, an outcome not in the registration, a sample size that stopped exactly when significance arrived, or a p-value that sits just under the line. Head and colleagues found that bump in the published literature. Corporate A/B tests that are checked daily and stopped when the winner appears have the same structure.

Real-world examples

  • Simmons, Nelson and Simonsohn (2011) reported a real experiment in which undergraduates who listened to a Beatles song were significantly younger than those who listened to a control track — a demonstration of what undisclosed flexibility can produce from noise.
  • In 2016 Brian Wansink of Cornell's Food and Brand Lab described, in a blog post celebrating a visiting student's diligence, how a null dataset was re-sliced until it yielded several publishable papers; the scrutiny that followed led to more than a dozen retractions and his resignation in 2019.
  • Bem's 2011 paper in the Journal of Personality and Social Psychology reporting evidence of precognition used analytic choices that were standard in the field; its failure to replicate became a founding case of the replication crisis and of the argument that the standards, not the author, were the problem.
  • Head and colleagues (2015) examined the distribution of reported p-values across scientific fields and found an excess of values just below 0.05, the signature of results pushed over the threshold.
  • Online experiments that are monitored continuously and declared when a variant reaches significance produce a large share of false winners; the practice is common enough that experimentation platforms sell “peeking-safe” sequential tests to prevent it.

Ethical guidelines

  • Pre-register the hypothesis, the outcome, the sample size and the analysis before seeing the data, or label the work exploratory and say so in the abstract, not the appendix.
  • Report every outcome measured, every condition run and every exclusion made, with the results under the alternatives.
  • A subgroup result found after the fact is a hypothesis for the next study, not a finding of this one.
  • For A/B tests, fix the sample size or use a sequential method designed for interim looks; do not stop when the number looks good.

How to defend against it

  • Ask whether the study was pre-registered and whether the reported outcome is the one registered; outcome switching is checkable in trial registries.
  • Count the comparisons. If a finding is one of twenty things measured, one significant result is what chance predicts.
  • Be suspicious of findings that hold only in a subgroup, only after excluding some participants, or only with one of several possible outcome measures.
  • Look for an independent replication with a fixed protocol before treating a surprising effect as real; replications are what the incentives of the original study did not supply.
  • Ask whether the sample size was decided in advance; “we collected data until the result was significant” is the confession that the p-value is meaningless.

From the Defense Playbook

Every playbook entry states how strong its evidence is and when not to use it. Browse the full playbook.

References

  1. Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366 · link
    The simulations showing false-positive rates rising to about 61 percent under combined researcher degrees of freedom, and the Beatles-song demonstration.
  2. Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460-465
    The garden of forking paths: data-dependent analysis choices invalidate p-values even without explicit multiple testing.
  3. Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124 · link
    The argument that low prior odds, small samples and analytic flexibility make a large share of published positive findings false.
  4. Head, M. L., Holman, L., Lanfear, R., Kahn, A. T., & Jennions, M. D. (2015). The extent and consequences of p-hacking in science. PLoS Biology, 13(3), e1002106
    Text-mining evidence of an excess of p-values just below 0.05 across disciplines.
Last reviewed
Suggest a correction

Detect P-Hacking in any text

Paste any message, email, or article into our free Manipulation Detector to see if P-Hacking or other techniques are being used on you.