LogicalMANIPULATIVE
P-Hacking
What it is
Trying many analytic choices — outcomes, subgroups, covariates, exclusions, stopping rules — and reporting the one that crossed the significance threshold as though it were the only test planned.
How it works
Real-world examples
- •Simmons, Nelson and Simonsohn (2011) reported a real experiment in which undergraduates who listened to a Beatles song were significantly younger than those who listened to a control track — a demonstration of what undisclosed flexibility can produce from noise.
- •In 2016 Brian Wansink of Cornell's Food and Brand Lab described, in a blog post celebrating a visiting student's diligence, how a null dataset was re-sliced until it yielded several publishable papers; the scrutiny that followed led to more than a dozen retractions and his resignation in 2019.
- •Bem's 2011 paper in the Journal of Personality and Social Psychology reporting evidence of precognition used analytic choices that were standard in the field; its failure to replicate became a founding case of the replication crisis and of the argument that the standards, not the author, were the problem.
- •Head and colleagues (2015) examined the distribution of reported p-values across scientific fields and found an excess of values just below 0.05, the signature of results pushed over the threshold.
- •Online experiments that are monitored continuously and declared when a variant reaches significance produce a large share of false winners; the practice is common enough that experimentation platforms sell “peeking-safe” sequential tests to prevent it.
Ethical guidelines
- ●Pre-register the hypothesis, the outcome, the sample size and the analysis before seeing the data, or label the work exploratory and say so in the abstract, not the appendix.
- ●Report every outcome measured, every condition run and every exclusion made, with the results under the alternatives.
- ●A subgroup result found after the fact is a hypothesis for the next study, not a finding of this one.
- ●For A/B tests, fix the sample size or use a sequential method designed for interim looks; do not stop when the number looks good.
How to defend against it
- ►Ask whether the study was pre-registered and whether the reported outcome is the one registered; outcome switching is checkable in trial registries.
- ►Count the comparisons. If a finding is one of twenty things measured, one significant result is what chance predicts.
- ►Be suspicious of findings that hold only in a subgroup, only after excluding some participants, or only with one of several possible outcome measures.
- ►Look for an independent replication with a fixed protocol before treating a surprising effect as real; replications are what the incentives of the original study did not supply.
- ►Ask whether the sample size was decided in advance; “we collected data until the result was significant” is the confession that the p-value is meaningless.
From the Defense Playbook
Every playbook entry states how strong its evidence is and when not to use it. Browse the full playbook.
References
- Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359-1366 · linkThe simulations showing false-positive rates rising to about 61 percent under combined researcher degrees of freedom, and the Beatles-song demonstration.
- Gelman, A., & Loken, E. (2014). The statistical crisis in science. American Scientist, 102(6), 460-465The garden of forking paths: data-dependent analysis choices invalidate p-values even without explicit multiple testing.
- Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124 · linkThe argument that low prior odds, small samples and analytic flexibility make a large share of published positive findings false.
- Head, M. L., Holman, L., Lanfear, R., Kahn, A. T., & Jennions, M. D. (2015). The extent and consequences of p-hacking in science. PLoS Biology, 13(3), e1002106Text-mining evidence of an excess of p-values just below 0.05 across disciplines.
Last reviewed
Suggest a correction