LogicalDUAL-USE
Base-Rate Neglect
What it is
Presenting individuating information — a test result, a profile, a vivid description — without the prior frequency of the thing in question, so the audience judges by resemblance and ignores how rare or common it actually is.
How it works
Real-world examples
- •Casscells, Schoenberger and Graboys (1978) put the 1-in-1,000 disease problem to staff and students at Harvard Medical School; fewer than one in five gave the correct answer of about 2 percent, and the most common answer was 95 percent.
- •Gigerenzer's surveys of gynecologists found most could not say what a positive mammogram meant for a woman with no symptoms; with roughly 1 percent prevalence, 90 percent sensitivity and a 9 percent false-positive rate, about one positive in ten is a cancer, yet many physicians answered 80 or 90 percent.
- •Arguments for profiling after 2001 ran from “most of the attackers were young Muslim men” to the risk posed by any young Muslim man; arguments on the other side run from “most mass shooters are white men” to the risk posed by any white man. Both drop the base rate — the millions in each group who never do anything of the kind — and both are the same inversion.
- •Workplace drug screening with a 1 percent prevalence and a test 95 percent specific produces far more false positives than true ones, which is why confirmatory testing exists; a positive result reported as “the test is 95 percent accurate” conceals that most positives are wrong.
- •Face-recognition scans of large crowds: even a system with one false match per million faces produces a false alarm for every real match when a million faces are scanned for a watchlist of one, and far more at realistic error rates.
Ethical guidelines
Where the line is
Describing a test's sensitivity, a case's features or a group's share of some outcome is legitimate when the prevalence is stated beside it and the reader is left able to compute the posterior; it becomes manipulation when the base rate is omitted or buried because it would shrink the conclusion the presenter wants drawn.
- ●Whenever you report a test's accuracy, a profile match or a probability of guilt, report the prevalence beside it; the audience cannot compute the answer without it.
- ●Use natural frequencies — “of 1,000 people like this, X have the condition and Y test positive” — because they are demonstrably understood and percentages demonstrably are not.
- ●Do not argue from “most of the people who did X were Y” to the risk posed by a given Y without stating how many Ys there are.
- ●Stating a hit rate without the base rate to make a screening tool, a product or a suspect look more certain is a decision to let the audience misread it.
How to defend against it
- ►Before believing a match, a positive or a prediction, ask how common the thing is in the first place — the base rate — and refuse to update without it.
- ►Rewrite the claim in people out of 1,000 or 100,000: how many have the condition, how many of those the test catches, how many without it test positive anyway. The answer usually appears in the rewriting.
- ►When a profile is offered as evidence (“fits the pattern of…”), ask how many people fit the pattern and never do the thing; that number is the denominator of the argument.
- ►Treat “99 percent accurate” as two numbers, not one: how often it catches real cases and how often it flags innocent ones, then apply both to the base rate.
- ►For any inverted conditional — from “most Xs are Ys” to “this Y is probably an X” — say the inversion out loud; the fallacy is usually visible once stated.
From the Defense Playbook
Every playbook entry states how strong its evidence is and when not to use it. Browse the full playbook.
References
- Kahneman, D., & Tversky, A. (1973). On the psychology of prediction. Psychological Review, 80(4), 237-251 · linkThe lawyer-engineer experiments: judgments followed the description and ignored the stated base rates.
- Bar-Hillel, M. (1980). The base-rate fallacy in probability judgments. Acta Psychologica, 44(3), 211-233The cab problem and the analysis of when base rates are used and ignored.
- Casscells, W., Schoenberger, A., & Graboys, T. B. (1978). Interpretation by physicians of clinical laboratory results. New England Journal of Medicine, 299(18), 999-1001The Harvard survey: the 1-in-1,000 problem and the modal answer of 95 percent against the correct 2 percent.
- Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review, 102(4), 684-704Natural-frequency formats let most people solve base-rate problems that they fail in probability format.
Last reviewed
Suggest a correction