Calibration Practice
Takes practiceRegularly attach a probability to your predictions, record them, and score them against what happened, so that "I am sure" comes to mean something and you can recognize false certainty in others.
How to do it
- 1Each week write down a handful of predictions with explicit probabilities and dates: "70 percent the contractor finishes by March", "90 percent this stock tip underperforms the index in a year".
- 2Use a scale you can act on: 50 percent means you would take either side of the bet; 90 percent means you would be surprised about one time in ten.
- 3When the outcome is known, score it. Your 90-percent claims should come true about nine times in ten; if they come true six times in ten, you are overconfident by that margin and should lower your numbers.
- 4Practice small updates: move 5 or 10 points on new evidence rather than swinging from certain to certain.
- 5Apply the scale to persuaders. Ask a forecaster, salesperson, or pundit for a number and a date, and check it later. People who never give numbers or never keep score are exempting themselves from being wrong.
What to say
- “How confident am I, as a number? And what would I need to see in a year to say I was wrong?”
- “You say this is certain. Would you put 90 percent on it, and can we check in six months?”
When to use it
- •Investment, career, and health decisions made on the strength of someone's confident forecast.
- •Consuming pundits, analysts, and influencers whose track record no one keeps.
- •Reviewing your own past judgments to learn which of your instincts deserve trust.
Counters
Evidence and how strong it is
Decades of calibration research (reviewed by Lichtenstein, Fischhoff & Phillips 1982) show that people who say they are 90 percent sure are typically right far less often, and that the gap shrinks with prompt, outcome-based feedback. Tetlock (2005) followed 284 experts across tens of thousands of forecasts over nearly two decades and found average expert accuracy barely better than simple extrapolation, with the most confident and most famous forecasters the least accurate. In the Good Judgment Project, an hour of probabilistic-reasoning training, keeping score, and small frequent updates produced measurable and lasting accuracy gains (Mellers et al. 2014), and the best-calibrated forecasters beat intelligence analysts with access to classified information (Tetlock & Gardner 2015). Evidence strength: strong for the existence of overconfidence and for feedback-based training in forecasting tournaments; the transfer of tournament calibration to everyday decisions is assumed rather than tested.
- Calibration is not accuracy. A weather forecaster who says 50 percent every day is perfectly calibrated and useless; you also need to be right more often than chance.
- Scoring requires outcomes you will actually learn. Predictions about things you will never check teach nothing and feel like practice.
- Numbers can launder guesses ("our model gives 87 percent"). A precise figure from someone who keeps no score is a rhetorical device, not a forecast.
- Lichtenstein, S., Fischhoff, B., & Phillips, L. D. (1982). Calibration of Probabilities: The State of the Art to 1980. In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under Uncertainty: Heuristics and Biases (pp. 306-334). Cambridge University PressReview establishing that untrained confidence is systematically miscalibrated and that outcome feedback improves it.
- Tetlock, P. E. (2005). Expert Political Judgment: How Good Is It? How Can We Know?. Princeton University PressThe long-run study of expert forecasts showing poor average accuracy and the inverse relationship between confidence, fame, and accuracy.
- Mellers, B., Ungar, L., Baron, J., et al. (2014). Psychological Strategies for Winning a Geopolitical Forecasting Tournament. Psychological Science, 25(5), 1106-1115Randomized evidence that brief probabilistic-reasoning training and teaming improved forecasting accuracy.
- Tetlock, P. E., & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. CrownThe account of the Good Judgment Project, including the habits of keeping score and updating in small increments.