PsychologicalDUAL-USE

Reward Prediction Error

What it is

The dopamine signal that fires when an outcome is better than expected and falls silent when it is exactly as expected. Anything that delivers rewards unpredictably keeps the signal alive, which is why uncertain payoffs hold attention far longer than reliable ones.

How it works

Wolfram Schultz recorded midbrain dopamine neurons in monkeys and, with Dayan and Montague, showed in 1997 that they do not encode reward itself but the difference between the reward received and the reward predicted. An unexpected drop of juice produces a burst; a fully predicted one produces nothing; a predicted drop that fails to arrive produces a dip. This is the teaching signal of reinforcement learning, and it is among the best-replicated findings in systems neuroscience. Fiorillo, Tobler and Schultz then found a second, sustained signal that peaks when reward probability is about fifty percent — maximal uncertainty. Two consequences follow for persuasion. First, a predictable reward stops teaching: the hundredth identical notification is ignored. Second, variable and uncertain rewards — Skinner's variable-ratio schedule — keep the signal firing and produce the highest, most extinction-resistant response rates. Slot machines, loot boxes, feeds, and dating apps are engineered around this. Lindström and colleagues showed in 2021 that posting behavior on social platforms follows a reward-learning model fed by likes. One caution: the phrase “dopamine hit” conflates this learning signal with pleasure; Berridge and Robinson separate “wanting” from “liking”, and it is wanting that the machines exploit.

Real-world examples

  • Schultz, Dayan and Montague (1997): dopamine neurons fired for unexpected juice, shifted to the predictive cue once it was learned, and dipped when a predicted reward was withheld.
  • Slot machines pay on a variable-ratio schedule and add engineered near-misses; Natasha Dow Schüll's ethnography of Las Vegas machine design documents an industry goal of “time on device” rather than big wins.
  • Loot boxes in video games deliver randomized items for money; the Belgian Gaming Commission ruled in 2018 that several implementations constituted gambling under Belgian law.
  • Lindström et al. (2021) fitted a reinforcement-learning model to more than a million posts across several platforms and found that posting frequency tracked the rate of likes as the model predicted, with an online experiment confirming the causal direction.
  • Email and messaging apps that reward checking unpredictably — most checks find nothing, some find something good — train compulsive checking more effectively than a reliable hourly digest would.

Ethical guidelines

Where the line is

Reward feedback is legitimate when it tracks real progress toward something the user chose and leaves clear stopping points; it becomes manipulation when randomness is added, odds are hidden, or rewards are timed so that the learning signal keeps a person engaged past the point they intended to stop.

  • Do not add randomness to a reward purely to increase compulsion; if variability exists, it should reflect the real nature of the activity, not a design goal of session length.
  • Disclose odds wherever a paid outcome is randomized, and never sell randomized rewards to minors.
  • Design for satiation: let users finish, batch notifications, and make stopping points visible.
  • Using progress and reward feedback to help someone reach a goal they chose is legitimate; using it to extend a session past the point they wanted to stop is not.

How to defend against it

  • Identify the schedule. If you cannot predict when the next reward comes, the product is using a variable schedule, and the pull you feel is engineered rather than a signal that the content is good.
  • Make rewards predictable on your side: check messages at fixed times, turn off unpredictable notifications, and let digests replace pings.
  • Remove the cue. Grayscale screens, moving apps off the home screen, and logging out between sessions blunt the predictive cue that triggers wanting.
  • Set a pre-committed budget of time or money before the session starts and decide the stopping rule before the first reward arrives — the design is built to make that decision for you afterward.
  • Distinguish wanting from liking: ask whether you actually enjoyed the last hour, not whether you felt pulled to continue.

References

  1. Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599 · link
    The founding demonstration that dopamine neurons encode reward prediction error rather than reward.
  2. Fiorillo, C. D., Tobler, P. N., & Schultz, W. (2003). Discrete coding of reward probability and uncertainty by dopamine neurons. Science, 299(5614), 1898-1902
    The sustained dopamine signal that peaks at maximal reward uncertainty.
  3. Ferster, C. B., & Skinner, B. F. (1957). Schedules of Reinforcement. Appleton-Century-Crofts
    Variable-ratio schedules produce the highest and most extinction-resistant response rates.
  4. Lindström, B., Bellander, M., Schultner, D. T., Chang, A., Tobler, P. N., & Olsson, A. (2021). A computational reward learning account of social media engagement. Nature Communications, 12, 1311
    Posting behavior on social platforms follows reward-learning dynamics driven by likes.
  5. Berridge, K. C., & Robinson, T. E. (1998). What is the role of dopamine in reward: Hedonic impact, reward learning, or incentive salience?. Brain Research Reviews, 28(3), 309-369
    The wanting-versus-liking distinction that undercuts the “dopamine equals pleasure” shorthand.
  6. Schüll, N. D. (2012). Addiction by Design: Machine Gambling in Las Vegas. Princeton University Press
    Ethnographic evidence that machine gambling is engineered for time on device using variable reward and near-miss design.
Last reviewed
Suggest a correction

Detect Reward Prediction Error in any text

Paste any message, email, or article into our free Manipulation Detector to see if Reward Prediction Error or other techniques are being used on you.