Crowd Fact-Checking

Minutes

Use the aggregated judgment of a small, politically mixed group, instead of your own reading or that of your like-minded friends, to assess whether a story is accurate; averaged ratings from a balanced crowd of laypeople track professional fact-checkers surprisingly well.

How to do it

  1. 1On platforms that have them, read community context notes before sharing, and check that the note cites a source you can open. Notes are designed to appear only when raters who usually disagree with each other both find them helpful.
  2. 2Build your own small crowd: a handful of people you know who differ in politics and background and who will answer "does this look right to you?" honestly. Ask before you share, not after.
  3. 3Collect judgments independently. Ask each person separately, or have everyone rate before anyone comments, so that the first confident voice does not anchor the rest.
  4. 4Average, do not pick. The value is in the aggregate; one person's rating, including your own, is noisy and partisan.
  5. 5Weight agreement across the divide most heavily. When people who usually disagree both think a story is false or misleading, take that seriously.
  6. 6Contribute properly if you rate or write notes yourself: rate claims that favour your side as strictly as those that do not, and cite primary sources.

What to say

  • Before I post this, can three of you who vote differently from me tell me whether it looks accurate? Do not discuss it first.
  • There is a context note on this one with a link to the original report. Read that before sharing.

When to use it

  • Moderating a community group, forum, or workplace channel without a professional fact-checker.
  • Judging a political or culturally charged story that you want to be true.
  • Fast-moving events where professional fact-checks do not yet exist.
  • Deciding how much weight to give platform context labels.

Counters

Evidence and how strong it is

Allen et al. (2021) had politically balanced groups of laypeople rate 207 articles after reading only the headline and lede. The average rating of a crowd of as few as ten to fifteen people correlated with the average of three professional fact-checkers about as well as the fact-checkers correlated with one another. The result depends on aggregation and balance. Allen, Martel & Rand (2022) analysed Twitter's Birdwatch programme and found that individual volunteers mostly flagged and rated along partisan lines, judging content from the other side more harshly. Wojcik et al. (2022) describe the "bridging" algorithm, which publishes a note only when raters with a history of disagreement both rate it helpful, and report that such notes reduced agreement with and resharing of misleading posts. Evidence strength: good experimental evidence that balanced, aggregated lay ratings are accurate; mixed observational evidence on deployed systems, where notes are often slow and many misleading posts never receive one.

Cautions
  • An unbalanced crowd is a mob with a scoreboard. Ratings from a like-minded group measure what the group likes, and coordinated raters can game open systems.
  • Crowds judge plausibility, not ground truth. They do well on claims that can be checked against common knowledge and poorly on technical questions and on claims that are surprising but true.
  • The absence of a note is not a verdict of accuracy. Most posts are never reviewed, and notes commonly arrive after the peak of sharing.
  • Do not turn it into a pile-on. The purpose is to decide what you will believe and share, not to organize a group against an individual.
  1. Allen, J., Arechar, A. A., Pennycook, G., & Rand, D. G. (2021). Scaling Up Fact-Checking Using the Wisdom of Crowds. Science Advances, 7(36), eabf4393
    The finding that small, politically balanced lay crowds agree with professional fact-checkers about as well as the professionals agree with each other.
  2. Allen, J., Martel, C., & Rand, D. G. (2022). Birds of a Feather Don't Fact-Check Each Other: Partisanship and the Evaluation of News in Twitter's Birdwatch Crowdsourced Fact-Checking Program. Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems
    Evidence that individual crowd raters act along partisan lines, the reason balance and aggregation are required.
  3. Wojcik, S., Hilgard, S., Judd, N., et al. (2022). Birdwatch: Crowd Wisdom and Bridging Algorithms Can Inform Understanding and Reduce the Spread of Misinformation. arXiv preprint arXiv:2210.15723
    Description and evaluation of the bridging-based ranking used for community notes. Authored by the platform's own team; not independently peer-reviewed.
Last reviewed

More in Online, media, and misinformation