DigitalMANIPULATIVE
LLM Grooming and Data-Poisoning Narratives
What it is
Publishing large volumes of low-quality or false web content with the aim, or the effect, of shaping what search engines and AI chatbots later repeat, together with the wider and often overstated claim that synthetic text is degrading the models themselves.
How it works
Real-world examples
- •NewsGuard's March 2025 audit tested ten leading chatbots on 15 false narratives spread by the Pravda network and reported that they repeated the claims 33 percent of the time, with seven citing Pravda articles as sources. NewsGuard sells ratings to AI companies, which readers should weigh.
- •Alyukov, Makhortykh, Voronovici and Sydorova, writing in the Harvard Kennedy School Misinformation Review in October 2025, audited chatbot answers and concluded that references to Kremlin-linked sources reflected information gaps rather than grooming.
- •DFRLab's April 2026 report Pravda in the Pipeline found Russian and Chinese state-linked articles in the Common Crawl archive and showed that an open-weight base model could complete some of them nearly verbatim, while stating that a model's training data cannot be precisely reconstructed this way.
- •Commercial actors do a version of the same thing: marketers openly sell generative engine optimization, publishing content built to be picked up in AI answers, which becomes deceptive when it uses fake reviews or fake independent sites.
Ethical guidelines
- ●Flooding the web with fabricated or laundered content to steer machine answers is deception at one remove; the reader never sees the source that misled the machine.
- ●Researchers and journalists should separate what has been shown (content exists, chatbots sometimes cite it) from what has not (intent, and measurable effects on users).
- ●AI developers owe users visible citations and source-quality controls, so that seeded content can be seen and challenged.
How to defend against it
- ►Treat a chatbot answer as a summary of sources, then look at the sources. If the citations are sites you have never heard of, read laterally about those sites before accepting the claim (Wineburg & McGrew).
- ►Be most careful on narrow, breaking, or obscure questions. Data voids are where low-quality sources dominate, because little reliable material exists yet.
- ►Ask the same question in a way that names a reliable source type, for example asking what major wire services or official statistics say, and compare.
- ►Do not swing to the opposite error. Claims that AI answers are wholly poisoned, or that models are collapsing, outrun the evidence; the measured rates are partial and topic-specific.
- ►For decisions that matter, confirm through a second channel that is not an AI system: a primary document, a professional, or an institution you can call.
References
- NewsGuard (2025). A well-funded Moscow-based global news network has infected Western artificial intelligence tools worldwide with Russian propaganda. NewsGuard Reality Check, March 2025 · linkThe audit of ten chatbots on 15 Pravda-network narratives, the 33 percent repetition rate, and the reference to the American Sunlight Project's LLM grooming hypothesis.
- Alyukov, M., Makhortykh, M., Voronovici, A., & Sydorova, M. (2025). LLMs grooming or data voids? LLM-powered chatbot references to Kremlin disinformation reflect information gaps, not manipulation. Harvard Kennedy School Misinformation Review, October 2025 · linkThe competing finding that chatbot citations of Kremlin-linked sources are infrequent and best explained by data voids.
- Brooking, E. T. (DFRLab) (2026). Pravda in the pipeline: Early evidence of state-adjacent propaganda in AI training data. Atlantic Council Digital Forensic Research Lab, April 8, 2026 · linkPresence of state-linked content in Common Crawl, near-verbatim reproduction by an open base model, and the stated limits of that evidence.
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755-759 · linkModel collapse as a finding about indiscriminate recursive training on generated data.
- Golebiewski, M., & boyd, d. (2019). Data Voids: Where Missing Data Can Easily Be Exploited. Data & Society Research InstituteThe data void concept used to explain why low-quality sources surface on sparse topics.
Last reviewed
Suggest a correctionRelated Articles
Dark Patterns in UX: How Apps Manipulate Your Behavior
Subscription traps, misleading interfaces, and engineered addiction. Understanding the persuasion techniques built into the apps you use every day.
7 min read
OSINT for Beginners: Open Source Intelligence Explained
Open Source Intelligence (OSINT) uses publicly available data to gather actionable insights. Here is a beginner-friendly guide to what OSINT is and how it is used.
10 min read