DigitalMANIPULATIVE

LLM Grooming and Data-Poisoning Narratives

What it is

Publishing large volumes of low-quality or false web content with the aim, or the effect, of shaping what search engines and AI chatbots later repeat, together with the wider and often overstated claim that synthetic text is degrading the models themselves.

How it works

Chatbots answer from two sources: training data scraped from the web and live search results. Both can be seeded. In February 2025 the American Sunlight Project argued that the pro-Kremlin Pravda network, which mass-republishes content across roughly 150 domains and draws little human traffic, looked designed for machines, and called this LLM grooming. A March 2025 NewsGuard audit found ten chatbots repeated the network's false narratives about a third of the time. The interpretation is contested. Alyukov and colleagues (2025) found that only about 8 percent of responses from four chatbots referenced Kremlin-linked sites, mainly where credible coverage was missing, which points to data voids (Golebiewski and boyd) rather than successful manipulation. In April 2026 DFRLab showed that state-linked text is present in Common Crawl and can be reproduced by an open model, while cautioning that this does not prove training influence. Model collapse is a separate laboratory finding (Shumailov et al., 2024) about recursive training, not evidence that today's chatbots are collapsing.

Real-world examples

  • NewsGuard's March 2025 audit tested ten leading chatbots on 15 false narratives spread by the Pravda network and reported that they repeated the claims 33 percent of the time, with seven citing Pravda articles as sources. NewsGuard sells ratings to AI companies, which readers should weigh.
  • Alyukov, Makhortykh, Voronovici and Sydorova, writing in the Harvard Kennedy School Misinformation Review in October 2025, audited chatbot answers and concluded that references to Kremlin-linked sources reflected information gaps rather than grooming.
  • DFRLab's April 2026 report Pravda in the Pipeline found Russian and Chinese state-linked articles in the Common Crawl archive and showed that an open-weight base model could complete some of them nearly verbatim, while stating that a model's training data cannot be precisely reconstructed this way.
  • Commercial actors do a version of the same thing: marketers openly sell generative engine optimization, publishing content built to be picked up in AI answers, which becomes deceptive when it uses fake reviews or fake independent sites.

Ethical guidelines

  • Flooding the web with fabricated or laundered content to steer machine answers is deception at one remove; the reader never sees the source that misled the machine.
  • Researchers and journalists should separate what has been shown (content exists, chatbots sometimes cite it) from what has not (intent, and measurable effects on users).
  • AI developers owe users visible citations and source-quality controls, so that seeded content can be seen and challenged.

How to defend against it

  • Treat a chatbot answer as a summary of sources, then look at the sources. If the citations are sites you have never heard of, read laterally about those sites before accepting the claim (Wineburg & McGrew).
  • Be most careful on narrow, breaking, or obscure questions. Data voids are where low-quality sources dominate, because little reliable material exists yet.
  • Ask the same question in a way that names a reliable source type, for example asking what major wire services or official statistics say, and compare.
  • Do not swing to the opposite error. Claims that AI answers are wholly poisoned, or that models are collapsing, outrun the evidence; the measured rates are partial and topic-specific.
  • For decisions that matter, confirm through a second channel that is not an AI system: a primary document, a professional, or an institution you can call.

References

  1. NewsGuard (2025). A well-funded Moscow-based global news network has infected Western artificial intelligence tools worldwide with Russian propaganda. NewsGuard Reality Check, March 2025 · link
    The audit of ten chatbots on 15 Pravda-network narratives, the 33 percent repetition rate, and the reference to the American Sunlight Project's LLM grooming hypothesis.
  2. Alyukov, M., Makhortykh, M., Voronovici, A., & Sydorova, M. (2025). LLMs grooming or data voids? LLM-powered chatbot references to Kremlin disinformation reflect information gaps, not manipulation. Harvard Kennedy School Misinformation Review, October 2025 · link
    The competing finding that chatbot citations of Kremlin-linked sources are infrequent and best explained by data voids.
  3. Brooking, E. T. (DFRLab) (2026). Pravda in the pipeline: Early evidence of state-adjacent propaganda in AI training data. Atlantic Council Digital Forensic Research Lab, April 8, 2026 · link
    Presence of state-linked content in Common Crawl, near-verbatim reproduction by an open base model, and the stated limits of that evidence.
  4. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755-759 · link
    Model collapse as a finding about indiscriminate recursive training on generated data.
  5. Golebiewski, M., & boyd, d. (2019). Data Voids: Where Missing Data Can Easily Be Exploited. Data & Society Research Institute
    The data void concept used to explain why low-quality sources surface on sparse topics.
Last reviewed
Suggest a correction

Detect LLM Grooming and Data-Poisoning Narratives in any text

Paste any message, email, or article into our free Manipulation Detector to see if LLM Grooming and Data-Poisoning Narratives or other techniques are being used on you.

Related Articles