llm-grooming-deliberate-training-data-poisoning
IN premise — entries/2026/06/21/wiki-Large_language_model-chunk-3.md
Created 2026-06-21T09:50:09+00:00
LLM grooming is the deliberate mass-publishing of web content to bias LLM training data and outputs, a term coined by the American Sunlight Project in 2025 (e.g., the Pravda network).
Summary
Organized groups are flooding the internet with large volumes of targeted content specifically to steer what large language models learn during training, meaning the "neutral" outputs of LLMs can be quietly biased by coordinated publishing operations rather than reflecting genuine human consensus. This matters because it turns LLMs into a target for information manipulation at scale, and any system that relies on LLM output for reasoning or fact-checking is vulnerable to being subtly poisoned at the source.
Dependents
These beliefs depend on this one:
- OUT data-scaling-paradigm-remains-safely-dominant — The data-volume-first scaling strategy — independently validated by Chinchilla scaling laws and Llama's compression evidence — remains the dominant and safe approach to capability improvement, with massive web-scale data ingestion as the primary scaling lever.
- IN llm-security-spans-training-and-inference-surfaces — LLM security threats operate across both training time (deliberate data poisoning via grooming) and inference time (prompt injection), requiring defense at multiple pipeline stages rather than a single chokepoint.
- IN memorization-creates-attack-surface-for-training-poisoning — LLM training data memorization (measured at 1-7% exact duplicates in GPT-2) is a plausible mechanism through which deliberate data poisoning (grooming) could materialize in model outputs — if poisoned content is memorized, it may be retrievable at inference time, suggesting that memorization rates could serve as one indicator of susceptibility to poisoning attacks.
- OUT uncontainable-innovation-diffusion-serves-collective-progress — The institutional uncontainability of LLM innovation — craft knowledge crossing organizational and disciplinary boundaries, with breakthrough techniques escaping their originators — serves collective capability progress by ensuring no single institution bottlenecks frontier advancement.