data-scaling-paradigm-remains-safely-dominant
OUT derived (depth 4)
Created 2026-06-21T11:44:45+00:00
The data-volume-first scaling strategy — independently validated by Chinchilla scaling laws and Llama's compression evidence — remains the dominant and safe approach to capability improvement, with massive web-scale data ingestion as the primary scaling lever.
Justifications
SL — Data-first scaling is validated by theory and practice unless deliberate data poisoning makes web-scale ingestion unsafe
Antecedents (all must be IN):
- IN data-scaling-outweighs-parameter-scaling — Empirical results consistently show data volume matters more than parameter count: Chinchilla demonstrated models were undertrained, Llama 1 13B beat GPT-3 175B, and Llama 3 8B continued improving at 75x Chinchilla-optimal data.
- IN optimal-scaling-validated-from-theory-and-compression — The optimal scaling strategy (MoE architecture + massive training data) is independently validated by two converging lines of evidence: Chinchilla scaling theory showing data matters more than parameters, and empirical compression results (DistilBERT, ALBERT, weight tying) showing models carry significant parameter redundancy — confirming from both theoretical and empirical directions that intelligent data/compute allocation dominates raw parameter count.
Unless (any of these IN defeats this justification):
- IN llm-grooming-deliberate-training-data-poisoning — LLM grooming is the deliberate mass-publishing of web content to bias LLM training data and outputs, a term coined by the American Sunlight Project in 2025 (e.g., the Pravda network).