compute-revolution-contaminates-own-data-supply
IN derived (depth 11)
Created 2026-06-21T14:12:26+00:00 · Reviewed 2026-06-21T15:37:01+00:00
The 300,000x compute increase that drove ML's capability revolution simultaneously creates the conditions for model collapse — massive compute enables training on internet-scale data that produces capable models, but those capable models flood the internet with synthetic content, contaminating the very data ecosystem that enabled the scaling in the first place.
Justifications
SL — The compute-enabled capability revolution poisons its own data supply through the models it produces
Antecedents (all must be IN):
- IN compute-scaling-quantifies-structure-displacement-rate — The 300,000x compute increase from AlexNet to AlphaZero (doubling every 3.4 months) provides a quantitative measure for the rate at which capacity growth has accompanied the displacement of structured mechanisms. The observed pattern — where increases in compute coincide with replacement of components like tree search, handcrafted features, and symbolic rules by neural capacity — suggests that structure displacement operates as an exponential process, though the precise relationship between each order of magnitude of compute and specific structural replacements is an observed correlation rather than a confirmed causal law.
- IN model-collapse-recursive-crisis-amplifier — Model collapse from synthetic data creates a recursive amplifier within ML's compounding reliability crisis — as capable models generate training data for next-generation models, reliability degradation is inherited and compounded across model generations, meaning capability scaling now directly poisons the data substrate on which future capability depends, adding a temporal feedback dimension to the crisis.
Dependents
These beliefs depend on this one:
- IN compute-scaling-self-undermining-on-two-fronts — ML's compute-driven capability scaling is self-undermining on two independent fronts — it displaces structured mechanisms with raw capacity (300,000x scaling systematically replacing search, theory, and domain expertise) while simultaneously contaminating the data ecosystem required for further scaling (capable models flood the internet with synthetic content, triggering model collapse) — meaning the compute revolution destroys both the intellectual and material substrates on which its own continuation depends.