compute-revolution-contaminates-own-data-supply

IN derived (depth 11)

Created 2026-06-21T14:12:26+00:00 · Reviewed 2026-06-21T15:37:01+00:00

The 300,000x compute increase that drove ML's capability revolution simultaneously creates the conditions for model collapse — massive compute enables training on internet-scale data that produces capable models, but those capable models flood the internet with synthetic content, contaminating the very data ecosystem that enabled the scaling in the first place.

Justifications

SL — The compute-enabled capability revolution poisons its own data supply through the models it produces

Antecedents (all must be IN):

  • IN compute-scaling-quantifies-structure-displacement-rate — The 300,000x compute increase from AlexNet to AlphaZero (doubling every 3.4 months) provides a quantitative measure for the rate at which capacity growth has accompanied the displacement of structured mechanisms. The observed pattern — where increases in compute coincide with replacement of components like tree search, handcrafted features, and symbolic rules by neural capacity — suggests that structure displacement operates as an exponential process, though the precise relationship between each order of magnitude of compute and specific structural replacements is an observed correlation rather than a confirmed causal law.
  • IN model-collapse-recursive-crisis-amplifier — Model collapse from synthetic data creates a recursive amplifier within ML's compounding reliability crisis — as capable models generate training data for next-generation models, reliability degradation is inherited and compounded across model generations, meaning capability scaling now directly poisons the data substrate on which future capability depends, adding a temporal feedback dimension to the crisis.

Dependents

These beliefs depend on this one: