superposition-unifies-data-and-representation-redundancy-v2

IN premise

Created 2026-08-25T04:33:46+00:00

Over-complete superposition is a primary structural factor consistent with both the data-side redundancy (5× correlated corpora at Spearman 0.87–0.97 yielding only marginal accuracy gains, suggesting 'relevant document count' overcounts unique information) and the representation-side redundancy (SAE dead features ranging from ~2% at 1M to ~65% at 34M, alongside systematic under-reconstruction): in both cases the observed patterns are consistent with available capacity exceeding the unique information to be encoded, though the antecedents establish these as descriptive correlations rather than isolating over-completeness as the sole causal mechanism.

Summary

The network has far more internal space to store information than the data actually contains, so most of that space sits empty: adding more highly similar documents yields almost no benefit, and a growing fraction of internal features (up to two-thirds at large scale) go completely unused. The practical implication is that the real bottleneck is the amount of unique information present, not the capacity or data volume available, so scaling either one further is unlikely to close the gap.

Dependents

These beliefs depend on this one: