vulnerability-persists-infinite-data-limit

IN premisesummaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-8.md

Created 2026-08-25T02:58:01+00:00

Superposition-induced adversarial vulnerability is present in the infinite-data limit and is a property of the optimal sparse representation, not a finite-sample or training artifact.

Summary

The vulnerability that lets an adversary exploit the model is baked into the very structure of its best possible representation, not caused by having too little data or imperfect training. This means no amount of extra data or longer training will make it go away; the fix has to come from changing how the representation works at a fundamental level.