merrill-2022-saturated-transformers-threshold-circuits
IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-sR-references.md
Created 2026-08-25T02:58:33+00:00
Merrill et al. (2022) showed that saturated (ReLU-activated) transformers are computationally equivalent to constant-depth threshold circuits, bounding their in-context computational power.
Summary
Transformers that use ReLU activations can only perform a very limited kind of computation, roughly equivalent to a shallow, fixed-depth circuit of threshold gates, meaning they hit a hard ceiling on how much in-context reasoning they can do no matter how many layers you stack. This matters because it sets a theoretical upper bound on what a standard transformer architecture can accomplish internally, pointing to where more expressive or iterative mechanisms would be needed to go beyond that limit.