state-space-models-linear-complexity-alternative
IN premise — entries/2026/06/21/wiki-Transformer_deep_learning_architecture-chunk-1.md
Created 2026-06-21T09:55:54+00:00
State-space models (e.g., Mamba) have emerged as alternatives to transformers that handle long sequences with linear rather than quadratic complexity.
Dependents
These beliefs depend on this one:
- IN quadratic-attention-drives-architectural-succession-pressure — The Transformer's quadratic attention cost creates permanent architectural succession pressure — just as RNNs were displaced by Transformers for failing to parallelize, Transformers face displacement pressure from linear-complexity alternatives (Mamba, RWKV, Reformer), confirming that hardware scalability determines paradigm survival applies reflexively to the currently dominant architecture.
- IN ssm-architecturally-validates-transformer-quadratic-limitation — State space models (Mamba, RWKV) achieving competitive performance with linear complexity architecturally validates that the Transformer's quadratic attention cost is a genuine limitation, not merely a theoretical concern — alternative architectures prove that sequence modeling does not inherently require quadratic computation.
- OUT ssm-breaks-transformer-hardware-lock-in — State space models would break the hardware lock-in that entrenches ML's crisis through the Transformer-TPU synergy — by achieving competitive performance with linear complexity, SSMs could redirect hardware co-evolution away from attention-optimized architectures, potentially reopening the material pathway that attention's hardware embedding has closed.