fast-weight-networks-1992-equivalent-linear-transformer

IN premisesummaries/2026-08-24/wiki-Transformer_deep_learning_architecture-chunk-1.md

Created 2026-08-24T17:11:24+00:00

Fast-weight/dynamic-link networks (1992) are mathematically equivalent to the unnormalized linear transformer, providing a theoretical predecessor to the 2017 architecture.

Summary

The 1992 fast-weight network and the 2017 linear transformer turn out to be the same mathematical object written in different notation, meaning the later architecture was not a genuinely new idea but a rediscovery of a connection already worked out more than two decades earlier. This matters because it reframes a widely cited 2017 contribution as a restatement of prior work and gives the 1992 formulation a stronger claim on historical priority.