pre-ln-removes-warmup-need
IN premise — entries/2026/06/21/wiki-Transformer_deep_learning_architecture-chunk-2.md
Created 2026-06-21T09:55:55+00:00
Pre-LN transformer (layer normalization before attention/FFN) stabilizes training and removes the need for learning rate warmup, unlike the original post-LN convention.