neural-history-compressor-1991-pretraining
IN premise — entries/2026/06/21/wiki-Deep_learning-chunk-2.md
Created 2026-06-21T09:55:50+00:00
Schmidhuber's neural history compressor (1991) used predictive coding and self-supervised pre-training with a hierarchy of RNNs; in 1993 it solved a task requiring >1000 unfolded layers
Dependents
These beliefs depend on this one:
- IN pretraining-30-year-delayed-adoption — Modern self-supervised pretraining has roots in Schmidhuber's 1991 neural history compressor, which used predictive coding and self-supervised pre-training decades before the paradigm became dominant in modern deep learning — a multi-decade gap between early work and widespread adoption that suggests hardware and ecosystem readiness may play a significant role in determining when theoretical ideas achieve industrial impact.