self-supervised-pretraining-1991-schmidhuber
IN premise — entries/2026/06/21/wiki-Neural_network_28machine_learning29-chunk-2.md
Created 2026-06-21T09:55:51+00:00
Self-supervised pre-training originated with Schmidhuber's neural history compressor in 1991, predating its use in GPT by decades
Dependents
These beliefs depend on this one:
- IN pretraining-30-year-delayed-adoption — Modern self-supervised pretraining has roots in Schmidhuber's 1991 neural history compressor, which used predictive coding and self-supervised pre-training decades before the paradigm became dominant in modern deep learning — a multi-decade gap between early work and widespread adoption that suggests hardware and ecosystem readiness may play a significant role in determining when theoretical ideas achieve industrial impact.