pretraining-30-year-delayed-adoption
IN derived (depth 1)
Created 2026-06-21T11:59:56+00:00 · Reviewed 2026-06-21T15:37:01+00:00
Modern self-supervised pretraining has roots in Schmidhuber's 1991 neural history compressor, which used predictive coding and self-supervised pre-training decades before the paradigm became dominant in modern deep learning — a multi-decade gap between early work and widespread adoption that suggests hardware and ecosystem readiness may play a significant role in determining when theoretical ideas achieve industrial impact.
Justifications
SL — 1991 invention to 2020s dominance parallels backprop's delayed adoption and confirms hardware-readiness as the adoption bottleneck
Antecedents (all must be IN):
- IN self-supervised-pretraining-1991-schmidhuber — Self-supervised pre-training originated with Schmidhuber's neural history compressor in 1991, predating its use in GPT by decades
- IN self-supervised-learning-dominant-pretraining-paradigm — Self-supervised learning is the dominant pre-training paradigm for modern deep learning, as opposed to supervised or unsupervised learning.
- IN neural-history-compressor-1991-pretraining — Schmidhuber's neural history compressor (1991) used predictive coding and self-supervised pre-training with a hierarchy of RNNs; in 1993 it solved a task requiring >1000 unfolded layers
Dependents
These beliefs depend on this one:
- OUT dormant-solutions-await-enabling-conditions — ML's pattern of multi-decade adoption latencies combined with the convergent discovery of genuine mathematical necessities across disconnected fields suggests that solutions to current reliability challenges may already exist in published research, awaiting the economic or hardware conditions that would make them viable.
- IN idea-latency-validates-economic-gating — ML ideas exhibit systematic multi-decade adoption latencies — transfer learning (invented 1976, adopted 2010s) and self-supervised pretraining (invented 1991, dominant 2018) were both available for decades before widespread use, providing independent evidence that ML progress is gated by economic and hardware readiness rather than idea availability.