self-supervised-learning-dominant-pretraining-paradigm
IN premise — entries/2026/06/21/wiki-Deep_learning-chunk-7.md
Created 2026-06-21T09:55:49+00:00
Self-supervised learning is the dominant pre-training paradigm for modern deep learning, as opposed to supervised or unsupervised learning.
Dependents
These beliefs depend on this one:
- OUT modern-llm-pipeline-sustainable — Modern LLM training pipelines (self-supervised pretraining → instruction tuning → RLHF) are a sustainable methodology — they dissolve classical paradigm boundaries by successfully combining all three ML paradigms, and self-supervised learning provides an effectively unlimited source of training signal.
- IN pretraining-30-year-delayed-adoption — Modern self-supervised pretraining has roots in Schmidhuber's 1991 neural history compressor, which used predictive coding and self-supervised pre-training decades before the paradigm became dominant in modern deep learning — a multi-decade gap between early work and widespread adoption that suggests hardware and ecosystem readiness may play a significant role in determining when theoretical ideas achieve industrial impact.
- IN pretraining-is-transfer-learning-at-scale — Modern self-supervised pretraining is transfer learning at industrial scale — the formal transfer learning framework (source domain D_S → target domain D_T) exactly describes the pretrain-then-finetune pipeline, unifying a 50-year-old theoretical concept with the dominant modern training methodology.
- OUT pretraining-universally-beneficial — The pretrain-then-finetune paradigm universally improves downstream task performance, as demonstrated by its adoption across BERT, GPT, and all modern LLMs as the standard training pipeline.
- IN self-supervised-paradigm-boundary-contested — Self-supervised learning occupies a contested taxonomic position — formally a subset of unsupervised learning, yet it has become the dominant pre-training paradigm, creating tension between its classification and its practical centrality.