modern-pretraining-dominant-but-transient
IN derived (depth 3)
Created 2026-06-21T10:23:12+00:00 · Reviewed 2026-06-21T15:37:01+00:00
Modern pretraining as industrial-scale transfer learning is simultaneously the most successful ML methodology and the most likely to be displaced — its dominance rests on empirically fragile foundations (pretraining can hurt), and the broader pattern of paradigm succession (GANs→diffusion) suggests today's self-supervised pipelines are transient.
Justifications
SL — the most successful paradigms are historically the most transient; pretraining's current dominance follows the same pattern as GAN dominance before displacement
Antecedents (all must be IN):
- IN pretraining-is-transfer-learning-at-scale — Modern self-supervised pretraining is transfer learning at industrial scale — the formal transfer learning framework (source domain D_S → target domain D_T) exactly describes the pretrain-then-finetune pipeline, unifying a 50-year-old theoretical concept with the dominant modern training methodology.
- IN dominant-paradigms-empirically-fragile-and-transient — The most successful ML paradigms are simultaneously dominant and fragile — pretrain-then-finetune is standard practice yet empirically hurtful in some transfer settings, GANs dominated generative modeling for years yet were displaced by diffusion — suggesting that current best practices are locally optimal recipes liable to succession rather than fundamental principles.
Dependents
These beliefs depend on this one:
- IN pretraining-dominance-hardware-contingent — Modern pretraining's dominance reflects hardware economics, not paradigm maturity — it is simultaneously the most successful ML methodology (transfer learning at industrial scale) and the most hardware-dependent (scaling selected it over theoretically superior alternatives), making it uniquely vulnerable to displacement by the next hardware transition, just as transformers' GPU synergy displaced RNN-based approaches.