modern-pretraining-dominant-but-transient

IN derived (depth 3)

Created 2026-06-21T10:23:12+00:00 · Reviewed 2026-06-21T15:37:01+00:00

Modern pretraining as industrial-scale transfer learning is simultaneously the most successful ML methodology and the most likely to be displaced — its dominance rests on empirically fragile foundations (pretraining can hurt), and the broader pattern of paradigm succession (GANs→diffusion) suggests today's self-supervised pipelines are transient.

Justifications

SL — the most successful paradigms are historically the most transient; pretraining's current dominance follows the same pattern as GAN dominance before displacement

Antecedents (all must be IN):

  • IN pretraining-is-transfer-learning-at-scale — Modern self-supervised pretraining is transfer learning at industrial scale — the formal transfer learning framework (source domain D_S → target domain D_T) exactly describes the pretrain-then-finetune pipeline, unifying a 50-year-old theoretical concept with the dominant modern training methodology.
  • IN dominant-paradigms-empirically-fragile-and-transient — The most successful ML paradigms are simultaneously dominant and fragile — pretrain-then-finetune is standard practice yet empirically hurtful in some transfer settings, GANs dominated generative modeling for years yet were displaced by diffusion — suggesting that current best practices are locally optimal recipes liable to succession rather than fundamental principles.

Dependents

These beliefs depend on this one: