pretraining-dominance-hardware-contingent

IN derived (depth 4)

Created 2026-06-21T10:27:01+00:00 · Reviewed 2026-06-21T15:37:01+00:00

Modern pretraining's dominance reflects hardware economics, not paradigm maturity — it is simultaneously the most successful ML methodology (transfer learning at industrial scale) and the most hardware-dependent (scaling selected it over theoretically superior alternatives), making it uniquely vulnerable to displacement by the next hardware transition, just as transformers' GPU synergy displaced RNN-based approaches.

Justifications

SL — Pretraining's transience (d3) explained by scalability-over-theory principle (d3) — dominance contingent on current hardware regime

Antecedents (all must be IN):

  • IN modern-pretraining-dominant-but-transient — Modern pretraining as industrial-scale transfer learning is simultaneously the most successful ML methodology and the most likely to be displaced — its dominance rests on empirically fragile foundations (pretraining can hurt), and the broader pattern of paradigm succession (GANs→diffusion) suggests today's self-supervised pipelines are transient.
  • IN scalability-trumps-elegance-in-ml — Hardware-architecture co-evolution favored architectures that could exploit parallelism (neural networks) over mathematically complete frameworks with limited parallelism benefits (SVMs). SVMs offered convex guarantees, kernel elegance, and sparse analytical solutions — a degree of mathematical closure few ML paradigms achieve — but neural networks' ability to scale with massive compute increases (300,000x from AlexNet to AlphaZero) was a significant factor in deep learning's dominance. This suggests engineering scalability became a major selection criterion for ML prominence, though the relative importance of compute scaling versus algorithmic innovation remains unestablished.

Dependents

These beliefs depend on this one: