modern-llm-pipeline-sustainable

OUT derived (depth 3)

Created 2026-06-21T10:13:05+00:00

Modern LLM training pipelines (self-supervised pretraining → instruction tuning → RLHF) are a sustainable methodology — they dissolve classical paradigm boundaries by successfully combining all three ML paradigms, and self-supervised learning provides an effectively unlimited source of training signal.

Justifications

SL — depth-3 gated — the pipeline's paradigm-dissolving power and self-supervised scaling support sustainability, BUT model collapse from recursive training on synthetic data threatens the foundation if AI-generated content contaminates future pretraining corpora

Antecedents (all must be IN):

  • IN modern-pipelines-dissolve-classical-paradigm-taxonomy — Modern LLM training pipelines dissolve the classical three-paradigm taxonomy — self-supervised pretraining blurs the supervised/unsupervised boundary (its taxonomic status is actively debated), and the full pipeline synthesizes all three paradigms sequentially, suggesting the taxonomy was always a pedagogical convenience rather than a natural partition of learning.
  • IN self-supervised-learning-dominant-pretraining-paradigm — Self-supervised learning is the dominant pre-training paradigm for modern deep learning, as opposed to supervised or unsupervised learning.

Unless (any of these IN defeats this justification):

  • IN ml-model-collapse-synthetic-data — Model collapse is the degradation that occurs when models train on uncurated synthetic data or outputs of prior model versions, also called 'model autophagy disorder (MAD)'