distillation-validated-across-full-scale-spectrum

IN derived (depth 1)

Created 2026-06-21T12:50:29+00:00 · Reviewed 2026-06-21T14:41:08+00:00

Knowledge distillation is validated as a scale-invariant capability across the full spectrum of language model sizes: from BERT-scale (DistilBERT retaining 95% performance at 60% of parameters) to frontier-scale (Llama 4 Maverick codistilled from the unreleased ~2T-parameter Behemoth), demonstrating that larger models reliably compress their capability into smaller ones regardless of absolute scale.

Justifications

SL — Distillation working at both ends of the scale spectrum (110M→66M and ~2T→400B) validates it as a general property of neural language models, not a small-model artifact

Antecedents (all must be IN):

Dependents

These beliefs depend on this one: