long-tail-as-coordinate-coverage-gap-not-capacity-gap

OUT derived (depth 6)

Created 2026-08-25T03:45:56+00:00 · Reviewed 2026-08-25T04:02:18+00:00

The long-tail knowledge failure is fundamentally a coordinate-coverage problem (the rank-one addressable subspace does not span rare fact directions) rather than a raw parameter-capacity problem, making the 10¹⁵-parameter scaling estimate a misleading reframing of a geometric coverage gap.

Justifications

SL — The addressability-failure claim identifies the long-tail as a geometric issue; the subspace-boundary claim pinpoints the exact geometric mechanism (rank-one updates cannot span the rare-fact directions). Together they reframe Kandpal's scaling law as a coordinate-coverage problem: adding parameters (more directions) is the wrong fix if the issue is that the existing directions don't span the target. Both are load-bearing—the failure without the mechanism is vague, the mechanism without the failure is abstract.

Antecedents (all must be IN):

  • OUT long-tail-as-geometric-addressability-failure — The long-tail knowledge problem (Kandpal's 10¹⁵-parameter estimate, 176B-model failure on rare facts) is fundamentally a geometric addressability failure rather than a data-scarcity or model-capacity issue: tail facts fail to acquire well-conditioned individual directions in the residual stream because their key vectors lie in the poorly-conditioned tail of the covariance spectrum, making them inaccessible to both parametric recall (no clean MLP slot) and rank-one editing (C⁻¹k* becomes ill-conditioned), and correctly routed to the contextual channel instead.
  • OUT parametric-write-subspace-boundary — The operational boundary between parametric recall and contextual retrieval is precisely the geometric boundary of the rank-one addressable subspace: facts whose subject-key projection aligns with the locally-stored key covariance are parametrically editable, while facts outside this subspace must be externally supplied.

Dependents

These beliefs depend on this one: