head-tail-divide-scale-validated

OUT derived (depth 5)

Created 2026-08-25T03:08:54+00:00

The head-tail geometric divide (parametric vs. contextual knowledge split) is a genuine architectural property that generalizes across model scales (334M→20B) for single-hop factual knowledge, but its extension to multi-hop compositional queries is not yet validated.

Justifications

SL — The geometric divide (depth-4) combined with scale invariance evidence establishes generality, but the multi-hop drop (7.4% on MQuAKE-CF) is a negative claim that, if IN, retracts the generalization claim.

Antecedents (all must be IN):

  • OUT head-tail-geometric-divide — The parametric/contextual knowledge split is geometrically grounded rather than merely frequency-driven: head-of-distribution facts occupy individually addressable directions in the covariance-whitened feature space (enabling rank-one editing), while long-tail facts are distributed across superposed features where no single direction isolates the knowledge, making external retrieval the only faithful access mechanism.
  • IN rome-scale-invariance-334m-to-20b — The two-site causal pattern (early MLP at last subject token + late attention) persists across GPT-2 Medium (334M), GPT-2 Large (774M), GPT-2 XL (1.5B), GPT-J (6B), and GPT-NeoX (20B), though peak layer indices shift.

Unless (any of these IN defeats this justification):