head-tail-divide-scale-validated
OUT derived (depth 5)
Created 2026-08-25T03:08:54+00:00
The head-tail geometric divide (parametric vs. contextual knowledge split) is a genuine architectural property that generalizes across model scales (334M→20B) for single-hop factual knowledge, but its extension to multi-hop compositional queries is not yet validated.
Justifications
SL — The geometric divide (depth-4) combined with scale invariance evidence establishes generality, but the multi-hop drop (7.4% on MQuAKE-CF) is a negative claim that, if IN, retracts the generalization claim.
Antecedents (all must be IN):
- OUT head-tail-geometric-divide — The parametric/contextual knowledge split is geometrically grounded rather than merely frequency-driven: head-of-distribution facts occupy individually addressable directions in the covariance-whitened feature space (enabling rank-one editing), while long-tail facts are distributed across superposed features where no single direction isolates the knowledge, making external retrieval the only faithful access mechanism.
- IN rome-scale-invariance-334m-to-20b — The two-site causal pattern (early MLP at last subject token + late attention) persists across GPT-2 Medium (334M), GPT-2 Large (774M), GPT-2 XL (1.5B), GPT-J (6B), and GPT-NeoX (20B), though peak layer indices shift.
Unless (any of these IN defeats this justification):
- IN rome-gptj-mquake-cf-multi-hop-drop — ROME-edited GPT-J answers only 7.4% of MQuAKE-CF multi-hop questions, down from 40.5% before editing