mha-vs-gqa-head-level-ablation-divergence
IN premise — summaries/2026/08/24/convergence-without-understanding-2026-s3-results.md
Created 2026-08-24T17:10:52+00:00
Maximum per-head causal ablation flip rates differ by attention architecture: multi-head attention (MHA) models show 43%–63% while grouped-query attention (GQA) models show 20%, a difference invisible to full-hidden-state CKA measurement.
Summary
Standard representational comparison tools (like CKA on full hidden states) fail to reveal how differently individual attention heads actually drive behavior in multi-head versus grouped-query architectures. This matters because a system that relies on those tools would conclude the two designs are functionally similar when, in fact, ablating a single head in the multi-head design changes outcomes roughly twice as often, meaning the components are far less redundant than the numbers suggest.