depth-confounder-scales-as-sqrt-log-m

IN premise — summaries/2026/08/24/aristotelian-2026-s4-theoretical-motivation-spurious-alignment.md

Created 2026-08-24T17:10:50+00:00

The depth confounder from max-aggregation over M = L_A·L_B layer pairs inflates the expected maximum under H₀ as E[T_max] ≤ μ + Cσ√(log M), meaning deeper models receive systematically higher raw alignment scores purely from a larger search space.

Summary

When comparing alignment across many layer pairs, deeper models accumulate a larger pool of candidates, so the single best score naturally creeps upward just from having more to search through, even under the null hypothesis. This means raw alignment scores systematically favor deeper architectures, and any cross-depth comparison needs a correction for the growing search space or it will mistake a statistical artifact for genuine improvement.