fv-max-aie-decreases-with-model-size

IN premise — summaries/2026/08/24/todd-2023-function-vectors-s2-function-vector-causal-effects-cannot-be-recovered-from-the--chunk-3.md

Created 2026-08-25T02:58:41+00:00

Maximum AIE decreases with model size: GPT-J ≈ 0.053, Llama 2 7B ≈ 0.047, Llama 2 70B ≈ 0.037, despite increasing total head count in larger models.

Summary

Larger language models show less of the measured adversarial influence effect, even though they have more attention heads to potentially amplify it. This matters because it means you cannot assume that scaling a model automatically increases its vulnerability to this type of interference; in fact, the opposite trend holds across the tested family.