cad-gain-scales-with-model-size

IN premise — summaries/2026/08/24/shi-2024-context-aware-decoding-s4-results.md

Created 2026-08-25T02:58:35+00:00

On knowledge-conflict tasks (MemoTrap, NQ-Swap), the relative performance gain from CAD increases with model size, indicating larger models rely more heavily on prior knowledge.

Summary

Bigger language models lean harder on what they already "know" from training, which means they are more likely to answer from stale memory rather than the information actually in front of them. This observation matters because it shows that CAD corrections become more important as you scale up models, not less, since the risk of a large model defaulting to its prior knowledge grows with size.