icl-prompt-ordering-sensitivity-10-40pct
IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s4-simulations.md
Created 2026-08-25T02:58:56+00:00
Permuting the same 4 in-context prompt examples on GINC yields 10–40% variation in accuracy, confirming that ICL accuracy is not permutation-invariant and is sensitive to example ordering.
Summary
The order in which you list your example prompts can swing model accuracy by as much as 40 percent, even when the exact same four examples are used. That means prompt performance is not a fixed property of the example set itself; it depends on sequencing, so any system relying on in-context learning needs to treat ordering as a real design variable rather than an incidental detail.