icl-non-distinguishable-error-bound-o-of-1-over-k

IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s3-theoretical-analysis.md

Created 2026-08-25T02:58:56+00:00

In the non-distinguishable regime, the excess 0-1 risk of the in-context predictor scales as O(1/k) (i.e., L_{0-1}(f_n) ≤ inf_f L_{0-1}(f) + g⁻¹(O(ε/k))), where k is the per-example length and g is the multiclass logistic calibration function.

Summary

When individual examples are hard to tell apart, the in-context learner's extra error over the theoretical best predictor shrinks in proportion to 1/k, so longer per-example sequences directly reduce the performance gap. This gives the system a concrete upper bound on how much in-context learning can suffer in that difficult regime, with the gap tied to the model's calibration curve and sequence length.