prompt-sensitivity-40-percent-accuracy-shift

IN premiseentries/2026/06/21/wiki-Prompt_engineering.md

Created 2026-06-21T09:50:10+00:00

LLM performance is highly sensitive to prompt design, with accuracy shifts of over 40 percentage points from minor changes such as reordering examples, and up to 76 accuracy points difference across formatting changes

Summary

Small changes to how a question is phrased — like rearranging example answers or tweaking formatting — can swing a language model's accuracy by more than 40 points, meaning prompt wording is not a cosmetic detail but a first-order variable that can flip a model from reliable to unreliable. This implies that any comparison between models, or any evaluation of a system's real-world capability, must treat prompt structure as a controlled input rather than an afterthought.

Dependents

These beliefs depend on this one: