kalai-vempala-calibrated-lm-must-hallucinate-monofacts-bound

IN premise — summaries/2026/08/24/kalai-2023-hallucination-inevitable-s1-introduction.md

Created 2026-08-24T17:10:59+00:00

Kalai and Vempala (2023, OpenAI/Georgia Tech) prove that any language model satisfying semantic-level calibration must hallucinate at a rate ≥ dMF − Miscalibration − 300·|Facts|/|Possible hallucinations| − 7/√n, even with i.i.d. training data, no factual errors, and no architectural constraints.

Summary

This result shows that hallucination in language models is not a solvable engineering bug but a mathematical floor: even a perfectly trained, perfectly calibrated model with clean data and no architectural flaws is forced to state false facts at some minimum rate. In practical terms, no amount of further tuning, data cleaning, or architectural redesign can drive a model's factual error rate to zero, which means downstream systems must always treat model output as potentially wrong and plan verification accordingly.