kalai-hallucination-good-turing-bound

IN premise — summaries/2026/08/24/kalai-2023-hallucination-inevitable-s0-abstract.md

Created 2026-08-24T17:10:58+00:00

For statistically calibrated LMs with bounded maximum fact probability, hallucination probability on singleton facts approximates the fraction of facts occurring exactly once in training data (Good-Turing estimate)

Summary

This says that when a language model hallucinates a rare one-off fact, the likelihood of doing so is roughly proportional to what fraction of its training material consists of facts it has only ever seen once. In practical terms, the model's hallucination rate on obscure facts is not some mysterious, unmeasurable quantity; it can be estimated directly from the composition of the training data using a classic statistical shortcut, giving engineers a concrete lever to reduce hallucinations by addressing how much single-occurrence content the model absorbed during training.