kalai-2023-dmf-is-tight-hallucination-lower-bound
IN premise — summaries/2026/08/24/kalai-2023-hallucination-inevitable-s7-upper-bounds-on-hallucination-rate.md
Created 2026-08-24T17:10:59+00:00
The monofact rate (dMF), the fraction of factoids appearing exactly once in training data, is the tight (optimal) lower bound on the hallucination rate of any calibrated language model — no algorithm can guarantee a hallucination rate significantly below dMF while maintaining calibration.
Summary
There is a hard floor on how rarely a language model can hallucinate, and that floor is set by how much of its training data consists of facts that showed up only once. No clever algorithm can push the error rate below that number without giving up calibration, meaning the one-off facts in the training set are the irreducible source of the model's confabulation.