kalai-2023-semantic-calibration-computationally-intractable

IN premise — summaries/2026/08/24/kalai-2023-hallucination-inevitable-s9-conclusions-limitations-and-future-work.md

Created 2026-08-24T17:11:00+00:00

The paper's calibration metric Mis_b(g,p) = ‖p_{V_b(g)} − g‖_TV is defined at the semantic/fact level rather than per-token probability, making it computationally intractable to evaluate on large models — a stated limitation distinct from standard token-level calibration used in classification.

Summary

The paper measures whether a model's confidence matches its actual chance of getting a full answer right, which is a more meaningful test than checking individual word probabilities, but it comes at a cost: computing this metric is so expensive that it becomes impractical for large models. This is a known, acknowledged limitation of the approach rather than a flaw in the theory.