xu2024-gpt4-faviq-32pct-inconsistency
IN premise — summaries/2026/08/24/xu-2024-knowledge-conflicts-survey-sR-references-chunk-3.md
Created 2026-08-25T02:59:04+00:00
GPT-4 exhibits a 32% inconsistency rate on the FaVIQ benchmark (Zhao et al. 2023b), indicating no model is fully immune to intra-memory conflicts.
Summary
Even the strongest current language model contradicts itself in nearly a third of cases when tested on a standard question-answering benchmark, which means the system cannot assume any single model's outputs are internally coherent and must actively check for and resolve self-contradictions rather than trusting them at face value.