gpt4-13pct-inconsistency-commonsense-qa

IN premise — summaries/2026/08/24/xu-2024-knowledge-conflicts-survey-s4-intra-memory-conflict.md

Created 2026-08-25T02:59:02+00:00

GPT-4 shows approximately a 13% self-inconsistency rate on Commonsense QA tasks (Zhao et al., 2023b).

Summary

GPT-4 contradicts its own answers on basic common-sense questions about one time in eight, so a single output from the model should not be treated as settled reasoning. This gives the system a concrete baseline for how much doubt to build in whenever it relies on the model's answers as evidence rather than ground truth.