conflict-dataset-sizes
IN premise — summaries/2026/08/24/xu-2024-knowledge-conflicts-survey-s1-introduction.md
Created 2026-08-25T02:59:01+00:00
Benchmark dataset sizes per the survey include KC (9,803 CM examples), KRE (11,684 CM), WikiContradiction (2,210 IC), Pan et al. 2023a (52,189 IC), and PARAREL (328 IM).
Summary
This records the current landscape of available benchmark data for detecting contradictions and inconsistencies across five datasets, ranging from a few hundred to over fifty thousand examples. It matters because the wide size spread and the fact that no single dataset dominates means any evaluation result needs to be interpreted relative to which dataset and task type it came from, and smaller datasets leave less room for reliable generalization.