entityquestions-82pct-usable
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s3-evaluation-setup.md
Created 2026-08-25T02:58:11+00:00
Only 82% of EntityQuestions questions are retained for evaluation because the remaining 18% lack unique Wikidata entity annotations.
Summary
The EntityQuestions benchmark loses roughly one in six questions because those questions can't be pinned to a single, distinct entity in Wikidata, so they're too ambiguous to score cleanly. Anyone using this dataset should treat its effective size as 82% of the listed count and expect the dropped questions to be the harder, multi-entity or poorly specified ones.