popqa-constructed-from-wikidata-triples
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s3-collect-popularity.md
Created 2026-08-25T02:58:10+00:00
POPQA is constructed by sampling knowledge triples from Wikidata, converting them to natural-language questions via manually written templates, and computing popularity scores via the Wikipedia API.
Summary
POPQA is a question-answering dataset where each question is generated by filling in a hand-written sentence template with a fact pulled from Wikidata, and which facts get selected is guided by how well-known they are according to Wikipedia. This means the questions have a predictable, templated structure and their coverage is limited to whatever is already encoded in Wikidata, so the dataset reflects both the shape of those templates and the popularity bias of Wikipedia rather than open-ended natural language.