chameleon-datasets-popqa-strategyqa
IN premise — summaries/2026/08/24/xie-2024-chameleon-sloth-s0-abstract-chunk-1.md
Created 2026-08-25T02:58:58+00:00
The LLM knowledge conflict experiments use POPQA (14K entity-centric QA from Wikidata) and STRATEGY QA (multi-step True/False reasoning from Wikipedia).
Summary
The knowledge conflict tests draw on two question sets with different shapes: one checks whether the model can correctly recall facts about specific entities (people, places, things) at scale, and the other requires chaining multiple facts together to reach a yes-or-no conclusion. Together they cover both simple factual lookup and multi-step reasoning, so the results show where the model's knowledge conflicts surface under different cognitive demands, not just one narrow skill.