rome-counterfact-dataset-size
IN premise — summaries/2026/08/24/meng-2022-rome-s3-interventions-on-weights-for-understanding-factual-associati.md
Created 2026-08-25T02:58:15+00:00
The COUNTERFACT evaluation dataset contains 21,919 counterfactual records including subjects, relations, paraphrases, neighborhood prompts, and generation prompts.
Summary
The COUNTERFACT benchmark is a sizeable evaluation set of roughly 22,000 what-if scenarios, each fully annotated with the subject, the relationship, rephrasings, and the specific prompts needed to test counterfactual generation. This gives the system a concrete, repeatable yardstick for measuring how well it handles hypothetical facts rather than just verified real-world ones.