popqa-accuracy-substring-match

IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s3-collect-popularity.md

Created 2026-08-25T02:58:10+00:00

In POPQA evaluation, a prediction is marked correct if any substring of the generated answer exactly matches any of the gold answers.

Summary

POPQA scores count a model's answer as correct even if the right option just appears as a fragment buried inside a longer or messier generation, not only when the model produces the exact option on its own. This means POPQA accuracy numbers tend to flatter a model's performance, so when comparing models or tracking progress, you should treat those scores as a generous floor rather than a tight measure of what the model actually "knows."