hella-swag-incorrect-options-sampled-from-lm
IN premise — summaries/2026/08/24/wiki-Large_language_model-chunk-3.md
Created 2026-08-24T17:11:16+00:00
HellaSwag's incorrect answer options in its multiple-choice video-description completion task are generated by sampling from a language model, making them plausible to LMs but trivial for humans to distinguish from the correct option
Summary
In HellaSwag, the wrong multiple-choice answers are crafted by another language model, so they sound convincing to AI systems even though a person could spot the correct option almost instantly. This means the benchmark is specifically calibrated to test whether a language model can do something that is trivial for a human but genuinely hard for a machine, making it a sensitive measure of model-level common sense rather than general knowledge.