zhou-2023-evaluation-two-tasks-three-datasets

IN premise — summaries/2026/08/24/zhou-2023-context-faithful-prompting-s5-conclusion.md

Created 2026-08-25T02:59:11+00:00

The evaluation spans two tasks—machine reading comprehension (knowledge conflict) and relation extraction (prediction with abstention)—on three datasets, using two model scales (175B InstructGPT and 7B LLaMA-2-chat) in both zero-shot and few-shot regimes.

Summary

The evaluation in this paper is deliberately broad, testing both a comprehension task where the model must notice it conflicts with its own background knowledge and a structured extraction task where it can choose to abstain, across three datasets and two very different model sizes. This matters because the range of tasks, data, and scales determines whether the paper's conclusions about model behavior can actually be generalized beyond a single narrow setup.