zhou-eval-datasets-nq-rtqa-retacred
IN premise — summaries/2026/08/24/zhou-2023-context-faithful-prompting-s1-introduction.md
Created 2026-08-25T02:59:09+00:00
Experiments are conducted on machine reading comprehension datasets (Natural Questions, RealTime QA, SQuAD 2.0, CoQA, QuAC) and relation extraction dataset Re-TACRED.
Summary
The evaluation spans two distinct tasks: answering questions from passages (using five well-known benchmarks like SQuAD and Natural Questions) and identifying relationships between entities in text (using Re-TACRED). This means any performance claims are grounded in results across these specific benchmarks, and the work is not limited to a single type of question-answering or a single language task.