zhou-eval-datasets-nq-rtqa-retacred

IN premise — summaries/2026/08/24/zhou-2023-context-faithful-prompting-s1-introduction.md

Created 2026-08-25T02:59:09+00:00

Experiments are conducted on machine reading comprehension datasets (Natural Questions, RealTime QA, SQuAD 2.0, CoQA, QuAC) and relation extraction dataset Re-TACRED.

Summary

The evaluation spans two distinct tasks: answering questions from passages (using five well-known benchmarks like SQuAD and Natural Questions) and identifying relationships between entities in text (using Re-TACRED). This means any performance claims are grounded in results across these specific benchmarks, and the work is not limited to a single type of question-answering or a single language task.