cad-benchmarks-used

IN premise — summaries/2026/08/24/shi-2024-context-aware-decoding-s5-related-work.md

Created 2026-08-25T02:58:36+00:00

CAD is evaluated on summarization benchmarks (XSUM, CNN-DM) and knowledge-conflict benchmarks (Memotrap, NQSWAP) across OPT models of varying sizes.

Summary

CAD's performance is being measured on two distinct kinds of tasks—condensing long texts into summaries and handling situations where the model's own knowledge clashes with the information in front of it—and the results are checked across several different model sizes. This matters because it tells us whether CAD's behavior is consistent and generalizable, rather than an artifact of a single task type or one particular model scale.