mello-8-demos-per-relation-icL-filtering

IN premise — summaries/2026/08/24/zhong-2023-mquake-sR-references-chunk-1.md

Created 2026-08-25T02:59:08+00:00

MQuAKE dataset construction uses 8 in-context demonstration examples per relation type with GPT-J to filter out facts the model cannot recall before inclusion in the benchmark.

Summary

The MQuAKE benchmark was built by showing a large language model eight example questions for each relationship type and then dropping any fact the model still couldn't produce, so the final test set only contains questions the model is near the edge of recalling. This matters because it means the dataset is pre-tuned to a specific model's recall boundary, which makes it harder to generalize results to other models or to claim the benchmark measures open-ended reasoning rather than familiarity.