mello-73-1-percent-conditional-accuracy

IN premise — summaries/2026/08/24/zhong-2023-mquake-sR-references-chunk-2.md

Created 2026-08-25T02:59:07+00:00

MeLLo answers 73.1% of multi-hop questions correctly when all associated edited facts are successfully retrieved, isolating the LLM reasoning ceiling from retrieval failures.

Summary

Even when every relevant fact is perfectly retrieved, the LLM itself only gets about 73% of multi-step questions right, which sets a hard ceiling on overall system performance. This matters because it lets you tell whether a wrong answer is a retrieval problem or a reasoning problem, and it tells you that no amount of retrieval tuning can push accuracy above that 73% mark.