mello-retrieval-accuracy-bottleneck

IN premise — summaries/2026/08/24/zhong-2023-mquake-sR-references.md

Created 2026-08-25T02:59:08+00:00

MeLLo's multi-hop accuracy drops from 52.1% (retrieval accuracy 88.1%, k=1) to 30.9% (retrieval accuracy 59.7%, k=3000), and reaches 73.1% when all associated facts are correctly retrieved.

Summary

Finding the right facts is the main thing holding MeLLo's multi-hop reasoning back: when it pulls from a small, well-targeted set it does far better than when it has to wade through thousands of candidates, and its ceiling (about 73%) is only reached when every needed fact is actually retrieved. This means the highest-leverage fix for improving overall performance is not the reasoning step itself but the retrieval step that feeds it.