kandpal-entity-linking-trillions-tokens

IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s0-abstract.md

Created 2026-08-25T02:58:05+00:00

Kandpal et al. apply a highly-parallelized entity linking pipeline to trillions of web-crawled tokens to count relevant documents per QA pair.

Summary

Kandpal and colleagues built a fast, massively parallel system that scans trillions of tokens of web text to identify named entities and uses that to count how many documents actually speak to a given question. In practice, this gives the system a concrete, large-scale measurement of how much of the crawled web is genuinely relevant to each QA pair, rather than guessing or sampling.