kandpal-2023-dataset-sizes-pile-roots-c4-owt
IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s2-identifying-relevant-pre-training-data.md
Created 2026-08-25T02:58:06+00:00
The entity-linked pre-training datasets used are: The Pile (825 GB), ROOTS English (490 GB), C4 (305 GB), OpenWebText (39 GB), and Wikipedia (December 2018).
Summary
This records exactly which five data sources were mixed together for entity-linked pre-training in the Kandpal 2023 work, along with their storage sizes, so that anyone checking the results knows the full recipe and relative scale of the input data. It matters because the quality and diversity of downstream entity linking depend heavily on this specific combination, and the size differences (The Pile at roughly eight times the size of OpenWebText) shape what the model actually learned.