kandpal-2023-gpt3-data-not-public-owt-proxy
IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s2-identifying-relevant-pre-training-data.md
Created 2026-08-25T02:58:06+00:00
GPT-3 pre-training data is not public, so the authors approximate GPT-3 relevant document counts by scaling OpenWebText counts, introducing acknowledged uncertainty.
Summary
Because no one outside the model's builders can see exactly what documents went into training GPT-3, the paper's document-count estimates are built on a stand-in dataset scaled to approximate the real corpus. This means every number downstream is an informed approximation with acknowledged error, not a direct measurement, so conclusions built on them should be treated as rough estimates rather than definitive counts.