kandpal-2023-gpt3-openwebtext-proxy

IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-sR-references.md

Created 2026-08-25T02:58:07+00:00

For GPT-3, relevant document counts were estimated using OpenWebText as a proxy corpus because GPT-3's actual training data is private.

Summary

Because OpenAI never published what documents actually went into GPT-3, researchers borrowed the size and composition of the public OpenWebText dataset as a stand-in to estimate how many relevant documents the model likely saw. This means any downstream claim that depends on GPT-3's corpus size is resting on an approximation, not a verified fact, and the margin of error is unknown.