bert-pretraining-corpora-bookcorpus-wikipedia
IN premise — summaries/2026/08/24/wiki-BERT_language_model-chunk-1.md
Created 2026-08-24T17:11:06+00:00
BERT was pre-trained on BookCorpus (800M words) and filtered English Wikipedia (2,500M words).
Summary
BERT's foundational training drew on roughly 3.3 billion words from published books and English Wikipedia articles, which means its language patterns, factual grounding, and stylistic biases are shaped by relatively formal, well-written prose rather than casual or domain-specific text. Any reasoning about what BERT knows, where its blind spots are, or why it favors certain phrasings should trace back to this specific corpus as the starting point.