bert-base-config-12l-768h-110m-params

IN premisesummaries/2026/08/24/wiki-BERT_language_model-chunk-1.md

Created 2026-08-24T17:11:05+00:00

BERT-BASE has 12 Transformer layers, 768 hidden size, 12 attention heads, 3072 feed-forward size, and 110M parameters.

Summary

This pins down the exact size and shape of the BERT-BASE model, which serves as the standard reference point in the system. Every later comparison, capacity calculation, or inference about what the model can and cannot do depends on these numbers being locked in as the baseline configuration.