bert-sota-2018-glue-squad-swag
IN premise — summaries/2026/08/24/wiki-BERT_language_model-chunk-1.md
Created 2026-08-24T17:11:06+00:00
BERT set new state-of-the-art results on GLUE (9 tasks), SQuAD v1.1, SQuAD v2.0, and SWAG benchmarks at its 2018 release.
Summary
When BERT arrived in 2018, it outperformed every prior method across a wide spread of language tasks, from reading comprehension to commonsense reasoning, showing that a single well-trained model could do what previously required many specialized systems. This became the benchmark the system must reference whenever evaluating whether a later architecture actually represents progress.