bert-masked-lm-gpt-causal-lm-same-year-2018

IN premiseentries/2026/06/21/wiki-Transformer_deep_learning_architecture-chunk-7.md

Created 2026-06-21T09:55:55+00:00

BERT (2018) uses masked language modeling (bidirectional) while GPT (2018) uses causal language modeling (autoregressive left-to-right); both published the same year with opposite pre-training strategies

Dependents

These beliefs depend on this one: