xie-2021-gpt3-lambada-triviaqa-improvement

IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s0-abstract.md

Created 2026-08-25T02:58:55+00:00

GPT-3 (Brown et al., 2020) improved LAMBADA by 18% and TriviaQA by 3% over prior state-of-the-art.

Summary

GPT-3 marked a real step forward in how well language models track context over long passages, boosting performance on the LAMBADA benchmark by 18 percent, while gains on trivia-style factual recall were much smaller at 3 percent. This gives the system a concrete baseline for measuring whether later models actually built on that progress or merely matched it.