llama1-13b-outperformed-gpt3-175b

IN premiseentries/2026/06/21/wiki-LLaMA-chunk-1.md

Created 2026-06-21T09:50:09+00:00

Llama 1 13B outperformed GPT-3 175B on most NLP benchmarks, demonstrating the value of data scaling over parameter scaling

Summary

A model with roughly one thirteenth the parameters of GPT-3 beat it on most language tasks, which means that feeding a model more and better training data can matter more than simply making the model bigger. This challenges the assumption that bigger models are automatically better and suggests that resource budgets should weigh data quality and volume as heavily as raw parameter count.

Dependents

These beliefs depend on this one: