mallen-2023-evaluates-10-lms-three-families

IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s2-related-work.md

Created 2026-08-25T02:58:10+00:00

Mallen et al. 2023 evaluates 10 language models across three families (GPT-Neo 125M–2.7B, OPT 125M–66B, GPT-3 instruct/curie/babbage/ada/davinci-003) using zero-shot or few-shot prompting without fine-tuning.

Summary

Mallen et al. ran a hands-on comparison of ten language models from three different model families, testing each one purely through prompting — asking directly or giving a few example demonstrations — rather than doing any custom training. This isolates what each model can genuinely do out of the box across different sizes and architectures, without the confounding factor of fine-tuning hiding or inflating real capability.