mallen-2023-models-evaluated
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s4-memorization-depends-on-popularity.md
Created 2026-08-25T02:58:11+00:00
Mallen et al. (2023) evaluate OPT (1.3B, 2.7B, 6.7B, 13B), GPT-Neo (1.3B, 2.7B, 6B, 20B), and GPT-3 (davinci-002, davinci-003) without fine-tuning.
Summary
Mallen and colleagues tested a set of large language models in their stock, off-the-shelf form, without any task-specific training or customization. This matters because it frames the results as reflecting each model's general-purpose ability rather than a tailored solution, which sets the baseline for any comparison or inference that builds on their findings.