zhou-2023-smaller-llm-opinion-degrades-abstention
IN premise — summaries/2026/08/24/zhou-2023-context-faithful-prompting-s4-experiments.md
Created 2026-08-25T02:59:10+00:00
For smaller LLMs (LLaMA-2-7B-chat), opinion-based prompting can degrade abstention performance (higher Brier score) because it converts uncertain-but-correct predictions to 'I don't know'.
Summary
For smaller language models, asking them to express opinions or confidence can backfire: instead of improving calibration, it pushes the model to withhold answers it would have gotten right, simply because it felt uncertain. In practice, this means that prompting strategies that work well for large models may actively hurt a 7-billion-parameter model's ability to be both accurate and honest about its limits.