zhou-2023-brier-score-reduction-gpt35

IN premise — summaries/2026/08/24/zhou-2023-context-faithful-prompting-s4-experiments.md

Created 2026-08-25T02:59:10+00:00

OPIN+INSTR reduced the Brier score by 24.2% in zero-shot and 7.8% in few-shot settings for GPT-3.5 (text-davinci-003) compared to the base prompt.

Summary

A particular way of phrasing the prompt — asking for the model's opinion while giving a clear instruction — made GPT-3.5's probability predictions meaningfully closer to reality, cutting its error by about 24% when no examples were shown and about 8% when a few were. The much larger gain in the zero-shot case suggests the technique is most helpful when the model has no examples to calibrate its confidence against.