gd-accuracy-rte-8-vs-512-demos

IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s6-related-work.md

Created 2026-08-25T02:58:33+00:00

Gradient descent fine-tuning with 8 demos achieves 0.36 accuracy on RTE, rising to 0.65 with 512 demos.

Summary

On the RTE entailment task, the number of example demonstrations used for fine-tuning makes a large practical difference: a handful of examples leaves the model barely above a third on accuracy, while five hundred examples push it past sixty percent. This sets an observed baseline showing that data volume, not just the choice of training procedure, is a major driver of how well the model performs on this task.