dai-2023-lr-grid-36-candidates-sst5-exception

IN premise — summaries/2026/08/24/dai-2023-icl-gradient-descent-sA-appendix.md

Created 2026-08-24T17:10:54+00:00

Dai et al. 2023 finetuning learning rate search uses a 36-candidate grid (9 base values {1-9} × 4 scales {0.1, 0.01, 0.001, 0.0001}), with one exception: GPT 1.3B on SST5 required a finer search yielding LR = 0.00016.

Summary

In the Dai et al. 2023 finetuning study, learning rates were picked from a fixed grid of 36 standard values spanning four orders of magnitude. The one case that fell outside that grid — GPT 1.3B on SST5, which needed a more precise search — is a reminder that a coarse grid can miss the optimal setting for particular model-and-task pairings.