dai-2023-agnews-highest-learning-rate

IN premise — summaries/2026/08/24/dai-2023-icl-gradient-descent-sA-appendix.md

Created 2026-08-24T17:10:55+00:00

In Dai et al. 2023 finetuning results, the AGNews dataset received the highest learning rate (0.2) in both GPT 1.3B and GPT 2.7B model sizes, compared to much lower rates for sentiment tasks (e.g., 0.0005 for SST2 at 1.3B).

Summary

In that 2023 paper, the authors trained the multi-class news classification task with a learning rate roughly 400 times higher than the binary sentiment task, meaning they took a far more aggressive optimization approach for the former. This matters because it shows the two datasets were tuned under very different conditions, so comparing their raw accuracy numbers directly could be misleading.