gd-learning-rates-tested-shen-2023

IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s8-icl-demonstrations.md

Created 2026-08-25T02:58:33+00:00

Learning rates tested for the GD/cGD simulation in Shen et al. (2023) were 5e-3, 1e-3, 5e-4, and 1e-4, with training up to ~175 epochs.

Summary

Shen et al. (2023) only tested four specific learning rates (0.005, 0.001, 0.0005, and 0.0001) and trained for roughly 175 epochs, so any conclusions from their gradient-descent simulation are bounded by that narrow search space. This matters because results outside those settings are not supported by their work, and anyone comparing or extending their findings needs to treat the untested range as unknown.