gd-construction-and-trained-weights-same-loss-basin
IN premise — summaries/2026/08/24/von-oswald-2023-icl-gd-sR-references.md
Created 2026-08-24T17:11:05+00:00
After correcting for scalar ambiguity in the W_K·W_Q and P·W_V products, a 50/50 linear interpolation between the analytical GD-construction weights and trained Transformer weights produces minimal loss increase, confirming both occupy the same loss basin (linear mode connectivity).
Summary
Mixing the analytically constructed weights with the trained Transformer weights half-and-half barely worsens the model's performance, which means the two sets of weights are essentially different coordinates for the same underlying solution rather than two different solutions. This matters because it validates that the construction method captures what training actually produces, giving confidence that the analytical approach is a faithful model of the learning process rather than just a toy approximation.