lora-best-weight-selection-wq-wv-rank4
IN premise — summaries/2026/08/24/hu-2021-lora-s7-u-nderstanding-the-low-rank-updates.md
Created 2026-08-24T17:10:55+00:00
On GPT-3 175B, adapting both Wq and Wv at rank 4 outperforms adapting a single weight type (Wq alone or Wk alone) at rank 8 under a fixed 18M parameter budget
Summary
When you have a fixed budget of about 18 million extra parameters for fine-tuning a large model, you get better results by splitting that budget across two parts of the attention mechanism (the query and value paths) rather than dumping it all into a single path. In practice, this means a model-adaptation strategy should favor broad, shallow coverage of multiple attention components over deep, narrow modification of just one.