dola-dynamic-layer-selection

IN premise — summaries/2026/08/24/xu-2024-knowledge-conflicts-survey-s5-challenges-and-future-directions.md

Created 2026-08-25T02:59:03+00:00

DoLa uses a dynamic layer-selection strategy to choose appropriate premature/mature layers per token, distinguishing it from static contrastive decoding with fixed layer splits.

Summary

DoLa picks which early and late layers of the model to compare on a word-by-word basis, rather than locking in one fixed pair for the whole generation. This matters because it lets the system adapt how much it leans on the model's "second thoughts" versus its "first impressions" depending on the specific token, which can improve output quality where a one-size-fits-all split would over- or under-correct.