prh-caption-density-improves-vision-language-alignment
IN premise — summaries/2026/08/24/huh-2024-prh-s6-counterexamples-and-limitations.md
Created 2026-08-24T17:10:56+00:00
Using the Densely-Captioned-Images dataset with caption lengths of 5, 10, 20, and 30 words (summarized with LLaMA3-8B-Instruct), higher caption density yields better mutual nearest-neighbor vision-language alignment scores, supporting the information-ceiling sub-hypothesis that more bijective mappings improve convergence.
Summary
When image descriptions are longer and more detailed, vision-language models lock onto each other more tightly, which means richer captions give the model more anchor points to connect what it sees with what it says. This matters practically because it points toward a clear design choice: investing in more thorough, one-to-one image descriptions during training should produce better multimodal alignment rather than just adding noise.