llama-4-first-natively-multimodal-in-line
IN premise — summaries/2026/08/24/wiki-LLaMA-chunk-3.md
Created 2026-08-24T17:11:15+00:00
Llama 4 (April 2025) is the first natively multimodal (text + vision) model in the Llama family; Llama 3.2 (September 2024) was the first to add vision via a separate vision encoder.
Summary
Llama 4 is the first in the Llama line built from the ground up to treat text and images as one unified input, rather than tacking on a separate vision module the way Llama 3.2 did. This sets a new architectural baseline: going forward, the family's default expectation is tight cross-modal integration rather than a text model with a vision adapter bolted on.