gpt4-multimodal-input-text-only-output-gpt4o-full-multimodal

IN premisesummaries/2026/08/24/wiki-Generative_pre-trained_transformer-chunk-1.md

Created 2026-08-24T17:11:10+00:00

GPT-4 (2023) accepts text and image input but produces text-only output; GPT-4o (2024) processes and generates text, images, and audio in both directions.

Summary

GPT-4 (2023) can understand text and images but can only reply in text, whereas GPT-4o (2024) handles text, images, and audio in both directions. This capability gap is a hard design constraint: any system that needs a model to produce speech, generate images, or respond in audio must use the newer architecture, and assumptions about what a given model can output should never cross this boundary.