whisper-trained-680k-hours-weak-supervision
IN premise — summaries/2026/08/24/wiki-Transformer_deep_learning_architecture-chunk-6-chunk-2.md
Created 2026-08-24T17:11:27+00:00
Whisper (Radford et al., 2022, arXiv:2212.04356) is an encoder-decoder Transformer trained on 680,000 hours of weakly supervised speech data for robust automatic speech recognition.
Summary
This records that OpenAI's Whisper model exists and works: it turns spoken audio into text using a Transformer architecture, trained on roughly 680,000 hours of real-world speech that wasn't perfectly labeled, which is why it handles noisy and diverse recordings well. It matters to the system as a foundation fact — anything that depends on automatic transcription, audio understanding, or multimodal processing is built on this capability.