Part of History of Language AI
Whisper applies encoder-decoder transformers to multilingual speech recognition, translation, and language identification across noisy audio conditions.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
2022: Whisper
OpenAI released Whisper in September 2022 as a family of encoder-decoder speech models. The models were trained on about 680,000 hours of audio paired with text collected from the web. One checkpoint could transcribe several languages, translate non-English speech into English, identify the spoken language, and predict timestamps.
The paper called the method weak supervision because the dataset was much larger and less tightly curated than a conventional speech corpus. Whisper was evaluated without fine-tuning on each benchmark. Its zero-shot results were competitive across many datasets, although accuracy varied sharply by language and recording condition.
Whisper used a familiar transformer architecture rather than a new speech-specific network. The notable choice was to put many supervised speech tasks into one token-prediction format and train at web scale. OpenAI released several model checkpoints and inference code, which made the system practical to run locally and straightforward to adapt.
The Problem
An ASR model can score well on the corpus it was tuned for and fail on a different microphone, accent, or speaking style. Training separate systems for every domain reduces that mismatch, but it multiplies data collection and maintenance work. The same problem becomes more expensive when each language has its own pipeline.
Labeled speech is also costly. Conventional corpora use carefully segmented recordings and checked transcripts. They provide clean supervision but cover only a small part of the audio found in actual use. Low-resource languages receive less coverage because suitable recordings and transcripts are scarce.
Speech translation added another pipeline. A system might first transcribe the source language and then pass the text to a translation model. Errors from the first stage become input to the second. Whisper instead learned direct speech-to-English translation in the same sequence model used for transcription.
The Solution
Whisper converted each task into conditional text generation. Audio entered the encoder, while control tokens told the decoder which language and task to use. The shared model could therefore switch behavior without loading a separate network for each request.
Architecture Design
Audio is resampled to 16 kHz and divided into 30-second windows. Whisper converts each window to an 80-channel log-Mel spectrogram. Two convolutional layers process the spectrogram before a transformer encoder produces the audio representation. A transformer decoder then predicts text tokens.
The decoder prompt contains control tokens. A language token records the detected or supplied language. A task token selects transcription or translation into English, and another token selects whether timestamps should be emitted. This framing lets one set of weights learn the related tasks together.
Training Data Collection
The 680,000-hour dataset consisted of audio-transcript pairs gathered from the internet. About 438,000 hours were English transcription data. The paper reported 117,000 hours of transcription in other languages and 125,000 hours in which non-English speech was paired with an English translation.
Web data brought varied recording conditions at a scale unavailable in hand-labeled benchmarks, but it also brought unreliable labels. The authors applied automated filters for language agreement, transcript quality, and duplicate content. They also tried to remove machine-generated transcripts. These filters reduced obvious errors without making the corpus equivalent to manually verified speech data.
Task-Specific Training
Training examples used the same decoder vocabulary but different token prefixes. An English transcription example and a French-to-English translation example therefore exercised the same network with different instructions. Timestamp tokens taught the decoder to align spans of text with positions in the 30-second audio window.
OpenAI trained model sizes from Tiny through Large. The family exposed a practical tradeoff: smaller checkpoints need less memory and run faster, while larger checkpoints generally produce lower error rates. English-only variants were also released for several sizes.
Applications and Impact
Because the weights and inference implementation were public, developers could run Whisper without sending recordings to a hosted transcription service. That mattered for offline workflows and for applications where audio could not leave a controlled environment. Local use still required enough compute for the selected checkpoint.
Whisper became a common component for captions, searchable audio archives, and meeting transcripts. Its timestamps supported subtitle generation, although long recordings first had to be divided into the model's 30-second processing windows. The language detector could choose a transcription route, and the translation task could produce English text directly from supported non-English speech.
The release also provided a strong reusable baseline for speech research. Teams could compare a specialized model against Whisper zero-shot or fine-tune a released checkpoint for a narrower domain. That accessibility was a direct consequence of publishing the weights, not evidence that every language received equal accuracy.
Limitations
Error rates were uneven across languages. The amount of training audio differed greatly, and the paper found a relationship between language-specific data volume and transcription quality. A headline language count therefore says little about whether a particular language or dialect is usable.
Whisper can hallucinate text that was not spoken, especially in silence, noise, or ambiguous audio. It may also repeat phrases. A fluent transcript is not proof of an accurate one, so high-consequence use requires checking the output against the recording.
The 30-second window and autoregressive decoder were designed for batch transcription rather than low-latency streaming. Processing long recordings requires windowing and timestamp heuristics. Boundary errors can omit speech or duplicate it across adjacent segments.
Compute remains another constraint. The largest original checkpoint has roughly 1.55 billion parameters. It can be slow on a CPU or a small device, while the faster checkpoints trade some accuracy for lower latency and memory use.
The web-sourced training set was not released, and its composition cannot be independently audited in full. This limits analysis of consent, copyrighted recordings, representational gaps, and label errors. Publishing model weights does not resolve those data questions.
Finally, Whisper's translation task targets English. It does not provide arbitrary speech-to-speech or speech-to-text translation between every pair of supported languages.
Legacy
Whisper showed what large, weakly supervised speech datasets could do with a conventional encoder-decoder transformer. Its cross-dataset zero-shot evaluation shifted attention away from a single in-domain score and toward performance under distribution change.
The task-token interface was equally useful. Transcription, English translation, language detection, and timestamp prediction became variants of one decoding problem. That design made the released checkpoints easier to integrate than a collection of unrelated speech pipelines.
The release also exposed the tradeoff behind web-scale supervision. Greater coverage improved transfer, but uneven data and noisy labels remained visible in language-level error rates and hallucinations. Whisper's lesson is not that scale removes the need for curation or evaluation. It is that scale changes where those problems appear.
Quiz
The following questions review Whisper's training data, architecture, task tokens, and deployment limits.
Whisper Speech Recognition Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore History of Language AIStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!