Back to whisper
OpenAI /

AI
Whisper

API OnlyAudio & SpeechaudiotextUpdated September 21, 2022

Whisper-1: OpenAI's Speech Recognition API Model

Model Overview

whisper-1 is the version of OpenAI's Whisper model available via the OpenAI API for transcription and translation tasks. Released as an API endpoint in March 2023, it provides access to a highly capable multilingual speech recognition system for automatic transcription, speech-to-text, and audio translation (to English). The API wraps Whisper's large model for easy integration without the need to self-host. It accepts audio files up to 25MB in formats including MP3, MP4, MPEG, MPGA, M4A, WAV, and WEBM.


✨ Key Features

Specification Table
FeatureDescription
Multilingual TranscriptionTranscribes audio in 100+ languages automatically
Speech TranslationTranslates speech from any supported language directly to English
Multiple Audio FormatsAccepts MP3, MP4, MPEG, MPGA, M4A, WAV, WEBM
TimestampsReturns word-level and segment-level timestamps
Prompt-GuidedAccepts optional prompt text to improve accuracy for domain-specific terms

🔗 Resources

Specification Table

📜 License & Access

Proprietary API — Available via OpenAI API. The underlying Whisper model weights are open-source (MIT).

You might also want to compare

Verified Sources

Model Specs

api-only

Parameters

Unknown

Context Window

Unknown

License

Proprietary

Deployment

api-only

Resources & Links

Lineage

Model Family

Part of the whisper family

Curator Notes

Bulk imported from OpenAI developer docs.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model