Whisper large-v3, Hosted
Whisper is OpenAI’s open-weight speech recognition model. We run Whisper large-v3 on our own GPUs. Running it yourself means a GPU, a queue, a warm deployment and somebody to own all three. This endpoint is the same model with none of that: you POST the URL of an audio file and read the transcript out of the response.
For transcription as a general API, with word timestamps, pricing and a quick start, see our Speech to Text API page.
Because the weights are open, this is not a lock-in decision. The same model you call here is the one you can run on your own hardware later, which is the main reason teams pick Whisper over a proprietary transcription API in the first place.