Unlimited Open Source Models

Get Plan
Skip to main content
Imagen

Whisper API — Pricing, Limits and Providers

Last updated · By ModelsLab Engineering

The ModelsLab Whisper API runs Whisper large-v3 on our own GPUs. POST an audio URL to /api/v6/whisper/transcribe and get the transcript with word-level timestamps. Each request costs $0.0047 from your plan, whatever the length of the audio. For comparison, OpenAI whisper-1 costs $0.006 per minute and Groq whisper-large-v3-turbo $0.04 per hour (checked September 23, 2026).

  • $0.0047 per request
  • Whisper large-v3 on our own GPUs
  • Unlimited on the $149 plan

What the Whisper API Does

Whisper large-v3, Hosted

Whisper is OpenAI’s open-weight speech recognition model. We run Whisper large-v3 on our own GPUs. Running it yourself means a GPU, a queue, a warm deployment and somebody to own all three. This endpoint is the same model with none of that: you POST the URL of an audio file and read the transcript out of the response.

For transcription as a general API, with word timestamps, pricing and a quick start, see our Speech to Text API page.

Because the weights are open, this is not a lock-in decision. The same model you call here is the one you can run on your own hardware later, which is the main reason teams pick Whisper over a proprietary transcription API in the first place.

Billed per Request, Not per Minute

Most transcription APIs meter by the minute or the hour of audio, so a long recording costs proportionally more. ModelsLab bills transcription per request, $0.0047 from your plan, so the length of the file does not change the price of the request. That changes the arithmetic for podcasts, lecture capture, call archives and anything else measured in hours rather than clips. For many short clips, a per-minute vendor such as Groq is cheaper.

  • Whisper large-v3 over plain HTTP — no SDK to install
  • Word-level or sentence-level timestamps for subtitles
  • Multilingual, with an optional language hint
  • Webhook delivery for long files
  • $0.0047 per request; no per-request charge on the $149 plan

Whisper API quick start

One POST with an audio URL. Python, JavaScript and cURL.

Python — transcribe a file

Python
1import requests
2
3response = requests.post(
4 "https://modelslab.com/api/v6/whisper/transcribe",
5 json={
6 "key": "YOUR_MODELSLAB_API_KEY",
7 "init_audio": "https://example.com/interview.mp3",
8 "language": "english",
9 "timestamp_level": "word",
10 },
11)
12
13print(response.json())

JavaScript — transcribe with a webhook

JavaScript
1const response = await fetch(
2 'https://modelslab.com/api/v6/whisper/transcribe',
3 {
4 method: 'POST',
5 headers: { 'Content-Type': 'application/json' },
6 body: JSON.stringify({
7 key: 'YOUR_MODELSLAB_API_KEY',
8 init_audio: 'https://example.com/episode-142.mp3',
9 // Long files queue; the webhook receives the finished transcript.
10 webhook: 'https://your-app.example.com/hooks/transcript',
11 track_id: 142,
12 }),
13 },
14);
15
16console.log(await response.json());

cURL — quick test from the terminal

bash
1curl -X POST 'https://modelslab.com/api/v6/whisper/transcribe' \
2 -H 'Content-Type: application/json' \
3 -d '{
4 "key": "YOUR_MODELSLAB_API_KEY",
5 "init_audio": "https://example.com/voicemail.wav",
6 "language": "english"
7 }'

Request and Response

The full contract of the transcription endpoint.

FieldTypeMeaning
init_audioURL (required)Audio file to transcribe
languagestring (optional)Source language; defaults to English
timestamp_levelstring (optional)word or sentence timings
webhookURL (optional)Where to POST the finished transcript
track_idstring (optional)Your own numeric reference id, echoed back
outputresponseThe transcript, with timings when asked

The same handler is registered at /api/v6/voice/speech_to_text; /api/v6/whisper/transcribe is the alias. A v7 route at /api/v7/voice/speech-to-text takes an explicit model_id.

What Transcription Costs

One ModelsLab subscription covers transcription, speech, image and video. A transcript is one request, not a per-minute charge.

PlanPrice / MonthTranscriptionWhisper included
Basic$21$0.0047 per requestYes
Standard$47$0.0047 per requestYes
Open Source Unlimited$149No per-request chargeYes

Plan prices from the live plans table, 2026-09-23. A transcription request draws $0.0047 from the plan's dollar allowance. Open Source Unlimited covers the open-weight speech models ModelsLab self-hosts, which includes Whisper large-v3.

Whisper API Providers Compared

Hosted Whisper and its closest rivals, priced per hour of audio. Every price read on the provider's own page on September 23, 2026.

ProviderModelListed pricePer hour of audio
Groqwhisper-large-v3-turbo$0.04 / hour (10 s minimum)$0.04
Together AIWhisper Large v3$0.0015 / minute$0.09
Groqwhisper-large-v3$0.111 / hour$0.111
AssemblyAIUniversal-2$0.15 / hour$0.15
OpenAIgpt-4o-mini-transcribe~$0.003 / minute~$0.18
ElevenLabsScribe v2$0.22 / hour$0.22
DeepgramNova-3 pre-recorded (pay as you go)$0.0043 / minute$0.258
OpenAIwhisper-1$0.006 / minute$0.36
ModelsLabWhisper large-v3$0.0047 / request$0.0047 per request, any length

Per-hour figures are the listed per-minute price times 60. ModelsLab bills per request rather than by audio length, so its per-hour cost falls as files get longer. The row shows the $0.0047 request price on Basic and Standard.

Which Whisper API is cheapest?

Per hour of audio, Groq is the cheapest hosted Whisper we found: $0.04 an hour for whisper-large-v3-turbo, with every request billed for at least 10 seconds. Together AI is next at $0.0015 a minute. OpenAI's whisper-1 is $0.006 a minute, nine times Groq's turbo rate.

ModelsLab is the cheaper choice when files are long: a request draws $0.0047 from your plan, so a one-hour file costs $0.0047 and a two-hour file costs the same. On the $149 Open Source Unlimited plan there is no per-request charge at all. For thousands of 30-second voice notes, a per-minute vendor wins. For live captions or speaker labels, Deepgram and AssemblyAI do what batch Whisper does not.

Whisper API Limits

File size, input type and minimum billing, where the provider publishes them.

ProviderInputSize limitMinimum billed
ModelsLabAudio URL (init_audio)No published hard limit; webhook for long filesOne request
OpenAIFile upload: mp3, mp4, mpeg, mpga, m4a, wav, webm25 MBPer minute
GroqFile upload or URL25 MB (base tier), 100 MB upload (dev tier), 25 MB by URL10 seconds

OpenAI and Groq limits from their speech-to-text docs, read 2026-09-23. ModelsLab does not publish a hard size or length limit for the shared endpoint; test your longest files before you depend on them, and use the webhook for long recordings.

Whisper price sources

Every competitor price and limit on this page, with the date we read it.

Transcribe in 3 Steps

From API key to a transcript.

STEP 01
STEP 01

Step 1: Get Your API Key

Create a ModelsLab account, subscribe to a plan from $21/month, and copy your key from the dashboard.

STEP 02
STEP 02

Step 2: POST the Audio URL

Send init_audio to /api/v6/whisper/transcribe. Pass language when you know it, and timestamp_level when you need word timings.

STEP 03
STEP 03

Step 3: Read the Transcript

Short files return the transcript directly. Long files return an id to fetch, or pass a webhook URL and receive the result when it is ready.

What Whisper Is Not Good At

Whisper is a batch model. It transcribes files well and it does not do live streaming, and it does not label who is speaking. If your product needs real-time captions on a live call, or speaker diarisation out of the box, a specialist streaming vendor will serve you better than this endpoint will — and it is cheaper to know that before you integrate than after.

Why Run Whisper Here

Key advantages that set us apart

Whisper large-v3 on our own GPUs
Billed per request, not per minute of audio
Word-level timestamps for subtitles
Multilingual, with an optional language hint
Webhook delivery for long recordings
Plain HTTP — no SDK, no streaming session
Plans from $21/month; unlimited on the $149 plan

Our Popular Use Cases

What teams build on hosted Whisper:

Transcribe a back catalogue at one API call per episode, then publish the text for search.

Podcast and Video Transcripts

Your Data is Secure: GDPR Compliant AI Services

ModelsLab GDPR Compliance Certification Badge

GDPR Compliant

Pricing That's Perfect

Choose plan as per your needs, cancel anytime.

100% refund policy on monthly & yearly plans — cancel anytime

Contact Sales
Best Value

Open Source Unlimited

Mission-Critical

$149/month

100% refund policy · cancel anytime

Unlimited Open Source Models
100% refund policy
24x7 Support
15 parallel generations
Access to all APIs
Unlimited generations on all open-source models
For mission critical workloads
Add Team Members
Priority GPU Clusters
Most Popular

Standard

Production

$47/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Priority Developer Support
10 concurrent API requests
For Production workloads
API access to all models
Prototype

Basic

Prototype

$21/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Developer Support via Discord/Email
5 concurrent API requests
API access to all models
Shared GPU

Get Expert Support in Seconds

We're Here to Help.

Want to know more? You can email us anytime at support@modelslab.com

View Docs

POST https://modelslab.com/api/v6/whisper/transcribe with `init_audio` set to the URL of the audio file. `language` is optional and defaults to English, `timestamp_level` takes `word` or `sentence`, and `webhook` receives the result for long files. The same handler is also registered at /api/v6/voice/speech_to_text.

Whisper large-v3, running on ModelsLab's own GPUs.

A transcription request draws $0.0047 from your plan whatever the length of the audio: Basic is $21/month, Standard $47/month, and the $149/month Open Source Unlimited plan has no per-request charge on self-hosted models including Whisper. For comparison, OpenAI whisper-1 costs $0.006 per minute and Groq whisper-large-v3-turbo $0.04 per hour (checked September 23, 2026).

It depends on volume and on whether you already own GPUs. Self-hosting Whisper large-v3 means a GPU, a queue and a deployment to keep warm; on a plan the same work is one HTTP call. Above about 31,700 transcripts a month ($149 ÷ $0.0047), the $149 Open Source Unlimited plan removes the per-request charge, which is usually cheaper than keeping your own GPU warm.

Yes. Set `timestamp_level` to `word` to get per-word timings rather than the sentence-level default, which is what subtitle and caption pipelines need.

Whisper is a multilingual model and the endpoint accepts a `language` hint rather than a fixed list, so pass it when you know the language and leave it out when the input varies. If a specific language matters to your product, test it on your own audio before committing — we would rather you measure it than take a number off a marketing page.

Yes. ModelsLab is a paid service; plans start at $21/month. Creating an account costs nothing, but transcription needs an active plan.

Per hour of audio, Groq: whisper-large-v3-turbo is $0.04 per hour with a 10-second minimum per request. Together AI charges $0.0015 per minute and OpenAI whisper-1 $0.006 per minute (checked September 23, 2026). ModelsLab bills $0.0047 per request from your plan, so it is cheaper for long files: a two-hour recording costs the same as a two-minute one.

OpenAI's transcription API accepts files up to 25 MB, and Groq accepts 25 MB (100 MB by upload on its dev tier), both read September 23, 2026. The ModelsLab endpoint fetches init_audio from a URL instead of taking an upload and does not publish a hard size limit. Use the webhook parameter for long recordings and test your longest files first.

Yes. OpenAI serves whisper-1 at $0.006 per minute, plus gpt-4o-mini-transcribe at about $0.003 per minute (checked September 23, 2026). Whisper's weights are open, so hosts such as ModelsLab, Groq and Together AI serve the same model family.

Send requests.post('https://modelslab.com/api/v6/whisper/transcribe', json={...}) with key and init_audio set to the audio URL, plus language and timestamp_level when you need them. The JSON response carries the transcript in output; long files return an id to fetch, or deliver to your webhook.
Plugins

Explore Plugins for Pro

Our plugins are designed to work with the most popular content creation software.

API

Build Apps with
ML
API

Use our API to build apps, generate AI art, create videos, and produce audio with ease.