Unlimited Open Source Models

Get Plan
Skip to main content
AudioGen

Speech to Text API

Send an audio file URL, get a transcript with word-level timestamps. Whisper large-v3 on our own GPUs, $0.0047 per request.

  • $0.0047 per request
  • Whisper large-v3 on our own GPUs
  • Unlimited on the $149 plan

Turn audio into text with one request

Send the URL of an audio file to POST /api/v6/voice/speech_to_text. The response is the transcript, with word-level timestamps when you ask for them.

Each request costs $0.0047 from your plan allowance. On the $149 Open Source Unlimited plan there is no per-request charge. One subscription and one API key cover image, video, speech and LLM APIs.

Transcription API with word timestamps

Set timestamp_level to word or sentence. You get start and end times with the text, which you can turn into SRT or VTT captions.

For long recordings, pass a webhook URL. We send the finished transcript to it.

FieldTypeWhat it does
init_audioURL (required)The mp3, wav, flac or opus file to transcribe
languagestring (optional)The language of the audio
timestamp_levelword or sentenceAdds timings to the transcript
webhookURL (optional)Receives the finished transcript
track_idnumber (optional)Your own reference id, sent back

Built on Whisper large-v3

Transcription runs on Whisper large-v3, the open-weight speech recognition model. We run it on our own GPUs, so you do not manage a GPU, a queue or a deployment.

For provider prices and limits side by side, see our Whisper API page.

Speech to text pricing

A transcription request costs $0.0047. On Basic ($21 a month) and Standard ($47 a month) it comes out of the dollar allowance your plan includes. On Open Source Unlimited ($149 a month) there is no per-request charge, because the model runs on our own GPUs.

PlanPrice per monthSpeech to textParallel generations
Basic$21$0.0047 per request5
Standard$47$0.0047 per request10
Open Source Unlimited$149No per-request charge15

Creating an account is free. API calls need a plan.

Speech to text API quick start

One POST with an audio URL. Python, JavaScript and cURL.

Python: transcribe with word timestamps

Python
1import requests
2
3response = requests.post(
4 "https://modelslab.com/api/v6/voice/speech_to_text",
5 json={
6 "key": "YOUR_MODELSLAB_API_KEY",
7 "init_audio": "https://example.com/interview.mp3",
8 "timestamp_level": "word",
9 },
10)
11
12print(response.json())

JavaScript: long file with a webhook

JavaScript
1const response = await fetch(
2 'https://modelslab.com/api/v6/voice/speech_to_text',
3 {
4 method: 'POST',
5 headers: { 'Content-Type': 'application/json' },
6 body: JSON.stringify({
7 key: 'YOUR_MODELSLAB_API_KEY',
8 init_audio: 'https://example.com/lecture.mp3',
9 timestamp_level: 'word',
10 // Long recordings queue; the webhook receives the transcript.
11 webhook: 'https://your-app.example.com/hooks/transcript',
12 }),
13 },
14);
15
16console.log(await response.json());

cURL: quick test from the terminal

bash
1curl -X POST 'https://modelslab.com/api/v6/voice/speech_to_text' \
2 -H 'Content-Type: application/json' \
3 -d '{
4 "key": "YOUR_MODELSLAB_API_KEY",
5 "init_audio": "https://example.com/voicemail.wav",
6 "timestamp_level": "word"
7 }'

Transcribe, translate and re-voice

Send the transcript to an LLM on the same key to translate or summarize it. LLMs are partner models billed per million tokens: from the plan allowance first on Basic and Standard, then the wallet, and from the wallet on the $149 plan. Then read the result back with our Text to Speech API in the voice you choose.

Speech to text by language

Language pages with transcription examples for each language.

Pricing That's Perfect

Choose plan as per your needs, cancel anytime.

100% refund policy on monthly & yearly plans — cancel anytime

Contact Sales
Best Value

Open Source Unlimited

Mission-Critical

$149/month

100% refund policy · cancel anytime

Unlimited Open Source Models
100% refund policy
24x7 Support
15 parallel generations
Access to all APIs
Unlimited generations on all open-source models
For mission critical workloads
Add Team Members
Priority GPU Clusters
Most Popular

Standard

Production

$47/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Priority Developer Support
10 concurrent API requests
For Production workloads
API access to all models
Prototype

Basic

Prototype

$21/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Developer Support via Discord/Email
5 concurrent API requests
API access to all models
Shared GPU

Start transcribing audio

Plans start at $21 a month. Each transcript is $0.0047, or unlimited on the $149 plan.

Get API key

Get Expert Support in Seconds

We're Here to Help.

Want to know more? You can email us anytime at support@modelslab.com

View Docs

POST https://modelslab.com/api/v6/voice/speech_to_text with `init_audio` set to the URL of the audio file. The response carries the transcript, or an id to fetch when the file is long enough to queue. Add `language` when you know it and `timestamp_level` when you need word timings.

Whisper large-v3, running on ModelsLab's own GPUs. The same handler is registered at /api/v6/voice/speech_to_text and /api/v6/whisper/transcribe.

Transcription costs $0.0047 per request whatever the length of the audio, paid from the usage included in your plan (from $21/month). Open Source Unlimited ($149/month) has no limit on the open-weight speech models ModelsLab self-hosts. Billing is per request, not per minute of audio, so one long recording costs the same as one short one.

Yes — it is a plain HTTP POST, so `requests.post()` with the JSON body is the whole integration. There is no SDK to install and no streaming session to manage; you send a URL and read the transcript from the response.

Yes. `timestamp_level` takes `word` or `sentence`, which is what you need to build SRT or VTT output.

ModelsLab is a paid service. Plans start at $21/month. Creating an account is free, but transcription requires an active plan.

Yes. The transcript can go straight into the LLM endpoint for translation and then into /api/v6/voice/text_to_audio to speak it — including in a cloned voice — and into /api/v7/video-fusion/lip-sync if the source was video. All of it runs on one API key.