Speech to Text API
Send an audio file URL, get a transcript with word-level timestamps. Whisper large-v3 on our own GPUs, $0.0047 per request.
- $0.0047 per request
- Whisper large-v3 on our own GPUs
- Unlimited on the $149 plan
Turn audio into text with one request
Send the URL of an audio file to POST /api/v6/voice/speech_to_text. The response is the transcript, with word-level timestamps when you ask for them.
Each request costs $0.0047 from your plan allowance. On the $149 Open Source Unlimited plan there is no per-request charge. One subscription and one API key cover image, video, speech and LLM APIs.
Transcription API with word timestamps
Set timestamp_level to word or sentence. You get start and end times with the text, which you can turn into SRT or VTT captions.
For long recordings, pass a webhook URL. We send the finished transcript to it.
| Field | Type | What it does |
|---|---|---|
| init_audio | URL (required) | The mp3, wav, flac or opus file to transcribe |
| language | string (optional) | The language of the audio |
| timestamp_level | word or sentence | Adds timings to the transcript |
| webhook | URL (optional) | Receives the finished transcript |
| track_id | number (optional) | Your own reference id, sent back |
Built on Whisper large-v3
Transcription runs on Whisper large-v3, the open-weight speech recognition model. We run it on our own GPUs, so you do not manage a GPU, a queue or a deployment.
For provider prices and limits side by side, see our Whisper API page.
Speech to text pricing
A transcription request costs $0.0047. On Basic ($21 a month) and Standard ($47 a month) it comes out of the dollar allowance your plan includes. On Open Source Unlimited ($149 a month) there is no per-request charge, because the model runs on our own GPUs.
| Plan | Price per month | Speech to text | Parallel generations |
|---|---|---|---|
| Basic | $21 | $0.0047 per request | 5 |
| Standard | $47 | $0.0047 per request | 10 |
| Open Source Unlimited | $149 | No per-request charge | 15 |
Creating an account is free. API calls need a plan.
Speech to text API quick start
One POST with an audio URL. Python, JavaScript and cURL.
Python: transcribe with word timestamps
Python1import requests23response = requests.post(4 "https://modelslab.com/api/v6/voice/speech_to_text",5 json={6 "key": "YOUR_MODELSLAB_API_KEY",7 "init_audio": "https://example.com/interview.mp3",8 "timestamp_level": "word",9 },10)1112print(response.json())
JavaScript: long file with a webhook
JavaScript1const response = await fetch(2 'https://modelslab.com/api/v6/voice/speech_to_text',3 {4 method: 'POST',5 headers: { 'Content-Type': 'application/json' },6 body: JSON.stringify({7 key: 'YOUR_MODELSLAB_API_KEY',8 init_audio: 'https://example.com/lecture.mp3',9 timestamp_level: 'word',10 // Long recordings queue; the webhook receives the transcript.11 webhook: 'https://your-app.example.com/hooks/transcript',12 }),13 },14);1516console.log(await response.json());
cURL: quick test from the terminal
bash1curl -X POST 'https://modelslab.com/api/v6/voice/speech_to_text' \2 -H 'Content-Type: application/json' \3 -d '{4 "key": "YOUR_MODELSLAB_API_KEY",5 "init_audio": "https://example.com/voicemail.wav",6 "timestamp_level": "word"7 }'
Transcribe, translate and re-voice
Send the transcript to an LLM on the same key to translate or summarize it. LLMs are partner models billed per million tokens: from the plan allowance first on Basic and Standard, then the wallet, and from the wallet on the $149 plan. Then read the result back with our Text to Speech API in the voice you choose.
Speech to text by language
Language pages with transcription examples for each language.
Pricing That's Perfect
Choose plan as per your needs, cancel anytime.
100% refund policy on monthly & yearly plans — cancel anytime
Open Source Unlimited
Mission-Critical
100% refund policy · cancel anytime
Standard
Production
100% refund policy · cancel anytime
Basic
Prototype
100% refund policy · cancel anytime
Start transcribing audio
Plans start at $21 a month. Each transcript is $0.0047, or unlimited on the $149 plan.
Get API keyRelated Speech APIs
- Whisper API
Whisper large-v3 on our own GPUs.
- Text to Speech API
48 languages, billed per second of audio.
- Voice Cloning API
Clone a voice from a short sample.
- Cheapest text to speech API
TTS cost per minute of audio, compared.
- Cheapest AI API
Image, video and speech prices side by side.
- Unlimited AI API
Every self-hosted open-source model for $149/mo.
Get Expert Support in Seconds
We're Here to Help.
Want to know more? You can email us anytime at support@modelslab.com