Unlimited Open Source Models

Get Plan
Skip to main content
AudioGen

Sound Effects API — Text to SFX

Last updated · By ModelsLab Engineering

The ModelsLab sound effects API turns a text prompt into a sound effect. It runs Stable Audio Open 1.0 on our own GPUs and costs $0.001 per second of audio ($0.0047 minimum), so a 10-second effect is $0.01. Clips are 3 to 15 seconds in mp3, wav or flac, and take about 13 seconds to generate. POST your prompt to /api/v6/voice/sfx.

  • $0.001 per second of audio
  • Stable Audio Open 1.0
  • No per-effect charge on the $149 plan

What the Sound Effects API Does

Text In, Sound Effect Out

Send a plain-English description — “glass bottle shattering on a concrete floor” — and the API returns an audio file of that sound. The model is Stable Audio Open 1.0 from Stability AI, run on ModelsLab’s own GPUs with 80 diffusion steps at 44.1 kHz stereo. You choose the length (3 to 15 seconds) and the format (mp3, wav or flac).

Billing is per second of audio: $0.001 a second with a $0.0047 minimum per effect, drawn from your plan. On the $149/month Open Source Unlimited plan there is no per-effect charge at all, which makes it a flat price for teams that generate sound libraries in bulk.

ElevenLabs Sound Effects is on the same API key too, through /api/v7/voice/sound-generation with model_id eleven_sound_effect at $0.06 per generation. Request errors logged per endpoint are on the API status page.

  • Stable Audio Open 1.0 over plain HTTP — no SDK to install
  • $0.001 per second, $0.0047 minimum; a 10-second effect is $0.01
  • 3 to 15 second clips in mp3, wav or flac, 44.1 kHz stereo
  • About 13 seconds of GPU time per effect
  • File URL in the response, fetch by id, or webhook delivery

Hear the Model

Generated on this endpoint on October 6, 2026 with duration 5, mp3 at 192k. Unedited: what you hear is what the API returned.

Door creak

Prompt: “Heavy wooden door creaking open slowly in a stone hallway”

Download the mp3

Thunderstorm

Prompt: “Thunderstorm with heavy rain and a close lightning strike”

Download the mp3

Laser shots

Prompt: “Sci-fi laser gun firing three quick shots”

Download the mp3

Glass shatter

Prompt: “Glass bottle shattering on a concrete floor”

Download the mp3

Sound effects API quick start

One POST with a prompt. cURL, Python and JavaScript.

cURL — one sound effect

bash
1curl -X POST 'https://modelslab.com/api/v6/voice/sfx' \
2 -H 'Content-Type: application/json' \
3 -d '{
4 "key": "YOUR_MODELSLAB_API_KEY",
5 "prompt": "Heavy wooden door creaking open slowly in a stone hallway",
6 "duration": 5,
7 "output_format": "mp3"
8 }'

Python — generate, then fetch if still rendering

Python
1import time
2import requests
3
4KEY = "YOUR_MODELSLAB_API_KEY"
5
6result = requests.post(
7 "https://modelslab.com/api/v6/voice/sfx",
8 json={
9 "key": KEY,
10 "prompt": "Glass bottle shattering on a concrete floor",
11 "duration": 5,
12 "output_format": "wav",
13 },
14 timeout=60,
15).json()
16
17# With a clear GPU queue the file comes back in this response. Otherwise the
18# status is "processing": keep the fetch URL and poll it until the status changes.
19fetch_url = result.get("fetch_result")
20for _ in range(60):
21 if result.get("status") != "processing":
22 break
23 time.sleep(10)
24 result = requests.post(fetch_url, json={"key": KEY}, timeout=60).json()
25
26if result.get("status") == "success":
27 print(result["output"][0])
28else:
29 print("Not ready:", result.get("status"), result.get("message"))

JavaScript — with a webhook

JavaScript
1const response = await fetch('https://modelslab.com/api/v6/voice/sfx', {
2 method: 'POST',
3 headers: { 'Content-Type': 'application/json' },
4 body: JSON.stringify({
5 key: 'YOUR_MODELSLAB_API_KEY',
6 prompt: 'Sci-fi laser gun firing three quick shots',
7 duration: 3,
8 // The finished file is POSTed here as well as returned or fetched.
9 webhook: 'https://your-app.example.com/hooks/sfx',
10 track_id: 'level-4-laser',
11 }),
12});
13
14console.log(await response.json());

Specs and Limits

Everything an integration needs to know before the first call.

SpecModelsLab sound effects API
EndpointPOST /api/v6/voice/sfx
ModelStable Audio Open 1.0 (Stability AI), 80 steps
Price$0.001 per second of audio, $0.0047 minimum per effect
10-second effect$0.01 (no per-effect charge on the $149 plan)
Clip length3 to 15 seconds, whole seconds, default 8
Formatsmp3 (default), wav, flac
Bitrate128k, 192k or 320k (default 320k)
Audio44.1 kHz, stereo
Typical speedAbout 13 s per effect (12.9–13.3 s measured)
InputText prompt only
DeliveryFile URL in the response, fetch by id, or webhook

Read from the running endpoint on October 6, 2026. Speed is GPU time per effect from the worker logs; add network time and any queue wait.

Request and Response

The full contract of POST /api/v6/voice/sfx.

FieldTypeMeaning
keystring (required)Your ModelsLab API key
promptstring (required)The sound to generate, in plain English
durationinteger (optional)Seconds of audio, 3 to 15; default 8
output_formatstring (optional)mp3, wav or flac; default mp3
bitratestring (optional)128k, 192k or 320k; default 320k
webhookURL (optional)Where to POST the result when it is ready
track_idstring (optional)Your own reference, echoed back
outputresponseArray with the URL of the audio file
id, fetch_resultresponseReturned with status processing; POST the fetch URL to get the file

The key goes in the JSON body. A response with status success carries the file URL in output. A response with status processing carries id, fetch_result and future_links; POST to fetch_result with your key, or wait for the webhook.

What Sound Effects Cost

One ModelsLab subscription covers sound effects, speech, image and video.

PlanPrice / monthSound effectsParallel requestsAlso covers
Basic$21$0.001 per second from included usage5Image, video, speech and LLM APIs
Standard$47$0.001 per second from included usage10Image, video, speech and LLM APIs
Open Source Unlimited$149No per-effect charge on Stable Audio Open15Every open-weight model ModelsLab self-hosts

Plan prices from the live plans table. On Basic and Standard, each effect draws $0.001 per second of audio, with a $0.0047 minimum, from the plan's included usage.

Sound Effects API Providers Compared

Fourteen text-to-sound-effects offers, sorted by what one 10-second effect costs.

ProviderModelListed price10-second effectMax length
Pika (dev.pika.art)Pika SFX$0.0002 / second + $10/month membership$0.00220 s
ReplicateMMAudio (zsxkib/mmaudio)≈ $0.0049 / run (L40S GPU time)≈ $0.0049Not published
ModelsLabStable Audio Open 1.0$0.001 / second, $0.0047 minimum; none on the $149 plan$0.0115 s
falMMAudio V2$0.001 / second$0.0130 s
falSonilo v1.1$0.0018 / second$0.018180 s
ElevenLabsSound Effects v2 (eleven_text_to_sound_v2)$0.12 / minute of audio$0.0230 s
falElevenLabs Sound Effects V2$0.002 / second$0.0222 s
falStable Audio 3 Small SFX$0.0206 / audio$0.0206120 s
WaveSpeedKling text-to-audio$0.035 / run$0.035 (max 10 s)10 s
SegmindElevenLabs sound generation$0.05 / generation$0.05Not published
ModelsLabElevenLabs Sound Effects (eleven_sound_effect)$0.06 / generation$0.06Not published
WaveSpeedMirelo SFX 1.6$0.01 / second$0.1060 s
Stability AIStable Audio 2.520 credits ($0.20) / generation$0.20190 s
falStable Audio 2.5$0.20 / audio$0.20190 s

Every price read on the provider's own page on October 6, 2026, in the unit the provider publishes. ElevenLabs lists Sound Effects at $0.12 per minute on its API pricing page, while its FAQ says sound effects are billed per generation and its docs say 40 credits per second when a duration is set. Replicate bills GPU time, so its cost per run varies. Pika's $10/month membership is charged on top of usage.

Which sound effects API should you use?

Per second, Pika SFX is the cheapest offer we found at $0.0002, but it adds a $10/month membership. MMAudio V2 on fal and this endpoint both cost $0.001 a second; ElevenLabs costs $0.12 a minute, which is $0.002 a second. At volume the arithmetic changes: above about 14,900 ten-second effects a month, the $149 Open Source Unlimited plan costs less than any $0.001-per-second API.

Price is not the whole answer. ElevenLabs Sound Effects v2 allows 30-second clips and loops. MMAudio can generate audio from a video as well as from text. Stable Audio 2.5 and Sonilo run past two minutes. This endpoint is text only and stops at 15 seconds — use it for libraries of short effects, game one-shots and app sounds, and pick one of the others when length, video sync or the last bit of realism matters.

Sound effects price sources

Every competitor price and limit on this page, with the date we read it.

Generate a Sound Effect in 3 Steps

From API key to an audio file.

STEP 01
STEP 01

Step 1: Get Your API Key

Create a ModelsLab account, subscribe to a plan from $21/month, and copy your key from the dashboard.

STEP 02
STEP 02

Step 2: POST a Prompt

Send key and prompt to /api/v6/voice/sfx. Add duration (3 to 15 seconds) and output_format (mp3, wav or flac) when the defaults do not fit.

STEP 03
STEP 03

Step 3: Download the Audio

With a clear GPU queue, the response carries the file URL in output. Otherwise POST to fetch_result with your key until the status changes, or pass a webhook and receive the file when it is ready.

Why Generate Sound Effects Here

Key advantages that set us apart

Stable Audio Open 1.0 on our own GPUs
$0.001 per second of audio, $0.0047 minimum
No per-effect charge on the $149 plan
3 to 15 second clips in mp3, wav or flac
ElevenLabs Sound Effects on the same key
One key for sound, speech, image, video and LLMs
Plans from $21/month

Our Popular Use Cases

What teams build with the sound effects API:

Generate footsteps, impacts, pickups and UI sounds in bulk, then keep the best variation of each.

Game Audio Libraries

Your Data is Secure: GDPR Compliant AI Services

ModelsLab GDPR Compliance Certification Badge

GDPR Compliant

Pricing That's Perfect

Choose plan as per your needs, cancel anytime.

100% refund policy on monthly & yearly plans — cancel anytime

Contact Sales
Best Value

Open Source Unlimited

Mission-Critical

$149/month

100% refund policy · cancel anytime

Unlimited Open Source Models
100% refund policy
24x7 Support
15 parallel generations
Access to all APIs
Unlimited generations on all open-source models
For mission critical workloads
Add Team Members
Priority GPU Clusters
Most Popular

Standard

Production

$47/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Priority Developer Support
10 concurrent API requests
For Production workloads
API access to all models
Prototype

Basic

Prototype

$21/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Developer Support via Discord/Email
5 concurrent API requests
API access to all models
Shared GPU
Testimonials

Trusted by Enterprise Teams Worldwide

Enterprise Success Stories

“

ModelsLab's Voice Cloning API has revolutionized how we approach character development in our games. It's like having a studio full of voice actors at our fingertips!

Alex Rivera
AR

Alex Rivera

Game Developer at TVC

“

The ease of creating lifelike voiceovers for our e-learning courses has dramatically increased engagement. A real breakthrough for educational content!

Priya Singh
PS

Priya Singh

Instructional Designer at TVC1

“

The LLM Chat API has dramatically helped me in how I approach chat integration. It's like giving an AI voice to my application, making it truly engaging. Thanks, ModelsLab!

John H.
JH

John H.

Developer Enthusiast at Mr

“

Voice Cloning from ModelsLab gave our marketing campaigns a unique edge with custom, realistic voiceovers. It's incredibly easy to use and effective.

Michael Chen
MC

Michael Chen

Digital Marketing Manager at TVC2

Get Expert Support in Seconds

We're Here to Help.

Want to know more? You can email us anytime at support@modelslab.com

View Docs

POST https://modelslab.com/api/v6/voice/sfx with your `key` and a `prompt`. Optional fields: `duration` (a whole number of seconds from 3 to 15, default 8), `output_format` (mp3, wav or flac, default mp3), `bitrate` (128k, 192k or 320k, default 320k), `webhook` and `track_id`. The response carries the audio URL in `output`, or an `id` to fetch from /api/v6/voice/fetch/{id} while the clip is still rendering.

Stable Audio Open 1.0 from Stability AI, running on ModelsLab’s own GPUs with 80 diffusion steps and 44.1 kHz stereo output. ElevenLabs Sound Effects is also available on the same API key through POST /api/v7/voice/sound-generation with model_id eleven_sound_effect.

$0.001 per second of audio, with a minimum of $0.0047 per effect, drawn from the usage included in your plan: a 10-second effect costs $0.01. Plans start at $21/month. On the $149/month Open Source Unlimited plan, Stable Audio Open effects carry no per-effect charge. ElevenLabs Sound Effects (eleven_sound_effect) costs $0.06 per generation.

About 13 seconds of GPU time per effect, whatever its length: the worker logged 12.9 to 13.3 seconds across 20 effects on October 6, 2026. When the GPU queue is clear, the API call returns the finished file. When it is not, the response has status processing, an id and the future file URL: POST to the fetch URL with your key until status is success, failed or error, or pass a webhook.

3 to 15 seconds per request, in whole seconds. For longer ambience, generate overlapping clips and crossfade them, or use a model built for long audio such as Stable Audio 2.5, which runs up to 190 seconds on Stability AI’s API.

ModelsLab charges no royalties or licence fees on the audio you generate. Stable Audio Open 1.0 is released by Stability AI under the Stability AI Community License; read it for the terms that apply to your own use.

modelslab.com/status lists every API endpoint, including /api/v6/voice/sfx, with the request errors logged over the last seven days. It does not yet count effects that time out in the GPU queue, so check your own responses for status failed or error as well.

No. The endpoint turns text into audio. For foley synced to footage, use a video-to-audio model such as MMAudio, Mirelo SFX or Hunyuan Video Foley.
Plugins

Explore Plugins for Pro

Our plugins are designed to work with the most popular content creation software.

API

Build Apps with
ML
API

Use our API to build apps, generate AI art, create videos, and produce audio with ease.