Unlimited Open Source Models

Get Plan
Skip to main content
Imagen

Best Text to Speech API in 2026

Text to speech plus voice cloning from a 4-second sample in 48 languages, billed at $0.001 per second of generated audio ($0.06 a minute) on plans from $21/month, and with no per-second charge on the $149/month Open Source Unlimited plan.

Last updated · By ModelsLab Engineering

Best text-to-speech API for developers in 2026: the short answer

For the most natural voices and the largest voice library, ElevenLabs: $0.05 per 1,000 characters on Flash and Turbo and $0.10 on Multilingual v2 and v3, about $0.045 to $0.09 a minute of speech. For the lowest price per minute, the cloud APIs: Google Cloud and Amazon Polly standard voices are $4 per million characters (about $0.004 a minute), and OpenAI tts-1, Azure neural voices and Fish Audio are about $15 per million (about $0.014 a minute).

ModelsLab bills self-hosted speech by the audio it returns: $0.001 per second, or $0.06 a minute, with voice cloning from a 4-second sample on every plan. That is more per minute than most per-character APIs. It costs less when you move to the $149/month Open Source Unlimited plan, which has no per-second charge: above about 28 hours of speech a month compared with ElevenLabs Multilingual, and above about 184 hours compared with OpenAI tts-1 or Fish Audio.

PlayHT is no longer an option: play.ht did not resolve on September 27, 2026. For voice cloning through an API without a sales call, the choices are ElevenLabs (from the $6/month Starter plan), Cartesia (from $5/month), Fish Audio, Hume and ModelsLab.

Text-to-Speech API Prices per Minute, September 2026

List prices from each provider's own pricing page. Per-minute figures assume about 900 characters per minute of speech.

ProviderModelList priceAbout per minuteVoice cloning via API
ModelsLabSelf-hosted TTS and voice cloning$0.001 per second of audio$0.06Yes, from a 4 s sample
ModelsLabSame, Open Source Unlimited$149/monthNo per-second chargeYes
ElevenLabsFlash / Turbo$0.05 per 1K characters$0.045Instant cloning from Starter ($6/month)
ElevenLabsMultilingual v2 / v3$0.10 per 1K characters$0.09Instant cloning from Starter ($6/month)
OpenAItts-1 / tts-1-hd$15 / $30 per 1M characters$0.014 / $0.027Eligible customers, through sales
Google CloudStandard, WaveNet / Neural2 / Chirp 3 HD$4 / $16 / $30 per 1M characters$0.004 / $0.014 / $0.027Instant Custom Voice, allow-listed ($60 per 1M)
Amazon PollyStandard / Neural / Generative$4 / $16 / $30 per 1M characters$0.004 / $0.014 / $0.027No
Azure AI SpeechNeural / Neural HD$15 / $22 per 1M characters$0.014 / $0.020Personal Voice, limited access
CartesiaSonic$38 to $65 per 1M credits, by plan$0.034 to $0.059From Pro ($5/month)
Fish Audios2-pro$15 per 1M UTF-8 bytes$0.014Yes
DeepgramAura-2$0.030 per 1K characters$0.027Not listed
HumeOctave$0.05 to $0.15 per 1K characters over plan$0.045 to $0.135Yes, every plan
PlayHT—No longer available——

Prices checked September 27, 2026 on elevenlabs.io/pricing/api, developers.openai.com/api/docs/pricing, cloud.google.com/text-to-speech/pricing, aws.amazon.com/polly/pricing, the Azure retail prices API (East US), cartesia.ai/pricing, docs.fish.audio (pricing and rate limits), deepgram.com/pricing and hume.ai/pricing. Per-minute figures are our arithmetic at 900 characters a minute; vendors quote 720 to 1,390. ModelsLab bills $0.001 per second of audio with a $0.0047 minimum per request, from the usage included in a plan from $21/month.

Where ModelsLab Fits Among Text to Speech APIs in 2026

The 2026 Text to Speech API Landscape

The text to speech API market in 2026 splits into three camps. ElevenLabs leads on studio-grade voice quality with subscription-plus-usage pricing. Cloud vendors — OpenAI TTS, Google Cloud TTS, Amazon Polly, Azure Speech — offer reliable preset voices billed per million characters, but voice cloning requires enterprise custom-voice programs. And unified AI platforms like ModelsLab bundle TTS, voice cloning, and music generation with the rest of a multimodal API stack.

ModelsLab stands out in 2026 on breadth: voice cloning from a 4-second sample, multilingual output from the same voice profile in 48 languages, and self-hosted speech at $0.001 per second of generated audio on plans from $21/month, with no per-second charge on the $149/month Open Source Unlimited plan, plus image, video and LLM APIs on the same key.

This guide evaluates the top TTS APIs across the criteria that matter to production teams: voice cloning, pricing, language coverage, latency, and integration experience.

What to Look for in a Text to Speech API

When evaluating text to speech APIs in 2026, prioritize these factors:

  • Voice cloning — Can you clone a custom voice from a short sample, or are you limited to preset voices?
  • Pricing basis — Per character, per second, or subscription tiers. Model the cost at your real monthly volume.
  • Language coverage — Multilingual generation, ideally from a single cloned voice profile.
  • Latency — Time to first audio matters for conversational and real-time use cases.
  • Audio stack breadth — Music generation, sound effects, and audio processing alongside TTS.
  • Integration experience — REST simplicity, webhooks, and predictable error handling.
  • Compliance — Consent-based cloning policies and GDPR-compliant data handling.

Trusted by

Google logo
Salesforce logo
Amazon logo
IBM logo
Adobe logo
Sony logo
Google logo
Salesforce logo
Amazon logo
IBM logo
Adobe logo
Sony logo
Google logo
Salesforce logo
Amazon logo
IBM logo
Adobe logo
Sony logo
Google logo
Salesforce logo
Amazon logo
IBM logo
Adobe logo
Sony logo
500M+

API Requests Processed

400K+

Registered Users

5K+

Discord Community Members

300+

Available AI APIs

500

GPUs in Our Own Datacenter

Best Text to Speech APIs Compared (2026)

Side-by-side comparison of the leading TTS and voice cloning API providers.

FeatureModelsLabElevenLabsOpenAI TTSGoogle Cloud TTSAmazon Polly
Voice Cloning from Short Sample10s samplePaid tiersNoCustom programCustom program
Pricing Basis$0.001/s of audio, plans from $21/moSubscription + usagePer 1M charactersPer 1M charactersPer 1M characters
Multilingual from One VoiceYesYesPreset voicesPreset voicesPreset voices
Music GenerationSame platformSFX onlyNoNoNo
Image + Video + LLM APIs (same key)YesNoSeparate pricingSeparate productsSeparate products
Free TierPaid, from $21/moLimitedPaid onlyTrial credits12 months
Speed and Emotion ControlPer requestYesSpeed onlySSMLSSML
Enterprise SLA99.9%EnterpriseYesYesYes

Pricing basis and voice cloning rows checked September 27, 2026 on each vendor’s pricing page; see the price table above.

Quick Start: Clone a Voice and Generate Speech

Get started with the best TTS API using simple REST calls.

Clone a voice and speak in one call (Python)

Python
1import requests
2
3# One call: clone from a short sample and generate speech
4url = "https://modelslab.com/api/v6/voice/text_to_audio"
5payload = {
6 "key": "YOUR_API_KEY",
7 "prompt": "Welcome to our platform. We are glad to have you here.",
8 "init_audio": "https://your-storage.com/voice-sample.wav",
9 "language": "english",
10 "speed": 1.0
11}
12
13response = requests.post(url, json=payload)
14audio_url = response.json()["output"][0]
15print(f"Generated audio: {audio_url}")

Generate speech with a library voice (Python)

Python
1# Generate speech with a voice from the voice library
2url = "https://modelslab.com/api/v6/voice/text_to_speech"
3payload = {
4 "key": "YOUR_API_KEY",
5 "prompt": "Welcome to our platform. We are glad to have you here.",
6 "voice_id": "scott",
7 "language": "english",
8 "output_format": "mp3"
9}
10
11response = requests.post(url, json=payload)
12audio_url = response.json()["output"][0]
13print(f"Generated audio: {audio_url}")

Generate speech with JavaScript (Node.js)

JavaScript
1const response = await fetch('https://modelslab.com/api/v6/voice/text_to_audio', {
2 method: 'POST',
3 headers: { 'Content-Type': 'application/json' },
4 body: JSON.stringify({
5 key: 'YOUR_API_KEY',
6 prompt: 'This is generated speech from the ModelsLab voice API.',
7 init_audio: 'https://your-storage.com/voice-sample.wav',
8 language: 'english'
9 })
10});
11
12const data = await response.json();
13console.log(`Audio: ${data.output[0]}`);

How to Get Started with the Best TTS API

From signup to your first generated audio in minutes.

STEP 01
STEP 01

Step 1: Create Your Account

Create a ModelsLab account, subscribe to a plan (from $21/month), and generate your API key from the dashboard. Every plan includes text to speech and voice cloning.

STEP 02
STEP 02

Step 2: Clone a Voice or Pick a Preset

Pass an audio sample of 4 to 10 seconds as init_audio to clone a voice (longer samples are trimmed to 10 seconds), or generate immediately with a voice from the voice library. Set language, speed, and emotion per request.

STEP 03
STEP 03

Step 3: Integrate and Scale

Wire the REST endpoints into your product, handle the returned audio URLs, and scale from Basic ($21/month) to Standard ($47/month, about 13 hours) or Open Source Unlimited ($149/month, no per-second charge).

2026 Text to Speech API Pricing Breakdown

ModelsLab bills self-hosted TTS at $0.001 per second of generated audio, with a $0.0047 minimum per request, paid from the usage included in your plan: $47 on Standard covers about 13 hours of speech a month. The $149/month Open Source Unlimited plan has no per-second charge on the speech models ModelsLab self-hosts.

By comparison, ElevenLabs bills subscription tiers plus usage-based credits, and cloud vendors like Google Cloud TTS and Amazon Polly charge per million characters with custom-voice programs gated behind enterprise contracts.

Voice Cloning That Ships in Minutes

The ModelsLab voice API clones in a single call: POST to text_to_audio with your text as prompt and a short reference sample as init_audio, and the response returns the generated audio URL. Cloning supports multilingual output, so one reference voice can speak every language your product ships in.

  • Clone from 4 to 10 seconds of reference audio
  • Multilingual generation from a single voice profile
  • Speed and emotion control per request
  • Audio URLs returned for direct playback or storage
  • Consent-based cloning policy and GDPR-compliant handling

Why Developers Choose ModelsLab for Voice in 2026

Key advantages that set us apart

Voice cloning from a 4-second sample
$0.001 per second of generated audio ($0.06 a minute)
Plans from $21/month, no per-second charge on the $149 plan
Multilingual output from one cloned voice
Speed and emotion control per request
One-call cloning: text_to_audio with prompt + init_audio
Music generation on the same platform
Image, video, and LLM APIs with the same key
$47/month Standard covers 10,000 API calls
99.9% uptime SLA for enterprise
GDPR-compliant with consent-based cloning
Basic plan at $21/month with 3,250 API calls

Our Popular Use Cases

What teams build with the best text to speech API:

Generate hours of consistent narration from a single cloned voice. The $149/month Open Source Unlimited plan has no per-second charge, so long-form narration does not grow the bill.

Audiobook and Long-Form Narration

Best Text to Speech API FAQ

It depends on what you optimise for. ElevenLabs has the most natural voices and the largest library ($0.05 to $0.10 per 1,000 characters). Google Cloud and Amazon Polly standard voices are the cheapest per minute ($4 per million characters). ModelsLab bills $0.001 per second of audio with voice cloning from a 4-second sample on every plan, and its $149/month Open Source Unlimited plan has no per-second charge, which makes it cheaper than OpenAI tts-1, Azure or Fish Audio above about 184 hours of speech a month (Google and Amazon standard voices stay cheaper). Prices checked 2026-09-27.

Per minute of speech (about 900 characters): Google Cloud and Amazon Polly standard voices about $0.004, OpenAI tts-1, Azure neural and Fish Audio about $0.014, ElevenLabs $0.045 (Flash) to $0.09 (Multilingual v2), and ModelsLab $0.06 ($0.001 per second of audio), or no per-second charge on the $149/month Open Source Unlimited plan. Prices checked 2026-09-27.

ModelsLab clones voices from 4 to 10 seconds of reference audio in a single call: POST to /api/v6/voice/text_to_audio with your text as prompt and the sample URL as init_audio. ElevenLabs offers instant cloning from its $6/month Starter plan, Cartesia from $5/month, and Fish Audio and Hume through their APIs. OpenAI custom voices and Google Instant Custom Voice are limited to approved customers, and Amazon Polly has no cloning.

Yes. ModelsLab voice profiles support multilingual generation via the language parameter, so a single cloned voice can narrate content in every language your product supports.

ModelsLab is a paid service — plans start at $21/month (Basic, 3,250 API calls) and include text to speech and voice cloning. Creating an account is free; API generation requires an active plan.

Yes. The ModelsLab audio stack covers text to speech, voice cloning, music generation, and audio processing — and the same API key unlocks image, video, and LLM endpoints for a full multimodal product.

Your Data is Secure: GDPR Compliant AI Services

ModelsLab GDPR Compliance Certification Badge

GDPR Compliant

Pricing That's Perfect

Choose plan as per your needs, cancel anytime.

100% refund policy on monthly & yearly plans — cancel anytime

Contact Sales
Best Value

Open Source Unlimited

Mission-Critical

$149/month

100% refund policy · cancel anytime

Unlimited Open Source Models
100% refund policy
24x7 Support
15 parallel generations
Access to all APIs
Unlimited generations on all open-source models
For mission critical workloads
Add Team Members
Priority GPU Clusters
Most Popular

Standard

Production

$47/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Priority Developer Support
10 concurrent API requests
For Production workloads
API access to all models
Prototype

Basic

Prototype

$21/month

100% refund policy · cancel anytime

Moderate Traffic
100% refund policy
Developer Support via Discord/Email
5 concurrent API requests
API access to all models
Shared GPU

Get Expert Support in Seconds

We're Here to Help.

Want to know more? You can email us anytime at support@modelslab.com

View Docs

It depends on what you optimise for. ElevenLabs has the most natural voices and the largest library, at $0.05 to $0.10 per 1,000 characters. Google Cloud and Amazon Polly standard voices are the cheapest per minute, at $4 per million characters. ModelsLab bills $0.001 per second of generated audio ($0.06 a minute) with voice cloning from a 4-second sample on every plan, and its $149/month Open Source Unlimited plan has no per-second charge, which makes it cheaper than OpenAI tts-1, Azure or Fish Audio above about 184 hours of speech a month; Google and Amazon standard voices stay cheaper. Prices checked 2026-09-27.

Per minute of speech (about 900 characters): Google Cloud and Amazon Polly standard voices about $0.004, OpenAI tts-1, Azure neural voices and Fish Audio about $0.014, ElevenLabs $0.045 (Flash) to $0.09 (Multilingual v2), and ModelsLab $0.06, from $0.001 per second of generated audio (at least $0.0047 a request) on plans from $21/month. The $149/month ModelsLab Open Source Unlimited plan has no per-second charge. Prices checked 2026-09-27.

ModelsLab clones a voice from a 4-to-10-second sample in a single call: POST to /api/v6/voice/text_to_audio with your text as prompt and the sample URL as init_audio. ElevenLabs offers instant cloning from its $6/month Starter plan, Cartesia from $5/month, and Fish Audio and Hume through their APIs. OpenAI custom voices and Google Instant Custom Voice are limited to approved customers, and Amazon Polly has no cloning.

Yes. Both text to speech and cloned voices support multilingual generation via the language parameter, so one cloned voice profile can speak multiple languages.

ModelsLab is a paid service — plans start at $21/month (Basic) and cover text to speech and voice cloning. Creating an account is free; generating audio requires an active plan.

Yes. The ModelsLab audio stack covers text to speech, voice cloning, music generation and audio processing, and the same API key also unlocks image, video and LLM endpoints — one integration for a full multimodal product.

No. play.ht did not resolve on September 27, 2026, and its API is no longer reachable. For text to speech with voice cloning through a public API, the alternatives are ElevenLabs, Cartesia, Fish Audio, Hume and ModelsLab.
Plugins

Explore Plugins for Pro

Our plugins are designed to work with the most popular content creation software.

API

Build Apps with
ML
API

Use our API to build apps, generate AI art, create videos, and produce audio with ease.