Text to Speech API — Batch, 30+ Languages, Pay per Character
Pricing
from $5.00 / 1,000 100 characters, standard voices
Text to Speech API — Batch, 30+ Languages, Pay per Character
Text-to-speech (TTS) API: turn one or many texts into MP3 or WAV voiceover files and get, for each text, a dataset item with the audio URL, duration and character count. English, French, Spanish, German and 30+ languages, auto-detected. No API key, pay per character.
Pricing
from $5.00 / 1,000 100 characters, standard voices
Rating
0.0
(0)
Developer
Sherwood
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Give it one text or a thousand, get one voice-over file per text (MP3 or WAV), one dataset item per text: the audio URL, its duration, the character count, or a clear error. English runs on Kokoro (fast, cheap); French, Spanish, German, Japanese and 30+ other languages run on MiniMax Speech 2.8 with natural voices. Language is detected automatically. No API key, no subscription. From $0.05 per 1,000 characters, failures are free.
Built for automation and AI agents: every input alias is accepted (texts, text, prompts, prompt, sentences), nothing is required, and the default input runs on its own so you can see the output shape before sending your own texts.
What you get
One item per text, in your dataset:
{"index": 1,"text": "Bienvenue dans notre boutique. Votre commande a été expédiée.","characters": 61,"outputUrl": "https://api.apify.com/v2/key-value-stores/abc123/records/speech-0001.mp3","durationMs": 4310,"format": "mp3","language": "French","voice": "Calm_Woman","model": "fal-ai/minimax/speech-2.8-turbo","processingMs": 3800,"status": "ok","error": null}
A text that fails (empty, over 5,000 characters, model error) comes back with "status": "error" and the reason in error. It is never charged.
Use cases
- Voice product descriptions, notifications or order updates in the customer's language.
- Turn articles, scripts or lessons into audio versions in bulk.
- Generate voice-overs for short videos from a content calendar.
- Give an AI agent a single call that turns a list of lines into a list of audio files.
Input
| Field | Type | Default | Notes |
|---|---|---|---|
texts | array | two sample sentences | Up to 5,000 characters each. For a single text, use text. |
text, prompts, prompt, sentences | aliases | — | Merged with texts. |
language | auto or a language name | auto | Detected per text when auto. English uses the standard model; other languages always use premium. |
voice | female, male, or a model voice id | female | Kokoro ids (af_heart, am_michael…) for English standard; MiniMax ids (Wise_Woman, Deep_Voice_Man, Friendly_Person…) for premium. |
quality | standard, premium | standard | premium forces MiniMax for every text, with pauses <#1.5#> and tags like (laughs). |
speed | 0.5 to 2.0 | 1.0 | |
outputFormat | mp3, wav | mp3 | Standard (Kokoro) always returns WAV; premium returns MP3. |
maxTexts | integer | 200 | Stop after this many texts. |
Minimal call:
{ "texts": ["Your package has been delivered.", "Votre colis a été livré."] }
Languages
36 languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Polish, Russian, Turkish, Arabic, Hindi, Japanese, Korean, Chinese (Mandarin), Cantonese, Indonesian, Malay, Vietnamese, Thai, Swedish, Danish, Norwegian, Finnish, Greek, Czech, Slovak, Romanian, Hungarian, Ukrainian, Bulgarian, Croatian, Slovenian, Catalan, Hebrew, Afrikaans. Each text is detected on its own when language is auto; Malay and Cantonese must be forced, because detection cannot tell them apart from Indonesian and Mandarin. A language detected outside this list is still voiced by the premium model, which then detects it itself. Close languages can be confused (Croatian is often detected as Slovenian): force language when you know it.
Pricing
Charged per started 100 characters of text actually voiced:
- Standard (English, Kokoro) — $0.005 per 100 characters, i.e. $0.05 per 1,000.
- Premium (any language, MiniMax) — $0.012 per 100 characters, i.e. $0.12 per 1,000.
- Failed texts: free.
Set a maximum cost per run in Apify and the Actor stops cleanly when it is reached; remaining texts are reported, not processed.
Via API, n8n, Make or MCP
Call it like any Apify Actor: POST https://api.apify.com/v2/actors/sherwood~text-to-speech-batch/run-sync-get-dataset-items with the JSON input above, or add it as a tool through the Apify MCP server. Results are plain dataset items, so get-dataset-items returns the URLs directly.
Limits
- 5,000 characters per text; split longer scripts into paragraphs (one item each).
- Generation is typically 3 to 8 seconds per text; the standard model can take longer on its first call of a run while it warms up. 4 texts are processed in parallel.
- Output files live in the run's key-value store and follow your account's data retention.