Text to Speech AI - 21 Natural Voices, 8 Languages avatar

Text to Speech AI - 21 Natural Voices, 8 Languages

Pricing

Pay per usage

Go to Apify Store
Text to Speech AI - 21 Natural Voices, 8 Languages

Text to Speech AI - 21 Natural Voices, 8 Languages

Convert text into natural speech with 21 AI voices across English, Spanish, French, Italian, Portuguese, Hindi, Japanese and Chinese. Speed control, long text support, MP3/WAV/Opus download. No rental fee - you pay Apify compute only.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Andrew Babo

Andrew Babo

Maintained by Community

Actor stats

0

Bookmarked

6

Total users

5

Monthly active users

11 days ago

Last modified

Categories

Share

Text to Speech AI — 21 Natural Voices, 8 Languages

Turn any text into natural-sounding speech in seconds. Paste your text, pick a voice, get a ready-to-use audio file (MP3, WAV or Opus) you can drop straight into a video, podcast, app or website.

Good for: YouTube / TikTok / Reels voiceovers, e-learning narration, product demos, audiobooks, IVR and phone prompts, accessibility read-aloud, app notifications.

Quick start

{
"text": "Hello! This audio was generated automatically.",
"voice": "af_heart",
"format": "mp3"
}

Run it and you get back an audio file URL in the run's storage — no setup, no API key, no external account.

What you get

  • 21 voices across English, Spanish, French, Italian, Portuguese, Hindi, Japanese and Chinese
  • Male and female voices, several accents per language
  • Long text supported — the Actor handles the splitting and joins the audio seamlessly
  • Speed control from slow narration to fast reads
  • MP3, WAV or Opus output, ready for editing software
  • Fast warm mode so repeated requests answer almost instantly

Input

FieldDefaultWhat it does
textrequiredThe text you want spoken
voiceaf_heartWhich voice to use
langautoForce a language if your text mixes several
speed1.00.5 = slower, 1.5 = faster
formatwavwav, mp3 or opus
precisionfp32Use int8 to trade a little quality for more speed

Output

Each run stores the audio file and adds one dataset row with the audio URL, duration, voice used and processing time — easy to plug into Make, Zapier, n8n or your own backend.

Pricing

No rental fee. You only pay Apify platform compute, which is typically a fraction of a cent per short clip.

Notes

  • For Vietnamese, use the Vietnamese Text to Speech Actor instead — this one is not trained on Vietnamese.
  • Generated voices are synthetic. Make sure your use complies with the rules of the platform you publish on.

For AI agents (MCP-ready)

This actor is built to be called by AI agents. It works out of the box with the Apify MCP Server — add it to Claude Desktop, Cursor or any MCP client: the agent can produce voiceovers on its own.

{
"mcpServers": {
"kokoro-tts": {
"url": "https://mcp.apify.com/?actors=andrew_babo/kokoro-tts",
"headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
}
}
}

Agent skill (paste into your agent's instructions)

Use the "kokoro-tts" tool to turn text into natural speech in English,
Spanish, French, Italian, Portuguese, Hindi, Japanese or Chinese — 21
preset voices. It is the fastest, cheapest default for plain narration.
For Vietnamese use "vieneu-tts"; for voice cloning use "chatterbox-tts"
(English) or "qwen3-tts".
HOW TO CALL
- { "text": "...", "voice": "af_heart", "format": "mp3" }
- "speed" 0.5-1.5; "format" wav | mp3 | opus (mp3 for publishing).
- Long scripts are split automatically — send the whole text in one call
rather than many small calls.
OUTPUT CONTRACT
- One row per run: audio_url, duration, voice, processing time.
- Give the user the audio_url; do not claim audio exists if
errorCode / NO_INPUT is returned (empty text).
- These are synthetic voices — mention that when the user publishes them.