Kokoro Text to Speech: $0.08/1K Chars, MP3, No API Key avatar

Kokoro Text to Speech: $0.08/1K Chars, MP3, No API Key

Pricing

from $5.60 / 1,000 100 characters spokens

Go to Apify Store
Kokoro Text to Speech: $0.08/1K Chars, MP3, No API Key

Kokoro Text to Speech: $0.08/1K Chars, MP3, No API Key

Convert text to speech with the open Kokoro-82M model, in bulk. $0.08 per 1,000 characters. 41 voices, MP3 or WAV plus optional SRT subtitles, one file per text, from a list or a dataset. Runs inside the Actor, no API key to buy. MCP-ready, no login.

Pricing

from $5.60 / 1,000 100 characters spokens

Rating

0.0

(0)

Developer

Don Mangu

Don Mangu

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

an hour ago

Last modified

Share

Kokoro Text to Speech

Kokoro Text to Speech turns text into natural-sounding speech with the open Kokoro-82M voice model. Paste one text or thousands, pick one of 41 voices, and get one MP3 or WAV file per text, with optional SRT subtitles. The model runs inside the Actor on CPU, so there is no API key to buy and your text is not sent to any outside speech service. Use it as a Kokoro text to speech API: send texts through the Apify API and read the audio links from the dataset.

What this Kokoro text to speech Actor does

  • Turns each text into one audio file (MP3 or 16-bit WAV, 24 kHz mono) in the run's key-value store.
  • Offers 41 voices in 7 languages: English (US and UK), Spanish, French, Hindi, Italian and Portuguese (Brazil).
  • Handles long texts. Text is split at sentence ends and spoken piece by piece, with short pauses between sentences and longer ones between paragraphs, so a 20-minute script works in one go.
  • Writes an optional SRT subtitle file per text, with one caption per sentence or line.
  • Reads texts from another Actor's dataset (for example product descriptions or article text), so you can chain it after a scraper.
  • Gives one dataset row per text with the audio link, length in seconds, file size and status.

How to use Kokoro Text to Speech

  1. Add your texts under Texts. Each entry becomes one audio file.
  2. Pick a Voice. The voice sets the language. Heart (English US) is a good default.
  3. Choose Audio format (MP3 or WAV) and a Speed (1.0 is normal).
  4. Turn on SRT subtitles if you need captions for a video.
  5. Click Start. When the run ends, open the Audio files table and click the links, or open the key-value store to download every file.

To read texts from a dataset, put its ID in Dataset with texts and the field name in Text field. Use ID field to carry your own item ID (for example a SKU) into each row.

How much does text to speech cost?

You pay $0.008 per started 100 characters of text that became audio, which is $0.08 per 1,000 characters (about $0.08 per minute of speech). The price covers the compute: speech is made on CPU inside the run, with no per-minute platform fee on top. Empty texts, texts over your limit and failed texts are free.

Example: 200 product descriptions of 450 characters each. Each text is 5 units of 100 characters, so 200 x 5 x $0.008 = $8.00 for 200 voice clips, roughly 85 minutes of audio in total. One 10,000-character script (about 10 minutes of speech) costs $0.80.

Set a maximum cost per run in the run options. The Actor stops before a text that would go over it and tells you which texts were left out.

Input example

{
"texts": [
"Welcome to our store. Every order ships within two days.",
"Chapter one. The morning was quiet, and the harbor was still."
],
"voice": "bf_emma",
"speed": 1.0,
"audioFormat": "mp3",
"subtitles": true
}

Output example

{
"index": 0,
"id": null,
"status": "ok",
"error": null,
"textPreview": "Welcome to our store. Every order ships within two days.",
"characters": 56,
"voice": "bf_emma",
"voiceName": "Emma",
"language": "English (UK)",
"speed": 1.0,
"audioFormat": "mp3",
"audioKey": "speech-0001.mp3",
"audioUrl": "https://api.apify.com/v2/key-value-stores/.../records/speech-0001.mp3",
"durationSeconds": 3.9,
"fileSizeBytes": 31920,
"subtitlesKey": "speech-0001.srt",
"subtitlesUrl": "https://api.apify.com/v2/key-value-stores/.../records/speech-0001.srt",
"chargedUnits": 1,
"createdAt": "2026-09-27T01:10:38Z"
}

The status field is ok, empty_text, too_long, no_speech, over_spending_limit, missing_text or error. The STATS record in the key-value store has totals for the run: texts, characters, audio minutes and units charged.

Voices

English (US): Heart, Bella, Nicole, Aoede, Kore, Sarah, Nova, Sky, Alloy, Jessica, River, Michael, Fenrir, Puck, Echo, Eric, Liam, Onyx, Adam, Santa. English (UK): Emma, Isabella, Alice, Lily, George, Fable, Lewis, Daniel. Spanish: Dora, Alex, Santa. French: Siwis. Hindi: Alpha, Beta, Omega, Psi. Italian: Sara, Nicola. Portuguese (Brazil): Dora, Alex, Santa. Heart, Bella and Emma are the most natural. Some other voices sound flatter.

FAQ

Can I use the audio commercially? The Kokoro-82M model is released under the Apache 2.0 license, which allows commercial use. You are responsible for the text you convert.

How long can a text be? Up to 20,000 characters by default and 50,000 at most per text (about 50 minutes of speech). Split longer books into chapters.

How fast is it? On the default 4 GB memory (one CPU core) it makes about one minute of speech in 45 to 60 seconds. More memory gives more cores and faster runs for long texts. The price per character stays the same.

Can it clone a voice? No. It uses the 41 built-in voices only.

Why does a number or abbreviation sound odd? The voice reads text as written. Spell out unusual abbreviations, symbols and units for the best result.

Is my text stored? Text is only used inside your run. The audio files and rows stay in your own Apify storage.