# Kokoro Text to Speech: $0.08/1K Chars, MP3, No API Key (`conserving_celerytop/kokoro-text-to-speech`) Actor

Convert text to speech with the open Kokoro-82M model, in bulk. $0.08 per 1,000 characters. 41 voices, MP3 or WAV plus optional SRT subtitles, one file per text, from a list or a dataset. Runs inside the Actor, no API key to buy. MCP-ready, no login.

- **URL**: https://apify.com/conserving\_celerytop/kokoro-text-to-speech.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.60 / 1,000 100 characters spokens

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Kokoro Text to Speech

Kokoro Text to Speech turns text into natural-sounding speech with the open Kokoro-82M voice model. Paste one text or thousands, pick one of 41 voices, and get one MP3 or WAV file per text, with optional SRT subtitles. The model runs inside the Actor on CPU, so there is no API key to buy and your text is not sent to any outside speech service. Use it as a Kokoro text to speech API: send texts through the Apify API and read the audio links from the dataset.

### What this Kokoro text to speech Actor does

- Turns each text into one audio file (MP3 or 16-bit WAV, 24 kHz mono) in the run's key-value store.
- Offers 41 voices in 7 languages: English (US and UK), Spanish, French, Hindi, Italian and Portuguese (Brazil).
- Handles long texts. Text is split at sentence ends and spoken piece by piece, with short pauses between sentences and longer ones between paragraphs, so a 20-minute script works in one go.
- Writes an optional SRT subtitle file per text, with one caption per sentence or line.
- Reads texts from another Actor's dataset (for example product descriptions or article text), so you can chain it after a scraper.
- Gives one dataset row per text with the audio link, length in seconds, file size and status.

### How to use Kokoro Text to Speech

1. Add your texts under **Texts**. Each entry becomes one audio file.
2. Pick a **Voice**. The voice sets the language. Heart (English US) is a good default.
3. Choose **Audio format** (MP3 or WAV) and a **Speed** (1.0 is normal).
4. Turn on **SRT subtitles** if you need captions for a video.
5. Click **Start**. When the run ends, open the **Audio files** table and click the links, or open the key-value store to download every file.

To read texts from a dataset, put its ID in **Dataset with texts** and the field name in **Text field**. Use **ID field** to carry your own item ID (for example a SKU) into each row.

### How much does text to speech cost?

You pay $0.008 per started 100 characters of text that became audio, which is $0.08 per 1,000 characters (about $0.08 per minute of speech). The price covers the compute: speech is made on CPU inside the run, with no per-minute platform fee on top. Empty texts, texts over your limit and failed texts are free.

Example: 200 product descriptions of 450 characters each. Each text is 5 units of 100 characters, so 200 x 5 x $0.008 = $8.00 for 200 voice clips, roughly 85 minutes of audio in total. One 10,000-character script (about 10 minutes of speech) costs $0.80.

Set a maximum cost per run in the run options. The Actor stops before a text that would go over it and tells you which texts were left out.

### Input example

```json
{
  "texts": [
    "Welcome to our store. Every order ships within two days.",
    "Chapter one. The morning was quiet, and the harbor was still."
  ],
  "voice": "bf_emma",
  "speed": 1.0,
  "audioFormat": "mp3",
  "subtitles": true
}
```

### Output example

```json
{
  "index": 0,
  "id": null,
  "status": "ok",
  "error": null,
  "textPreview": "Welcome to our store. Every order ships within two days.",
  "characters": 56,
  "voice": "bf_emma",
  "voiceName": "Emma",
  "language": "English (UK)",
  "speed": 1.0,
  "audioFormat": "mp3",
  "audioKey": "speech-0001.mp3",
  "audioUrl": "https://api.apify.com/v2/key-value-stores/.../records/speech-0001.mp3",
  "durationSeconds": 3.9,
  "fileSizeBytes": 31920,
  "subtitlesKey": "speech-0001.srt",
  "subtitlesUrl": "https://api.apify.com/v2/key-value-stores/.../records/speech-0001.srt",
  "chargedUnits": 1,
  "createdAt": "2026-09-27T01:10:38Z"
}
```

The `status` field is `ok`, `empty_text`, `too_long`, `no_speech`, `over_spending_limit`, `missing_text` or `error`. The `STATS` record in the key-value store has totals for the run: texts, characters, audio minutes and units charged.

### Voices

English (US): Heart, Bella, Nicole, Aoede, Kore, Sarah, Nova, Sky, Alloy, Jessica, River, Michael, Fenrir, Puck, Echo, Eric, Liam, Onyx, Adam, Santa. English (UK): Emma, Isabella, Alice, Lily, George, Fable, Lewis, Daniel. Spanish: Dora, Alex, Santa. French: Siwis. Hindi: Alpha, Beta, Omega, Psi. Italian: Sara, Nicola. Portuguese (Brazil): Dora, Alex, Santa. Heart, Bella and Emma are the most natural. Some other voices sound flatter.

### Related Actors

- [Speech to Text & Audio Transcription](https://apify.com/conserving_celerytop/audio-podcast-transcription): Use it to transcribe audio, video and podcast episodes to text with timestamps and subtitles.
- [Article Extractor](https://apify.com/conserving_celerytop/article-extractor): Use it to pull clean article text, dates and metadata from news and blog pages.
- [Web Page to Markdown for AI](https://apify.com/conserving_celerytop/web-page-to-markdown): Use it to turn web pages into clean Markdown for LLMs, RAG and AI agents.
- [Podcast Scraper API](https://apify.com/conserving_celerytop/podcast-episode-scraper): Use it to list every episode of a podcast with audio file URLs from its RSS feed.

### FAQ

**Can I use the audio commercially?** The Kokoro-82M model is released under the Apache 2.0 license, which allows commercial use. You are responsible for the text you convert.

**How long can a text be?** Up to 20,000 characters by default and 50,000 at most per text (about 50 minutes of speech). Split longer books into chapters.

**How fast is it?** On the default 4 GB memory (one CPU core) it makes about one minute of speech in 45 to 60 seconds. More memory gives more cores and faster runs for long texts. The price per character stays the same.

**Can it clone a voice?** No. It uses the 41 built-in voices only.

**Why does a number or abbreviation sound odd?** The voice reads text as written. Spell out unusual abbreviations, symbols and units for the best result.

**Is my text stored?** Text is only used inside your run. The audio files and rows stay in your own Apify storage.

# Actor input Schema

## `texts` (type: `array`):

Enter the texts to turn into speech. Each text becomes one audio file. Line breaks inside a text become short pauses.

## `voice` (type: `string`):

Pick the voice. The voice also sets the language: English (US or UK), Spanish, French, Hindi, Italian or Portuguese (Brazil).

## `speed` (type: `number`):

Set the speaking speed. 1.0 is normal, 0.8 is slower, 1.2 is faster.

## `audioFormat` (type: `string`):

Pick the file type. MP3 is small and plays everywhere. WAV is uncompressed 16-bit audio for editing.

## `subtitles` (type: `boolean`):

Save an SRT subtitle file next to each audio file, with one timed caption per sentence group.

## `datasetId` (type: `string`):

Enter the ID of an Apify dataset. Each item with text in the text field becomes one audio file, after the texts typed above.

## `textField` (type: `string`):

Name of the dataset field that holds the text.

## `idField` (type: `string`):

Name of a dataset field to copy into the id column of each row, so you can match audio files to your items.

## `maxCharactersPerText` (type: `integer`):

Longer texts get a free too\_long row instead of audio. 20,000 characters is about 20 minutes of speech.

## `maxTexts` (type: `integer`):

Stop after this many texts.

## `fileNamePrefix` (type: `string`):

Start of each file name in the key-value store, for example speech-0001.mp3.

## Actor input object example

```json
{
  "texts": [
    "Hello! This is a short test of Kokoro text to speech. Each text you add here becomes one audio file."
  ],
  "voice": "af_heart",
  "speed": 1,
  "audioFormat": "mp3",
  "subtitles": false,
  "textField": "text",
  "maxCharactersPerText": 20000,
  "maxTexts": 1000,
  "fileNamePrefix": "speech"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `files` (type: `string`):

No description

## `stats` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "texts": [
        "Hello! This is a short test of Kokoro text to speech. Each text you add here becomes one audio file."
    ],
    "voice": "af_heart"
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/kokoro-text-to-speech").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "texts": ["Hello! This is a short test of Kokoro text to speech. Each text you add here becomes one audio file."],
    "voice": "af_heart",
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/kokoro-text-to-speech").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "texts": [
    "Hello! This is a short test of Kokoro text to speech. Each text you add here becomes one audio file."
  ],
  "voice": "af_heart"
}' |
apify call conserving_celerytop/kokoro-text-to-speech --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/kokoro-text-to-speech"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/h6mQTL3xUcaEGj9p0/builds/MlHVN1nX3NDPwelSh/openapi.json
