# Text-to-Speech API — Local Voices, Six Audio Formats (`dropin-apis/tts-compat`) Actor

English text-to-speech API with a compatible /v1/audio/speech endpoint. Independent Kokoro voices, six audio formats and speed control. Standby $0.009; stored batch $0.12 per started 1,000 characters. No voice cloning or affiliation.

- **URL**: https://apify.com/dropin-apis/tts-compat.md
- **Developed by:** [drop-in apis](https://apify.com/dropin-apis) (community)
- **Categories:** AI, Developer tools
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $9.00 / 1,000 1,000 input characters

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Text-to-Speech API — Local Voices, Six Audio Formats

Convert English text into MP3, Opus, AAC, FLAC, WAV or PCM through a compatible `POST /v1/audio/speech` endpoint, using independent local Kokoro voices at **$0.009 per started 1,000 input characters in Standby**, or **$0.12 per started 1,000 in a normal run that stores a downloadable file**.

Standby streams MP3, Opus, AAC and PCM progressively for up to 4,096 characters. WAV and FLAC remain buffered and accept up to 500 characters over HTTP; normal Store/MCP runs support 4,096 characters in all six formats.

This is an independent speech service. Compatibility covers the request shape and audio response; it does **not** mean identical voices, pronunciation, prosody, model quality or provider affiliation. English only in this version. No voice cloning.

### Quick start

Use the Standby host shown in this Actor's **Endpoints** tab. Authentication is an Apify API token in a Bearer header. Agents can also use normal Store runs or Apify MCP `call-actor`: provide `input` text and optionally `model`, `voice`, `response_format` and `speed`. That route returns one dataset item with the audio URL and metadata, and uses the separate batch price shown below.

```bash
curl --fail-with-body \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  --data '{"model":"tts-1","input":"Hello, world.","voice":"amber","response_format":"mp3"}' \
  'https://dropin-apis--tts-compat.apify.actor/v1/audio/speech' \
  --output speech.mp3
```

The hostname above is the Actor-level host; confirm the actual deployed host in **Endpoints**. Do not put your token in a URL. A cold instance takes longer than a warm one; allow up to five minutes for the platform's first response and a longer total/read timeout for a long audio stream. MP3, Opus, AAC and PCM use chunked binary streaming, with one continuous codec stream. WAV and FLAC remain buffered.

Python clients with a compatible speech interface can use this pattern:

```python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["APIFY_TOKEN"],
    base_url="https://dropin-apis--tts-compat.apify.actor/v1",
    timeout=1500,  # Total/read budget for a long stream; first audio is earlier.
    max_retries=0,  # Explicit caller-controlled retries; no idempotency guarantee.
)
with client.audio.speech.with_streaming_response.create(
    model="tts-1", input="Hello, world.", voice="amber", response_format="wav"
) as response:
    response.stream_to_file("speech.wav")
```

This is a technical client example only. MP3 audio begins after the first synthesis window and acknowledgement of its first character event. Client version 3.24.0 was tested with the real local engine, an MP3 decode and an instruction-error response. Cloud proxy compatibility and measured latency are recorded in `VERIFY.md`.

### Store and MCP batch quick start

Use this Actor's **Start** button or an Apify MCP `call-actor` tool with this input:

```json
{"input":"Hello, world.","voice":"amber","response_format":"mp3"}
```

Exactly one output row gives the audio download URL/key, format, duration, input character count, local voice, provider, billed units and processing time. The audio is stored in the default key-value store. HTTP requests still require `model`, `input` and `voice`; normal runs supply the short defaults for missing fields. Batch runs allow up to 4,096 characters and use 4 GB by default.

### Request contract

| Parameter | Values / behavior |
| --- | --- |
| `model` | `tts-1`, `tts-1-hd`, `gpt-4o-mini-tts`. These are compatibility identifiers: **all three use the same local model**. No HD/model-quality distinction. |
| `input` | 1–4,096 Unicode characters of English text. Whitespace counts toward billing; whitespace is normalized for pronunciation. Unsupported scripts and unknown lexicon words are rejected rather than silently skipped. |
| `voice` | One of the local names below or its compatibility alias. Required for HTTP; normal runs default to `amber`. |
| `response_format` | `mp3` (default), `opus`, `aac`, `flac`, `wav`, `pcm`. |
| `speed` | Number from 0.25 to 4.0; default 1. Pitch-preserving tempo conversion after synthesis. |
| `instructions` | Omit or use an empty string. Nonempty instruction-controlled prosody returns HTTP 400. |
| `stream_format` | Omit or use `audio`. SSE is unsupported. |

POST is available at `/v1/audio/speech` and `/audio/speech`. GET `/v1/models` and `/models` list the accepted identifiers. GET `/v1/voices` and `/voices` expose the exact local voice inventory and alias map. A pronunciation token longer than the model window is rejected before the character charge. Output is capped at 20 minutes: an overflow discovered after streaming starts interrupts the paid stream.

### Independent voices and disclosure

| Compatibility parameter value | Local voice | Upstream tensor | Accent |
| --- | --- | --- | --- |
| `alloy` | amber | af\_heart | American English |
| `echo` | birch | am\_fenrir | American English |
| `fable` | cove | bf\_emma | British English |
| `onyx` | dusk | am\_michael | American English |
| `nova` | ember | af\_bella | American English |
| `shimmer` | fern | af\_nicole | American English |
| `ash` | grove | am\_puck | American English |
| `ballad` | harbor | bm\_george | British English |
| `coral` | iris | af\_sarah | American English |
| `sage` | jade | af\_kore | American English |
| `verse` | lake | af\_aoede | American English |
| `marin` | meadow | bf\_isabella | British English |
| `cedar` | north | bm\_lewis | British English |

The left column consists solely of accepted **API parameter values**, not advertised voice identities. Each alias maps to a distinct local voice. Every successful speech response has an `X-Voice-Provider` disclosure header and `X-Local-Voice`. Pronunciation and timbre differ from any prior provider. Long inputs are split into bounded pronunciation windows, so phrase boundaries can affect prosody.

### Audio formats

| Format | Content type | Container / decoding |
| --- | --- | --- |
| MP3 | `audio/mpeg` | MP3, mono, 24 kHz, 128 kbps |
| Opus | `audio/ogg` | Opus in Ogg, mono, 48 kHz decoder rate |
| AAC | `audio/aac` | AAC with ADTS framing, mono, 24 kHz |
| FLAC | `audio/flac` | Buffered HTTP; FLAC, mono, 24 kHz; the non-seekable encoder header can omit total duration |
| WAV | `audio/wav` | PCM16 inside WAV, mono, 24 kHz |
| PCM | `audio/pcm` | **Headerless signed 16-bit little-endian**, mono, 24 kHz |

Decode PCM with explicit parameters, for example `ffplay -f s16le -ar 24000 -ac 1 speech.pcm`. Clients must not treat PCM bytes as a WAV file.

Streamed formats set `X-Audio-Streaming: true` and `X-Billing-Mode: incremental`, and omit `Content-Length` and final-duration/processing-time headers, because those values are not known at the first byte. `X-Billed-Units` states the full-input count for a completed request; an interrupted stream may charge fewer units. WAV/FLAC set streaming to `false` and include final duration. Operational logs contain only timing, character/byte counts, charge milestones and error codes; they never contain the input text or audio.

### Pricing and error handling

- Standby event `characters-1k`: `ceil(len(input) / 1000) × $0.009`, measured in Unicode code points, including spaces. 250 or 1,000 characters cost $0.009; 1,001 cost $0.018; 4,096 cost $0.045.
- Normal Store/MCP event `batch-characters-1k`: `ceil(len(input) / 1000) × $0.12`. The batch price covers its different compute/storage economics. A 250- or 1,000-character batch costs $0.12; 4,096 costs $0.60. The two character events are mutually exclusive for a request.
- Apify additionally applies its synthetic Actor-start event, normally $0.00005 per allocated GB at each instance start. At 8 GB Standby this is $0.0004; at the normal 4 GB batch default it is $0.0002. It is per instance start, not necessarily per speech request.
- Validation, whole-input pronunciation and first-window synthesis/encoding failures do not trigger the character event. Later failures after streaming begins can leave charged partial audio. The platform's start event can still apply to a failed request.
- A **full-budget precheck runs before synthesis**, so a spending cap unable to cover `ceil(len(input) / 1000)` returns **402 without work or audio**. Streaming charges one `characters-1k` event before the first audio byte, then each further event before releasing audio that can cross the next 1,000-character boundary in the original text. Whole-input pronunciation validation and each new unit's first synthesis/encoding preflight happen before its charge. Failure before a new unit starts does not charge that unit; failure within an already-started unit retains that unit.
- Words are kept whole: a word spanning a billing boundary needs the next event before its audio, slightly before the exact character boundary. Leading/trailing whitespace still counts; silent trailing characters attach to the last real audio window, so completed requests always total `ceil(len(input) / 1000)` events. Unusual whitespace or one word spanning multiple units can require more than one event before the same packet.
- Before the first audio byte, a missing/zero acknowledgement or billing exception returns **503 without audio**; an insufficient/exhausted spending cap returns **402 without audio**. During streaming, a failed charge, changed cap, synthesis or encoding failure truncates the connection without JSON or a successful EOF. Only started units remain charged; later units do not. Batch and buffered formats complete generation/encoding before billing.
- One native synthesis per instance; additional requests return **429** if they reach the same busy instance. The platform can start more instances. Errors use an `error` JSON object with `message`, `type`, `param` and `code`.

There is no atomic transaction spanning billing and network delivery. A lost connection after an acknowledged event can still prevent the client receiving that window; incremental charging cannot guarantee receipt, and there is no automatic refund. Malformed/uncertain SDK acknowledgements also fail closed and cannot be treated as proof of no platform charge. Do not retry blindly on network failures: repeated successful calls are separately charged. No idempotency key is implemented in this version. The platform's separate Actor-start event is not a character charge and can apply before speech succeeds.

### Privacy and runtime

Standby does not log input text or audio, does not write either into Actor storage, and uses pipes rather than request-created temporary files. Audio is retained in memory only while handling the request. The cloud platform still transports the request; this is not a promise that no infrastructure metadata is collected.

Normal runs are supported for customers and agents. Apify retains their input in run storage; generated audio is saved in the default key-value store, and exactly one dataset row contains `audioUrl`, `audioKey`, `format`, `duration`, `characters`, local `voice` and `provider`. Clear the run/storage when you no longer need them. An empty input object uses `Hello, world.` with local voice `amber` and MP3; missing model defaults to `tts-1`. Supplied invalid values are rejected. The fixed health endpoint returns only synthesis metadata, never audio or caller-controlled speech, and caches one successful self-test per instance.

Kokoro-82M v1.0 runs through quantized ONNX on CPU. Assets are fetched at image build time from a pinned revision and verified with SHA-256. Misaki English lexicon/POS G2P runs with no espeak-ng or GPL phonemizer fallback. FFmpeg is a separate encoding executable; the Debian build includes GPL components, disclosed in [LICENSES.md](./LICENSES.md). Every bundled voice's original filename, license and hash is recorded in `assets.json`. Source/model provenance is described in [Kokoro's model card](https://huggingface.co/hexgrad/Kokoro-82M) and the [ONNX export](https://huggingface.co/onnx-community/Kokoro-82M-v1.0-ONNX).

### Limits

- **English only** in this version; an unsupported script returns HTTP 400. No voice cloning.
- Standby streams MP3, Opus, AAC and PCM for up to **4,096 characters**; WAV and FLAC are buffered and limited to **500 characters** over HTTP (longer input returns 413 `use_batch_mode` before any charge). A normal Store/MCP batch run supports all six formats up to 4,096 characters.
- Non-empty `instructions` and SSE `stream_format` are not supported and return HTTP 400; the `model` field is a compatibility identifier only (all models select the same local voice model).
- Generation runs on CPU with no real-time guarantee; see the measured latency in the FAQ below.

### Frequently asked questions

#### Can I change only the base URL of an existing speech client?

Yes for the supported JSON endpoint, model identifiers, voices, formats and speed. Use your Apify token as the client key. Re-test English pronunciation and voice selection; acoustic output is different. Remove nonempty instructions and SSE options.

#### Does an accepted model identifier select a different speech model?

No. Every accepted identifier selects Kokoro-82M v1.0. There is no claim of equivalent quality, HD behavior or instruction following.

#### Does this text-to-speech API support Arabic or other languages?

This version supports English only. Unsupported alphabetic scripts return HTTP 400. Multilingual model weights alone do not make the chosen permissive G2P pipeline multilingual.

#### Why can a rare name or acronym return 400?

The no-espeak English pronunciation lexicon does not cover every spelling. Unknown pronunciations are rejected so that words are not silently lost. Rephrase or spell out a supported name. This is a substantive compatibility limit.

#### Can an AI agent use Apify MCP call-actor to synthesize speech?

Yes. Normal `call-actor` runs synthesize your text and return one dataset row with a downloadable audio URL. Use the batch price ($0.12 per started 1,000 characters) for this route. An HTTP tool can use the cheaper Standby endpoint ($0.009). The OpenAPI schema describes all HTTP routes. No x402 eligibility is claimed for this Standby-enabled Actor.

#### What happens when billing fails or I reach my spending limit?

The whole-input budget is checked before synthesis; character events are then acknowledged one at a time before their staged audio is released. Before streaming, billing failure returns 503 and an insufficient/exhausted cap returns 402, with no audio. If a charge, synthesis, encoding or delivery fails after streaming begins, the connection is truncated: started units remain charged, later units do not, and there is no automatic refund or replacement JSON. A client must reject an interrupted HTTP stream rather than treating partial audio as the completed result. Normal batch runs still charge only after complete encoding.

#### How fast is CPU speech generation?

It depends on input length, voice, allocated CPU and cold-start overhead. Streaming changes time to first audio, not the total work required. There is no real-time guarantee. Standby uses 8 GB/two ONNX intra-op threads; normal runs default to 4 GB/two threads. Although Apify documents one/two nominal cores, measured cloud cgroup limits were 0.75/1.5 cores. A controlled 1,000-character cloud comparison improved from 171.91 seconds at 4 GB/one thread to 98.68 seconds at 8 GB/two threads (42.6% faster).

The latest **incrementally billed 4,096-character MP3 stream** delivered its first bytes after **10.50 seconds**, completed after **482.40 seconds**, and fully decoded **264.70 seconds of audio**. Telemetry acknowledged its five events separately, before audio began and at later stream milestones. Historical streaming measurements before incremental billing were 8.52/119.34 seconds for 1,000 characters and 7.31/503.00 seconds for 4,096. All proxy timings include a readiness GET and are observations, not isolated cold POST startup or a speedup guarantee. First audio has a 270-second application deadline; a stream has a further 1,200-second completion deadline. Buffered WAV/FLAC requests longer than 500 characters return 413 with batch guidance before synthesis or charging. See `VERIFY.md` for workflow links and exact conditions.

#### Are the voices or API affiliated with another provider?

No. The voices are independent Kokoro tensors, disclosed in headers and the voice inventory. Compatibility identifiers are accepted only as parameter values.

### Related tools

[Drop-in Labs](https://alidaram99.github.io/) · [API migration guides](https://alidaram99.github.io/api-alternatives/) · [Published Actor portfolio](https://apify.com/dropin-apis)

# Actor input Schema

## `model` (type: `string`):

Compatibility identifier; all three select the same independent local engine.

## `input` (type: `string`):

English text only. Unknown pronunciations are rejected. Whitespace counts toward billing.

## `voice` (type: `string`):

Local voice or compatibility alias. Aliases select distinct local voices, not original-provider voices.

## `response_format` (type: `string`):

Standby streams MP3/Opus/AAC/PCM up to 4096 chars. Buffered WAV/FLAC is limited to 500 chars in Standby; batch supports 4096 for all formats. Opus is Ogg; PCM is signed 16-bit mono 24 kHz.

## `speed` (type: `number`):

Pitch-preserving playback speed multiplier, from 0.25 to 4.0.

## `instructions` (type: `string`):

Nonempty instructions are rejected with HTTP 400; controlled prosody is unsupported.

## `stream_format` (type: `string`):

Binary audio streaming for MP3/Opus/AAC/PCM; WAV/FLAC buffered. SSE is unsupported.

## Actor input object example

```json
{
  "model": "tts-1",
  "input": "Hello, world.",
  "voice": "amber",
  "response_format": "mp3",
  "speed": 1
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "model": "tts-1",
    "input": "Hello, world.",
    "voice": "amber"
};

// Run the Actor and wait for it to finish
const run = await client.actor("dropin-apis/tts-compat").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "model": "tts-1",
    "input": "Hello, world.",
    "voice": "amber",
}

# Run the Actor and wait for it to finish
run = client.actor("dropin-apis/tts-compat").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "model": "tts-1",
  "input": "Hello, world.",
  "voice": "amber"
}' |
apify call dropin-apis/tts-compat --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dropin-apis/tts-compat"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iJ0c74wOD3zOePq5R/builds/orv7lV6S6EOE2c4RA/openapi.json
