Text-to-Speech API — Local Voices, Six Audio Formats avatar

Text-to-Speech API — Local Voices, Six Audio Formats

Pricing

from $9.00 / 1,000 1,000 input characters

Go to Apify Store
Text-to-Speech API — Local Voices, Six Audio Formats

Text-to-Speech API — Local Voices, Six Audio Formats

English text-to-speech API with a compatible /v1/audio/speech endpoint. Independent Kokoro voices, six audio formats and speed control. Standby $0.009; stored batch $0.12 per started 1,000 characters. No voice cloning or affiliation.

Pricing

from $9.00 / 1,000 1,000 input characters

Rating

0.0

(0)

Developer

drop-in apis

drop-in apis

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

4 hours ago

Last modified

Share

Convert English text into MP3, Opus, AAC, FLAC, WAV or PCM through a compatible POST /v1/audio/speech endpoint, using independent local Kokoro voices at $0.009 per started 1,000 input characters in Standby, or $0.12 per started 1,000 in a normal run that stores a downloadable file.

Standby streams MP3, Opus, AAC and PCM progressively for up to 4,096 characters. WAV and FLAC remain buffered and accept up to 500 characters over HTTP; normal Store/MCP runs support 4,096 characters in all six formats.

This is an independent speech service. Compatibility covers the request shape and audio response; it does not mean identical voices, pronunciation, prosody, model quality or provider affiliation. English only in this version. No voice cloning.

Quick start

Use the Standby host shown in this Actor's Endpoints tab. Authentication is an Apify API token in a Bearer header. Agents can also use normal Store runs or Apify MCP call-actor: provide input text and optionally model, voice, response_format and speed. That route returns one dataset item with the audio URL and metadata, and uses the separate batch price shown below.

curl --fail-with-body \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H 'Content-Type: application/json' \
--data '{"model":"tts-1","input":"Hello, world.","voice":"amber","response_format":"mp3"}' \
'https://dropin-apis--tts-compat.apify.actor/v1/audio/speech' \
--output speech.mp3

The hostname above is the Actor-level host; confirm the actual deployed host in Endpoints. Do not put your token in a URL. A cold instance takes longer than a warm one; allow up to five minutes for the platform's first response and a longer total/read timeout for a long audio stream. MP3, Opus, AAC and PCM use chunked binary streaming, with one continuous codec stream. WAV and FLAC remain buffered.

Python clients with a compatible speech interface can use this pattern:

import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["APIFY_TOKEN"],
base_url="https://dropin-apis--tts-compat.apify.actor/v1",
timeout=1500, # Total/read budget for a long stream; first audio is earlier.
max_retries=0, # Explicit caller-controlled retries; no idempotency guarantee.
)
with client.audio.speech.with_streaming_response.create(
model="tts-1", input="Hello, world.", voice="amber", response_format="wav"
) as response:
response.stream_to_file("speech.wav")

This is a technical client example only. MP3 audio begins after the first synthesis window and acknowledgement of its first character event. Client version 3.24.0 was tested with the real local engine, an MP3 decode and an instruction-error response. Cloud proxy compatibility and measured latency are recorded in VERIFY.md.

Store and MCP batch quick start

Use this Actor's Start button or an Apify MCP call-actor tool with this input:

{"input":"Hello, world.","voice":"amber","response_format":"mp3"}

Exactly one output row gives the audio download URL/key, format, duration, input character count, local voice, provider, billed units and processing time. The audio is stored in the default key-value store. HTTP requests still require model, input and voice; normal runs supply the short defaults for missing fields. Batch runs allow up to 4,096 characters and use 4 GB by default.

Request contract

ParameterValues / behavior
modeltts-1, tts-1-hd, gpt-4o-mini-tts. These are compatibility identifiers: all three use the same local model. No HD/model-quality distinction.
input1–4,096 Unicode characters of English text. Whitespace counts toward billing; whitespace is normalized for pronunciation. Unsupported scripts and unknown lexicon words are rejected rather than silently skipped.
voiceOne of the local names below or its compatibility alias. Required for HTTP; normal runs default to amber.
response_formatmp3 (default), opus, aac, flac, wav, pcm.
speedNumber from 0.25 to 4.0; default 1. Pitch-preserving tempo conversion after synthesis.
instructionsOmit or use an empty string. Nonempty instruction-controlled prosody returns HTTP 400.
stream_formatOmit or use audio. SSE is unsupported.

POST is available at /v1/audio/speech and /audio/speech. GET /v1/models and /models list the accepted identifiers. GET /v1/voices and /voices expose the exact local voice inventory and alias map. A pronunciation token longer than the model window is rejected before the character charge. Output is capped at 20 minutes: an overflow discovered after streaming starts interrupts the paid stream.

Independent voices and disclosure

Compatibility parameter valueLocal voiceUpstream tensorAccent
alloyamberaf_heartAmerican English
echobircham_fenrirAmerican English
fablecovebf_emmaBritish English
onyxduskam_michaelAmerican English
novaemberaf_bellaAmerican English
shimmerfernaf_nicoleAmerican English
ashgroveam_puckAmerican English
balladharborbm_georgeBritish English
coralirisaf_sarahAmerican English
sagejadeaf_koreAmerican English
verselakeaf_aoedeAmerican English
marinmeadowbf_isabellaBritish English
cedarnorthbm_lewisBritish English

The left column consists solely of accepted API parameter values, not advertised voice identities. Each alias maps to a distinct local voice. Every successful speech response has an X-Voice-Provider disclosure header and X-Local-Voice. Pronunciation and timbre differ from any prior provider. Long inputs are split into bounded pronunciation windows, so phrase boundaries can affect prosody.

Audio formats

FormatContent typeContainer / decoding
MP3audio/mpegMP3, mono, 24 kHz, 128 kbps
Opusaudio/oggOpus in Ogg, mono, 48 kHz decoder rate
AACaudio/aacAAC with ADTS framing, mono, 24 kHz
FLACaudio/flacBuffered HTTP; FLAC, mono, 24 kHz; the non-seekable encoder header can omit total duration
WAVaudio/wavPCM16 inside WAV, mono, 24 kHz
PCMaudio/pcmHeaderless signed 16-bit little-endian, mono, 24 kHz

Decode PCM with explicit parameters, for example ffplay -f s16le -ar 24000 -ac 1 speech.pcm. Clients must not treat PCM bytes as a WAV file.

Streamed formats set X-Audio-Streaming: true and X-Billing-Mode: incremental, and omit Content-Length and final-duration/processing-time headers, because those values are not known at the first byte. X-Billed-Units states the full-input count for a completed request; an interrupted stream may charge fewer units. WAV/FLAC set streaming to false and include final duration. Operational logs contain only timing, character/byte counts, charge milestones and error codes; they never contain the input text or audio.

Pricing and error handling

  • Standby event characters-1k: ceil(len(input) / 1000) × $0.009, measured in Unicode code points, including spaces. 250 or 1,000 characters cost $0.009; 1,001 cost $0.018; 4,096 cost $0.045.
  • Normal Store/MCP event batch-characters-1k: ceil(len(input) / 1000) × $0.12. The batch price covers its different compute/storage economics. A 250- or 1,000-character batch costs $0.12; 4,096 costs $0.60. The two character events are mutually exclusive for a request.
  • Apify additionally applies its synthetic Actor-start event, normally $0.00005 per allocated GB at each instance start. At 8 GB Standby this is $0.0004; at the normal 4 GB batch default it is $0.0002. It is per instance start, not necessarily per speech request.
  • Validation, whole-input pronunciation and first-window synthesis/encoding failures do not trigger the character event. Later failures after streaming begins can leave charged partial audio. The platform's start event can still apply to a failed request.
  • A full-budget precheck runs before synthesis, so a spending cap unable to cover ceil(len(input) / 1000) returns 402 without work or audio. Streaming charges one characters-1k event before the first audio byte, then each further event before releasing audio that can cross the next 1,000-character boundary in the original text. Whole-input pronunciation validation and each new unit's first synthesis/encoding preflight happen before its charge. Failure before a new unit starts does not charge that unit; failure within an already-started unit retains that unit.
  • Words are kept whole: a word spanning a billing boundary needs the next event before its audio, slightly before the exact character boundary. Leading/trailing whitespace still counts; silent trailing characters attach to the last real audio window, so completed requests always total ceil(len(input) / 1000) events. Unusual whitespace or one word spanning multiple units can require more than one event before the same packet.
  • Before the first audio byte, a missing/zero acknowledgement or billing exception returns 503 without audio; an insufficient/exhausted spending cap returns 402 without audio. During streaming, a failed charge, changed cap, synthesis or encoding failure truncates the connection without JSON or a successful EOF. Only started units remain charged; later units do not. Batch and buffered formats complete generation/encoding before billing.
  • One native synthesis per instance; additional requests return 429 if they reach the same busy instance. The platform can start more instances. Errors use an error JSON object with message, type, param and code.

There is no atomic transaction spanning billing and network delivery. A lost connection after an acknowledged event can still prevent the client receiving that window; incremental charging cannot guarantee receipt, and there is no automatic refund. Malformed/uncertain SDK acknowledgements also fail closed and cannot be treated as proof of no platform charge. Do not retry blindly on network failures: repeated successful calls are separately charged. No idempotency key is implemented in this version. The platform's separate Actor-start event is not a character charge and can apply before speech succeeds.

Privacy and runtime

Standby does not log input text or audio, does not write either into Actor storage, and uses pipes rather than request-created temporary files. Audio is retained in memory only while handling the request. The cloud platform still transports the request; this is not a promise that no infrastructure metadata is collected.

Normal runs are supported for customers and agents. Apify retains their input in run storage; generated audio is saved in the default key-value store, and exactly one dataset row contains audioUrl, audioKey, format, duration, characters, local voice and provider. Clear the run/storage when you no longer need them. An empty input object uses Hello, world. with local voice amber and MP3; missing model defaults to tts-1. Supplied invalid values are rejected. The fixed health endpoint returns only synthesis metadata, never audio or caller-controlled speech, and caches one successful self-test per instance.

Kokoro-82M v1.0 runs through quantized ONNX on CPU. Assets are fetched at image build time from a pinned revision and verified with SHA-256. Misaki English lexicon/POS G2P runs with no espeak-ng or GPL phonemizer fallback. FFmpeg is a separate encoding executable; the Debian build includes GPL components, disclosed in ./LICENSES.md. Every bundled voice's original filename, license and hash is recorded in assets.json. Source/model provenance is described in Kokoro's model card and the ONNX export.

Limits

  • English only in this version; an unsupported script returns HTTP 400. No voice cloning.
  • Standby streams MP3, Opus, AAC and PCM for up to 4,096 characters; WAV and FLAC are buffered and limited to 500 characters over HTTP (longer input returns 413 use_batch_mode before any charge). A normal Store/MCP batch run supports all six formats up to 4,096 characters.
  • Non-empty instructions and SSE stream_format are not supported and return HTTP 400; the model field is a compatibility identifier only (all models select the same local voice model).
  • Generation runs on CPU with no real-time guarantee; see the measured latency in the FAQ below.

Frequently asked questions

Can I change only the base URL of an existing speech client?

Yes for the supported JSON endpoint, model identifiers, voices, formats and speed. Use your Apify token as the client key. Re-test English pronunciation and voice selection; acoustic output is different. Remove nonempty instructions and SSE options.

Does an accepted model identifier select a different speech model?

No. Every accepted identifier selects Kokoro-82M v1.0. There is no claim of equivalent quality, HD behavior or instruction following.

Does this text-to-speech API support Arabic or other languages?

This version supports English only. Unsupported alphabetic scripts return HTTP 400. Multilingual model weights alone do not make the chosen permissive G2P pipeline multilingual.

Why can a rare name or acronym return 400?

The no-espeak English pronunciation lexicon does not cover every spelling. Unknown pronunciations are rejected so that words are not silently lost. Rephrase or spell out a supported name. This is a substantive compatibility limit.

Can an AI agent use Apify MCP call-actor to synthesize speech?

Yes. Normal call-actor runs synthesize your text and return one dataset row with a downloadable audio URL. Use the batch price ($0.12 per started 1,000 characters) for this route. An HTTP tool can use the cheaper Standby endpoint ($0.009). The OpenAPI schema describes all HTTP routes. No x402 eligibility is claimed for this Standby-enabled Actor.

What happens when billing fails or I reach my spending limit?

The whole-input budget is checked before synthesis; character events are then acknowledged one at a time before their staged audio is released. Before streaming, billing failure returns 503 and an insufficient/exhausted cap returns 402, with no audio. If a charge, synthesis, encoding or delivery fails after streaming begins, the connection is truncated: started units remain charged, later units do not, and there is no automatic refund or replacement JSON. A client must reject an interrupted HTTP stream rather than treating partial audio as the completed result. Normal batch runs still charge only after complete encoding.

How fast is CPU speech generation?

It depends on input length, voice, allocated CPU and cold-start overhead. Streaming changes time to first audio, not the total work required. There is no real-time guarantee. Standby uses 8 GB/two ONNX intra-op threads; normal runs default to 4 GB/two threads. Although Apify documents one/two nominal cores, measured cloud cgroup limits were 0.75/1.5 cores. A controlled 1,000-character cloud comparison improved from 171.91 seconds at 4 GB/one thread to 98.68 seconds at 8 GB/two threads (42.6% faster).

The latest incrementally billed 4,096-character MP3 stream delivered its first bytes after 10.50 seconds, completed after 482.40 seconds, and fully decoded 264.70 seconds of audio. Telemetry acknowledged its five events separately, before audio began and at later stream milestones. Historical streaming measurements before incremental billing were 8.52/119.34 seconds for 1,000 characters and 7.31/503.00 seconds for 4,096. All proxy timings include a readiness GET and are observations, not isolated cold POST startup or a speedup guarantee. First audio has a 270-second application deadline; a stream has a further 1,200-second completion deadline. Buffered WAV/FLAC requests longer than 500 characters return 413 with batch guidance before synthesis or charging. See VERIFY.md for workflow links and exact conditions.

Are the voices or API affiliated with another provider?

No. The voices are independent Kokoro tensors, disclosed in headers and the voice inventory. Compatibility identifiers are accepted only as parameter values.

Drop-in Labs · API migration guides · Published Actor portfolio