AI Voice Cloning Text to Speech - English avatar

AI Voice Cloning Text to Speech - English

Pricing

Pay per usage

Go to Apify Store
AI Voice Cloning Text to Speech - English

AI Voice Cloning Text to Speech - English

Clone a voice from one short English recording and make it read any script with natural expression, including laughs and sighs. Built-in default voice, long text support, MP3/WAV/Opus output. No rental fee - you pay Apify compute only.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Andrew Babo

Andrew Babo

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

11 days ago

Last modified

Categories

Share

AI Voice Cloning Text to Speech — Expressive English Voices

Clone a voice from one short recording and make it read any English script — with natural expression, including laughs and sighs. This is the quality-first option when the voice has to sound convincingly human.

Good for: branded narration in your own voice, character voices for stories and games, podcast and YouTube voiceovers, ads, audiobook samples, dubbing tests.

Quick start

{
"text": "Oh, that is hilarious! [laugh] Anyway, how are you today?",
"ref_audio_url": "https://example.com/my-voice.wav",
"format": "mp3"
}

No reference clip? Leave ref_audio_url empty and a built-in voice is used.

What you get

  • Voice cloning from a clean 5–15 second English recording
  • Expressive delivery with inline cues such as [laugh], [chuckle], [cough]
  • Built-in default voice when you don't supply a reference
  • Long text supported — split and joined with smooth transitions
  • Speed control, silence trimming, MP3 / WAV / Opus output
  • One dataset row per run with the audio URL and timings

Input

FieldDefaultWhat it does
textrequiredThe script to read
ref_audio_url—Link to the voice you want to clone
speed1.0Playback speed
dtypefp32Use a quantized option to trade a little quality for speed
formatwavwav, mp3 or opus

Pricing

No rental fee — you only pay Apify compute. This Actor prioritises quality, so a long script takes longer than the fast voice Actors; for bulk jobs use the Fast Text to Speech Actor instead.

Notes

  • English only. For Vietnamese use the Vietnamese Text to Speech Actor.
  • Only clone voices you own or have permission to use.

For AI agents (MCP-ready)

This actor is built to be called by AI agents. It works out of the box with the Apify MCP Server — add it to Claude Desktop, Cursor or any MCP client: the agent can clone an English voice on its own.

{
"mcpServers": {
"chatterbox-tts": {
"url": "https://mcp.apify.com/?actors=andrew_babo/chatterbox-tts",
"headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
}
}
}

Agent skill (paste into your agent's instructions)

Use the "chatterbox-tts" tool when ENGLISH speech must sound convincingly
human, or must be read in a specific cloned voice, including expressive
cues like [laugh] or [sigh]. It is the quality-first option and slower than
"kokoro-tts" / "piper-tts" — use those for bulk or draft work.
HOW TO CALL
- { "text": "Oh, that is hilarious! [laugh] Anyway...",
"ref_audio_url": "https://.../my-voice.wav", "format": "mp3" }
- No reference clip: omit "ref_audio_url" and a built-in voice is used.
- "dtype": a quantized option trades a little quality for speed.
OUTPUT CONTRACT
- One row per run: audio_url, duration, processing time.
- English only — route Vietnamese to "vieneu-tts", other languages to
"qwen3-tts" or "kokoro-tts".
- errorCode / noResults means empty input.
- Only clone voices the user owns or has permission to use.