AI Voice Cloning Text to Speech - English
Pricing
Pay per usage
AI Voice Cloning Text to Speech - English
Clone a voice from one short English recording and make it read any script with natural expression, including laughs and sighs. Built-in default voice, long text support, MP3/WAV/Opus output. No rental fee - you pay Apify compute only.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Andrew Babo
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
11 days ago
Last modified
Categories
Share
AI Voice Cloning Text to Speech — Expressive English Voices
Clone a voice from one short recording and make it read any English script — with natural expression, including laughs and sighs. This is the quality-first option when the voice has to sound convincingly human.
Good for: branded narration in your own voice, character voices for stories and games, podcast and YouTube voiceovers, ads, audiobook samples, dubbing tests.
Quick start
{"text": "Oh, that is hilarious! [laugh] Anyway, how are you today?","ref_audio_url": "https://example.com/my-voice.wav","format": "mp3"}
No reference clip? Leave ref_audio_url empty and a built-in voice is used.
What you get
- Voice cloning from a clean 5–15 second English recording
- Expressive delivery with inline cues such as
[laugh],[chuckle],[cough] - Built-in default voice when you don't supply a reference
- Long text supported — split and joined with smooth transitions
- Speed control, silence trimming, MP3 / WAV / Opus output
- One dataset row per run with the audio URL and timings
Input
| Field | Default | What it does |
|---|---|---|
text | required | The script to read |
ref_audio_url | — | Link to the voice you want to clone |
speed | 1.0 | Playback speed |
dtype | fp32 | Use a quantized option to trade a little quality for speed |
format | wav | wav, mp3 or opus |
Pricing
No rental fee — you only pay Apify compute. This Actor prioritises quality, so a long script takes longer than the fast voice Actors; for bulk jobs use the Fast Text to Speech Actor instead.
Notes
- English only. For Vietnamese use the Vietnamese Text to Speech Actor.
- Only clone voices you own or have permission to use.
For AI agents (MCP-ready)
This actor is built to be called by AI agents. It works out of the box with the Apify MCP Server — add it to Claude Desktop, Cursor or any MCP client: the agent can clone an English voice on its own.
{"mcpServers": {"chatterbox-tts": {"url": "https://mcp.apify.com/?actors=andrew_babo/chatterbox-tts","headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }}}}
Agent skill (paste into your agent's instructions)
Use the "chatterbox-tts" tool when ENGLISH speech must sound convincinglyhuman, or must be read in a specific cloned voice, including expressivecues like [laugh] or [sigh]. It is the quality-first option and slower than"kokoro-tts" / "piper-tts" — use those for bulk or draft work.HOW TO CALL- { "text": "Oh, that is hilarious! [laugh] Anyway...","ref_audio_url": "https://.../my-voice.wav", "format": "mp3" }- No reference clip: omit "ref_audio_url" and a built-in voice is used.- "dtype": a quantized option trades a little quality for speed.OUTPUT CONTRACT- One row per run: audio_url, duration, processing time.- English only — route Vietnamese to "vieneu-tts", other languages to"qwen3-tts" or "kokoro-tts".- errorCode / noResults means empty input.- Only clone voices the user owns or has permission to use.