Text To Speech (NO API)
Pricing
from $0.04 / 1,000 text converted (per character)s
Text To Speech (NO API)
Convert text to natural speech in 43 curated voices across 8 languages. Clone any voice from a 15-second sample. Long text splits automatically. No API key, no subscription โ pay only for what you convert.
Pricing
from $0.04 / 1,000 text converted (per character)s
Rating
0.0
(0)
Developer
Dead
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Text to Speech
Turn text into natural-sounding audio. Paste your text, pick a voice, run it โ you get MP3 files back.
No API key to set up, no subscription, no monthly credits to manage. You pay only for what you convert.
๐ Listen to samples of all 43 voices
What it does
- Converts text to speech in 43 curated voices across 8 languages
- Clones a voice from a short audio sample you provide
- Handles long text automatically โ paste a whole article, no length juggling
- Returns MP3, WAV or Opus files with direct download links
Quick start
- Paste your text into Text to convert โ one line per audio file
- Pick a Voice from the dropdown
- Click Start
That's it. Your audio files appear in the run's output, each with a link.
Voices
43 voices, hand-picked and checked. No celebrity impressions, no cartoon characters โ just clean, general-purpose voices you can use without worrying about whose likeness you're borrowing.
| Language | Voices |
|---|---|
| English | 20 (6 female, 13 male, 1 neutral) |
| Spanish | 6 (2 female, 4 male) |
| Russian | 4 (3 female, 1 male) |
| Arabic | 3 (3 female) |
| French | 3 (3 female) |
| Japanese | 3 (3 female) |
| Portuguese | 3 (1 female, 2 male) |
| Chinese | 1 (1 male) |
Each is labelled by language, gender and style โ for example "English, female โ Laura, confident narrator" โ so you can pick without guessing.
Hear them first: the samples page linked above has a play button for every voice saying the same sentence. Worth two minutes before your first real run.
Expression tags
Drop these anywhere in your text to shape the delivery:
[excited] I can't believe it! [pause] [whisper] Come closer.
Available: [whisper] [excited] [pause] [emphasis] [laughing] [sigh] [angry] [sad] [shouting]
Voice cloning
Want it to sound like a specific person? Two fields under Voice cloning:
Clone a voice from an audio sample โ a direct link to a 10โ30 second recording. It has to be a link that downloads the file, not a page with a player on it.
- โ
Works: raw GitHub links, S3 URLs, Dropbox links ending in
?dl=1 - โ Doesn't work: Google Drive share pages, SoundCloud pages, YouTube links
What is said in that sample โ type out the exact words spoken in the recording. Optional, but it noticeably improves how closely the clone matches, because the words can be lined up against the audio.
When you fill these in, the sample overrides whichever voice you picked from the dropdown.
Only clone voices you have permission to use.
How long text is handled
You don't have to think about this. Paste an entire article and it just works.
Behind the scenes there's a 10,000-byte limit per audio file. If a line goes over, it's split at a sentence boundary โ never mid-word โ and you get two files instead of one. Nothing is lost or cut off.
For reference, 10,000 bytes is roughly 1,500 English words. Longer than that and you'll simply get more than one file back.
Pricing
Two charges:
| Price | |
|---|---|
| Each audio file created | $0.015 |
| Each byte of text | $0.00004 |
That works out to $0.04 per 1,000 characters of English. Billing is per byte with no rounding, so a short line is never charged like a long one. Failed lines aren't charged at all.
What "bytes" means
For English, bytes and characters are the same thing. If your word processor says 4,000 characters, that's 4,000 bytes. Nothing to convert.
It only differs for other scripts, because those characters take more space:
| Script | Bytes per character |
|---|---|
| English, Spanish, Portuguese, French | 1 |
| Russian, Arabic | 2 |
| Chinese, Japanese, Korean | 3 |
So 1,000 bytes is about 1,000 English characters, or about 330 Chinese characters.
Real examples
| What you're converting | Cost |
|---|---|
| A tweet (150 characters) | $0.02 |
| A product description (400 characters) | $0.03 |
| A short blog post (1,000 characters) | $0.055 |
| A video script (4,500 characters) | $0.20 |
| A long article (13,000 characters, 2 files) | $0.56 |
A tip worth knowing
You're charged $0.015 for each audio file, so ten short sentences as ten separate lines cost $0.15 in file fees, while the same text on one line costs $0.015.
Combining short pieces into one line also produces better audio, since a single continuous generation keeps the voice more consistent than ten separate ones.
How it compares
Prices for the same text through this Actor and through ElevenLabs' published API rates:
| What you're converting | This Actor | ElevenLabs Flash | ElevenLabs Multilingual v2 |
|---|---|---|---|
| Tweet (150 characters) | $0.02 | $0.01 | $0.02 |
| Short description (400 characters) | $0.03 | $0.02 | $0.04 |
| Blog post (1,000 characters) | $0.055 | $0.05 | $0.10 |
| Video script (4,500 characters) | $0.20 | $0.23 | $0.45 |
| Long article (13,000 characters) | $0.56 | $0.66 | $1.33 |
Per hour of generated audio:
| Cost per hour | |
|---|---|
| This Actor | $2.55 |
| ElevenLabs Flash | $3.00 |
| ElevenLabs Multilingual v2 | $6.00 |
The short version
- 15% cheaper than ElevenLabs Flash
- 57% cheaper than ElevenLabs Multilingual v2
Those margins hold for anything paragraph-length or longer.
Where it doesn't win: very short text. Each audio file carries a $0.015 fee, which is nothing against a 13,000-character article but triples the price of a tweet. The break-even point against ElevenLabs Flash is around 1,500 characters โ below that, a per-character service works out cheaper.
So if you're converting articles, scripts, chapters or product copy, this is the cheaper option. If you're generating thousands of one-line notifications, it isn't.
Beyond price
- No subscription โ pay per job, not per month
- No credits to buy in advance or watch expire
- No API key to create or manage
- Runs in your Apify workflows and schedules alongside everything else
- Voice cloning included, not gated behind a higher tier
ElevenLabs figures are their published API rates as of early 2026 โ check their current pricing before relying on these. Note also that their API rates sit on top of a monthly subscription, so the real cost to an existing subscriber differs from the table above.
Output
Each line produces one row:
{"index": 0,"text": "The morning began quietly as the first rays of sunlight...","audioUrl": "https://api.apify.com/v2/key-value-stores/.../audio-0000.mp3","format": "mp3","bytes": 616906,"status": "ok","error": null}
- audioUrl โ direct link to the MP3
- bytes โ size of the audio file
- status โ
okorfailed - error โ what went wrong, if anything
Files are also in the run's Storage tab, named audio-0000.mp3, audio-0001.mp3 and so on.
Limits
| Lines per run | 200 |
| Bytes per audio file | 10,000 (longer text is split automatically) |
| Timeout per file | 300 seconds, adjustable up to 1,800 |
Tips
Paste whole paragraphs, not single sentences. Cheaper, and the voice stays consistent across the whole piece.
Try a short line first. One sentence in your chosen voice costs about $0.02 and tells you whether it's the right one before you run a hundred pages.
Use [pause] if it feels rushed. Placing it between sentences gives the delivery room to breathe.
Check the seams on split files. Very long text becomes multiple files, and the voice can shift slightly between them.
Common problems
"The voice sample could not be downloaded" Your link opens a page rather than downloading a file. Use a raw or direct-download URL.
"The service quota has been used up" Temporary capacity limit. Wait a few minutes and try again.
"This text could not be processed" Usually unusual characters or formatting. Try simplifying the text.
A file stops earlier than expected Check the log โ long lines get split into multiple files, so the rest is in the next one.