SRT Subtitle Generator โ Subtitles from Any Video
Pricing
from $6.30 / 1,000 minute transcribeds
SRT Subtitle Generator โ Subtitles from Any Video
Give it a YouTube, TikTok, Instagram or Vimeo link, or a video file, and get a finished SRT or WebVTT subtitle file with the timings already aligned. 90+ languages, translation, speaker labels. Export data, run via API, schedule runs, or integrate with AI workflows.
Pricing
from $6.30 / 1,000 minute transcribeds
Rating
5.0
(2)
Developer
Matvey
Maintained by CommunityActor stats
0
Bookmarked
14
Total users
10
Monthly active users
3 days ago
Last modified
Categories
Share
SRT Subtitle Generator takes a video or an audio file and hands back a finished subtitle file โ SRT or WebVTT, timings already aligned, ready to drop next to the video or upload to YouTube. 90+ languages, optional translation to English, optional speaker labels, and no API key to set up.
Tested head-to-head โ 26 September 2026
Same 11-second speech clip through eight transcription Actors on the Apify Store. Four of the others put the whole 108-character sentence into a single 11-second caption โ a wall of text on screen. This Actor cuts captions to the usual reading norm: at most two lines of 42 characters, split at commas and sentence ends, so the same sentence becomes two readable captions. Price for the clip: $0.011 (one started minute plus the file).
What it is for
Subtitling videos for social media and YouTube, making footage accessible, preparing translation-ready caption files, and adding captions in bulk to a library of recordings. Give it a list of URLs and every file comes back with its own subtitle track.
| What you give it | What you get back |
|---|---|
| A video URL | An SRT file with aligned timings |
subtitleFormat: vtt | A WebVTT file instead |
translateToEnglish: true | English subtitles for a foreign-language video |
diarize: true | Speaker labels inside the captions |
| A three-hour file | One continuous subtitle file, timings unbroken |
Formats: MP4, MOV, WEBM, MP3, M4A, WAV, FLAC, OGG โ from a public URL or an upload.
What data does it return?
| Field | Example |
|---|---|
subtitles | the complete SRT or WebVTT file, as text |
source, fileName | https://example.com/clip.mp4 ยท clip.mp4 |
durationSeconds, durationMinutes | 461.05 ยท 7.68 |
language, model, translatedToEnglish | English ยท accurate ยท false |
transcript, wordCount, charCount | the plain text behind the captions |
segments | [{"start": 0, "duration": 4.56, "text": "โฆ", "speaker": "Speaker 1"}] |
segmentCount, speakerCount, speakers | 101 ยท 3 ยท ["Speaker 1", โฆ] |
status, errorCode, errorMessage | ok, or why a file failed |
The subtitle file arrives as a field on the row, so you can save it straight to disk from the API or from an integration.
How much does it cost?
| Event | Price | When it is charged |
|---|---|---|
| Minute transcribed | $0.009 | Per started minute, using the built-in key |
| Minute with your own key | $0.004 | Per started minute when you supply a Groq key |
| Add-on: Speaker labels | $0.006 | Per started minute, only when speaker labels are on |
| File processed | $0.002 | Per file downloaded and prepared |
| Video link resolved | $0.09 | Per YouTube, TikTok, Instagram, X, Vimeo or other video page whose audio has to be pulled through a residential exit node. Direct file links never pay it |
A ten-minute video costs about $0.30. A failed file is an error row and is never charged.
Bulk export
This Actor is built for bulk jobs โ put hundreds of links into one run, or call it from the API on a schedule. There is no fee per run: you pay per started minute of audio plus $0.002 per file, and $0.09 only for a video page (not a direct file link) whose audio has to be fetched through a residential exit node. Files that fail are returned as error rows and are free.
โฌ๏ธ Input
{"urls": ["https://example.com/episode-12.mp3"],"quality": "accurate","includeSegments": true}
Speaker labels for a two-person interview:
{ "urls": ["https://example.com/interview.m4a"], "diarize": true, "speakerCount": 2 }
Every input field
| Field | What it does |
|---|---|
urls | Public links to audio or video files. Several at a time is fine. |
file | An uploaded file instead of a link. |
quality | fast for speed, accurate for difficult audio and non-English speech. |
language | Name the spoken language instead of letting the model guess โ faster and safer on short clips. |
translateToEnglish | Write the text in English whatever was spoken. |
diarize, speakerCount | Split the transcript by speaker. Give the count when you know it. |
vocabularyHint | Names, brands and jargon, so they come out spelled your way. |
includeSegments | Time-stamped passages alongside the full text. |
subtitleFormat | srt or vtt to get a ready subtitle file. |
apiKey | Your own Groq key โ cuts the per-minute price by more than half. |
maxConcurrency | Files processed in parallel. |
maxFileSizeMb, timeoutPerFileSecs | Guard rails for very large or very slow downloads. |
โฌ๏ธ Output
{"source": "https://example.com/episode-12.mp3","fileName": "episode-12.mp3","durationMinutes": 7.68,"language": "English","transcript": "Welcome back to the showโฆ","wordCount": 1665,"segments": [{ "start": 0, "duration": 4.56, "text": "Welcome back to the show" }],"status": "ok"}
Use cases
Subtitles for video
A finished .srt you can drop next to the file, or a .vtt for a web player. Captions follow the usual reading norm: at most two lines of 42 characters each, split at commas and sentence ends.
Accessibility
Captions that make a video usable without sound, which is how most of it is watched on a phone anyway.
Translated captions
Turn on translation and the cues come out in English whatever language was spoken.
Interviews and panels
With speaker labels on, each cue carries who is talking โ the difference between readable and confusing captions.
Batch work
A back catalogue of videos subtitled in one run, one file per item.
๐ค For AI agents and LLM apps
Compact reference for agents calling this Actor through the Apify MCP server or the Apify API (lergassy/srt-subtitle-generator).
Purpose: produce a ready SRT or WebVTT subtitle file from an audio or video URL.
Minimal input:
{ "urls": ["https://example.com/audio.mp3"] }
Behaviours an agent should know:
- Billing is per started minute of audio, not per row. A 40-minute file costs the same whether you keep the segments or not.
durationMinuteson the row tells you what the run actually cost.- Speaker labels are an extra per-minute charge. Turn
diarizeon only when who-said-what matters. - Passing
speakerCountwhen you know it makes the split noticeably cleaner. languageis detected automatically, but naming it removes the guess on short or noisy clips.- A file that cannot be downloaded or decoded returns a
status: "error"row with the reason, free of charge โ checkstatusbefore readingtranscript. - Video files work: the audio track is extracted and the picture discarded.
Use via API
Run this Actor from your own code or pipeline. Get your token in Apify Console โ Settings โ API & Integrations. Every input field below has the same name as in the JSON tab of the input form.
Python
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("lergassy/srt-subtitle-generator").call(run_input={"urls": ["https://www.tiktok.com/@tiktok/video/7670293981833612574","https://raw.githubusercontent.com/lergassy/apify-actor-assets/main/audio-samples/sample-1min.m4a"],"quality": "fast","translateToEnglish": False,"diarize": False,"includeSegments": True,"subtitleFormat": "srt","maxConcurrency": 3,"maxFileSizeMb": 500,"timeoutPerFileSecs": 600})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });const run = await client.actor('lergassy/srt-subtitle-generator').call({"urls": ["https://www.tiktok.com/@tiktok/video/7670293981833612574","https://raw.githubusercontent.com/lergassy/apify-actor-assets/main/audio-samples/sample-1min.m4a"],"quality": "fast","translateToEnglish": false,"diarize": false,"includeSegments": true,"subtitleFormat": "srt","maxConcurrency": 3,"maxFileSizeMb": 500,"timeoutPerFileSecs": 600});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
cURL (runs the Actor and returns the results in one call)
curl -X POST "https://api.apify.com/v2/acts/lergassy~srt-subtitle-generator/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \-H "Content-Type: application/json" \-d '{"urls": ["https://www.tiktok.com/@tiktok/video/7670293981833612574", "https://raw.githubusercontent.com/lergassy/apify-actor-assets/main/audio-samples/sample-1min.m4a"], "quality": "fast", "translateToEnglish": false, "diarize": false, "includeSegments": true, "subtitleFormat": "srt", "maxConcurrency": 3, "maxFileSizeMb": 500, "timeoutPerFileSecs": 600}'
Limitations, stated plainly
- Audio quality decides accuracy. Overlapping speech, heavy accents and phone-line compression cost accuracy no model recovers fully.
- Speaker labels are labels, not identities. You get "Speaker 1" and "Speaker 2", never names โ mapping them to people is your step.
- Timings are passage-level, not word-level.
- Very large files need time. Raise
timeoutPerFileSecsfor multi-hour recordings, and expect a long run. - Links must be public. A file behind a login or an expiring signed URL cannot be fetched.
Integrations
Runs from the Apify API and the Python and JavaScript clients, and from n8n, Make or Zapier through the Apify app. Point a webhook at the dataset to push transcripts into Notion, Google Docs or your own database, or let an agent call it through the Apify MCP server.
โ FAQ
Do I need an API key or an account anywhere?
No. Recognition runs with a built-in key. Supply your own Groq key only if you want the cheaper per-minute rate.
Which languages are supported?
Around 90, including Russian, Indonesian, Spanish, German, French, Japanese and Chinese. Use accurate for non-English audio.
Can it handle a three-hour file?
Yes. Long files are processed in parts and the timings stay continuous across them.
Does it work with video?
Yes. Give it an MP4, MOV or WEBM link and the audio track is pulled out for you.
What happens if a file is unreachable?
You get an error row with the reason and no charge for that file. The other files in the run are still transcribed.
How accurate is it?
It runs Whisper large v3. On clean speech it is close to a careful human first pass; on overlapping or noisy audio it is not.
Can I get speaker names?
No โ you get "Speaker 1", "Speaker 2" and so on. Naming them is your step.
How do I make it spell product names correctly?
Put them in vocabularyHint. The hint is given to the model before it starts.
You might also like
| Actor | What it does |
|---|---|
| Speech to Text | The same engine, framed as a speech-to-text API |
| SRT Subtitle Generator | Ready subtitle files from audio or video |
| YouTube Transcript Scraper | Captions straight from YouTube, no transcription needed |
| Document Text Extractor | The same idea for PDFs and Word files |
Also known as
People look for this Actor as an SRT generator, subtitle generator, automatic captions, video to SRT converter, WebVTT generator and an AI subtitle maker.
Notes
Recognition is Whisper large v3 with a built-in key โ nothing to sign up for. Timings come from the model's own segment boundaries, so captions break where the speaker pauses rather than at a fixed character count.