Auto Caption Burner - Viral Subtitles Burned In
Pricing
Pay per event
Auto Caption Burner - Viral Subtitles Burned In
Captions are the #1 retention hack, and this burns them straight into your MP4. Send a video URL, get the finished video back with word-by-word animated subtitles, plus the full transcript and word timings. 5 viral presets. $0.04 per 30-second caption block, $0.08 per minute of video.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Dami's Studio
Maintained by CommunityActor stats
1
Bookmarked
1
Total users
0
Monthly active users
4 days ago
Last modified
Categories
Share
Auto Caption & Subtitle Burner
Give it a public video URL and it returns the same video with animated word-by-word captions burned in, plus the matching SRT file and word-level timestamps. It is built for short-form vertical content (TikTok, Reels, Shorts) but works on anything with speech: podcasts, course lessons, ads, interviews. No login, no watermark on the output.
How it works
The actor downloads your video, pulls the audio with ffmpeg, runs it through Whisper for word-level timestamps, builds an ASS subtitle track in the preset you picked, and re-encodes the video with the captions burned in using x264. The active word is highlighted as it is spoken.
Input
| Field | Required | Notes |
|---|---|---|
videoUrl | yes | Public direct URL to the source video (.mp4 / .mov / .webm). You host it (S3, a CDN, a Drive direct link). Nothing is scraped. |
preset | no | Visual style of the captions. One of hormozi, beast, tiktok, clean, karaoke. Defaults to hormozi. |
language | no | ISO-639-1 code of the spoken language (en, es, fr, ...). Leave as auto to detect it. |
allCaps | no | Force captions to uppercase, overriding the preset. |
fontSize | no | Override the preset font size, in px at 1080 width (24-200). |
primaryColor | no | Hex color of inactive words, e.g. #FFFFFF. |
highlightColor | no | Hex color of the currently-spoken word, e.g. #FFE000. |
position | no | Where captions sit: top, center, or bottom. |
wordsPerGroup | no | How many words are on screen at once (1-8). |
crf | no | x264 quality. Lower is better and bigger. 18 is visually lossless, 23 is smaller. Defaults to 18. |
openaiApiKey | no | Your OpenAI key for Whisper. Required unless the actor owner has set a shared key. Stored as a secret. |
Every preset sets a sane font, size, outline, and highlight color. The style override fields only change the parts you specify and leave the rest of the preset alone.
Output
Each run pushes one dataset record with the source dimensions, duration, word count, and the full transcript text. The rendered files land in the run's key-value store: the captioned MP4, an SRT file, and a JSON file of per-word timestamps (word, start, end). The dataset record's output object holds the keys and public URLs for the MP4 and SRT.
Example
{"videoUrl": "https://example.com/my-short.mp4","preset": "hormozi","language": "auto","allCaps": true,"highlightColor": "#FFE000"}
Pricing
$0.04 per 30-second caption block, which works out at $0.08 per minute of video, plus $0.0000125 each time a run starts. Flat rate — no volume tiers, no plan gates — and you are charged only for blocks actually burned in.
Blocks are counted from the finished video, rounded up, so a 45-second short is two blocks ($0.08) and a 3-minute video is six ($0.24). There is no separate render fee, no per-preset fee and no charge for the SRT or the word-timing JSON — they come with the same block charge. A run that fails before anything is rendered charges nothing but the run-start fee.
Whisper transcription runs on your own OpenAI key, billed to you by OpenAI and not marked up here.
Notes
Transcription runs on Whisper, so you need an OpenAI key in openaiApiKey unless the owner has configured a shared one. Word timing is only as good as the audio: clean speech maps cleanly, heavy background noise or music can drift the highlights. The source has to be reachable at a direct URL, since the actor downloads the file rather than scraping a page.