Video Transcriber: TikTok, Instagram, Facebook, YouTube Shorts
Pricing
from $1.00 / 1,000 seconds of video transcribeds
Video Transcriber: TikTok, Instagram, Facebook, YouTube Shorts
Turn TikTok, Instagram Reels, Facebook, YouTube Shorts and X video links into timestamped transcripts and SRT or VTT subtitles. Video to text in 28 languages with auto detection and speaker labels. Bulk links, no login. Failed and silent videos are never charged. Works via API and MCP.
Pricing
from $1.00 / 1,000 seconds of video transcribeds
Rating
0.0
(0)
Developer
The Mine Works
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors.
Paste public video links from TikTok, Instagram, Facebook, YouTube or X (Twitter) and get back what was said: the full transcript, a timed segment for every sentence, ready to use SRT and VTT subtitles, the detected language and, if you want them, speaker labels. The actor downloads only the audio and runs it through a dedicated speech recognition engine, so it works on Reels and TikToks that have no captions at all. No login, no cookies and no API key of your own.
Why choose this actor?
- A batch of Reels in minutes, not an afternoon. On 26 September a run transcribed 55 Instagram Reels (61.6 minutes of speech) in 266 seconds, 55 of 55 successful, each with timed segments and SRT and VTT subtitles. A single 11 second Reel on the current build takes about 8 to 10 seconds.
- You pay for seconds actually transcribed. Billing is per video plus per second of audio, measured from the audio file and rounded to the nearest second. Links that fail to download, private or deleted videos, unsupported links, videos over your length limit and clips with no speech are never charged.
- Subtitles and speaker labels at no extra cost. SRT and WebVTT text come on every successful row by default, split into two line captions. Turn on speaker labels and every sentence is tagged with who said it.
Part of The Mine Works Social media and video family: Threads Scraper, Reddit Scraper, Threads Search Scraper, Instagram Profile Scraper, Instagram Followers & Following, Reddit Search Scraper.
Try it in one minute
Paste this into the JSON tab of the input page and press Start:
{"urls": ["https://www.instagram.com/reel/DcHyP0GsROe"]}
You get one row with the transcript of an 11 second Reel about ad formats, its 12 timed segments, SRT and VTT subtitles and the detected language, in about 10 seconds. It costs a few cents on any plan.
Give the videos in urls, one public link per line. Accepted links:
- TikTok:
tiktok.com/@user/video/<id>, plus short links fromvm.tiktok.comandvt.tiktok.com. - Instagram:
instagram.com/reel/<code>,/reels/<code>,/p/<code>(video posts) and/tv/<code>. - Facebook:
facebook.com/reel/<id>,facebook.com/<page>/videos/<id>,facebook.com/watch/?v=<id>andfb.watch/<code>. - YouTube:
youtube.com/shorts/<id>,youtube.com/watch?v=<id>andyoutu.be/<id>. - X (Twitter):
x.com/<user>/status/<id>andtwitter.com/<user>/status/<id>.
Repeated links are processed once. Older integrations that send a single link in start_urls still work.
Apify's free plan includes $5 of credit every month, which covers about 80 transcripts of 30 second Reels at this actor's Free plan price ($0.008 per video plus $0.0018 per second, and the $0.005 start fee once per run).
Copy to your AI assistant
themineworks/instagram-tiktok-video-transcript on Apify. Transcribes public TikTok, Instagram, Facebook, YouTube and X (Twitter) videos from their audio and returns one row per link with the transcript, timed sentence segments, SRT and VTT subtitles, detected language, optional speaker labels and the seconds billed. Call ApifyClient("TOKEN").actor("themineworks/instagram-tiktok-video-transcript").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Required: urls (string[] of public video links). Optional: include_timestamps (default true; adds time markers inside transcript), include_subtitles (default true; adds srt and vtt), enable_diarization (default false; speaker labels), language (default "auto"; "multi" for mixed language speech, or a code such as "en", "hi", "es", "pt-BR"), maxVideoMinutes (default 60, max 240; longer videos skipped, not charged), maxConcurrency (default 3, max 8). Rows with status "failed" carry an error and are never charged; the row with _type "info" is a run note, never billed. For more than about 50 short videos, raise the run timeout above the 300 second default. Full spec: GET https://api.apify.com/v2/acts/themineworks~instagram-tiktok-video-transcript/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations?fpr=ymnoit&utm_source=apify-readme&utm_medium=referral
Key features
- Five platforms, one input box. TikTok, Instagram, Facebook, YouTube and X links can be mixed in one run. Each row records its
platformand avideoIdtaken from the link, so you can join results back to your list. - Listens to the audio. Captions are not needed. The actor fetches only the audio track, decodes it to clean 16 kHz mono sound and sends it to a dedicated speech recognition engine, not a chat model, so sentences are written down as spoken rather than summarised or tidied away.
- 28 languages with automatic detection.
autodetects each video's language on its own, so one run can mix English, Hindi and Spanish videos.multifollows speakers who switch language mid sentence (for example Hinglish or Spanglish). Pick an exact code when you know it. Transcripts are always in the spoken language; nothing is translated. - Timed output in three shapes. A
transcriptstring (with optional time markers and speaker labels), asegmentsarray withstart,endandtextfor every sentence, andsrtandvttsubtitle text with captions of at most two lines of about 42 characters. - Parallel and budget safe. Up to 8 videos are processed at once. Before each video the actor checks that your run's maximum charge can still pay for it, and skips it with a clear row if not.
- Blocked links retried for you. Each link is fetched through Apify's datacenter proxy first and retried through residential proxy if the platform blocks it, at no extra charge to you. In our test runs, two of three YouTube Shorts and a 3 minute X video were downloaded on that second try.
How to use it
Basic: a few Reels or TikToks
{"urls": ["https://www.instagram.com/reel/DcHyP0GsROe","https://www.tiktok.com/@duolingo/video/7192679543415639338"]}
Timestamps and subtitles are on by default. Leave language on auto.
Plain text only, for an AI or search pipeline
{"urls": ["https://www.instagram.com/reel/DcHyP0GsROe"],"include_timestamps": false,"include_subtitles": false}
The transcript is then plain sentences, ready to embed or summarise, and the rows are smaller. The segments array still carries the timing.
Interviews and podcasts with speaker labels
{"urls": ["https://x.com/NASA/status/1491475671058681863/video/1"],"enable_diarization": true,"language": "en"}
Each segment gets a speaker number, the transcript and subtitles are labelled Speaker 0, Speaker 1 and so on, and speakers_detected reports how many people spoke. On this 3 minute 25 second NASA video, two speakers were found.
A creator's last 100 videos, for hook research
{"urls": ["<paste up to a few hundred links, one per line>"],"include_subtitles": false,"maxVideoMinutes": 3,"maxConcurrency": 6}
Take the first segment of every transcript to compare opening lines across a niche. maxVideoMinutes: 3 skips anything longer without charging for it. Get the links from a profile scraper first, then pass them here.
Set the run timeout in the run options to match the batch. The default is 300 seconds, and 55 Reels took 176 to 266 seconds at 4 at a time; if a run reaches its timeout, it stops 15 seconds early and keeps every transcript already delivered.
Subtitles for editing in Premiere Pro, CapCut or YouTube Studio
{"urls": ["https://www.youtube.com/shorts/Gc4dtEp2fcw"],"include_subtitles": true,"language": "en"}
Save the srt field as a .srt file or the vtt field as a .vtt file and import it. For mixed Hindi and English speech, set language to multi.
Daily transcripts of a brand's new videos
Save a task with the day's links and schedule it, or let an automation (Make, Zapier, n8n) collect new video links each morning and start a run with them. Each video is billed only once per run it appears in, so send only new links.
Input parameters
| Parameter | Type | Default | What it does |
|---|---|---|---|
urls | array of strings | required (form prefill: one Instagram Reel) | Public video links, one per line. TikTok, Instagram, Facebook, YouTube and X (Twitter). Duplicates are removed. |
include_timestamps | boolean | true | Puts a start and end time marker in square brackets before each sentence inside transcript. The segments array has timings either way. |
include_subtitles | boolean | true | Adds srt and vtt subtitle text to each successful row. No extra cost. |
enable_diarization | boolean | false | Speaker labels: Speaker 0, Speaker 1 and so on in the transcript, a speaker number on each segment, voice tags in the VTT, and speakers_detected. No extra cost. |
language | string | auto | auto detects each video's language; multi handles speech that switches language mid sentence; or pick one of 35 language codes (below). |
maxVideoMinutes | integer (1 to 240) | 60 | Videos longer than this are skipped with a failed row and never charged. |
maxConcurrency | integer (1 to 8) | 3 | How many videos are processed at the same time. Higher finishes big lists faster; lower is gentler on platforms that rate limit. |
start_urls | string | none (hidden) | A single link, kept so older integrations keep working. Use urls instead. |
Language codes: en, en-US, en-GB, en-AU, en-IN, es, es-419, fr, fr-CA, de, nl, pt, pt-BR, it, ja, ko, zh, hi, id, ms, ru, uk, pl, sv, da, fi, no, tr, th, vi, el, cs, sk, ro and hu: 28 languages in all, with regional variants for English, Spanish, French and Portuguese. Setting the right language is the single biggest help on accents, background music and brand names.
The run itself uses 1 GB of memory and a 300 second timeout by default. For more than about 50 short videos, raise the timeout in the run options.
What data do you get?
One row per link, in the order the videos finish. Successful and failed links both get a row, so nothing on your list disappears silently.
Identity: sourceUrl (the link as you gave it), videoId (the platform's ID taken from the link), platform (tiktok, instagram, facebook, youtube or x), status (success or failed).
The words: transcript (the full text; with time markers and speaker labels when those options are on), segments (one object per sentence with start and end in seconds, text, and speaker when speaker labels are on), srt and vtt (subtitle text, when subtitles are on).
About the audio: durationSec (length measured from the downloaded audio itself), detected_language (a code such as en), speakers_detected (a number when speaker labels are on, otherwise null), billed_seconds (the seconds charged for this video), timestamp (when the row was produced).
Failed rows carry sourceUrl, videoId, platform, status: "failed", an error in plain words and timestamp. The reasons you can see: the video is private, deleted, age restricted or needs a login; no downloadable audio at the link; blocked or rate limited by the platform; the video is longer than maxVideoMinutes (the row then also has durationSec); the link is not from a supported site; or the run's charge budget is used up. None of these are charged.
A video with music or silence only comes back as a success row with an empty transcript, an empty segments array and billed_seconds of 0. It is not charged.
The dataset also ends with one _type: "info" row, a note about the run that is never billed. The run's totals (transcribed, failed, skipped for budget, seconds billed) are written to the run's OUTPUT record in its key value store.
Stable fields for automations
These 9 fields were present in every successful row we sampled (490 rows from 26 of our runs between 26 September and 2 October 2026):
| Field | What it holds |
|---|---|
sourceUrl | The link exactly as you gave it |
videoId | The platform's video ID from the link, or the link itself if none is found |
platform | tiktok, instagram, facebook, youtube or x |
status | success or failed |
durationSec | Audio length in seconds, to two decimals |
transcript | The transcript text |
segments | Array of start, end, text (and speaker) per sentence |
billed_seconds | Seconds charged for this video |
timestamp | ISO time the row was produced |
We will not rename these fields. New fields may be added over time; existing ones keep their names.
Output examples
Real rows from our own runs, trimmed for length.
A Reel with time markers and subtitles turned off (live build, 2 October 2026; 12 segments, 3 shown):
{"sourceUrl": "https://www.instagram.com/reel/DcHyP0GsROe","videoId": "DcHyP0GsROe","platform": "instagram","status": "success","durationSec": 11.05,"transcript": "Before and after ad. Two. UGC problem solution ad. Three. Because I think there's gonna be something better. Founder ad. One. Static testimonial ad. Five. Static offer. LED ad. Four.","segments": [{ "start": 0, "end": 1.28, "text": "Before and after ad." },{ "start": 1.28, "end": 1.6, "text": "Two." },{ "start": 3.92, "end": 5.6, "text": "Because I think there's gonna be something better." }],"detected_language": "en","billed_seconds": 11,"timestamp": "2026-10-02T11:15:11.683Z"}
The subtitle fields of the same Reel with the defaults on (26 September, first cues shown). The full row also carries the transcript with time markers and every segment:
{"videoId": "DcHyP0GsROe","srt": "1\n00:00:00,000 --> 00:00:01,280\nBefore and after ad.\n\n2\n00:00:01,280 --> 00:00:01,600\nTwo.\n\n3\n00:00:01,600 --> 00:00:03,520\nUGC problem solution ad.\n","vtt": "WEBVTT\n\n00:00:00.000 --> 00:00:01.280\nBefore and after ad.\n\n00:00:01.280 --> 00:00:01.600\nTwo.\n","detected_language": "en","billed_seconds": 11}
A NASA video on X with speaker labels on (26 September; 8 segments, 3 shown). The first 49 seconds are music, which is why the first sentence starts at 49.86 seconds. Silence and music are left out of the text, but billing is on the full audio length:
{"sourceUrl": "https://x.com/NASA/status/1491475671058681863/video/1","videoId": "1491475671058681863","platform": "x","status": "success","durationSec": 204.89,"segments": [{ "start": 49.86, "end": 53.53, "text": "It's thrilling to be able to see something that's never been seen before.", "speaker": 0 },{ "start": 53.62, "end": 56.18, "text": "This emission that we're seeing is thermal emission.", "speaker": 0 },{ "start": 56.59, "end": 65.39, "text": "Even on the night side, the surface of Venus is so hot that it's it's glowing, faintly at very red wavelengths.", "speaker": 0 }],"vtt": "WEBVTT\n\n00:00:49.860 --> 00:00:53.530\n<v Speaker 0>It's thrilling to be able to see\nsomething that's never been seen before.\n","detected_language": "en","speakers_detected": 2,"billed_seconds": 205,"timestamp": "2026-09-26T11:02:36.164Z"}
The note row at the end of every run (never charged):
{"_type": "info","delivered": 1,"message": "1 transcripts delivered. This row is informational: it is never billed.","scraped_at": "2026-09-26T13:12:01.723Z"}
Pricing
Pay per event: a fee for each video transcribed and delivered, a fee per second of its audio, and a small start fee per run.
| Event | Free | Bronze | Silver | Gold and above |
|---|---|---|---|---|
video-transcribed, per video | $0.008 | $0.005 | $0.005 | $0.005 |
transcription-second, per second of audio | $0.0018 | $0.0015 | $0.0012 | $0.001 |
| Same, per minute of audio | $0.108 | $0.09 | $0.072 | $0.06 |
apify-actor-start, per run | $0.005 per GB of run memory, minimum one event | same | same | same |
The start fee, exactly. Apify's apify-actor-start event is charged once when a run starts, at $0.005 for each GB of memory the run uses, with a minimum of one event. This actor runs on 1 GB by default, so a default run pays $0.005. Put all your links in one run and you pay it once.
How seconds are counted. From the downloaded audio file itself, not from the platform's metadata, rounded to the nearest second with a minimum of one. Music and silence inside a video count as part of its length.
What real jobs cost, start fee included:
| Job | Free | Bronze | Silver | Gold and above |
|---|---|---|---|---|
| One 30 second Reel | $0.067 | $0.055 | $0.046 | $0.040 |
| 100 Reels of 45 seconds, one run | $8.91 | $7.26 | $5.91 | $5.01 |
| 1,000 TikToks of 20 seconds, one run | $44.01 | $35.01 | $29.01 | $25.01 |
| One 10 minute YouTube video | $1.09 | $0.91 | $0.73 | $0.61 |
Our 26 September proof run (55 Reels, 3,698 billed seconds) would cost $7.10 on the Free plan and $3.98 on Gold.
Never charged: links that fail to download; private, deleted, age restricted or login only videos; unsupported links; videos longer than maxVideoMinutes; videos with no speech; videos skipped because the run's budget was used up; a transcription that fails on our side; and the info row. Proxy traffic, including the residential retry, is included in the price.
There is no scheduled price change for this actor. The Pricing tab on this page always shows the rate for your own plan; if it and this table ever differ, the Pricing tab is right.
FAQ
Which sites does it support? TikTok, Instagram (Reels, video posts and IGTV links), Facebook (Reels, page videos, Watch and fb.watch links), YouTube (Shorts, regular videos and youtu.be links) and X (Twitter) videos. In our own runs, Instagram and X videos were transcribed end to end, and audio from TikTok videos, YouTube Shorts and Facebook Reels downloaded through the same path. Links from other sites get a failed row and no charge.
How many videos can I send in one run? There is no fixed cap. The practical limit is the run timeout: 300 seconds by default, and 55 Reels took 176 to 266 seconds at 4 at a time. For bigger lists, raise the timeout in the run options or split the list. A run that reaches its timeout stops 15 seconds early and keeps every transcript delivered so far.
How accurate is it?
The transcripts come from a speech recognition engine built for the job, not from a general chat model, so it writes down what was said instead of deciding what is worth keeping. Clear speech, voiceovers and talking head videos come back close to word for word. Accuracy drops, as with any transcriber, on loud background music, people talking over each other, strong accents with the wrong language picked, and slang or brand names the engine has not heard. Setting the right language helps most.
Do I need a TikTok or Instagram login, cookies or an API key? No. Everything runs on Apify's servers without any account, cookie or key from you, so nothing of yours can be blocked or banned.
Does it translate?
No. The transcript is in the language that is spoken. Use multi for speakers who switch languages mid sentence.
What happens with private, deleted or blocked videos?
They come back as a row with status: "failed" and a plain error, and they are not charged. A link the platform blocks is retried once through residential proxy before it is marked failed. Links that can never work (private, deleted, age restricted or login only) fail at once instead of keeping you waiting.
What about videos with only music? They return a success row with an empty transcript and zero billed seconds, and cost nothing.
Can it stay inside my budget? Yes. Before each video, the actor checks that the run's maximum charge can still pay for it, and skips it with a clear row if not. Nothing you did not budget for is transcribed.
How do I export the results?
From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API. To get subtitle files, save the srt or vtt field of each row as a .srt or .vtt file.
Can I run it on a schedule?
Yes. Save your input as a task, then in Apify Console open Schedules, choose Create new and pick a time or a cron expression such as 0 7 * * *. The task transcribes the same links each time, so for new videos have Make, Zapier or n8n collect fresh links and start the run with them.
How is this different from YouTube caption scrapers? Caption scrapers copy subtitles that already exist, so they only work when the uploader or the platform made some. Most TikToks, Reels and Facebook videos have none. This actor listens to the audio, so it works on any public video with speech and gives you timings and subtitles either way.
Can I use it from Claude, ChatGPT or another AI assistant?
- Connector URL:
https://mcp.apify.com/?tools=themineworks/instagram-tiktok-video-transcript. - Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
- ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
- Cursor or VS Code: add it as an HTTP MCP server with that URL.
- Claude Code:
claude mcp add -t http instagram-tiktok-video-transcript "https://mcp.apify.com/?tools=themineworks/instagram-tiktok-video-transcript".
Is it legal to transcribe public social media videos? The actor only reads videos that anyone can watch without logging in. Transcribing public videos for research, accessibility, analysis and search is common practice, but the words in someone else's video can be protected by copyright, so check the platform's terms and the creator's rights before you republish a transcript. Videos can contain personal data, and you are responsible for handling it under the laws that apply to you, such as GDPR and CCPA. This is general information, not legal advice.
Integrations
- Google Sheets: export a run to a sheet, one row per video with its transcript.
- Make, Zapier and n8n: use the Apify app or node to start a run with new links and send transcripts to Notion, Slack, a CMS or a vector database.
- Webhooks: have Apify call your URL when a run succeeds, then fetch the dataset.
- API and client libraries: start runs and read datasets from Python, JavaScript or any HTTP client. The "Copy to your AI assistant" block above has the exact call. For a handful of links,
run-sync-get-dataset-itemsreturns the transcripts in one HTTP call. - MCP clients: Claude, ChatGPT, Cursor, VS Code and other MCP clients can call the actor through
https://mcp.apify.com.
More from The Mine Works
Social media and video
- Threads Scraper
- Reddit Scraper
- Threads Search Scraper
- Instagram Profile Scraper
- Instagram Followers & Following
- Reddit Search Scraper
- Twitter / X Scraper
- YouTube Transcript
- Xiaohongshu (RED) Scraper
- Telegram Channel Scraper
- Telegram Channel Finder
- Pinterest Profile Scraper
Leads and business directories
Marketing, SEO and reviews
Real estate
Science, health and government data
Jobs and hiring
Company and business data
E-commerce and marketplaces
Food and local services
Developer and AI tools
More tools
Support
A link failed that should have worked, or you need a field we do not return yet? Open an issue on the Issues tab of this page with the link and we will reply there. To ask for a new platform or data source, email dmineworks@gmail.com.
Video Transcriber turns public TikTok, Instagram, Facebook, YouTube and X videos into timed transcripts and SRT and VTT subtitles, billed by the second of audio and never for a link that fails.

