Video Transcriber: TikTok, Instagram, Facebook, YouTube Shorts avatar

Video Transcriber: TikTok, Instagram, Facebook, YouTube Shorts

Pricing

from $1.00 / 1,000 seconds of video transcribeds

Go to Apify Store
Video Transcriber: TikTok, Instagram, Facebook, YouTube Shorts

Video Transcriber: TikTok, Instagram, Facebook, YouTube Shorts

Turn TikTok, Instagram Reels, Facebook, YouTube Shorts and X video links into timestamped transcripts and SRT or VTT subtitles. Video to text in 28 languages with auto detection and speaker labels. Bulk links, no login. Failed and silent videos are never charged. Works via API and MCP.

Pricing

from $1.00 / 1,000 seconds of video transcribeds

Rating

0.0

(0)

Developer

The Mine Works

The Mine Works

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

4 days ago

Last modified

Share

Video Transcriber: TikTok, Instagram Reels, Facebook, YouTube Shorts and X

Video Transcriber turns public TikTok, Instagram Reels, Facebook, YouTube Shorts and X (Twitter) video links into accurate, timestamped transcripts and ready to use SRT and VTT subtitles. Paste one link or a few thousand and get clean JSON with sentence level segments, speaker labels and the detected language. No login, no cookies and no API keys of your own.

It listens to the audio itself, so it works on Reels and TikToks that have no captions at all, which is most of them.

At a glance

  • Platforms: TikTok, Instagram (Reels, video posts, IGTV), Facebook (Reels, Watch, fb.watch), YouTube (Shorts and regular videos) and X / Twitter videos
  • Output: transcript text, timestamped segments, SRT subtitles, WebVTT subtitles, detected language, speaker count, exact billed seconds
  • Languages: 28 languages with automatic detection, plus a multilingual mode for videos that switch language mid sentence (Hinglish, Spanglish, Taglish)
  • Bulk: any number of links per run, processed in parallel
  • Fair billing: failed links, private videos and clips with no speech are never charged
  • Works on Apify's free plan. The free monthly credit covers about 80 transcripts of 30 second Reels.

What does Video Transcriber do?

You give it public video URLs. For each one it downloads only the audio track, runs it through a dedicated speech recognition model and returns:

  1. The full transcript, with optional [start - end] timestamps and [Speaker N] labels inline
  2. A segments array with one entry per sentence: start time, end time, text and speaker
  3. SRT and WebVTT subtitle files as text fields, split into two line captions that fit on screen
  4. The detected language, the number of speakers, the video length and the exact seconds you were billed for

Every link gets its own output row, including the ones that fail, so you can always join results back to your list.

PlatformLinks that work
TikToktiktok.com/@user/video/..., short links vm.tiktok.com/... and vt.tiktok.com/...
Instagraminstagram.com/reel/..., /reels/..., /p/... (video posts), /tv/...
Facebookfacebook.com/reel/..., facebook.com/.../videos/..., facebook.com/watch/?v=..., fb.watch/...
YouTubeyoutube.com/shorts/..., youtube.com/watch?v=..., youtu.be/...
X (Twitter)x.com/user/status/..., twitter.com/user/status/...

Only public videos can be transcribed. Anything else is returned as a failed row with a clear reason and is not charged.

How do I transcribe TikTok videos or Instagram Reels in bulk?

  1. Click Try for free and paste your links into Video URLs, one per line.
  2. Leave Language on auto, or pick the language if you know it.
  3. Turn on Speaker labels for interviews, podcasts and street interviews.
  4. Click Start. Results appear in the Output tab as they finish, and you can download them as JSON, CSV, Excel or HTML.

Input example:

{
"urls": [
"https://www.instagram.com/reel/DcHyP0GsROe",
"https://www.tiktok.com/@user/video/7412345678901234567",
"https://www.youtube.com/shorts/abcDEF12345",
"https://x.com/NASA/status/1491475671058681863/video/1"
],
"include_timestamps": true,
"include_subtitles": true,
"enable_diarization": false,
"language": "auto",
"maxVideoMinutes": 60,
"maxConcurrency": 3
}

Only urls is required. Everything else has a sensible default.

What does the output look like?

A real record from a NASA video on X, shortened to its first four sentences so it fits here. A full record carries every sentence in transcript, segments, srt and vtt.

{
"sourceUrl": "https://x.com/NASA/status/1491475671058681863/video/1",
"videoId": "1491475671058681863",
"platform": "x",
"status": "success",
"durationSec": 204.89,
"transcript": "[49.86s - 53.53s] [Speaker 0] It's thrilling to be able to see something that's never been seen before. [53.62s - 56.18s] [Speaker 0] This emission that we're seeing is thermal emission. [56.59s - 65.39s] [Speaker 0] Even on the night side, the surface of Venus is so hot that it's it's glowing, faintly at very red wavelengths. [72.28s - 83.81s] [Speaker 1] These whisper images, I think, are really exciting because they provide a new window into the lower atmosphere and surface region of Venus where these extreme conditions exist.",
"segments": [
{ "start": 49.86, "end": 53.53, "text": "It's thrilling to be able to see something that's never been seen before.", "speaker": 0 },
{ "start": 53.62, "end": 56.18, "text": "This emission that we're seeing is thermal emission.", "speaker": 0 },
{ "start": 56.59, "end": 65.39, "text": "Even on the night side, the surface of Venus is so hot that it's it's glowing, faintly at very red wavelengths.", "speaker": 0 },
{ "start": 72.28, "end": 83.81, "text": "These whisper images, I think, are really exciting because they provide a new window into the lower atmosphere and surface region of Venus where these extreme conditions exist.", "speaker": 1 }
],
"srt": "1\n00:00:49,860 --> 00:00:53,530\n[Speaker 0] It's thrilling to be able to see\nsomething that's never been seen before.\n\n2\n00:00:53,620 --> 00:00:56,180\n[Speaker 0] This emission that we're\nseeing is thermal emission.\n\n3\n00:00:56,590 --> 00:01:02,830\n[Speaker 0] Even on the night side, the surface of\nVenus is so hot that it's it's glowing,\n\n4\n00:01:02,830 --> 00:01:05,390\n[Speaker 0] faintly at very red wavelengths.\n",
"vtt": "WEBVTT\n\n00:00:49.860 --> 00:00:53.530\n<v Speaker 0>It's thrilling to be able to see\nsomething that's never been seen before.\n\n00:00:53.620 --> 00:00:56.180\n<v Speaker 0>This emission that we're\nseeing is thermal emission.\n\n00:00:56.590 --> 00:01:02.830\n<v Speaker 0>Even on the night side, the surface of\nVenus is so hot that it's it's glowing,\n\n00:01:02.830 --> 00:01:05.390\n<v Speaker 0>faintly at very red wavelengths.\n",
"detected_language": "en",
"speakers_detected": 2,
"billed_seconds": 205,
"timestamp": "2026-09-26T11:02:36.713Z"
}

The first 49 seconds of that video are music, which is why the first sentence starts at 49.86s. Silence and music are skipped in the text, but you are billed on the full audio length, rounded to the nearest second.

A link that cannot be transcribed looks like this, and costs nothing:

{
"sourceUrl": "https://www.instagram.com/reel/THISDOESNOTEXIST123/",
"videoId": "THISDOESNOTEXIST123",
"platform": "instagram",
"status": "failed",
"error": "Could not download this video. It may be private, deleted, or temporarily blocked by the platform",
"timestamp": "2026-09-26T09:33:46.303Z"
}

Output fields

FieldWhat it holds
sourceUrlThe link you gave, exactly as given
videoIdThe platform's id for the video, taken from the link
platformtiktok, instagram, facebook, youtube or x
statussuccess or failed
durationSecAudio length in seconds, measured from the audio file itself
transcriptFull text, with [start - end] timestamps and [Speaker N] labels when those options are on
segmentsOne object per sentence: start, end, text, and speaker when speaker labels are on
srtSubRip subtitles, ready to save as a .srt file
vttWebVTT subtitles, ready to save as a .vtt file, with speakers as voice tags
detected_languageLanguage code, for example en, hi, es
speakers_detectedNumber of distinct speakers when speaker labels are on, otherwise null
no_speech_detectedtrue when the video has music or silence only (not charged)
billed_secondsThe seconds you were charged for this video
errorWhy a failed row failed
timestampWhen the row was produced

How much does it cost to transcribe a video?

You pay for three things, all listed on the Pricing tab:

ChargeFree planStarter (Bronze)Scale (Silver)Business (Gold) and up
Per second of video transcribed$0.0018$0.0015$0.0012$0.0010
Per video transcribed$0.008$0.005$0.005$0.005
Per run started$0.005$0.005$0.005$0.005

Compute, proxies and storage are included. There is nothing else on the bill.

What real jobs cost:

JobFree planStarterScaleBusiness
One 30 second Reel$0.067$0.055$0.046$0.040
100 Reels of 45 seconds, one run$8.91$7.26$5.91$5.01
1,000 TikToks of 20 seconds, one run$44.01$35.01$29.01$25.01
One 10 minute YouTube video$1.09$0.91$0.73$0.61

Charged: each video that was transcribed and delivered, its length in seconds (rounded to the nearest second, minimum one), and one start charge per run at the default 1 GB of memory.

Never charged: links that fail to download, private, deleted or login only videos, unsupported links, videos longer than your maxVideoMinutes limit, and videos with no speech in them.

Batch your links into one run: the start charge is paid once per run, not once per video.

Will it go over my budget?

No. Before transcribing each video, the actor checks that your run's maximum cost can still cover it. A video that would push the run over your limit is skipped with a clear message instead of being charged, and the run stops cleanly when the budget is used up.

Can I get SRT or VTT subtitle files?

Yes, on by default. Every successful row carries an srt and a vtt field. Save either one as a file and it drops straight into Premiere Pro, DaVinci Resolve, CapCut, YouTube Studio or any video player. Long sentences are split into captions of at most two lines of about 42 characters, the length broadcast captioning guidelines recommend, and the timing is shared across the split. Turn Include SRT and VTT subtitles off if you only need the text.

Does it detect the language automatically?

Yes. Leave Language on auto and each video's language is detected on its own, so one run can mix English, Hindi and Spanish videos. Pick multi for videos where speakers switch language mid sentence. Picking the exact language improves accuracy when you know it.

Supported languages: English (US, UK, Australia, India), Spanish (Spain and Latin America), French (France and Canada), German, Dutch, Portuguese (Portugal and Brazil), Italian, Japanese, Korean, Chinese, Hindi, Indonesian, Malay, Russian, Ukrainian, Polish, Swedish, Danish, Finnish, Norwegian, Turkish, Thai, Vietnamese, Greek, Czech, Slovak, Romanian and Hungarian.

The transcript is always in the language that is spoken. This actor does not translate.

Can it tell different speakers apart?

Yes. Turn on Speaker labels and every sentence is tagged [Speaker 0], [Speaker 1] and so on, speakers_detected reports how many people spoke, and the VTT file carries speakers as voice tags. It costs nothing extra.

How accurate is the transcription?

Transcripts come from a dedicated speech recognition model, not from a general AI chatbot. In our own test on 55 Instagram Reels (26 Sep 2026), every passage of speech was captured with nothing dropped or repeated, and about 97% of words matched another commercial transcription service word for word. The remaining differences were at word level, for example "wanna" written as "want to", or a brand name spelled differently. That matters: language models asked to transcribe tend to tidy up, summarise or quietly drop sentences they judge unimportant, and you cannot see the gap. A speech model writes down what was said.

Clear speech, voiceovers and talking head videos come back close to word perfect. Accuracy drops, as it does for any transcriber, on heavy background music, several people talking over each other, strong accents in a language picked wrongly, and slang or brand names the model has not heard. Setting the correct language helps most.

Do I need a TikTok or Instagram login, cookies or an API key?

No. Everything runs from Apify's servers with no account, cookie or key from you, so there is nothing of yours that can be blocked or banned.

What happens when a video fails?

Each link is first fetched through fast datacenter servers. If the platform rate limits or blocks that request, it is retried automatically through residential servers at no extra cost to you. Links that can never work (private, deleted, age restricted, login only) fail immediately with a clear error so you are not kept waiting. None of these are charged.

How is this different from YouTube caption scrapers?

Caption scrapers copy subtitles the uploader or YouTube already made, so they only work when captions exist. Most TikToks, Reels and Facebook videos have none. Video Transcriber listens to the audio, so it works on any public video with speech, on any of the five platforms, and gives you timestamps and subtitles either way.

How do I use it from the API, Python or JavaScript?

Python:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("themineworks/instagram-tiktok-video-transcript").call(run_input={
"urls": ["https://www.instagram.com/reel/DcHyP0GsROe"],
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row["status"], row.get("transcript"))

JavaScript:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('themineworks/instagram-tiktok-video-transcript').call({
urls: ['https://www.tiktok.com/@user/video/7412345678901234567'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => r.transcript));

One HTTP call that waits and returns the transcripts, good for a handful of links:

curl -X POST "https://api.apify.com/v2/acts/themineworks~instagram-tiktok-video-transcript/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://www.youtube.com/shorts/abcDEF12345"]}'

It also works from n8n, Make, Zapier, Google Sheets, Airbyte and webhooks through Apify's integrations, and on a schedule if you want new videos transcribed every day.

Can AI agents use it through MCP?

Yes. Add Apify's MCP server to Claude, ChatGPT, Cursor or any MCP client with this actor enabled:

https://mcp.apify.com/?tools=themineworks/instagram-tiktok-video-transcript

Your agent can then transcribe any video link it finds and reason over the text, for example "summarise what these five creators said about our product this week".

What is it used for?

  • Content repurposing: turn Reels and TikToks into blog posts, captions, threads and newsletters
  • Subtitles and accessibility: burn in captions or upload SRT files so videos work on mute and for deaf and hard of hearing viewers
  • Competitor and trend research: read what competitors and creators say across hundreds of videos in minutes, and search it
  • Hook analysis: pull the first sentence of every top performing video in a niche and compare what works
  • Brand and influencer monitoring: check what creators actually say about a product before you pay them
  • AI and RAG pipelines: feed clean, timestamped text from short video into search indexes, embeddings and LLM analysis
  • Research and journalism: keep a searchable, time coded record of public statements made on video

Limitations

  • Public videos only. Private accounts, close friends stories, age restricted and login only videos fail and are not charged.
  • Photo posts and carousels without video have no audio to transcribe.
  • Music only or silent videos return an empty transcript with no_speech_detected: true and are not charged.
  • Videos longer than maxVideoMinutes (default 60, maximum 240) are skipped and not charged.
  • Platforms change and rate limit without warning. Retries handle most of it, but a small share of links can fail on any given day.
  • Transcripts are in the spoken language. There is no translation.

Transcribing publicly available videos for research, analysis, accessibility and search is common practice, and this actor only accesses content anyone can watch without logging in. The words in someone else's video can still be protected by copyright, so check the platform's terms and the creator's rights before you republish a transcript, and handle any personal data in line with the laws that apply to you. This is not legal advice.

Need something else?

Open an issue on the Issues tab with the link that failed or the feature you need. Issues are read and answered.