# Video Transcriber: TikTok, Instagram, Facebook, YouTube Shorts (`themineworks/instagram-tiktok-video-transcript`) Actor

Turn TikTok, Instagram Reels, Facebook, YouTube Shorts and X video links into timestamped transcripts and SRT or VTT subtitles. Video to text in 28 languages with auto detection and speaker labels. Bulk links, no login. Failed and silent videos are never charged. Works via API and MCP.

- **URL**: https://apify.com/themineworks/instagram-tiktok-video-transcript.md
- **Developed by:** [The Mine Works](https://apify.com/themineworks) (community)
- **Categories:** Videos, AI, Social media
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 seconds of video transcribeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Video Transcriber: TikTok, Instagram Reels, Facebook, YouTube Shorts and X

Video Transcriber turns public **TikTok, Instagram Reels, Facebook, YouTube Shorts and X (Twitter)** video links into accurate, timestamped **transcripts** and ready to use **SRT and VTT subtitles**. Paste one link or a few thousand and get clean JSON with sentence level segments, speaker labels and the detected language. No login, no cookies and no API keys of your own.

It listens to the audio itself, so it works on Reels and TikToks that have no captions at all, which is most of them.

**At a glance**

- **Platforms:** TikTok, Instagram (Reels, video posts, IGTV), Facebook (Reels, Watch, fb.watch), YouTube (Shorts and regular videos) and X / Twitter videos
- **Output:** transcript text, timestamped segments, SRT subtitles, WebVTT subtitles, detected language, speaker count, exact billed seconds
- **Languages:** 28 languages with automatic detection, plus a multilingual mode for videos that switch language mid sentence (Hinglish, Spanglish, Taglish)
- **Bulk:** any number of links per run, processed in parallel
- **Fair billing:** failed links, private videos and clips with no speech are never charged
- **Works on Apify's free plan.** The free monthly credit covers about 80 transcripts of 30 second Reels.

### What does Video Transcriber do?

You give it public video URLs. For each one it downloads only the audio track, runs it through a dedicated speech recognition model and returns:

1. The full **transcript**, with optional `[start - end]` timestamps and `[Speaker N]` labels inline
2. A **segments** array with one entry per sentence: start time, end time, text and speaker
3. **SRT** and **WebVTT** subtitle files as text fields, split into two line captions that fit on screen
4. The **detected language**, the **number of speakers**, the video length and the exact seconds you were billed for

Every link gets its own output row, including the ones that fail, so you can always join results back to your list.

### Which platforms and links are supported?

| Platform | Links that work |
|---|---|
| TikTok | `tiktok.com/@user/video/...`, short links `vm.tiktok.com/...` and `vt.tiktok.com/...` |
| Instagram | `instagram.com/reel/...`, `/reels/...`, `/p/...` (video posts), `/tv/...` |
| Facebook | `facebook.com/reel/...`, `facebook.com/.../videos/...`, `facebook.com/watch/?v=...`, `fb.watch/...` |
| YouTube | `youtube.com/shorts/...`, `youtube.com/watch?v=...`, `youtu.be/...` |
| X (Twitter) | `x.com/user/status/...`, `twitter.com/user/status/...` |

Only **public** videos can be transcribed. Anything else is returned as a failed row with a clear reason and is not charged.

### How do I transcribe TikTok videos or Instagram Reels in bulk?

1. Click **Try for free** and paste your links into **Video URLs**, one per line.
2. Leave **Language** on `auto`, or pick the language if you know it.
3. Turn on **Speaker labels** for interviews, podcasts and street interviews.
4. Click **Start**. Results appear in the **Output** tab as they finish, and you can download them as JSON, CSV, Excel or HTML.

Input example:

```json
{
  "urls": [
    "https://www.instagram.com/reel/DcHyP0GsROe",
    "https://www.tiktok.com/@user/video/7412345678901234567",
    "https://www.youtube.com/shorts/abcDEF12345",
    "https://x.com/NASA/status/1491475671058681863/video/1"
  ],
  "include_timestamps": true,
  "include_subtitles": true,
  "enable_diarization": false,
  "language": "auto",
  "maxVideoMinutes": 60,
  "maxConcurrency": 3
}
```

Only `urls` is required. Everything else has a sensible default.

### What does the output look like?

A real record from a NASA video on X, shortened to its first four sentences so it fits here. A full record carries every sentence in `transcript`, `segments`, `srt` and `vtt`.

```json
{
  "sourceUrl": "https://x.com/NASA/status/1491475671058681863/video/1",
  "videoId": "1491475671058681863",
  "platform": "x",
  "status": "success",
  "durationSec": 204.89,
  "transcript": "[49.86s - 53.53s] [Speaker 0] It's thrilling to be able to see something that's never been seen before. [53.62s - 56.18s] [Speaker 0] This emission that we're seeing is thermal emission. [56.59s - 65.39s] [Speaker 0] Even on the night side, the surface of Venus is so hot that it's it's glowing, faintly at very red wavelengths. [72.28s - 83.81s] [Speaker 1] These whisper images, I think, are really exciting because they provide a new window into the lower atmosphere and surface region of Venus where these extreme conditions exist.",
  "segments": [
    { "start": 49.86, "end": 53.53, "text": "It's thrilling to be able to see something that's never been seen before.", "speaker": 0 },
    { "start": 53.62, "end": 56.18, "text": "This emission that we're seeing is thermal emission.", "speaker": 0 },
    { "start": 56.59, "end": 65.39, "text": "Even on the night side, the surface of Venus is so hot that it's it's glowing, faintly at very red wavelengths.", "speaker": 0 },
    { "start": 72.28, "end": 83.81, "text": "These whisper images, I think, are really exciting because they provide a new window into the lower atmosphere and surface region of Venus where these extreme conditions exist.", "speaker": 1 }
  ],
  "srt": "1\n00:00:49,860 --> 00:00:53,530\n[Speaker 0] It's thrilling to be able to see\nsomething that's never been seen before.\n\n2\n00:00:53,620 --> 00:00:56,180\n[Speaker 0] This emission that we're\nseeing is thermal emission.\n\n3\n00:00:56,590 --> 00:01:02,830\n[Speaker 0] Even on the night side, the surface of\nVenus is so hot that it's it's glowing,\n\n4\n00:01:02,830 --> 00:01:05,390\n[Speaker 0] faintly at very red wavelengths.\n",
  "vtt": "WEBVTT\n\n00:00:49.860 --> 00:00:53.530\n<v Speaker 0>It's thrilling to be able to see\nsomething that's never been seen before.\n\n00:00:53.620 --> 00:00:56.180\n<v Speaker 0>This emission that we're\nseeing is thermal emission.\n\n00:00:56.590 --> 00:01:02.830\n<v Speaker 0>Even on the night side, the surface of\nVenus is so hot that it's it's glowing,\n\n00:01:02.830 --> 00:01:05.390\n<v Speaker 0>faintly at very red wavelengths.\n",
  "detected_language": "en",
  "speakers_detected": 2,
  "billed_seconds": 205,
  "timestamp": "2026-09-26T11:02:36.713Z"
}
```

The first 49 seconds of that video are music, which is why the first sentence starts at 49.86s. Silence and music are skipped in the text, but you are billed on the full audio length, rounded to the nearest second.

A link that cannot be transcribed looks like this, and costs nothing:

```json
{
  "sourceUrl": "https://www.instagram.com/reel/THISDOESNOTEXIST123/",
  "videoId": "THISDOESNOTEXIST123",
  "platform": "instagram",
  "status": "failed",
  "error": "Could not download this video. It may be private, deleted, or temporarily blocked by the platform",
  "timestamp": "2026-09-26T09:33:46.303Z"
}
```

#### Output fields

| Field | What it holds |
|---|---|
| `sourceUrl` | The link you gave, exactly as given |
| `videoId` | The platform's id for the video, taken from the link |
| `platform` | `tiktok`, `instagram`, `facebook`, `youtube` or `x` |
| `status` | `success` or `failed` |
| `durationSec` | Audio length in seconds, measured from the audio file itself |
| `transcript` | Full text, with `[start - end]` timestamps and `[Speaker N]` labels when those options are on |
| `segments` | One object per sentence: `start`, `end`, `text`, and `speaker` when speaker labels are on |
| `srt` | SubRip subtitles, ready to save as a `.srt` file |
| `vtt` | WebVTT subtitles, ready to save as a `.vtt` file, with speakers as voice tags |
| `detected_language` | Language code, for example `en`, `hi`, `es` |
| `speakers_detected` | Number of distinct speakers when speaker labels are on, otherwise `null` |
| `no_speech_detected` | `true` when the video has music or silence only (not charged) |
| `billed_seconds` | The seconds you were charged for this video |
| `error` | Why a failed row failed |
| `timestamp` | When the row was produced |

### How much does it cost to transcribe a video?

You pay for three things, all listed on the **Pricing** tab:

| Charge | Free plan | Starter (Bronze) | Scale (Silver) | Business (Gold) and up |
|---|---|---|---|---|
| Per second of video transcribed | $0.0018 | $0.0015 | $0.0012 | $0.0010 |
| Per video transcribed | $0.008 | $0.005 | $0.005 | $0.005 |
| Per run started | $0.005 | $0.005 | $0.005 | $0.005 |

Compute, proxies and storage are included. There is nothing else on the bill.

What real jobs cost:

| Job | Free plan | Starter | Scale | Business |
|---|---|---|---|---|
| One 30 second Reel | $0.067 | $0.055 | $0.046 | $0.040 |
| 100 Reels of 45 seconds, one run | $8.91 | $7.26 | $5.91 | $5.01 |
| 1,000 TikToks of 20 seconds, one run | $44.01 | $35.01 | $29.01 | $25.01 |
| One 10 minute YouTube video | $1.09 | $0.91 | $0.73 | $0.61 |

**Charged:** each video that was transcribed and delivered, its length in seconds (rounded to the nearest second, minimum one), and one start charge per run at the default 1 GB of memory.

**Never charged:** links that fail to download, private, deleted or login only videos, unsupported links, videos longer than your `maxVideoMinutes` limit, and videos with no speech in them.

Batch your links into one run: the start charge is paid once per run, not once per video.

### Will it go over my budget?

No. Before transcribing each video, the actor checks that your run's **maximum cost** can still cover it. A video that would push the run over your limit is skipped with a clear message instead of being charged, and the run stops cleanly when the budget is used up.

### Can I get SRT or VTT subtitle files?

Yes, on by default. Every successful row carries an `srt` and a `vtt` field. Save either one as a file and it drops straight into Premiere Pro, DaVinci Resolve, CapCut, YouTube Studio or any video player. Long sentences are split into captions of at most two lines of about 42 characters, the length broadcast captioning guidelines recommend, and the timing is shared across the split. Turn **Include SRT and VTT subtitles** off if you only need the text.

### Does it detect the language automatically?

Yes. Leave **Language** on `auto` and each video's language is detected on its own, so one run can mix English, Hindi and Spanish videos. Pick `multi` for videos where speakers switch language mid sentence. Picking the exact language improves accuracy when you know it.

Supported languages: English (US, UK, Australia, India), Spanish (Spain and Latin America), French (France and Canada), German, Dutch, Portuguese (Portugal and Brazil), Italian, Japanese, Korean, Chinese, Hindi, Indonesian, Malay, Russian, Ukrainian, Polish, Swedish, Danish, Finnish, Norwegian, Turkish, Thai, Vietnamese, Greek, Czech, Slovak, Romanian and Hungarian.

The transcript is always in the language that is spoken. This actor does not translate.

### Can it tell different speakers apart?

Yes. Turn on **Speaker labels** and every sentence is tagged `[Speaker 0]`, `[Speaker 1]` and so on, `speakers_detected` reports how many people spoke, and the VTT file carries speakers as voice tags. It costs nothing extra.

### How accurate is the transcription?

Transcripts come from a dedicated speech recognition model, not from a general AI chatbot. In our own test on 55 Instagram Reels (26 Sep 2026), every passage of speech was captured with nothing dropped or repeated, and about 97% of words matched another commercial transcription service word for word. The remaining differences were at word level, for example "wanna" written as "want to", or a brand name spelled differently. That matters: language models asked to transcribe tend to tidy up, summarise or quietly drop sentences they judge unimportant, and you cannot see the gap. A speech model writes down what was said.

Clear speech, voiceovers and talking head videos come back close to word perfect. Accuracy drops, as it does for any transcriber, on heavy background music, several people talking over each other, strong accents in a language picked wrongly, and slang or brand names the model has not heard. Setting the correct language helps most.

### Do I need a TikTok or Instagram login, cookies or an API key?

No. Everything runs from Apify's servers with no account, cookie or key from you, so there is nothing of yours that can be blocked or banned.

### What happens when a video fails?

Each link is first fetched through fast datacenter servers. If the platform rate limits or blocks that request, it is retried automatically through residential servers at no extra cost to you. Links that can never work (private, deleted, age restricted, login only) fail immediately with a clear `error` so you are not kept waiting. None of these are charged.

### How is this different from YouTube caption scrapers?

Caption scrapers copy subtitles the uploader or YouTube already made, so they only work when captions exist. Most TikToks, Reels and Facebook videos have none. Video Transcriber listens to the audio, so it works on any public video with speech, on any of the five platforms, and gives you timestamps and subtitles either way.

### How do I use it from the API, Python or JavaScript?

Python:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("themineworks/instagram-tiktok-video-transcript").call(run_input={
    "urls": ["https://www.instagram.com/reel/DcHyP0GsROe"],
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(row["status"], row.get("transcript"))
```

JavaScript:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('themineworks/instagram-tiktok-video-transcript').call({
    urls: ['https://www.tiktok.com/@user/video/7412345678901234567'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => r.transcript));
```

One HTTP call that waits and returns the transcripts, good for a handful of links:

```bash
curl -X POST "https://api.apify.com/v2/acts/themineworks~instagram-tiktok-video-transcript/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://www.youtube.com/shorts/abcDEF12345"]}'
```

It also works from **n8n, Make, Zapier, Google Sheets, Airbyte and webhooks** through Apify's integrations, and on a **schedule** if you want new videos transcribed every day.

### Can AI agents use it through MCP?

Yes. Add Apify's MCP server to Claude, ChatGPT, Cursor or any MCP client with this actor enabled:

```
https://mcp.apify.com/?tools=themineworks/instagram-tiktok-video-transcript
```

Your agent can then transcribe any video link it finds and reason over the text, for example "summarise what these five creators said about our product this week".

### What is it used for?

- **Content repurposing:** turn Reels and TikToks into blog posts, captions, threads and newsletters
- **Subtitles and accessibility:** burn in captions or upload SRT files so videos work on mute and for deaf and hard of hearing viewers
- **Competitor and trend research:** read what competitors and creators say across hundreds of videos in minutes, and search it
- **Hook analysis:** pull the first sentence of every top performing video in a niche and compare what works
- **Brand and influencer monitoring:** check what creators actually say about a product before you pay them
- **AI and RAG pipelines:** feed clean, timestamped text from short video into search indexes, embeddings and LLM analysis
- **Research and journalism:** keep a searchable, time coded record of public statements made on video

### Limitations

- Public videos only. Private accounts, close friends stories, age restricted and login only videos fail and are not charged.
- Photo posts and carousels without video have no audio to transcribe.
- Music only or silent videos return an empty transcript with `no_speech_detected: true` and are not charged.
- Videos longer than `maxVideoMinutes` (default 60, maximum 240) are skipped and not charged.
- Platforms change and rate limit without warning. Retries handle most of it, but a small share of links can fail on any given day.
- Transcripts are in the spoken language. There is no translation.

### Is it legal to transcribe public social media videos?

Transcribing publicly available videos for research, analysis, accessibility and search is common practice, and this actor only accesses content anyone can watch without logging in. The words in someone else's video can still be protected by copyright, so check the platform's terms and the creator's rights before you republish a transcript, and handle any personal data in line with the laws that apply to you. This is not legal advice.

### Need something else?

Open an issue on the **Issues** tab with the link that failed or the feature you need. Issues are read and answered.

# Actor input Schema

## `urls` (type: `array`):

Public video links to transcribe, one per line: TikTok, Instagram Reels and video posts, Facebook Reels and Watch, YouTube Shorts and videos, X (Twitter) videos. Short links like vm.tiktok.com and fb.watch work. Private, deleted or login only videos fail and are never charged.

## `include_timestamps` (type: `boolean`):

Adds \[start - end] markers inside the transcript text, for example "\[0.00s - 2.50s] Welcome to my channel!". The segments array always has timestamps either way.

## `include_subtitles` (type: `boolean`):

Adds ready to use srt and vtt subtitle text to each result, split into two line captions of about 42 characters. Save either field as a file for Premiere Pro, CapCut, DaVinci Resolve or YouTube Studio. No extra cost.

## `enable_diarization` (type: `boolean`):

Labels who is speaking (\[Speaker 0], \[Speaker 1] ...) in the transcript, segments and subtitles, and reports speakers\_detected. Useful for interviews, podcasts and street interviews. No extra cost.

## `language` (type: `string`):

auto detects each video's language on its own, so one run can mix languages. multi handles speakers who switch language mid sentence (Hinglish, Spanglish). Pick the exact language for the best accuracy when you know it. Transcripts are in the spoken language; nothing is translated.

## `maxVideoMinutes` (type: `integer`):

Videos longer than this are skipped and never charged. Protects your budget from an accidental two hour upload in a list of short clips.

## `maxConcurrency` (type: `integer`):

How many videos are processed at the same time. Higher finishes big lists faster; lower is gentler on platforms that rate limit.

## `start_urls` (type: `string`):

Kept so older integrations keep working. Use urls instead.

## Actor input object example

```json
{
  "urls": [
    "https://www.instagram.com/reel/DcHyP0GsROe"
  ],
  "include_timestamps": true,
  "include_subtitles": true,
  "enable_diarization": false,
  "language": "auto",
  "maxVideoMinutes": 60,
  "maxConcurrency": 3
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.instagram.com/reel/DcHyP0GsROe"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("themineworks/instagram-tiktok-video-transcript").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://www.instagram.com/reel/DcHyP0GsROe"] }

# Run the Actor and wait for it to finish
run = client.actor("themineworks/instagram-tiktok-video-transcript").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.instagram.com/reel/DcHyP0GsROe"
  ]
}' |
apify call themineworks/instagram-tiktok-video-transcript --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,themineworks/instagram-tiktok-video-transcript"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/M06LX584YCty8SHHj/builds/ZXfYLItBGFrbw1vLO/openapi.json
