# TikTok Transcript Scraper — Video to Text, Subtitles, No Login (`blitzdata/tiktok-transcript-scraper`) Actor

TikTok transcripts from the link alone — no cookies, no login. Instagram Reels, YouTube and Shorts too. Transcript, timed segments and ready-made SRT/WebVTT subtitles, optional translation, hook and summary. A voice-activity gate means music and silence cost nothing. $0.005 per audio minute.

- **URL**: https://apify.com/blitzdata/tiktok-transcript-scraper.md
- **Developed by:** [⚡ Blitzdata](https://apify.com/blitzdata) (community)
- **Categories:** Videos, Social media, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 audio minute transcribeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## TikTok Transcript Scraper

Paste a link, get the words. This Actor transcribes a **TikTok video, Instagram
Reel, YouTube video or YouTube Short** to text: the full transcript, timed
segments, ready-made **SRT and WebVTT subtitles**, and optionally a
**translation** and a one-line **hook and summary**. No cookies, no login, no
API key of your own.

It listens to the audio. It does not scrape captions, so it works on the TikToks
and Reels that never had any — the ones a caption scraper returns empty.

**TikTok to text · TikTok transcript · Instagram Reel transcript · YouTube
transcript · audio to text · speech to text · subtitles in 99+ languages.**

### What this Actor does

- **Transcript** of the speech in an Instagram Reel, TikTok, YouTube video or Short
- **Timed segments** (`start`, `end`, `text`) and the same blocks as **SRT** and **WebVTT** files
- **Translation** into any language you name, from the same model, at no extra charge
- **Hook and summary**: a one-line hook and a one-sentence summary, when you ask for them
- **Who said what**: turn on `diarize` and every segment and subtitle line is labelled Speaker 1, Speaker 2 — for interviews, podcasts and panels, at no extra charge
- **Whole channels**: paste a YouTube channel, playlist or TikTok profile and set how many latest posts to transcribe
- **No charge for music or silence**: a voice-activity gate runs first, so a clip without speech costs nothing and returns no invented text
- **Direct media too**: an `.mp3` or `.mp4` URL is transcribed like any platform link

All three platforms are verified from a datacenter IP, the same kind of address
this Actor runs on. Vimeo is not supported: it serves no media to logged-out
clients (0 of 4 public videos). X and Facebook are untested and therefore not
claimed.

### Music and silence cost nothing

Reels are full of clips with no speech at all: a music-only edit, a montage, a
silent product shot. Most transcribers run the model anyway, bill you for it,
and hand back a sentence the model invented over the soundtrack.

A voice-activity gate runs first here. No speech means no text and **no charge**.

Measured against a leading competitor on six real clips, five of them without
speech:

| | This Actor | Competitor |
|---|---|---|
| Billed | **1 of 6** | 6 of 6 |
| On a music-only clip | *(empty)* | `шум` — Russian for "noise" |

### How accurate is the transcript?

It runs on Voxtral Small 24B, not Whisper. Over a 120-clip corpus in 12
languages it scores a word error rate of **6.9** against **9.7** for
whisper-large-v3 (lower is better: 6.9 means about seven wrong words in a
hundred). With music under the speech, the normal case for a Reel, the gap
widens: **9.3 against 43.9** on the six speech-with-music clips, and Whisper
dropped one of them entirely.

Per language, on the same corpus (word error rate; character error rate for
zh, ja, ko):

| en | nl | de | fr | es | it | pt | hi | ar | zh | ja | ko |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 7.7 | 6.8 | 2.9 | 9.4 | 1.7 | 2.2 | 5.1 | 10.1 | 16.4 | 7.4 | 5.4 | 8.1 |

These are read-speech benchmark clips. A Reel with a beat under the voice, a
phone microphone and street noise will score worse than this; the
speech-with-music number above is the better guide for that case. Languages
outside the twelve are handled by the same model but are not measured here.

On clean speech the transcript is level with the best competitors: two NASA
clips came back word-for-word equal. The audio is cut at silences and the
pieces are transcribed in parallel, at roughly 5 to 9 times realtime:
transcribing a 24-minute video takes about three minutes and a 69-minute
podcast about nine, plus the download.

#### Pass the language if you know it

Speech models decode better when they are told the language, and the
difference is not small. On the same 120-clip benchmark:

| | word error rate |
|---|---|
| `language` set to the right code | **6.8** |
| `language: auto` | 7.6 |

`auto` gets part of the way there by transcribing one chunk first, reading the
language off it, and running the rest with that. Each row reports what
happened in `language_source`: `given`, `probe`, or `none`.

If you are scraping one creator, one market or one campaign, you already know
the language. Passing it is the cheapest accuracy you will get anywhere.

### SRT and WebVTT subtitles from any Reel, TikTok or YouTube video

`segments` gives `{start, end, text}`; `srt` and `vtt` are the same blocks in
the two subtitle formats, ready to upload.

```
1
00:00:00,390 --> 00:00:04,270
NASA is building a moon base, a place where astronauts will live, work, and conduct

2
00:00:04,270 --> 00:00:08,530
science on and around the moon. Now building a moon base won't happen all at once.
```

Cue boundaries come from the voice-activity gate, so a cue starts and ends
where someone is actually speaking. Inside a cue the words are spread by
length, because the model gives no per-word timestamps and inventing them
would look more precise while being less true. Blocks stay within two lines of
42 characters and six seconds, and break on sentence ends where they can.
Japanese and Chinese get shorter lines, as subtitles in those scripts should.

### How to get a transcript from an Instagram Reel, TikTok or YouTube video

1. Open [TikTok Transcript Scraper](https://apify.com/blitzdata/tiktok-transcript-scraper) on Apify and click **Try for free**.
2. Paste one or more video links in **Video URLs**. Reels, TikToks, YouTube videos and Shorts can be mixed in one run.
3. Set **Language** to the ISO code if you know it (`en`, `nl`, `es`, ...). Leave `auto` if you do not.
4. Tick **Add hook + summary** or fill **Translate to** if you want them. They cost nothing extra.
5. Click **Start**. One dataset row per video appears with the transcript, segments, SRT and VTT.
6. Export the dataset as JSON, CSV or Excel, or read it through the API, the Python or JavaScript client, or the MCP server (see the **API** tab).

The same input works as a JSON body through the API:

```json
{
  "urls": [
    "https://www.instagram.com/nasa/reel/CvNoLm8Ouaa/",
    "https://www.tiktok.com/@nasa/video/7685115221458930957",
    "https://www.youtube.com/watch?v=aircAruvnKk"
  ],
  "language": "en",
  "include_hook": true
}
```

### Input

| Field | Type | Default | Meaning |
|---|---|---|---|
| `urls` | array | — | One or more video URLs, or channel/profile URLs (required) |
| `posts_per_profile` | int | `0` | Transcribe this many latest posts per channel or profile URL. `0` disables it |
| `language` | string | `auto` | ISO code, or `auto`. Pass the code when you know it — see above |
| `include_hook` | bool | `false` | Also return a one-line hook and a one-sentence summary |
| `diarize` | bool | `false` | Work out who speaks when; labels every segment and subtitle line |
| `num_speakers` | int | — | Exact number of speakers, when you know it. An interview is 2 |
| `translate_to` | string | — | Translate the transcript into this language |
| `use_proxy` | string | `auto` | `auto` routes YouTube through a residential pool (it refuses datacenter IPs) and retries anything else through it if the direct attempt is refused; `always` / `never` override |
| `proxy_region` | string | — | ISO-2 country for the residential exit (NL, US, …) |
| `max_duration_seconds` | int | `7200` | Skip longer audio so one upload cannot run up the bill |

Duplicate URLs in one run, including a video that arrives both on its own and
through a profile, are transcribed and billed once.

### Output

One row per video:

| Field | What |
|---|---|
| `text` | the full transcript |
| `language` | ISO code, read off the transcript |
| `language_source` | `given`, `probe` or `none` — how the language was decided |
| `segments` | `{start, end, text}` blocks |
| `srt`, `vtt` | the same blocks as subtitle files |
| `translation`, `translation_language` | when `translate_to` was set |
| `hook`, `summary` | when `include_hook` was set |
| `speakers`, `speaker_turns` | who spoke and when, with `speaker` on every segment — when `diarize` was set |
| `duration_seconds`, `speech_seconds` | audio length, and how much of it was speech |
| `speech_spans` | the raw voice-activity windows |
| `platform`, `video_id`, `title`, `uploader`, `upload_date`, `view_count`, `like_count`, `webpage_url` | source metadata, as far as the platform exposes it |
| `billed_audio_minutes` | what this row cost, in minutes |
| `from_profile` | the channel or profile a video was expanded from |
| `proxy_used` | whether the download went through the residential pool |
| `error` | set when one URL fails — a broken URL never stops the run |

#### Output example

A TikTok video, shortened:

```json
{
  "url": "https://www.tiktok.com/@nasa/video/7685115221458930957",
  "platform": "TikTok",
  "video_id": "7685115221458930957",
  "title": "Another week of building the future at NASA. 🚀 🎓 U.S. Space Academy...",
  "uploader": "nasa",
  "upload_date": "20260913",
  "view_count": 156500,
  "like_count": 6317,
  "text": "NASA's future depends on the people, missions, and the networks that connect them. This week had all three. Here's what's new in your NASA Minute. Just two weeks after President Trump signed the executive order, ...",
  "language": "en",
  "language_source": "probe",
  "duration_seconds": 82.91,
  "speech_seconds": 70.22,
  "segments": [
    { "start": 0.32, "end": 3.78, "text": "NASA's future depends on the people, missions, and the networks that connect them." },
    { "start": 3.78, "end": 6.61, "text": "This week had all three. Here's what's new in your NASA Minute." }
  ],
  "srt": "1\n00:00:00,320 --> 00:00:03,780\nNASA's future depends on the people, missions, and the networks that\n\n2\n...",
  "vtt": "WEBVTT\n\n00:00:00.320 --> 00:00:03.780\nNASA's future depends on the people, ...",
  "translation": null,
  "hook": null,
  "summary": null,
  "billed_audio_minutes": 2,
  "error": null
}
```

### Transcribe a whole YouTube channel or TikTok profile

Paste `https://www.youtube.com/@NASA/videos` with `posts_per_profile: 3` and the
three latest videos come back as three rows, each carrying `from_profile`.
Channel URLs, playlists and `ytsearch5:...` all work.

TikTok profiles work when TikTok feels like it. Instagram profiles cannot be
listed at all, so give Reel URLs there — you get a row saying so rather than an
empty result.

### How much does it cost to transcribe a Reel, TikTok or YouTube video?

**$0.005 per audio minute**, rounded up per video, and only when speech was
found. Everything is included: the download, the residential proxy YouTube
needs, the transcription, the subtitles, the translation and the hook. There is
no start fee and no per-result fee.

| Video | Billed | Cost |
|---|---|---|
| A 30-second Reel | 1 minute | $0.005 |
| An 82-second TikTok | 2 minutes | $0.01 |
| A 19-minute YouTube video | 19 minutes | $0.095 |
| A 70-minute podcast | 70 minutes | $0.35 |
| 1,000 Reels under a minute each | 1,000 minutes | $5 |

A clip with no speech costs nothing. A URL that fails costs nothing. Apify's
free plan includes $5 of usage a month, which covers up to 1,000 audio minutes
here without a card. Set a maximum charge on the run if you want a hard cap;
the Actor stops at the limit and tells you which URLs it did not reach.

### Use cases for a video transcript API

- **Content research**: pull the scripts of a creator's last fifty Reels and see what they actually say
- **Ad and competitor monitoring**: turn TikTok and Reel campaigns into searchable text, with the hook already extracted
- **Subtitles and repurposing**: SRT and VTT for re-uploads, translated in the same call
- **LLM and RAG pipelines**: clean, timed text from short video, one JSON row per clip
- **Social listening and brand safety**: what is being said about a brand in video, not just in captions

### FAQ

**Does it work without an Instagram or TikTok login?**
Yes. No cookies, no session ID, no account. Public content only.

**Does it transcribe YouTube Shorts?**
Yes. Shorts, regular videos, playlists, channels and `ytsearch` queries all work. YouTube refuses datacenter IPs, so those downloads go through a residential pool; that is included in the price.

**Which languages are supported?**
The model is multilingual; twelve languages are measured above (English, Dutch, German, French, Spanish, Italian, Portuguese, Hindi, Arabic, Chinese, Japanese, Korean). Pass `language` when you know it.

**What happens on a video with no speech?**
The row comes back with `speech: false`, an empty `text` and `billed_audio_minutes: 0`. No charge, and no invented sentence.

**How long can a video be?**
Up to two hours per video by default (`max_duration_seconds`). Long audio is cut at silences and transcribed in parallel; a 69-minute podcast returned 13,757 words, 1,111 timed segments and a 111 KB SRT.

**How is the translation billed?**
It is not. Translation and the hook are text passes on the same model and are included in the per-minute price.

**Are the segment timestamps word-accurate?**
Cue starts and ends sit on measured speech. Within a cue, words are spread by length, so expect them to be close rather than frame-exact. See the subtitles section above.

**What comes back for a private, deleted or mistyped URL?**
An error row that says so, not an empty transcript, so the two are easy to tell apart. The rest of the run continues and the failed URL is not billed.

**Can I call it from Python, JavaScript, curl or an AI agent?**
Yes. The **API** tab shows ready-made calls for the Apify API, the Python and JavaScript clients, the CLI and the MCP server. Runs can be scheduled and results delivered through webhooks and integrations like any Apify Actor.

**Does it support Vimeo, X or Facebook?**
Vimeo no: it serves no media to logged-out clients. X and Facebook are untested and are not claimed. Direct `.mp3` and `.mp4` URLs work.

### Limits

- Instagram profiles cannot be expanded to their latest posts; give Reel URLs.
- TikTok profile expansion depends on TikTok and is not guaranteed.
- One video is billed on its full audio length, rounded up to the minute, once speech is found anywhere in it.
- Korean is the weakest of the measured languages (8.1 character error rate); Arabic the weakest in word error rate (16.4).

### Is it legal to transcribe Instagram Reels, TikToks and YouTube videos?

This Actor reads publicly available video and turns speech into text. It does
not bypass logins and does not access private content. The audio is deleted
after transcription and is not used for training. You are responsible for using the output in line with
the platforms' terms, copyright and the privacy law that applies to you, in
particular when the transcripts contain personal data.

### Support

Something wrong, a platform that stopped working, a field you need? Open an
issue on the **Issues** tab of this Actor and it will be looked at.

# Actor input Schema

## `urls` (type: `array`):

Instagram Reels, TikTok and YouTube (incl. Shorts) — one dataset row per video, with transcript, timed segments and ready-made SRT and WebVTT subtitles. No cookies or login needed for public content. A channel or profile URL works too when 'Posts per profile' is set. Direct .mp3/.mp4 URLs work as well.

## `posts_per_profile` (type: `integer`):

Paste a channel or profile URL instead of single videos and this many of its latest posts are transcribed. 0 disables it, so a profile URL then yields one video. Works for YouTube channels, playlists and search; TikTok profiles are hit-and-miss; Instagram profiles cannot be listed, so give reel URLs there.

## `language` (type: `string`):

ISO code (en, nl, es, ...) or 'auto'. Passing the code is worth it when you know it: on a 120-clip benchmark the model scored 6.8 word error rate with the language given against 7.6 on 'auto', which establishes the language from one chunk first and would be 8.0 without that. The row reports which happened in 'language\_source'.

## `include_hook` (type: `boolean`):

Also return a one-line hook and a one-sentence summary, written from the transcript by the same model.

## `diarize` (type: `boolean`):

Work out who is speaking when, and label every segment and subtitle line with Speaker 1, Speaker 2 and so on. Useful for interviews, podcasts and panels. Costs nothing extra.

## `num_speakers` (type: `integer`):

Tell it the exact number of speakers if you know it — an interview is 2. Leave empty and it works the number out itself, which is slightly less reliable. Only used when Label the speakers is on.

## `translate_to` (type: `string`):

Language to translate the transcript into (e.g. English, Dutch, Spanish, or an ISO code). Leave empty to skip. The same speech model does the translation, so there is no second vendor and no extra fee.

## `max_duration_seconds` (type: `integer`):

Skip anything longer, so one long upload cannot run up the bill. Up to 2 hours is supported: the audio is cut at silences and the pieces are transcribed in parallel at roughly 5-9x realtime, so a 2-hour file takes minutes.

## `use_proxy` (type: `string`):

YouTube refuses datacenter IPs with a bot check, so 'auto' routes YouTube through a residential pool. Instagram and TikTok download fine without one and are left direct. Any other refused download is retried through the pool as well. 'always' / 'never' override.

## `proxy_region` (type: `string`):

ISO-2 country for the residential exit (e.g. NL, US). Leave empty for the default.

## Actor input object example

```json
{
  "urls": [
    "https://www.tiktok.com/@nasa/video/7685115221458930957",
    "https://www.instagram.com/nasa/reel/CvNoLm8Ouaa/"
  ],
  "posts_per_profile": 0,
  "language": "auto",
  "include_hook": false,
  "diarize": false,
  "translate_to": "",
  "max_duration_seconds": 7200,
  "use_proxy": "auto",
  "proxy_region": ""
}
```

# Actor output Schema

## `transcripts` (type: `string`):

One row per video: text, timed segments, SRT and WebVTT, optional translation, hook and summary, plus source metadata.

## `summary` (type: `string`):

How many URLs were given, how many videos were found after expanding any profiles, and how many succeeded.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.tiktok.com/@nasa/video/7685115221458930957",
        "https://www.instagram.com/nasa/reel/CvNoLm8Ouaa/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("blitzdata/tiktok-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://www.tiktok.com/@nasa/video/7685115221458930957",
        "https://www.instagram.com/nasa/reel/CvNoLm8Ouaa/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("blitzdata/tiktok-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.tiktok.com/@nasa/video/7685115221458930957",
    "https://www.instagram.com/nasa/reel/CvNoLm8Ouaa/"
  ]
}' |
apify call blitzdata/tiktok-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,blitzdata/tiktok-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/e9aLSdpf4YH0Uu57H/builds/ZDb0xcUxSEvvd6Ztc/openapi.json
