# TikTok Transcript Scraper (`inovaflow/tiktok-transcript-scraper`) Actor

Paste TikTok video URLs, get what is said in each video as text: the full transcript, timed segments and language, from TikTok's own captions, plus caption, author and stats. Optional speech-to-text for videos without captions. For content research, repurposing and AI. No login.

- **URL**: https://apify.com/inovaflow/tiktok-transcript-scraper.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** Social media, Videos, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 transcripts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Paste TikTok video links, get what is said in each video as clean text — with timestamps and language, straight from TikTok's own captions.**

If you study what works on TikTok, you know the problem. The value of a video is in what the creator says — the hook, the script, the product claim — and none of it is in the caption. So you watch, pause, rewind and type. For ten videos that's an afternoon; for a competitor's last 200 videos it never happens. And pasting a link into a generic transcription tool means waiting for audio uploads and paying per minute.

We built this for our own content research first, and now we're sharing it. Most talking videos on TikTok already carry an automatic caption track that TikTok generated itself. This Actor reads that track directly — no audio download, no AI guessing — so a transcript costs a tenth of a cent and arrives in about a second. You get the full text, every line with its start and end time, the language, and the video's caption, creator and stats alongside.

### Who it's for

- Content marketers and creators studying hooks and scripts — *"what do the top 50 videos in my niche actually say in the first 3 seconds?"*
- Agencies repurposing TikToks into blogs, newsletters and shorts — *"give me the text of our client's last 30 videos."*
- Brand safety and compliance teams — *"did the creator say the required disclosure out loud?"*
- Researchers and AI builders — transcripts as clean JSON for summaries, topic tagging, sentiment or RAG.

### What you get per video

- `transcript` — the full spoken text; `wordCount`.
- `segments` — every caption line with `start` / `end` in seconds (turn off with *Include timed segments*).
- `language` (TikTok's tag, e.g. `eng-US`), `languageCode` (`en`), `isAutoGenerated`, `transcriptSource` (`tiktok-captions` or `speech-to-text`), `availableLanguages`.
- The video: URL, creator, caption, hashtags, post date, length, views, likes, comments, shares, sound.

### Options

- Video URLs — full links, share links (`vm.tiktok.com`, `vt.tiktok.com`, `tiktok.com/t/…`) or bare video IDs, one per line. Up to 5,000 per run.
- Preferred language — optional (`en`, `es`, `de`…). If TikTok has captions in that language for a video you get them, otherwise the original language; the row tells you which.
- Include timed segments — on by default.
- Transcribe videos without captions — speech-to-text (Whisper) for videos TikTok didn't caption, when available for this Actor; with a length limit you set. Off, or not available: those videos come back as `no-captions` and cost nothing. If the speech-to-text service is busy (rate limit), the video comes back as `stt-rate-limited`, free — just run it again later.

### How to set it up

1. Paste your video links into **Video URLs**.
2. Click **Start**. A hundred videos take about a minute.
3. Open the **Transcripts** view, export as CSV / Excel / JSON, or pull the rows into your tool or LLM pipeline through the API.

And that's it.

### Pricing

- **$0.002 per transcript** from TikTok's captions — $2 per 1,000 videos.
- **$0.005 per speech-to-text transcript** (per started 5 minutes of audio), only for videos without captions and only when a transcript is delivered.
- Plus Apify's small start fee per run. Videos without captions (when speech-to-text is off or unavailable), videos without speech, deleted or private videos, invalid links and duplicates are free.

> Tip: TikTok captions most videos where someone talks. Music-only videos, dances and very old videos usually have no captions — keep speech-to-text off if you only want the cheap ones.

### Notes

- Transcripts are TikTok's own automatic captions, so they read the way TikTok's captions read on screen: very good for clear speech, weaker for heavy slang, music over speech or several people talking at once.
- Public videos only. No login, nothing liked or followed.
- Respect creators' rights when you reuse their words, and TikTok's terms.

Found a bug or need another field? Open an issue on the **Issues** tab. Need transcripts for a whole account or hashtag? Ask there too.

# Actor input Schema

## `videoUrls` (type: `array`):

TikTok video links, one per line — `https://www.tiktok.com/@creator/video/7114641064446676267`, a share link like `https://vm.tiktok.com/ZMabc123/`, or just the video ID. Up to 5,000 per run.

## `language` (type: `string`):

Optional. A language code like `en`, `es` or `de`. When TikTok has captions in that language for a video you get them; otherwise you get the video's original-language captions (the row says whether your language was found). Leave empty for the original language.

## `includeTimestamps` (type: `boolean`):

Adds `segments`: every caption line with its start and end time in seconds. Off = plain transcript text only (smaller output).

## `speechToText` (type: `boolean`):

When a video has no TikTok captions, transcribe its audio with Whisper speech-to-text (when available for this Actor). Off = such videos come back as `no-captions` and cost nothing.

## `maxSpeechToTextMinutes` (type: `integer`):

Videos longer than this are not sent to speech-to-text (they come back as `no-captions`, free).

## `maxConcurrency` (type: `integer`):

How many videos are read at the same time.

## `residentialFallback` (type: `boolean`):

Videos are read through fast datacenter proxies. If TikTok refuses one several times, it is retried through a residential proxy. Turn off to never use residential traffic.

## `proxyConfiguration` (type: `object`):

Apify datacenter proxy by default — TikTok's public video pages load fine there.

## Actor input object example

```json
{
  "videoUrls": [
    "https://www.tiktok.com/@garyvee/video/7114641064446676267",
    "https://vm.tiktok.com/ZMabc123/"
  ],
  "language": "en",
  "includeTimestamps": true,
  "speechToText": true,
  "maxSpeechToTextMinutes": 10,
  "maxConcurrency": 5,
  "residentialFallback": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `transcripts` (type: `string`):

One row per video: transcript text, timed segments, language, source, caption, author and stats.

## `summary` (type: `string`):

Transcripts delivered (from TikTok captions / speech-to-text), videos without captions, words, languages and what could not be read.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://www.tiktok.com/@garyvee/video/7114641064446676267",
        "https://www.tiktok.com/@washingtonpost/video/7609177768793787679"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/tiktok-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "videoUrls": [
        "https://www.tiktok.com/@garyvee/video/7114641064446676267",
        "https://www.tiktok.com/@washingtonpost/video/7609177768793787679",
    ],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/tiktok-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://www.tiktok.com/@garyvee/video/7114641064446676267",
    "https://www.tiktok.com/@washingtonpost/video/7609177768793787679"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call inovaflow/tiktok-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/tiktok-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CYcbhTCseQXJ8wu8F/builds/b9zINcHkDDg9fHYOD/openapi.json
