# TikTok Transcript Scraper: Videos & Accounts to Text (`fguiraud/tiktok-transcript-scraper`) Actor

TikTok transcript scraper: the spoken words of any TikTok video, or the newest videos of whole accounts, as text with views, likes and date, plus SRT subtitles. Open-source Whisper, no API key. $0.009 per video.

- **URL**: https://apify.com/fguiraud/tiktok-transcript-scraper.md
- **Developed by:** [Fernando Guiraud](https://apify.com/fguiraud) (community)
- **Categories:** Videos, AI, Developer tools
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $9.00 / 1,000 videos

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does TikTok Transcript Scraper do?

**TikTok Transcript Scraper** turns **TikTok videos into text**: the full **transcript** of what is said, with the video's **author, date, views, likes and comments**, and optional **SRT/VTT subtitles**. Paste video links, or give **whole TikTok accounts** (`@username`) and get the newest videos of each one, one row per video.

It runs **open-source Whisper** on the audio, so it works on **every video with speech, not only those with captions**. No TikTok login, no API key and no subscription: you pay **$0.009 per video** (first minute included), and videos with only music, private videos and failed downloads are **never billed**.

It runs on the Apify platform, so you also get an API, **scheduling**, integrations (Google Sheets, Make, Zapier, n8n) and access for **AI agents through the [Apify MCP server](https://mcp.apify.com)**.

![Sample output: real rows from a run of this Actor](https://fernandoguiraud16-coder.github.io/data-tools/assets/outputs/output-tiktok-transcript-scraper.png?v=3)

> **Independent tool**, not affiliated with or endorsed by TikTok or ByteDance. TikTok is a trademark of ByteDance Ltd.; it is named only to describe the public data this Actor collects. It only accesses publicly available content.

### Why use it?

- 🎣 **Hooks and scripts**: what viral videos say in their first sentence, ranked by views.
- 📺 **Track an account every week**: schedule it with **Only new videos** and get only what was posted since the last run.
- 🕵️ **Competitor and creator research**: a searchable log of what brands and creators say.
- ✍️ **Repurposing**: turn your own TikToks into captions, posts and newsletters.
- 🤖 **AI pipelines**: clean text ready for ChatGPT, Claude or a vector database.

### How to get the transcript of a TikTok video

1. Click **Try for free**.
2. Paste **TikTok video links**, or add accounts (`@nasa`) in **TikTok accounts**.
3. Optionally choose the **spoken language** (or `auto`) and extra outputs (SRT, segments, Markdown).
4. Click **Start**, then download the results as JSON, CSV or Excel.

### Output

One record per video. Real result (October 2026):

```json
{
  "source": "https://www.tiktok.com/@tiktok/video/7689568478570843422",
  "status": "ok",
  "language": "en",
  "durationSeconds": 52.72,
  "billedMinutes": 1,
  "text": "You three look so good. I'm getting serious 2000 flashbacks. Thanks. I have this disposable camera. …",
  "video": {
    "platform": "tiktok",
    "author": "tiktok",
    "uploadDate": "2026-09-25",
    "views": 327000,
    "likes": 4477,
    "comments": 1576
  }
}
```

#### A whole account, every week

```json
{
  "profiles": ["@nasa"],
  "maxVideosPerProfile": 10,
  "onlyNewVideos": true
}
```

Real run (October 2026): of the 5 newest `@nasa` videos, 3 had speech and were transcribed (243,300, 163,800 and 105,900 views); the other 2 had only music and were not billed. With `onlyNewVideos`, a second run skips them all and costs nothing.

### Input

| Field | Description | Default |
|---|---|---|
| `sources` | TikTok video links (Instagram, X, Facebook and MP4 links also work) | `sources` or `profiles` |
| `profiles` | TikTok accounts (`@username` or profile link) | - |
| `maxVideosPerProfile` | Newest videos per account | 10 |
| `onlyNewVideos` | Skip videos transcribed by earlier runs | on |
| `language` | Spoken language code, or `auto` | `auto` |
| `model` | `base` (recommended), `small` (most accurate) or `tiny` | `base` |
| `outputs` | `text`, `markdown`, `segments`, `srt`, `vtt`, `chunks` | `text` |

### How much does it cost?

| Event | Price |
|---|---|
| **Video** (first minute included) | **$0.009** |
| Each extra minute (base or tiny model) | $0.006 |
| Each extra minute (small model) | $0.012 |
| Run start (per GB of memory, default 4 GB) | $0.0005 |
| AI insights per video (optional; Claude usage billed to your own key) | $0.01 |

**100 TikToks cost about $0.90.** A 3-minute video costs $0.021. Music-only videos, private videos and failed downloads are free; each result explains why.

### Use it with AI agents (MCP)

Add `https://mcp.apify.com?tools=fguiraud/tiktok-transcript-scraper` to Claude, Cursor or any MCP client and ask: *"Transcribe the 5 newest videos of @nasa and summarize what they talk about."*

**Agent-friendly input**: common field names such as `url`, `urls`, `startUrls`, `query` or `keywords` are accepted, and an empty input runs a small sample from the input form (with a note in the results) instead of failing.

### FAQ and limitations

- **Videos without speech** (music only) come back as "no speech" and are not billed.
- **Private videos and videos that need a login** cannot be read; they are not billed.
- **Is it legal?** The Actor reads public videos. Respect the rights of the creators when you reuse their content.
- Found a problem or need a feature? Open an issue on the **Issues** tab. Replies within 48 hours.

### Related tools

- [Instagram Reels Transcript Scraper](https://apify.com/fguiraud/instagram-reels-transcript-scraper): Instagram Reels to text.
- [Video to Text Transcriber](https://apify.com/fguiraud/video-to-text-transcriber): any video file or social link, priced per minute (better for long videos).
- [YouTube Transcript Scraper](https://apify.com/fguiraud/youtube-transcript-scraper): YouTube videos, channels and playlists.

# Changelog

This Actor's version history is a separate document: https://apify.com/fguiraud/tiktok-transcript-scraper/changelog.md

# Actor input Schema

## `sources` (type: `array`):

Links to public TikTok videos (tiktok.com/@user/video/...). Instagram, X and Facebook video links and direct MP4 links work too. For whole accounts use 'TikTok accounts'.

## `profiles` (type: `array`):

TikTok accounts to transcribe, as '@username' or a profile link (https://www.tiktok.com/@username). The newest videos of each account are transcribed (see 'Videos per account'). Instagram profile links (https://www.instagram.com/username/) work too, up to their 12 newest reels.

## `maxVideosPerProfile` (type: `integer`):

How many of the newest videos to transcribe from each TikTok account.

## `onlyNewVideos` (type: `boolean`):

Skip TikTok videos already transcribed by previous runs of this Actor in your account (remembered in a key-value store named 'transcriber-state-tiktok'). Ideal for a weekly schedule: you only pay for new videos.

## `base64Files` (type: `array`):

Short media files without a URL, e.g. from an AI agent: \[{"fileName": "memo.m4a", "content": "<base64>"}]. Keep the total input under ~9 MB; use URLs for longer recordings.

## `model` (type: `string`):

'base': good accuracy, fast (recommended). 'small': best accuracy, especially for accents, noisy audio and non-English speech; slower and billed at a higher per-minute price. 'tiny': fastest draft quality.

## `language` (type: `string`):

ISO code of the spoken language (en, es, de, fr, pt, it, ja, zh, ...) or 'auto' to detect it. Setting it avoids misdetection on short clips.

## `task` (type: `string`):

'transcribe': text in the spoken language. 'translate': translate the speech to English text.

## `vocabulary` (type: `array`):

Words the speech recognition should favour: people and company names, product names, technical terms (e.g. 'Kubernetes', 'Dr. Nguyen', 'Apify'). Improves spelling of rare words.

## `outputs` (type: `array`):

'text': full transcript split into paragraphs at pauses. 'segments': timestamped segments. 'srt' / 'vtt': ready-to-use subtitle files. 'chunks': ~chunkSize-character passages with start/end times and a token estimate, ready for vector databases. 'markdown': paragraphs prefixed with their start time, e.g. '**\[00:01:23]** ...'.

## `subtitleMaxChars` (type: `integer`):

Characters per subtitle line in SRT/VTT (42 is the broadcast and YouTube standard; 32-37 for vertical video). 0 keeps Whisper's long raw segments.

## `subtitleMaxLines` (type: `integer`):

Maximum lines shown at once in each subtitle (when line length is set).

## `subtitleMaxDuration` (type: `number`):

Longest time a single subtitle stays on screen (when line length is set).

## `aiInsights` (type: `boolean`):

Analyse each transcript with Claude: title, summary, key points, chapters with start times, action items and topics (in 'insights'). Requires your Anthropic API key; Claude usage is billed to your Anthropic account, plus one small 'AI insights' event per file.

## `anthropicApiKey` (type: `string`):

Your key from console.anthropic.com. Stored as a secret input; used only to call Claude for this run.

## `insightsModel` (type: `string`):

'claude-opus-5': best quality (default). 'claude-sonnet-5': cheaper, great for meetings and podcasts. 'claude-haiku-4-5': cheapest (transcripts up to ~2 hours).

## `insightsInstructions` (type: `string`):

Optional, e.g. 'Summarise in Spanish', 'Focus on decisions and owners', 'Chapters every ~10 minutes'.

## `saveFiles` (type: `boolean`):

Save the transcript (.txt) and subtitles (.srt / .vtt, if selected in outputs) as files in the run's key-value store; the result includes their download links.

## `wordTimestamps` (type: `boolean`):

Add start/end times for every word inside each segment (for karaoke-style captions or precise search). Slightly slower.

## `chunkSize` (type: `integer`):

Target size of 'RAG chunks'.

## `skipSilence` (type: `boolean`):

Detect speech first and skip silent parts. Faster and reduces hallucinated text in long pauses.

## `maxDurationMinutes` (type: `integer`):

Only the first N minutes of each file are transcribed (and billed).

## `maxFileSizeMb` (type: `integer`):

Larger files are skipped (not billed).

## `failOnError` (type: `boolean`):

Mark the run as FAILED when a file cannot be transcribed. Useful for pipelines and monitoring.

## Actor input object example

```json
{
  "sources": [
    {
      "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422"
    }
  ],
  "maxVideosPerProfile": 10,
  "onlyNewVideos": true,
  "model": "base",
  "language": "auto",
  "task": "transcribe",
  "outputs": [
    "text"
  ],
  "subtitleMaxChars": 42,
  "subtitleMaxLines": 2,
  "subtitleMaxDuration": 6,
  "aiInsights": false,
  "insightsModel": "claude-opus-5",
  "saveFiles": true,
  "wordTimestamps": false,
  "chunkSize": 1000,
  "skipSilence": true,
  "maxDurationMinutes": 240,
  "maxFileSizeMb": 1000,
  "failOnError": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `files` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sources": [
        {
            "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422"
        }
    ],
    "outputs": [
        "text"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fguiraud/tiktok-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sources": [{ "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422" }],
    "outputs": ["text"],
}

# Run the Actor and wait for it to finish
run = client.actor("fguiraud/tiktok-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sources": [
    {
      "url": "https://www.tiktok.com/@tiktok/video/7689568478570843422"
    }
  ],
  "outputs": [
    "text"
  ]
}' |
apify call fguiraud/tiktok-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fguiraud/tiktok-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HoQvAwYHv41gvy2At/builds/WBU3C1Kv6kBvWUevO/openapi.json
