# TikTok Transcript Scraper (`tokfluence/tiktok-transcript-scraper`) Actor

Get plain-text transcripts of TikTok videos by URL, taken from TikTok's own subtitles. One row per video that has subtitles; videos without subtitles return no transcript, and there is no AI speech-to-text.

- **URL**: https://apify.com/tokfluence/tiktok-transcript-scraper.md
- **Developed by:** [Tokfluence Tiktok API](https://apify.com/tokfluence) (community)
- **Categories:** Social media, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.90 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## TikTok Transcript Scraper

Give it TikTok video URLs and it returns the plain-text transcript of each video that has TikTok subtitles, one row per video. Rows use the field names of Clockworks' TikTok transcript row where we have the data; the transcript text itself sits under `tokfluence.transcript` (see below for what that means for an existing pipeline).

### Read this first

- **Only videos with TikTok subtitles have a transcript.** Some TikTok videos carry none. Those have no transcript here, and no run can change that.
- **There is no AI transcription.** We do not run speech-to-text on the audio. A video without TikTok subtitles gives no row, not an empty or guessed transcript, and you are not charged for it. The run's status message names those videos.
- **Plain text, one language.** The transcript is one block of text taken from one of the video's subtitle tracks. Segments, timestamps and subtitle files (VTT/SRT) are not kept, so the row has no timing fields.

### What it returns

One row per video that has a transcript. Filled fields:

- `id`: the TikTok video id
- `webVideoUrl`: the video's URL
- `submittedVideoUrl`: the URL as you gave it in the input, so you can match rows back to your list

Fields Tokfluence adds, under a `tokfluence` object:

- `tokfluence.transcript`: the transcript as plain text
- `tokfluence.language`: the subtitle language as TikTok labels it, for example `eng-US`
- `tokfluence.source`: where the text came from; today always `TIKTOK_AUTO_CAPTION`, TikTok's own captions
- `tokfluence.summary`: a short summary of the transcript written by a language model, when Tokfluence has made one, otherwise null

**If you are moving a pipeline from another transcript actor:** Clockworks' transcript row does not carry the text in the row. It links to files, in `videoMeta.subtitleLinks[].downloadLink` (timed subtitles) and `videoMeta.transcriptionLink` (AI transcripts). We have no such files, so both are null here, and your pipeline should read `tokfluence.transcript` instead of downloading a link.

### Input

- `postURLs`: TikTok video URLs in the full `https://www.tiktok.com/@user/video/<id>` form. Short `vm.tiktok.com` links are not accepted and are named in the status message. Up to 50 per run.
- `maxResults`: stops once this many rows are in the dataset.
- `mode`, under Advanced: see below.

There is no option to transcribe videos without subtitles, because we do not do that.

### Fresh or fast: the Mode setting

Under **Advanced**, `mode` trades freshness for speed:

- `auto` (default): serves the transcript Tokfluence already stores, and scrapes TikTok live for videos we do not hold or hold without a transcript. Most runs want this.
- `database`: fastest. Never scrapes, so any video we have not captured, or captured without a transcript, is missing from the results.
- `live`: always scrapes TikTok now, and slower. Tokfluence keeps the first transcript it captures for a video, so for a video we already hold a transcript for, `live` returns the same text as `auto`; `auto` already goes live for videos we hold without one.

Live requests run on Tokfluence's shared scraping workers, one video after another, and wait behind other live requests, so a live run can take a while.

If a live scrape outlasts the run's timeout, the run fails with the request id. The scrape keeps going and is stored when it finishes, so a `database` run shortly afterwards returns those transcripts without scraping again.

### Fields that can be null

This actor never fills a gap itself: when we do not have a field it is `null`, not `0`, `false` or an empty string. Always null in this actor:

- `videoMeta.subtitleLinks` and `videoMeta.transcriptionLink`: we store plain text, not subtitle or transcript files
- `text`: in Clockworks' row this is the video caption, not the transcript; the transcript response does not include the caption
- `mediaUrls`: we do not download video files
- `url`, `error`, `errorCode`, `invalidUrls`: videos we could not serve are reported in the run's status message and log rather than as error rows

Often null:

- `tokfluence.summary`: only some transcripts have a summary

### Zero or short results

When a run returns fewer rows than the videos you gave it, it still succeeds, but it logs a warning and sets the run's status message to the likely cause, naming the videos:

- videos with no TikTok subtitles, so no transcript exists (not charged); in `database` mode, videos we hold without a transcript yet
- URLs we could not find, or could not read (for example short links)
- `database` mode, which never scrapes

The actor uses one Tokfluence account for every run. If that account is out of API credits, or Tokfluence's scrape service is down, the run fails with a message saying so. You are charged only for rows already in the dataset.

### What Tokfluence keeps from a run

Tokfluence logs every request the actor makes: the URLs, the mode, where each row was served from, and a reference made of a one-way hash of your Apify user id plus the run id. Every TikTok handle in your URLs and in the returned rows is recorded as a candidate for Tokfluence's creator discovery. Every video scraped live is stored in Tokfluence like any other scrape, and later runs, anyone's, can be served that stored copy.

### Memory

The actor defaults to 256 MB and allows up to 512 MB. It makes one API call, maps the rows and writes them to the dataset in batches of 100.

### Pricing

PLACEHOLDER (Daniel): pay per result, one `result` event per dataset row. Price to be set in Console at publish.

Only videos with a transcript become rows, so only they are charged. If your plan's remaining budget covers fewer rows than you asked for, the actor collects up to that ceiling, says so in the log, and stops cleanly rather than returning a silently short list.

### Notes

Transcripts come from subtitles TikTok publishes on public videos, captured and stored by Tokfluence. You are the data controller for anything you export; follow GDPR and any local rules that apply to you. More at [Tokfluence](https://tokfluence.com).

# Actor input Schema

## `postURLs` (type: `array`):

TikTok video URLs to get transcripts for, in the full https://www.tiktok.com/@user/video/<id> form (short vm.tiktok.com links are not accepted). Up to 50 per run. Only videos with TikTok subtitles have a transcript.

## `maxResults` (type: `integer`):

Stops once this many rows are in the dataset. Each row is one video's transcript.

## `mode` (type: `string`):

Trades freshness for speed. Auto serves what Tokfluence already stores when it is fresh enough and scrapes live for the rest. Database only is fastest and never scrapes, so anything we do not already hold is missing from the results. Live always scrapes TikTok now: freshest, and slower.

## Actor input object example

```json
{
  "postURLs": [
    "https://www.tiktok.com/@apifylife/video/7338085038258457889"
  ],
  "maxResults": 100,
  "mode": "auto"
}
```

# Actor output Schema

## `results` (type: `string`):

Rows written to the run's dataset, one per video with a transcript, in Clockworks field names.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "postURLs": [
        "https://www.tiktok.com/@apifylife/video/7338085038258457889"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tokfluence/tiktok-transcript-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "postURLs": ["https://www.tiktok.com/@apifylife/video/7338085038258457889"] }

# Run the Actor and wait for it to finish
run = client.actor("tokfluence/tiktok-transcript-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "postURLs": [
    "https://www.tiktok.com/@apifylife/video/7338085038258457889"
  ]
}' |
apify call tokfluence/tiktok-transcript-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tokfluence/tiktok-transcript-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/EFo1jYwHQZIYZff6e/builds/gQKsI9v2Seqq9Bfgr/openapi.json
