# TikTok Video Scraper: Post Stats, Transcripts (`automation_craft/tiktok-video-scraper`) Actor

Scrape TikTok videos by URL, one row per link in your order, no login: play, like, comment, share, save and repost counts, caption, hashtags, mentions, music, author stats at post time, content categories, related searches, media URLs with expiry and the ASR transcript inline. Pay per post.

- **URL**: https://apify.com/automation\_craft/tiktok-video-scraper.md
- **Developed by:** [Automation Craft](https://apify.com/automation_craft) (community)
- **Categories:** Social media, Marketing, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### TikTok Video Scraper: Post Stats, Transcripts

**TikTok Video Scraper**: paste TikTok video URLs, photo slideshow URLs, short links or post ids and get one row per link, in the order you gave them, with no login, no cookies and no API key. Each row carries the play, like, comment, share, save and repost counts, the caption with its hashtags and mentions, the music, the author and the author's follower and like counts at the time of the run, TikTok's own content categories and related searches, the media URLs with the time they expire, and the transcript of the video from TikTok's automatic caption track, inline and free.

You pay per post delivered. A link that could not be served gets a free status row in its place that says why (not found, private, blocked, invalid, duplicate, filtered), and the run summary is free.

### Quick start

1. Open the Actor and paste your links into **TikTok post URLs, short links or post ids**, one per line. Full `www.tiktok.com/@user/video/...` and `/photo/...` URLs, `vm.tiktok.com` and `vt.tiktok.com` short links, `m.tiktok.com/v/....html` links and bare post ids all work.
2. Leave **Include the transcript** on to get the spoken text of each video at no extra cost.
3. Set **Maximum posts** if you want a cap (0 means every link).
4. Run it. Download the **Posts** view as JSON, CSV or Excel, or read the dataset through the API. Row 0 is your first link, row 1 your second, and so on.
5. For a list you re-check on a schedule, give the run a **Memory name**: posts already delivered under that name are not fetched or charged again.

### What you get

Every row has `position` (the index of your link), `status` and `input` (your link as given). A delivered post (`type: "post"`, `status: "ok"`) has the fields below. Fill rates are measured on the Apify platform on 174 live posts from brand and creator accounts:

| Field | Filled in the 174 post sample (run 98WgGaP12B6r8TfZf) |
|---|---|
| `playCount`, `diggCount`, `commentCount`, `shareCount`, `collectCount` | 174 of 174 |
| `authorMeta.fans`, `authorMeta.heart` (author at post time) | 174 of 174 |
| `categoryType` | 174 of 174 |
| `textLanguage` | 174 of 174 |
| `mediaExpiresAt` | 174 of 174 |
| `commentsEnabled` | 174 of 174 |
| `videoMeta.downloadAddr`, `vqScore`, `loudness`, `dynamicCoverUrl` | 158 of 158 videos |
| `musicMeta.playUrl` | 164 of 174 |
| `suggestedWords` (related searches) | 145 of 174 |
| `diversificationLabels` (content categories) | 140 of 174 |
| `hashtags` | 111 of 174 |
| `transcript` | 82 of 158 videos |
| `mentions` | 47 of 174 |
| `musicMeta.dspLinks` | 54 of 174 |
| `isAd` true | 27 of 174 |
| `poi` (tagged place) | 11 of 174 |

| Field | What it is |
|---|---|
| `id`, `webVideoUrl`, `resolvedUrl` | TikTok's post id (a string), the canonical URL (`/photo/` for slideshows) and, for short links, where the link led |
| `text`, `textLanguage`, `hashtags`, `mentions`, `detailedMentions` | caption, its language, hashtags with ids, and mentioned accounts with their ids |
| `createTime`, `createTimeISO` | when the post was published |
| `playCount`, `diggCount`, `commentCount`, `shareCount`, `collectCount`, `repostCount` | engagement counts from TikTok's statsV2 block (see the note below) |
| `authorName`, `authorMeta` | the author as TikTok serves it with the post (never taken from your URL), with follower, following, like and video counts, verified, bio, bio link and account settings |
| `musicMeta` | song id, title, artist, original or not, copyright and commerce flags, the audio URL and streaming links when TikTok has them |
| `videoMeta` | duration, size, cover URLs, `downloadAddr`, the largest rendition (or every rendition), TikTok's video quality score, loudness, caption track links and TikTok's caption flags |
| `imagePost`, `slideshowImageLinks`, `isSlideshow` | for photo posts: every image with its size, a second host and its expiry |
| `diversificationLabels`, `categoryType`, `suggestedWords` | TikTok's content categories and the related searches TikTok shows under the post |
| `isAd`, `isAigc`, `commentsEnabled`, `duetEnabled`, `stitchEnabled`, `penaltyStatus`, `warnInfo` | flags TikTok sets on the post |
| `transcript`, `transcriptSegments`, `transcriptLanguage`, `transcriptSource`, `transcriptReason` | the caption track as plain text and as timed segments in seconds, or the reason there is none |
| `mediaUrls`, `mediaExpiresAt` | the signed media URLs and the earliest moment one of them stops working |

**About the counts.** On its web page TikTok prints play, like, comment and share counts to 4 significant digits once they reach 10,000 (212,000, not 212,113); below 10,000 they are exact. Saves (`collectCount`) are exact at every size. The Actor returns TikTok's figures as numbers and `null`, never 0, when TikTok does not give one.

**About the transcript.** It is TikTok's own automatic speech to text track (`transcriptSource: "ASR"`), downloaded and returned inline. TikTok had a track for 82 of 158 videos (52 percent) in the sample above and for 13 of 36 videos (36 percent) in a sample of hashtag feed posts. Every other post says why in `transcriptReason`: `no_captions` (TikTok made no track, 71 videos in the sample, mostly music without speech), `captions_disabled` (captions switched off, 5 videos, with TikTok's own code in `transcriptReasonDetail`), `slideshow` (photo posts have no speech track), `download_failed` or `disabled_by_input`. Set **Transcript language** to prefer a language when TikTok offers more than one track.

**Status rows** (`type: "status"`, free) sit in the same place as the link they answer:

| `status` | Meaning |
|---|---|
| `not_found` | TikTok says the post does not exist. TikTok answers a removed post and an id that never existed the same way, so the message says "removed or never existed". An unknown or expired short link also ends here; a short link TikTok refused on both the platform IP and the proxy is a `blocked` row instead. |
| `private`, `removed`, `unavailable` | TikTok answered with another code; `statusCode` and `statusMsg` echo it. A backend timeout on TikTok's side (statusCode 100004, seen once in 174 proxied loads) is fetched twice more before it becomes `unavailable` |
| `blocked` | TikTok did not serve the page after every retry on fresh exits |
| `invalid` | not a TikTok post link or id (a profile or hashtag URL, for example) |
| `duplicate` | the same post appeared earlier in your list; `duplicateOf` gives that position |
| `already_known` | delivered before under your memory name |
| `filtered` | removed by a free filter (`excludeAds`, `slideshowsOnly`) |
| `skipped` | not fetched because the run stopped first (`maxPosts` reached or your spending limit); `stoppedBy` says which |
| `failed` | an unexpected error; the message says what |

The last row is the run summary (`type: "summary"`) with counts per status, transcript outcomes, request and retry counts, the stop reason and billing counters.

### How much does it cost to scrape TikTok videos?

| Event | FREE and BRONZE | SILVER | GOLD | PLATINUM | DIAMOND | When it is charged |
|---|---|---|---|---|---|---|
| Post | $0.50 per 1,000 | $0.45 | $0.40 | $0.40 | $0.28 | once per post delivered with status ok |
| Actor start | $0.002 per run per GB of memory | same | same | same | same | once per run; the default 512 MB run pays it once |

The Store card shows "$0.50 / 1,000" on the Post row: one post costs a twentieth of a cent. A list of 1,000 links costs $0.50 plus $0.002 for the start. The transcript, the filters, the memory, status rows and the summary cost nothing. Measured on Apify, a run of 174 posts took 44 seconds, about 0.25 seconds per post with 4 parallel requests (about 230 posts per minute).

### Input

| Field | Type | Notes |
|---|---|---|
| `postUrls` | array | post URLs, short links or post ids; one row per value in your order |
| `maxPosts` | integer | stop after this many posts; 0 = no cap |
| `includeTranscript` | boolean | the inline transcript, free, on by default |
| `transcriptLanguage` | string | preferred caption language such as `en` or `es` |
| `includeRenditions` | boolean | every video rendition instead of the largest one |
| `excludeAds`, `slideshowsOnly` | boolean | free filters |
| `memoryName`, `resetMemory` | string, boolean | pay once per post across scheduled runs |
| `tryDirectFirst` | boolean | first request from the Apify platform IP, retries through the proxy |
| `proxyConfiguration` | object | the retry proxy; Apify datacenter proxy by default |
| `maxConcurrency` | integer | 1 to 8 parallel requests, 4 by default |

```json
{
  "postUrls": [
    "https://www.tiktok.com/@nike/video/7689162344139640077",
    "https://www.tiktok.com/@nike/video/7401458428457127211"
  ],
  "maxPosts": 10,
  "includeTranscript": true
}
```

On the Apify platform the first request for each post went straight from the platform IP and was served on the first try for 174 of 174 in each of five runs; with every request through the datacenter proxy 508 of 522 (three runs of 174) were served on the first try and the runs took 58 to 160 seconds instead of 44. No residential proxy is used.

### FAQ

#### Do I need a TikTok account, cookies or an API key?

No. Every post page is read the way a signed out visitor reads it, over plain HTTP: the first request goes from the Apify platform IP and any retry goes through a datacenter proxy. Nothing in the input asks for a credential.

#### Does it keep the order of my URLs?

Yes. The dataset carries exactly one row per URL in the order you gave it, with a position field, and a URL that could not be served gets a free status row in the same place saying why (not found, removed, private, blocked or invalid).

#### Does it return the transcript of a TikTok video?

When TikTok has an automatic caption track for the post, the Actor downloads it and returns the text inline plus timed segments, at no extra charge. Posts without a track carry a transcriptReason instead of an empty string.

#### Does it return comments?

No. Comment text is served only after a slide puzzle and a login on the web, so this Actor returns the comment count and whether comments are enabled, not the comment bodies.

#### Are the view and like counts exact?

Saves are exact at every size. Plays, likes, comments and shares are exact below 10,000; above that TikTok's page gives them to 4 significant digits (212,000) and the Actor returns TikTok's figure. The author's follower and like counts on a post are rounded the same way.

#### Can I download the video or the cover?

The row carries the signed media URLs TikTok serves (video, covers, music) and a mediaExpiresAt timestamp parsed from them. Fetch what you want to keep before that time; the Actor does not store files.

#### Why does this Actor run with limited permissions?

Least privilege. It reads and writes only its own run storages, and the optional cross run memory is a named key value store it creates itself.

### What this Actor does NOT do

- It does not return comment text. TikTok serves comments only after a slide puzzle and a login on the web; rows carry `commentCount` and `commentsEnabled` instead.
- It does not list a profile's posts or a hashtag's posts. Give it post links; the sibling Actors below cover profiles and hashtags.
- It does not store videos or images. Media URLs are signed by TikTok and expire, about 6 hours after the run for video and music and about 47 hours for slideshow images; `mediaExpiresAt` gives the exact time. Download what you want to keep before then.
- It does not transcribe audio itself. The transcript is TikTok's own caption track, so a post without one has no transcript.
- It does not tell a removed post from one that never existed: TikTok answers both the same way.

### API example

```bash
curl -X POST "https://api.apify.com/v2/acts/automation_craft~tiktok-video-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"postUrls":["https://www.tiktok.com/@nike/video/7689162344139640077"],"includeTranscript":true}'
```

Rows are about 11 KB each; the largest in the sample was 21 KB (a 9 image slideshow).

### More data tools by Automation Craft

- [TikTok Profile Scraper: Exact Stats, No Login](https://apify.com/automation_craft/tiktok-profile-scraper)
- [TikTok Hashtag Scraper: Posts, Free Date Filter](https://apify.com/automation_craft/tiktok-hashtag-scraper)
- [TikTok Ads Library Scraper: EU Ads, No Login](https://apify.com/automation_craft/tiktok-ads-library-scraper)
- [Google Trends Scraper - Compare and Trending Now](https://apify.com/automation_craft/google-trends-scraper)
- [Meta Ad Library Scraper - All Placements, Filters](https://apify.com/automation_craft/meta-ads-library-scraper)
- [Website Screenshot Scraper: Full Page, Bulk](https://apify.com/automation_craft/website-screenshot-scraper)

# Actor input Schema

## `postUrls` (type: `array`):

One value per line. Accepts www.tiktok.com/@user/video/ID and /photo/ID URLs, vm.tiktok.com and vt.tiktok.com short links, m.tiktok.com/v/ID.html links and bare numeric post ids. Every value gets exactly one row, in the order given; duplicates, profile URLs and anything that is not a post get a free status row.

## `maxPosts` (type: `integer`):

Stop after this many posts are delivered. Values after the cap are not fetched and get a free status row with status skipped. 0 means no cap.

## `includeRenditions` (type: `boolean`):

Free. When on, videoMeta.renditions lists every rendition TikTok serves (quality, bitrate, codec, size and URL). When off, only the largest rendition is kept so rows stay small; videoMeta.renditionCount still says how many exist.

## `includeTranscript` (type: `boolean`):

Free. Downloads the automatic caption track TikTok serves with the post and returns the text inline plus timed segments. Posts without a track get a transcriptReason instead (no\_captions, captions\_disabled, slideshow, download\_failed).

## `transcriptLanguage` (type: `string`):

Optional language code such as en, es or pt. When the post has a caption track in that language it is used; otherwise the original language track is returned and transcriptLanguageRequested records what you asked for. Leave empty for the original language.

## `excludeAds` (type: `boolean`):

Free filter. Posts TikTok marks as ads get a free status row with status filtered instead of a paid row.

## `slideshowsOnly` (type: `boolean`):

Free filter. Deliver only photo slideshow posts; videos get a free status row with status filtered.

## `memoryName` (type: `string`):

Optional. A name for a memory kept between runs: a post already delivered under this name is not fetched again and gets a free status row with status already\_known. If the memory cannot be opened or saved, posts are delivered free and a status row says so.

## `resetMemory` (type: `boolean`):

Forget every post stored under the memory name before this run starts.

## `tryDirectFirst` (type: `boolean`):

When on, the first request for each post goes straight from the Apify platform and every retry goes through the proxy below on a fresh session. Measured on Apify: the platform IP answered every post page on the first try and twice as fast. Turn off to send every request through the proxy.

## `proxyConfiguration` (type: `object`):

The proxy used for retries (and for every request when the option above is off). Apify datacenter proxy by default; no residential proxy is needed.

## `maxConcurrency` (type: `integer`):

How many posts are fetched at the same time, 1 to 8. Rows are still written in your order.

## Actor input object example

```json
{
  "postUrls": [
    "https://www.tiktok.com/@nike/video/7689162344139640077",
    "https://www.tiktok.com/@nike/video/7401458428457127211"
  ],
  "maxPosts": 10,
  "includeTranscript": true,
  "transcriptLanguage": "en",
  "memoryName": "weekly-campaign-posts",
  "tryDirectFirst": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "BUYPROXIES94952"
    ]
  },
  "maxConcurrency": 4
}
```

# Actor output Schema

## `rows` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "postUrls": [
        "https://www.tiktok.com/@nike/video/7689162344139640077",
        "https://www.tiktok.com/@nike/video/7401458428457127211"
    ],
    "maxPosts": 10,
    "includeRenditions": false,
    "includeTranscript": true,
    "excludeAds": false,
    "slideshowsOnly": false,
    "resetMemory": false,
    "tryDirectFirst": true,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "BUYPROXIES94952"
        ]
    },
    "maxConcurrency": 4
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation_craft/tiktok-video-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "postUrls": [
        "https://www.tiktok.com/@nike/video/7689162344139640077",
        "https://www.tiktok.com/@nike/video/7401458428457127211",
    ],
    "maxPosts": 10,
    "includeRenditions": False,
    "includeTranscript": True,
    "excludeAds": False,
    "slideshowsOnly": False,
    "resetMemory": False,
    "tryDirectFirst": True,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["BUYPROXIES94952"],
    },
    "maxConcurrency": 4,
}

# Run the Actor and wait for it to finish
run = client.actor("automation_craft/tiktok-video-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "postUrls": [
    "https://www.tiktok.com/@nike/video/7689162344139640077",
    "https://www.tiktok.com/@nike/video/7401458428457127211"
  ],
  "maxPosts": 10,
  "includeRenditions": false,
  "includeTranscript": true,
  "excludeAds": false,
  "slideshowsOnly": false,
  "resetMemory": false,
  "tryDirectFirst": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "BUYPROXIES94952"
    ]
  },
  "maxConcurrency": 4
}' |
apify call automation_craft/tiktok-video-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation_craft/tiktok-video-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QnyS5qA8xBI8LorwN/builds/QOvDSfDAX86rF5FPx/openapi.json
