# Douyin (TikTok China) Scraper — Videos, Creators, Comments (`edgy_dock/douyin-tiktok-cn-scraper`) Actor

Search Douyin by keyword and export videos, creator profiles and comments.

- **URL**: https://apify.com/edgy\_dock/douyin-tiktok-cn-scraper.md
- **Developed by:** [wangyan](https://apify.com/edgy_dock) (community)
- **Categories:** Social media, E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.005 / actor start

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Douyin (TikTok China) Scraper — Videos, Creators & Comments

Export Douyin data as JSON, CSV or Excel. Search by keyword, pull full video details, profile
creators, collect comments — without an API key, a Chinese business entity, or an agency.

Douyin has **700M+ daily active users** and is where Chinese consumers discover products,
music, food trends and entertainment. It is also almost completely closed to outsiders: no
public API, cryptographically-signed web endpoints, and the official Open Platform requires a
mainland Chinese company. This Actor gives you the data layer without any of that.

***

### Quick start (3 minutes)

1. **Grab a cookie** — open <https://www.douyin.com/> in your browser (log in for best
   results), press `F12` → Console, run `copy(document.cookie)`.
2. **Click "Try for free"** at the top of this page (or paste the cookie into the `cookie`
   field of a run).
3. **Hit "Start"** — results appear in the Dataset tab within ~45 seconds.

Minimal input that works:

```json
{
  "searchKeywords": ["美食"],
  "cookie": "ttwid=...; sessionid_ss=...;",
  "maxItems": 20
}
```

***

### What you can do with it

| You are | You use it to |
|---|---|
| **A brand entering China** | Measure share of voice, find which product claims resonate, track competitors |
| **An agency** | Build KOL / influencer shortlists from real engagement numbers, not media kits |
| **A dropshipper / sourcing operator** | See what is trending in China 3–6 months before it reaches Western marketplaces |
| **A music / entertainment team** | Spot breakout songs, hashtags, dance challenges by region |
| **A trend / AI team** | Feed a genuinely hard-to-obtain Chinese-language video dataset into your models |

***

### Three ways to use it

#### 1 · Search by keyword

```json
{
  "searchKeywords": ["火锅", "宠物"],
  "cookie": "...",
  "maxItems": 200,
  "sort": "most_liked",
  "searchType": "all",
  "fetchVideoDetails": true,
  "maxCommentsPerVideo": 20
}
```

#### 2 · Deep-scrape specific videos

```json
{
  "videoUrls": [
    "https://www.douyin.com/video/7421123456789012345"
  ],
  "cookie": "...",
  "maxCommentsPerVideo": 100
}
```

#### 3 · Pull a creator's recent posts

```json
{
  "creatorUrls": [
    "https://www.douyin.com/user/MS4wLjABAAAAyourSecUidHere"
  ],
  "cookie": "...",
  "maxPostsPerCreator": 50
}
```

The user URL path is a `sec_uid` (base64-url string starting with `MS4wLjABAAAA`), **not** the
pretty name you may see on the profile page.

***

### Input reference

| Field | Default | Notes |
|---|---|---|
| `searchKeywords` | `[]` | Chinese works ×10 better than English |
| `videoUrls` | `[]` | `https://www.douyin.com/video/<aweme_id>` |
| `creatorUrls` | `[]` | `https://www.douyin.com/user/<sec_uid>` |
| `cookie` | `""` | **Required in practice.** See [Cookie walkthrough](#cookie-walkthrough) |
| `maxItems` | `100` | Hard cap. You only pay for records actually delivered |
| `sort` | `general` | `general` / `latest` / `most_liked` |
| `searchType` | `all` | `all` / `video` / `user` |
| `fetchVideoDetails` | `false` | Open each video for full desc + music + IP + stats |
| `maxCommentsPerVideo` | `0` | Set > 0 to also pull comments (charged separately) |
| `maxPostsPerCreator` | `20` | Recent posts to collect per creator |
| `proxyConfiguration` | `{useApifyProxy: true, apifyProxyGroups: ["RESIDENTIAL"]}` | **CN proxy strongly recommended** |

***

### Cookie walkthrough

```
1. Open https://www.douyin.com/ in Chrome / Edge / Firefox
2. Log in — scan the QR with the Douyin mobile app (strongly recommended)
3. Press F12 → Developer Tools opens
4. Click the "Console" tab
5. Paste this and press Enter:
      copy(document.cookie)
6. Clipboard now holds the cookie string
7. Paste into the "cookie" field of the Actor input
```

The cookie looks like:

```
ttwid=xxx; odin_tt=xxx; sessionid_ss=xxx; msToken=xxx; s_v_web_id=xxx; ...
```

**Critical cookies:** `ttwid` and `sessionid_ss`. If either is missing, log in again.

**Cookies die fast on Douyin.** Expect to refresh every 1–3 days under heavy use. The Actor
logs a warning when the cookie looks poisoned.

***

### CN proxy — required in practice

Douyin blocks or heavily throttles most non-CN IPs. **The Actor usually returns zero without
a CN residential proxy.**

```json
"proxyConfiguration": {
  "useApifyProxy": true,
  "apifyProxyGroups": ["RESIDENTIAL"],
  "apifyProxyCountry": "CN"
}
```

Bring-your-own works too:

```json
"proxyConfiguration": {
  "proxyUrls": ["http://user:pass@cn-proxy.example.com:8080"]
}
```

***

### Output

Each record is a flat JSON object. Chinese counts like `1.2w` / `3.4万` / `2亿` are normalised
to integers.

#### Video (from search)

```json
{
  "source": "search",
  "searchKeyword": "美食",
  "awemeId": "7421123456789012345",
  "awemeType": 0,
  "title": "撸猫日常",
  "description": "一只小猫咪在玩",
  "awemeUrl": "https://www.douyin.com/video/7421123456789012345",
  "coverUrl": "https://p.douyin.com/cover.jpg",
  "videoUrl": "https://aweme.snssdk.com/video.mp4",
  "duration": 15,
  "diggCount": 12000,
  "commentCount": 350,
  "shareCount": 42,
  "collectCount": 12000,
  "playCount": 500000,
  "createdAt": 1727500000,
  "authorId": "9876543210",
  "authorSecUid": "MS4wLjABAAAA...",
  "authorNickname": "撸猫小队长",
  "authorAvatar": "...",
  "authorUrl": "https://www.douyin.com/user/MS4wLjABAAAA..."
}
```

#### Video detail (with `fetchVideoDetails: true`)

Adds `ipLocation`, `downloadCount`, `imageUrls[]` (photo-mode posts), `tags[]`, and
`music.{id,title,author}`.

#### Creator profile

```json
{
  "source": "creator",
  "secUid": "MS4wLjABAAAA...",
  "userId": "9876543210",
  "nickname": "撸猫小队长",
  "douyinId": "catmaster",
  "bio": "每天一只猫",
  "gender": 2,
  "ipLocation": "北京",
  "followingCount": 128,
  "followerCount": 125000,
  "totalFavorited": 34000,
  "awemeCount": 342,
  "verified": true,
  "verifyReason": "宠物博主"
}
```

#### Comment

```json
{
  "source": "comment",
  "awemeId": "7421...",
  "commentId": "7421000000000000001",
  "content": "太可爱了",
  "diggCount": 11000,
  "replyCount": 23,
  "createdAt": 1727500100,
  "ipLocation": "上海",
  "authorId": "111",
  "authorSecUid": "MS4wLjABAAAA...",
  "authorNickname": "路人甲"
}
```

***

### Pricing

Pay-per-event — you are only charged for records actually delivered.

| Event | Price (USD) | What triggers it |
|---|---|---|
| `apify-actor-start` | $0.005 | Once per run (per GB RAM, min 1) |
| `video-scraped` | $0.0015 | Each video delivered |
| `creator-scraped` | $0.004 | Each creator profile |
| `comment-scraped` | $0.0004 | Each comment |

#### Cost estimator

| Job | Setting | Approx cost |
|---|---|---|
| Quick keyword scan | 100 videos, no details | **$0.16** |
| Deep hashtag sweep | 1 000 videos + details + 20 comments | **$9.50** |
| 10-KOL shortlist | 10 creators × 50 posts each | **$0.80** |
| Daily trend monitor | 500 videos/day for 30 days | **$22.50 / month** |

Platform margin: Apify takes 20%, you keep 80%. Prices above are what you pay.

***

### Common errors

| Error / symptom | Cause | Fix |
|---|---|---|
| `No videos captured for 'X'` | Cookie expired or no CN proxy | Re-copy cookie + enable CN residential proxy |
| Zero results on every run | IP blocked by Douyin | Switch Apify proxy group, wait 10 min |
| `BrowserType.launch` timeout | Memory too low | Raise `minMemoryMbytes` to 4096 |
| Login page in screenshots | Cookie wasn't logged-in | Scan QR with Douyin mobile app, re-copy |
| `aweme_id` not extracted from URL | Wrong URL format | Must be `/video/<numeric_id>`, not shortened share link |
| Only comments, no videos | `maxItems` consumed by comments | Separate the runs or raise `maxItems` |

***

### Why this exists

Douyin signs every web API call with `X-Bogus` and `msToken` headers derived from obfuscated
browser JS that ByteDance updates without notice. Re-implementing is a losing maintenance
race.

Instead the Actor opens a real Chromium page in the Apify runtime, lets the page's own JS
make the real signed calls, and reads JSON off the wire. ByteDance can rotate signatures all
they want — the browser catches up for us.

***

### Known limits

- **Cookies expire within days.** Users are expected to refresh their own cookie.
- **Search depth is capped by Douyin's own paging** — a few hundred videos per keyword.
- **English keywords mostly return nothing.** The platform is Chinese-first.
- **Non-CN IPs are usually blocked.** CN residential proxy is practically required.
- **This is a scraper, not a CRM.** Data is read-only, best-effort.

***

### Changelog

- **0.1.1** · Initial public release

# Actor input Schema

## `searchKeywords` (type: `array`):

Keywords to search on Douyin. Chinese keywords return far more results than English — try 美食 instead of food.

## `videoUrls` (type: `array`):

Individual Douyin video URLs to scrape in full detail, e.g. https://www.douyin.com/video/7421123456789012345

## `creatorUrls` (type: `array`):

Creator profile URLs, e.g. https://www.douyin.com/user/MS4wLjABAAAA...

## `cookie` (type: `string`):

Required for reliable results. Open douyin.com in your browser, press F12 → Console → type document.cookie and paste the result here. See the README for a 30-second walkthrough.

## `maxItems` (type: `integer`):

Hard cap on records written to the dataset. You are only charged for records actually delivered.

## `sort` (type: `string`):

How search results are ordered.

## `searchType` (type: `string`):

Filter search results by content type.

## `fetchVideoDetails` (type: `boolean`):

Open every search hit to collect full description, music, IP location and more stats. Slower and uses more compute but returns much richer records.

## `maxCommentsPerVideo` (type: `integer`):

Set above 0 to also collect comments. Comments are charged separately and count towards Max items.

## `maxPostsPerCreator` (type: `integer`):

How many recent posts to collect from each creator profile.

## `proxyConfiguration` (type: `object`):

Douyin blocks most non-CN IPs. A CN residential proxy is strongly recommended — the Actor will likely return nothing without one.

## Actor input object example

```json
{
  "searchKeywords": [
    "美食",
    "穿搭",
    "宠物"
  ],
  "maxItems": 100,
  "sort": "general",
  "searchType": "all",
  "fetchVideoDetails": false,
  "maxCommentsPerVideo": 0,
  "maxPostsPerCreator": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `datasetUrl` (type: `string`):

Browse the full scraped dataset in Apify Console.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchKeywords": [
        "美食"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("edgy_dock/douyin-tiktok-cn-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchKeywords": ["美食"] }

# Run the Actor and wait for it to finish
run = client.actor("edgy_dock/douyin-tiktok-cn-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchKeywords": [
    "美食"
  ]
}' |
apify call edgy_dock/douyin-tiktok-cn-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,edgy_dock/douyin-tiktok-cn-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/F45h13DhVqI699PeT/builds/cEY4SdW7BM87Jut7h/openapi.json
