# Weibo Scraper: Posts, Search, User Profiles & Comments (`changefeeds/weibo-scraper`) Actor

Export Weibo posts by keyword, a user's profile and posts, single posts by URL, and their comments, from the logged-out m.weibo.cn and weibo.com web APIs. Full text for long posts, pictures, video URLs. No Weibo login or cookies needed.

- **URL**: https://apify.com/changefeeds/weibo-scraper.md
- **Developed by:** [Changefeeds Tools](https://apify.com/changefeeds) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$4.00 / 1,000 item returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Weibo Scraper: Posts, Search, User Profiles & Comments

A Weibo scraper from Changefeeds Tools. Search Weibo (新浪微博) by keyword,
export a user's profile and posts, fetch single posts by URL, and pull the
comments on any of them. It reads the same `m.weibo.cn` and `weibo.com` web
APIs a logged-out browser uses, with the anonymous visitor cookie Weibo hands
every visitor. No Weibo account, no cookies from you, no login.

Who it's for: brand and market researchers tracking Chinese social media,
analysts watching a public figure's or company's account, and anyone who needs
Weibo posts as clean JSON without maintaining a Weibo client of their own.

### What you get

- **Posts** with full text (long posts that Weibo truncates are expanded with
  one extra request), the original HTML, repost/comment/like counts, picture
  URLs, the video URL and cover, the author, a summary of the reposted post,
  the posting client and the region Weibo shows ("发布于 上海").
- **User profiles** with exact follower/following/post counts, verification
  and its reason, gender, location, description and avatar.
- **Comments** (top-level, Weibo's hot order) with author, likes, reply count,
  floor number and region.

### Input

You can mix the three kinds of input in one run.

| Field | Type | Default | What it does |
| --- | --- | --- | --- |
| `searchKeywords` | list of strings | none | Each keyword returns up to `maxItems` posts from Weibo search. |
| `users` | list of strings | none | Each user returns one profile row plus up to `maxItems` of their posts. Accepts `1669879400`, `weibo.com/u/<uid>`, `weibo.com/<uid>`, `weibo.com/<custom domain>`, `weibo.com/n/<screen name>`, `m.weibo.cn/u/<uid>`, `m.weibo.cn/profile/<uid>`, `@<screen name>`. |
| `postUrls` | list of strings | none | Each returns one post row. Accepts `weibo.com/<uid>/<bid>`, `m.weibo.cn/detail/<mid>`, `m.weibo.cn/status/<bid>`, `weibo.com/detail/<mid>`, or a bare mid (`5348645916642028`) or bid (`Rkpv9rRTu`). The two id forms are converted into each other for you. |
| `includeComments` | boolean | `false` | Also return comments for every post the run returns. |
| `maxItems` | integer | `50` | Max posts per keyword and per user (1 to 1,000). |
| `maxComments` | integer | `20` | Max comments per post (1 to 1,000). |
| `proxy` | proxy settings | Apify residential | See Limits. Proxy traffic is included in the price. |

```json
{ "users": ["1669879400"], "maxItems": 10 }
```

```json
{ "searchKeywords": ["华为"], "maxItems": 100 }
```

```json
{ "postUrls": ["https://weibo.com/1669879400/Rkpv9rRTu"], "includeComments": true, "maxComments": 45 }
```

### Sample output

From a live run on 2026-09-30 (long text and URLs shortened here). A `post`
row:

```json
{
  "type": "post",
  "id": "5348645916642028",
  "mid": "5348645916642028",
  "bid": "Rkpv9rRTu",
  "url": "https://weibo.com/1669879400/Rkpv9rRTu",
  "createdAt": "2026-09-29T16:40:22.000Z",
  "text": "#烈士纪念日#，铭记历史，缅怀先烈，吾辈自强！",
  "textHtml": "<a  href=\"https://m.weibo.cn/search?containerid=231522type...",
  "isLongText": false,
  "fullTextFetched": false,
  "reposts": 11291,
  "comments": 13213,
  "likes": 87099,
  "countsCapped": false,
  "pics": [],
  "videoUrl": null,
  "videoCover": null,
  "user": { "id": "1669879400", "screenName": "Dear-迪丽热巴", "followers": null, "followersText": "8255万", "verified": true },
  "retweeted": {
    "id": "5348560276032588",
    "bid": "Rknh1pjlO",
    "url": "https://weibo.com/2803301701/Rknh1pjlO",
    "createdAt": "2026-09-29T11:00:03.000Z",
    "text": "【#烈士纪念日#，发条微博，缅怀先烈】9月30日，#国庆节的前一天是烈士纪念日#。...",
    "user": { "id": "2803301701", "screenName": "人民日报" },
    "reposts": 523038,
    "comments": 5884,
    "likes": 34643
  },
  "source": null,
  "region": "法国",
  "input": "https://weibo.com/1669879400/Rkpv9rRTu",
  "inputType": "post",
  "scrapedAt": "2026-09-30T05:25:53.969Z"
}
```

A `user` row:

```json
{
  "type": "user",
  "id": "1669879400",
  "screenName": "Dear-迪丽热巴",
  "url": "https://weibo.com/u/1669879400",
  "description": "一只喜欢默默表演的小透明。工作联系jaywalk@jaywalk.com.cn 🍒",
  "followers": 82550167,
  "followersText": "8255万",
  "following": 294,
  "statusesCount": 1934,
  "verified": true,
  "verifiedType": 0,
  "verifiedReason": "嘉行传媒签约演员",
  "gender": "f",
  "location": "上海",
  "avatar": "https://tvax1.sinaimg.cn/crop.0.0.1080.1080.1024/63885668ly8geyrcrw0zjj20u00u0mz6.jpg?...",
  "customDomain": null,
  "input": "1669879400",
  "scrapedAt": "2026-09-30T05:23:12.391Z"
}
```

A `comment` row:

```json
{
  "type": "comment",
  "id": "5348546372437370",
  "postId": "5348478804824358",
  "postBid": "Rkl9Ckf2e",
  "createdAt": "2026-09-29T10:04:49.000Z",
  "text": "我怎么我今天去线下，promax是12g开头啊🙀",
  "textHtml": "我怎么我今天去线下，promax是12g开头啊🙀",
  "likes": 1,
  "replies": 5,
  "floor": 11,
  "region": "浙江",
  "user": { "id": "7875740762", "screenName": "共鸣198108", "followers": 73, "verified": false },
  "scrapedAt": "2026-09-30T05:24:40.125Z"
}
```

A failed input gets one free `status` row instead, and the run carries on:

```json
{ "type": "status", "input": "Qa1b2c3d4", "inputType": "post", "status": "not_found", "error": "Post not found (deleted, private, or never existed).", "checkedAt": "2026-09-30T05:26:04.091Z" }
```

`status` is `not_found`, `invalid` (not a uid, URL or id this actor
recognises), `blocked` (Weibo still asked for a login after a fresh visitor
cookie and every proxy rotation), or `error` (any other failure).

The key-value store record `OUTPUT` summarises the run: rows by type, the
charged count, `stopped_reason`, HTTP requests, proxy rotations, response
bytes, and a `truncated` list naming every keyword, user or post that returned
fewer rows than you asked for, with the reason (`logged_out_limit`,
`max_total_charge_reached`, `repeated_cursor` or `error`). Every run leaves at
least one dataset row.

A few field notes:

- `user.followers` on post rows is `null` when Weibo only gives a rounded
  figure there ("8255万"); the rounded string is in `followersText`. The
  `user` profile row has the exact count.
- `countsCapped: true` means Weibo reported reposts or comments as exactly
  1,000,000, its "100万+" display cap: the real count is at least that.
- Picture and avatar URLs are Weibo's CDN links; video URLs are signed and
  expire after a few hours, so download promptly if you need the files.

### Pricing

Pay per event, nothing else: **$0.004 per item returned** ($4 per 1,000
post, user or comment rows). Proxy traffic is included in the price.
Failed inputs (`not_found`, `invalid`, `blocked`, `error`) are never charged.
If you set a maximum total charge for the run, the actor stops before fetching
pages it could not sell, and says so in `OUTPUT.stopped_reason`
(`"max_total_charge_reached"`).

### Limits, stated plainly

- **Logged-out Weibo only shows so much, and this actor never logs in.** Weibo
  serves the first page of each feed to visitors and asks for a login from
  page 2 on. So:
  - **Search** returns at most what the first page of six search tabs (top,
    latest, hot, pictures, videos, one more) contain, merged and de-duplicated.
    For "华为" that was 87 distinct posts; niche keywords give fewer.
  - **User posts** come from the first page of the user's timeline and profile
    tabs (roughly the newest 10 to 25 posts of any kind), then from their
    video wall, which does paginate logged-out. Beyond the first couple of
    dozen posts you therefore get **video posts only**; older text-only and
    photo posts are not reachable without a login. The run reports this in
    `OUTPUT.truncated` as `logged_out_limit`.
  - **Comments** come from weibo.com's paginated comments endpoint (top-level
    comments, hot order). Replies to comments are not returned as rows (the
    `replies` field counts them). If that endpoint fails, the actor falls back
    to the mobile hot-comments feed, which gives only its first page.
- **Reposted posts** are summarised from what Weibo sends with the repost; a
  long original is not expanded.
- **Rate limits and login walls.** Weibo issues its visitor cookie per IP and
  rate-limits by IP. When it answers "login required" (`ok: -100`, a redirect
  to passport.weibo.com, "请登录后使用") or HTTP 418, the actor gets a fresh
  visitor cookie, then moves to a new proxy session with a new cookie, up to
  8 times for that request, and only then gives up that input with a free
  `blocked` row. Residential proxy is the default for that reason.
- **Counts are a snapshot**, and Weibo itself caps displayed reposts and
  comments at 1,000,000 (see `countsCapped`).
- **Polite by design:** about 400 ms between requests, at most two inputs in
  flight, HTTP 429/5xx retried with backoff (Retry-After honoured up to 60 s,
  3 retries), 20 s timeouts, and a normal desktop Chrome User-Agent with the
  site's own Referer. No CAPTCHA solving, no accounts.
- **Public data only.** Private accounts, deleted posts and anything that
  needs a login are out of scope.

### FAQ

**Do I need a Weibo account or cookie?** No. The actor gets Weibo's anonymous
visitor cookie itself, the same way a browser does on its first visit.

**Why did a keyword return fewer posts than `maxItems`?** Logged-out Weibo
search only serves the first page of each tab. `OUTPUT.truncated` says so.

**Can I get a user's full history?** Their newest posts plus their full video
history, yes. Older text and photo posts need a logged-in session, which this
actor does not use.

**What is the difference between `mid` and `bid`?** The same post id in two
encodings: `mid` is numeric (`5348645916642028`), `bid` is the base62 form in
weibo.com URLs (`Rkpv9rRTu`). Both are on every post row.

### Local development

```bash
pnpm --filter @mmnm/weibo test        # unit tests, no network
pnpm --filter @mmnm/weibo build
```

`node src/main.ts` runs the actor locally with Apify's local storage
(`APIFY_LOCAL_STORAGE_DIR`); `ACTOR_TEST_PAY_PER_EVENT=true
ACTOR_MAX_TOTAL_CHARGE_USD=1` exercises the charging path.

# Actor input Schema

## `searchKeywords` (type: `array`):

Keywords to search on Weibo. Each keyword returns up to `maxItems` posts, merged and de-duplicated from the first page of six Weibo search tabs (top, latest, hot, pictures, videos, one more). Logged-out Weibo serves only that first page per tab, so a keyword yields at most around 100 posts (fewer for niche terms).

## `users` (type: `array`):

Weibo users: a numeric uid (1669879400), weibo.com/u/<uid>, weibo.com/<uid>, weibo.com/<custom domain>, weibo.com/n/<screen name>, m.weibo.cn/u/<uid>, or @<screen name>. Each user returns one profile row plus up to `maxItems` of their posts.

## `postUrls` (type: `array`):

Single posts: weibo.com/<uid>/<bid>, m.weibo.cn/detail/<mid>, m.weibo.cn/status/<bid>, or a bare mid / bid. Each returns one post row.

## `includeComments` (type: `boolean`):

Also return the top-level comments (hot order) of every returned post, up to `maxComments` per post. Each comment is a billed row.

## `maxItems` (type: `integer`):

Maximum posts returned per search keyword and per user (the user's profile row is extra). 1 to 1,000.

## `maxComments` (type: `integer`):

Maximum comments returned per post when `includeComments` is on. 1 to 1,000.

## `proxy` (type: `object`):

Weibo issues its visitor cookie per IP and rate-limits by IP, so requests go through Apify residential proxy by default. When Weibo asks for a login or rate-limits, the actor gets a fresh visitor cookie, then moves to a new proxy session (up to 8 times per request). Proxy traffic is included in the price.

## Actor input object example

```json
{
  "users": [
    "1669879400"
  ],
  "includeComments": false,
  "maxItems": 10,
  "maxComments": 20,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "users": [
        "1669879400"
    ],
    "maxItems": 10,
    "proxy": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("changefeeds/weibo-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "users": ["1669879400"],
    "maxItems": 10,
    "proxy": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("changefeeds/weibo-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "users": [
    "1669879400"
  ],
  "maxItems": 10,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call changefeeds/weibo-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,changefeeds/weibo-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6hfDhJsdrSXrKrcvH/builds/jK5oo68bAfg8vsh7R/openapi.json
