# Weibo Scraper: Search Posts by Keyword or User (微博) (`datagleaner/weibo-scraper`) Actor

Weibo scraper for public posts by keyword search or user timeline: full text, ISO dates, likes, reposts, comments, images, video, poster region and author. No Weibo login, no browser, no cookies. $3 per 1,000 posts; export to JSON, CSV or Excel.

- **URL**: https://apify.com/datagleaner/weibo-scraper.md
- **Developed by:** [Data Gleaner](https://apify.com/datagleaner) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Weibo Scraper: Posts, Comments, Hot Search (微博热搜)

**Weibo Scraper** pulls public Sina Weibo (微博) posts by keyword search, user timeline or post URL, plus comments and the hot-search board, into JSON, CSV or Excel, with no login, cookies or API key. Each post comes with full text, an ISO 8601 timestamp, like / repost / comment counts, images, video, the poster's region and the author. It uses the anonymous guest session any first-time visitor gets, so there is no browser to run and no account to set up, and it costs $3 per 1,000 posts.

### What you can do with it

- **Brand and product monitoring.** Track every mention of your brand, competitors or a product launch (`瑞幸`, `iPhone 17`, `Nike`) as it happens. Results come newest first.
- **China social listening and market research.** See what Chinese consumers say about a topic, with the poster's province and engagement numbers.
- **Academic and media research.** Collect public discourse on an event or hashtag for analysis.
- **LLM and RAG pipelines.** `text` is clean plain text, one field per post, ready for embedding, summarising or translation.
- **Trend tracking.** Turn on `hotSearch` for the current 微博热搜 board (rank, word, heat, label).
- **Influencer scouting.** Turn on `includeUserProfile` to get each author's follower count, bio and verification status.

### Input

| Field | Type | What it does |
|---|---|---|
| `searchQueries` | list of text | Keywords, Chinese or English. One search per keyword. |
| `userIds` | list of text | Numeric user IDs or profile URLs (`https://weibo.com/u/1669879400`, `https://m.weibo.cn/u/1669879400`). Returns the user's latest posts. |
| `postUrls` | list of text | Single posts: `https://weibo.com/1785807165/RlqrBy6x0`, `https://m.weibo.cn/detail/5351065438128262` or a bare post ID. One request per post; deleted or private posts are skipped with a warning. |
| `hotSearch` | true / false | Default false. Adds the current 微博热搜 board, one row per trend (`type` is `hotSearch`: `rank`, `word`, `heat`, `label`, `category`, `isTopic`, `url`). Ads are skipped. |
| `maxHotSearchItems` | number | Trends to keep from the top of the board. Default 50, max 60. |
| `maxItemsPerQuery` | number | Cap per keyword and per user. Default 100, max 1000. |
| `sinceDate` | text | Optional. `2026-10-01` or full ISO 8601 (Beijing time if no timezone). Older posts are skipped; the run stops once a page holds nothing newer. |
| `untilDate` | text | Optional. Skips posts newer than this; a bare date means through the end of that day. For keyword search Weibo searches back from this moment, so together with `sinceDate` it reaches older date ranges, still up to 1,000 posts per keyword. `sinceDate` must be earlier than `untilDate`. |
| `includeComments` | true / false | Default false. Adds a `comments` list to each post: top-level comments, hottest first, with the first replies Weibo shows. Costs extra requests (about one per 20 comments). |
| `maxCommentsPerPost` | number | Top-level comments per post when comments are on. Default 20, max 1000. |
| `includeUserProfile` | true / false | Adds the exact follower count, bio, verified reason, following count, post count, gender and avatar to each author. One extra request per distinct author, so runs are slower. |
| `requestDelaySeconds` | number | Pause between requests, default 1.2 s (randomised up to 1.5x). |
| `proxyConfiguration` | object | Optional Apify Proxy. Default is no proxy. |

Give at least one of `searchQueries`, `userIds` or `postUrls`, or turn on `hotSearch`. With none of them, the Actor runs a built-in example (keyword 咖啡, up to 5 posts from about a day ago) instead of failing.

```json
{
  "searchQueries": ["咖啡", "luckin coffee"],
  "userIds": ["https://weibo.com/u/1669879400"],
  "maxItemsPerQuery": 200,
  "sinceDate": "2026-10-01",
  "untilDate": "2026-10-07",
  "hotSearch": true,
  "includeUserProfile": false
}
```

### Output

One dataset item per post, plus one per trend when `hotSearch` is on. Export as JSON, CSV, Excel or via the API.

| Field | Meaning |
|---|---|
| `type` | `post` for posts, `hotSearch` for hot-search rows (which carry `rank`, `word`, `heat`, `label`, `category`, `isTopic`, `url` and `scrapedAt` instead of the post fields). |
| `id`, `mid` | Weibo post ID (string). |
| `url` | Link to the post on weibo.com. |
| `createdAt` | ISO 8601 with timezone (Weibo shows Beijing time, `+08:00`). |
| `text` | Post text with HTML removed. Emoji appear as `[name]`, line breaks as `\n`. |
| `textRaw` | Weibo's markup for the text with zero-width characters removed. On long posts it is the markup preview unless Weibo returns the full text as markup. |
| `isLongText` | `true` if Weibo had truncated the post; `text` then holds the full text. |
| `repostsCount`, `commentsCount`, `likesCount` | Engagement counts. |
| `topics` | Hashtags without the `#` marks, in order of appearance. |
| `mentions` | `@` handles mentioned in the text. |
| `pics` | Image URLs, largest size available. |
| `picCount` | Number of images Weibo reports for the post. |
| `video` | `{url, duration (seconds), playCount, cover}` or `null`. |
| `region` | Poster's province / country as Weibo shows it ("发布于 江苏" becomes `江苏`). |
| `source` | Client the post was sent from (for example `iPhone客户端`). |
| `author` | `{id, screenName, profileUrl, followersCount, verified, verifiedReason, description}`; with `includeUserProfile` also `followingCount`, `postsCount`, `gender`, `avatarUrl`. |
| `retweetedStatus` | The original post when this one is a repost, same shape; otherwise `null`. |
| `comments` | Only with `includeComments`: `{id, postId, createdAt, text, likesCount, repliesCount, region, author, replies}`. |
| `query` / `userId` / `postUrl` | The keyword, user or post URL the post was found through. |
| `scrapedAt` | UTC time of collection. |

```json
{
  "type": "post",
  "id": "5351532104777774",
  "mid": "5351532104777774",
  "url": "https://weibo.com/7564024690/RlCAik2US",
  "createdAt": "2026-10-07T23:49:02+08:00",
  "text": "#宋亚轩古茗咖啡全球代言人# ... 暮歌丨宋亚轩的微博视频 @时代少年团-宋亚轩",
  "textRaw": "<a href=\"//s.weibo.com/weibo?q=...\" target=\"_blank\">#宋亚轩古茗<span style=\"color: red;\">咖啡</span>全球代言人#</a> ...",
  "isLongText": false,
  "repostsCount": 0,
  "commentsCount": 0,
  "likesCount": 0,
  "topics": ["宋亚轩古茗咖啡全球代言人"],
  "mentions": ["时代少年团-宋亚轩"],
  "pics": [],
  "picCount": 0,
  "video": {"url": "http://f.video.weibocdn.com/o0/9faVhSlilx08zIWo7ciI01041203E6Ci0E020.mp4", "duration": 154.0, "playCount": 174000, "cover": "https://wx2.sinaimg.cn/orj480/008gImyHly1ifr8cel3rhj30u01hcju1.jpg"},
  "region": "安徽",
  "source": "Android客户端",
  "author": {"id": "7564024690", "screenName": "椿岛轩巷", "followersCount": null, "verified": true, "verifiedReason": null, "description": null, "profileUrl": "https://weibo.com/u/7564024690"},
  "retweetedStatus": null,
  "query": "咖啡",
  "scrapedAt": "2026-10-07T15:49:05+00:00"
}
```

Keyword search does not carry the author's follower count, bio or verified reason, so those are `null` unless you turn on `includeUserProfile`. Posts found through `userIds` always include them.

### How much does it cost to scrape Weibo?

Pay per event: **$3 per 1,000 posts** (one `post` event, $0.003, for each post saved to the dataset). With `includeComments` on, each comment is a `comment` event at **$2 per 1,000** ($0.002). With `hotSearch` on, each hot-search board entry is a `trend` event at **$2 per 1,000** ($0.002). You pay only for posts you receive: nothing for failed pages or retries, and the run stops by itself when your **maximum total charge** is reached. Platform usage is included in the price.

Example: 5 keywords at 200 posts each is 1,000 posts, so $3. A daily brand-monitoring run of 100 new posts costs about $0.30.

### Limits

- **About 1,000 posts per keyword.** Weibo serves roughly 25 pages of keyword search results (about 1,000 to 1,200 posts), however large the topic. To go deeper, split a topic into several narrower keywords or run it daily with `sinceDate`, or slice the period with `sinceDate` and `untilDate`.
- **User timelines return the latest posts only (about 10).** Without a login, Weibo shows guests the first page of a user's timeline and asks for a login on page 2. The Actor stops cleanly there. To follow a person or brand over time, run it on a schedule and dedupe on `id`, or search for their name or hashtag.
- **Public data only, no login.** Private, followers-only and age-restricted posts and follower lists are not available. A few posts are hidden from guests; they come back truncated rather than failing the run.
- Weibo can change or throttle its guest access at any time. The Actor refreshes its guest session, backs off and retries on errors, but a run may return fewer posts than requested when Weibo limits it. If that happens, raise `requestDelaySeconds` or enable a proxy.

### Use with Python

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("datagleaner/weibo-scraper").call(
    run_input={"searchQueries": ["瑞幸"], "maxItemsPerQuery": 10}
)
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["createdAt"], item["likesCount"], item["text"][:80])
```

Ten posts cost $0.03.

### Use with JavaScript / Node.js

```js
// npm install apify-client
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagleaner/weibo-scraper').call({
    searchQueries: ['瑞幸'],
    maxItemsPerQuery: 10,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const post of items) console.log(post.createdAt, post.likesCount, post.text.slice(0, 80));
```

### Use it from n8n, Make, Zapier or an AI agent

Actor ID: `datagleaner/weibo-scraper`

Minimal input:

```
{"searchQueries": ["瑞幸"], "maxItemsPerQuery": 10}
```

Each tool below runs this Actor with your own Apify API token.

- **n8n:** add the **Apify** node (`@apify/n8n-nodes-apify`). On n8n Cloud you install it from the community node registry. Choose **Run an Actor and get dataset**, set Actor to `datagleaner/weibo-scraper` and paste the input above.
- **Make:** use the Apify app's **Run an Actor** module, then **Get Dataset Items** to read the results. **Watch Actor Runs** can trigger a scenario when a run finishes.
- **Zapier:** use the Apify action **Run Actor**, then the search **Fetch dataset items**. The trigger **Finished Actor run** starts a Zap when a run ends.
- **AI agents (MCP):** connect to `https://mcp.apify.com?tools=datagleaner/weibo-scraper`. Without a token, an agent can still find Actors with the `search-actors` tool, but running this one needs your Apify account. In Claude Code:

  ```
    claude mcp add --transport http apify "https://mcp.apify.com?tools=datagleaner/weibo-scraper"
  ```

  Then run `/mcp` to sign in to Apify in your browser. Then ask in plain language, for example: "Use the Weibo scraper to get the 50 newest Weibo posts mentioning 瑞幸 since 2026-10-01, and summarise in English what people complain about most, with the like counts."
- **LangChain (Python):**

```python
## pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/weibo-scraper")
result = tool.invoke({"run_input": json.loads('{"searchQueries": ["瑞幸"], "maxItemsPerQuery": 10}')})
```

### FAQ

**How do I scrape Weibo posts by keyword?** Put one or more keywords in `searchQueries` and set `maxItemsPerQuery`. Each keyword runs a Weibo search and returns the newest public posts first, up to about 1,000 per keyword. English keywords work, but Chinese keywords return far more.

**Can I scrape a single post or go back to older dates?** Put post links in `postUrls`. For older posts, set `untilDate` (with `sinceDate`) on a keyword search to reach earlier date ranges, up to 1,000 posts per range.

**Can I scrape a Weibo user's posts?** Yes, put a numeric user ID or profile link in `userIds`. Without a login Weibo shows guests only the latest page (about 10 posts); see Limits.

**Do I need a Weibo API key, account or cookies?** No. The Actor uses the anonymous guest session any visitor gets, so there is no Weibo developer app, login or cookie to set up. You only need an Apify account to run it.

**Can I get Weibo hot search, comments or follower lists?** Hot search and comments yes: set `hotSearch` for the 微博热搜 board and `includeComments` for top-level comments with their first replies (each is billed as described under pricing). Follower lists are not included.

**Why did my run return fewer posts than `maxItemsPerQuery`?** The topic has fewer public matches, `sinceDate` cut the list, Weibo's result cap was reached, or Weibo throttled the session. The run log says which; raising `requestDelaySeconds` or enabling a proxy helps with throttling.

### Responsible use

This Actor collects only publicly visible posts. Respect Weibo's terms and keep the request rate low (the default pacing is deliberately polite). Posts and profiles are personal data: **you** are responsible for having a lawful basis and meeting your obligations under GDPR, China's PIPL and any other law that applies to your use, including retention, deletion requests and cross-border transfer. This is not legal advice.

### 中文说明

**微博抓取器：** 无需登录、无需 Cookie、无需浏览器，按关键词、用户或帖子链接抓取微博公开帖子，并可获取评论和微博热搜榜，输出干净的 JSON：正文（已去除 HTML）、ISO 8601 时间、转发/评论/点赞数、图片、视频、发布地区、来源、作者信息、被转发的原微博。

- **用途：** 品牌监测、舆情与社交聆听、学术研究、喂给大模型的数据管线、达人筛选。
- **输入：** `searchQueries`（关键词，中英文均可）、`userIds`（数字 UID 或主页链接）、`maxItemsPerQuery`（每个关键词/用户的上限，默认 100）、`sinceDate` / `untilDate`（起止日期）、`postUrls`（单条帖子）、`includeComments` / `maxCommentsPerPost`（评论）、`hotSearch`（热搜榜）、`includeUserProfile`（补全作者粉丝数、简介、认证信息）。
- **价格：** 按条计费，帖子每 1,000 条 3 美元；开启评论时每 1,000 条评论 2 美元，开启热搜时每 1,000 条热搜 2 美元。只为实际保存的数据付费，达到你设置的最大花费后自动停止。
- **限制：** 每个关键词最多约 1,000 条（微博搜索的上限）；游客状态下用户主页只能获取最新约 10 条；只抓取公开数据。
- **合规：** 请遵守微博服务条款并保持较低的请求频率；个人信息的处理须符合 GDPR、《个人信息保护法》等适用法律，责任由使用者承担。

# Actor input Schema

## `searchQueries` (type: `array`):

Keywords to search on Weibo, Chinese or English (for example 咖啡, 瑞幸, iPhone 17). Each keyword returns at most 1,000 posts, a limit set by Weibo.

## `userIds` (type: `array`):

Numeric Weibo user IDs or profile URLs such as https://weibo.com/u/1669879400 or https://m.weibo.cn/u/1669879400. Returns the user's posts, newest first.

## `postUrls` (type: `array`):

Single posts to fetch, such as https://weibo.com/1785807165/RlqrBy6x0, https://m.weibo.cn/detail/5351065438128262 or a bare post ID. One request per post; deleted or private posts are skipped with a warning.

## `hotSearch` (type: `boolean`):

Adds the current 微博热搜 board: one row per trend with rank, word, heat, label and a search link (about 50 rows, ads skipped). Rows have type 'hotSearch'; posts have type 'post'.

## `maxHotSearchItems` (type: `integer`):

How many trends to keep from the top of the board.

## `maxItemsPerQuery` (type: `integer`):

Upper bound on posts for each keyword and each user. Search stops at 1,000 whatever you enter.

## `sinceDate` (type: `string`):

Skip posts older than this date (YYYY-MM-DD or ISO 8601, Beijing time if no timezone). For users the run stops at the first older post; for keyword search older posts are filtered out.

## `untilDate` (type: `string`):

Skip posts newer than this date (YYYY-MM-DD means through the end of that day, or ISO 8601; Beijing time if no timezone). For keyword search Weibo searches back from this moment, so with sinceDate it reaches older date ranges, still up to 1,000 posts per keyword.

## `includeUserProfile` (type: `boolean`):

Adds followingCount, postsCount, gender and avatarUrl to each author, and the exact followers count and bio. Costs one extra request per distinct author, so runs are slower.

## `includeComments` (type: `boolean`):

Adds a `comments` list to each post: top-level comments, hottest first, each with text, likes, reply count, region, author and the first replies Weibo shows. Costs extra requests (about one per 20 comments).

## `maxCommentsPerPost` (type: `integer`):

Upper bound on top-level comments for each post when comments are on.

## `requestDelaySeconds` (type: `number`):

Base pause between requests; the actual pause is randomised up to 1.5x. Raise it if you see throttling.

## `proxyConfiguration` (type: `object`):

Optional. Works without a proxy from Apify datacenter IPs; enable if Weibo starts blocking your runs.

## Actor input object example

```json
{
  "searchQueries": [
    "咖啡"
  ],
  "hotSearch": false,
  "maxHotSearchItems": 50,
  "maxItemsPerQuery": 100,
  "includeUserProfile": false,
  "includeComments": false,
  "maxCommentsPerPost": 20,
  "requestDelaySeconds": 1.2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "咖啡"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagleaner/weibo-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["咖啡"],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("datagleaner/weibo-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "咖啡"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call datagleaner/weibo-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagleaner/weibo-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/f2o6dx3t14cqNIxUI/builds/ptjQ4maxgVMjccUOV/openapi.json
