# Weibo Scraper (`reportable_broth/weibo-scraper`) Actor

Weibo (微博) scraper: keyword search, user posts, fans and following lists, comments, hot search and hot feed. Add your own cookie for deep results, or run without login. Monitor mode returns only new posts. $4.49 per 1,000 items.

- **URL**: https://apify.com/reportable\_broth/weibo-scraper.md
- **Developed by:** [Quiet Harvest](https://apify.com/reportable_broth) (community)
- **Categories:** Social media, Lead generation, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.49 / 1,000 item (your own login or no login)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Weibo Scraper

Scrape Weibo (微博, Sina Weibo) without a browser. Get keyword search results, a user's posts, fans and following lists, comments, the hot search list (热搜) and the hot feed.

**Pick how to log in:**

| Login | Price per 1,000 items | What you get |
|---|---|---|
| **Your own cookie** | $4.49 | Deep results through your account |
| **No login** | $4.49 | First page of search, user posts, fans and following. Comments, hot search and the hot feed work in full |
| **Our Weibo accounts** | Coming soon | Deep results with nothing to set up |

Posts, comments, users and hot-search topics all count as one item each.

***

### What you can get

| | No login | With your cookie |
|---|---|---|
| Keyword search | First page (~20 newest posts) | Hundreds of posts per keyword |
| A user's posts | Latest ~10 | Page after page, newest first |
| Fans / following | First 20 | Hundreds |
| Comments on a post | ✅ Deep, sorted by likes or newest | ✅ |
| User profile | ✅ | ✅ |
| Hot search (热搜), hot feed (热门) | ✅ | ✅ |
| Single posts, full text of long posts | ✅ | ✅ |

**Why the first page only without login?** Since 2025, Weibo asks anonymous visitors to log in from page 2 onward, for search, user timelines and follower lists. Every Weibo scraper runs into this. Add your own cookie to get deep results at the same $4.49 per 1,000.

**Monitoring doesn't need login.** Page one of search and of a user's timeline is always the newest content. Turn on **Only new items** and schedule the Actor every 15–30 minutes, and you capture every new post for a keyword or account without logging in.

***

### Quick start

Monitor keywords (no login, run on a schedule):

```json
{
  "searchKeywords": ["新能源汽车", "#比亚迪#"],
  "onlyNew": true,
  "maxCommentsPerFoundPost": 20
}
```

A brand account with its fans (without login you get the first page):

```json
{
  "users": ["https://weibo.com/u/2803301701"],
  "maxPostsPerUser": 200,
  "maxFans": 500
}
```

The same with your own account (deep results):

```json
{ "users": ["https://weibo.com/u/2803301701"], "maxPostsPerUser": 200, "accountMode": "own", "cookie": "SUB=...; SUBP=...; ..." }
```

All comments on a post:

```json
{ "postUrls": ["https://weibo.com/2803301701/RjAABq8CG"], "maxCommentsPerPostUrl": 2000, "commentsSort": "latest" }
```

Trending now:

```json
{ "hotSearch": true, "hotFeedPosts": 50 }
```

### Output

Every item has a `type`:

**`post`**: `id`, `bid`, `url`, `createdAt` (ISO 8601 UTC), `text` (plain text, full text for long posts), `likes`, `comments`, `reposts`, `source` (device or app), `region` (IP location shown by Weibo), `topics` (#hashtags#), `mentions`, `images`, `video` (`url`, `cover`, `title`, `durationSec`), `isPinned`, `repostOf` (the original post if this is a repost), `user` (`userId`, `name`, `verified`, `verifiedReason`, `followers`, `following`, `posts`, `gender`, `avatar`, `url`), and `keyword` or `sourceUserId` for where it was found.

**`comment`**: `id`, `postId`, `createdAt`, `text`, `likes`, `replies`, `region`, `floor`, `topReplies` (up to 3 replies), `user`.

**`user`**: the profile fields above plus `description`, `location`, `customUrl`, `memberLevel`, `verifiedType`.

**`fan`** / **`following`**: user fields plus `description`, `ofUserId`, `ofUserName` (whose fan or following it is).

**`hotSearch`**: `rank`, `word`, `heat`, `label` (e.g. 新, 热, 沸), `url`.

### How to copy your Weibo cookie (only for 'Your own cookie')

1. In Chrome, open **https://m.weibo.cn** and log in (you can scan the QR code with the Weibo app).
2. Press **F12**, open the **Network** tab and press **F5**.
3. Click the first request (`m.weibo.cn`), and under **Request Headers** copy the whole value of **Cookie**.
4. Set **Login** to *Use my own cookie* and paste it into the cookie field.

The field is encrypted by Apify. The cookie is only sent to Weibo.

### Safe by design

- Logged-in requests keep one IP for the whole run and wait 0.8 s between requests (adjustable). That is about a quarter of the speed at which Weibo started limiting in our tests.
- If Weibo rate-limits the run, the Actor switches to a new IP once and slows down. If it happens again, **the run stops right away**. You pay only for items already delivered, and your account is not pushed further.
- If a login expires, the run continues without login and says so in the log.

### Pricing

| Item | Price |
|---|---|
| Any item, with **your own cookie** or **no login** | $4.49 per 1,000 |

Logging in with our accounts (no setup) is coming soon.

With **Only new items** on, you only pay for what's new since the last run.

### Good to know

- Long posts are cut off with "...全文" in lists. With **Full text of long posts** on (default), the Actor fetches the complete text.
- Hidden fan lists and deleted posts return nothing.
- Weibo shows some counts rounded (e.g. "1.2万"). They are converted to numbers.

### Questions

Found a bug or want a field added? Open an issue on the Issues tab.

# Actor input Schema

## `searchKeywords` (type: `array`):

Chinese, English or hashtags. One search per line.

## `maxPostsPerKeyword` (type: `integer`):

Without a login cookie Weibo returns about 20 posts per keyword (the first page).

## `searchSort` (type: `string`):

Latest is best for monitoring.

## `users` (type: `array`):

User ID (2803301701), profile URL (weibo.com/u/2803301701 or weibo.com/rmrb) or @name.

## `includeProfile` (type: `boolean`):

Followers, following, post count, verification, bio, location.

## `maxPostsPerUser` (type: `integer`):

Newest first. Without a login cookie you get the latest ~10.

## `maxFans` (type: `integer`):

Accounts that follow the user. Without a login cookie you get the first 20.

## `maxFollowing` (type: `integer`):

Accounts the user follows. Without a login cookie you get the first 20.

## `postUrls` (type: `array`):

weibo.com/<uid>/<id>, m.weibo.cn/detail/<id>, or the post ID. Returns the post and its comments.

## `maxCommentsPerPostUrl` (type: `integer`):

Comments work without login and can go deep.

## `maxCommentsPerFoundPost` (type: `integer`):

Also fetch this many comments for every post found by search or user timelines. 0 = off.

## `commentsSort` (type: `string`):

Order of comments.

## `hotSearch` (type: `boolean`):

The ~50 real-time trending topics with heat values.

## `hotFeedPosts` (type: `integer`):

Posts from Weibo's trending feed. 0 = off.

## `fullText` (type: `boolean`):

Long posts are cut off with '...全文'. This fetches the full text (one extra request per long post).

## `onlyNew` (type: `boolean`):

Put the Actor on a schedule: each run returns only posts, comments and fans that earlier runs haven't returned. You only pay for new items.

## `monitorStoreName` (type: `string`):

Runs with the same inputs share memory automatically. Set a name to share it on purpose.

## `accountMode` (type: `string`):

Weibo only shows page 2+ of search, user posts, fans and following to logged-in users. Add your cookie below for deep results. Using our accounts (no setup): coming soon.

## `cookie` (type: `string`):

Unlocks page 2+ of search, user posts, fans and following. Log in at m.weibo.cn, copy the Cookie request header (see README). Stored encrypted; only sent to Weibo.

## `requestGapSecs` (type: `number`):

Slower is safer for your account.

## `proxyConfiguration` (type: `object`):

Anonymous requests rotate IPs. Logged-in requests keep one IP for the whole run.

## Actor input object example

```json
{
  "searchKeywords": [
    "新能源汽车"
  ],
  "maxPostsPerKeyword": 50,
  "searchSort": "latest",
  "users": [],
  "includeProfile": true,
  "maxPostsPerUser": 20,
  "maxFans": 0,
  "maxFollowing": 0,
  "postUrls": [],
  "maxCommentsPerPostUrl": 100,
  "maxCommentsPerFoundPost": 0,
  "commentsSort": "hot",
  "hotSearch": false,
  "hotFeedPosts": 0,
  "fullText": true,
  "onlyNew": false,
  "accountMode": "own",
  "requestGapSecs": 0.8,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchKeywords": [
        "新能源汽车"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("reportable_broth/weibo-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchKeywords": ["新能源汽车"],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("reportable_broth/weibo-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchKeywords": [
    "新能源汽车"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call reportable_broth/weibo-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,reportable_broth/weibo-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/MQL0GYzf055uYkSqN/builds/Ek7BZVVlizDDfPkrq/openapi.json
