Weibo Scraper — Hot Search, Posts, Comments & Creator Feeds avatar

Weibo Scraper — Hot Search, Posts, Comments & Creator Feeds

Pricing

from $35.00 / 1,000 item scrapeds

Go to Apify Store
Weibo Scraper — Hot Search, Posts, Comments & Creator Feeds

Weibo Scraper — Hot Search, Posts, Comments & Creator Feeds

Sina Weibo (微博) scraper API, no login: keyword search, trending posts, hot-search board (微博热搜), comments, user profiles and post details, plus Black Cat (黑猫投诉) consumer complaints. Creator timelines (user_posts) need your cookie. Nine modes for social listening and China market research.

Pricing

from $35.00 / 1,000 item scrapeds

Rating

1.0

(1)

Developer

Sami

Sami

Maintained by Community

Actor stats

3

Bookmarked

571

Total users

252

Monthly active users

2 days ago

Last modified

Share

Weibo Scraper - Chinese Social Intelligence

Extract Chinese public opinion, trending topics, and real-time consumer sentiment from Weibo (微博) — China's dominant microblog with 580M+ monthly users producing a dense public-opinion signal. Built for AI training corpora, Chinese consumer equity research alt-data, brand monitoring agencies, and academic NLP teams. No API key, no VPN, and no login except user_posts (your cookie).

💡 Tracking a brand, not just scraping one platform? Pair this with two recurring monitors that watch the whole picture for you: Chinese Brand Monitor — Weibo + RedNote + Bilibili + Douban + Xueqiu mentions in one scheduled, deduped, sentiment-tagged feed — and AI Brand Visibility Monitor — how the AI engines (DeepSeek, Qwen, Kimi, GLM + ChatGPT/Gemini/Claude) answer about your brand.

How to scrape Weibo in 3 easy steps

  1. Go to the Weibo Scraper page on Apify and click "Try for free"
  2. Configure your input — choose a mode (hot_timeline, hot_search, hot_search_delta, post_comments, post_details, search, user_posts, user_profile, or complaints for Black Cat consumer complaints), enter your keywords, post IDs or URLs, user IDs or company IDs, and set the number of results
  3. Click "Run", wait for the scraper to finish, then download your data in JSON, CSV, or Excel format

No coding required. No API key. Works with Apify's free plan.

📅 Need a specific period, or a deeper search? In search mode, set searchDateFrom / searchDateTo (Beijing time, +08:00) and the run pages back from searchDateTo through that period, up to your maxResults; with or without dates, a search keeps paging past Weibo's first result window (about 1,000 posts) toward your maxResults (up to 5,000 per run; this deeper walk is off in deltaMode). Need the accounts rather than their posts? user_profile returns followers, following, post count, verification, bio and location per user ID, no login. Tracking specific posts? post_details returns each post's current text, reposts, comments and likes from its URL, no login. And complaints pulls Black Cat (黑猫投诉) consumer complaints about the companies you list, also no login. → Date window details

Part of the Chinese Digital Intelligence Suite

Built by Zhorex, who maintains a suite of Chinese-platform scrapers — built specifically for AI training data buyers, equity research analysts covering Chinese consumer stocks, and brand monitoring teams:

  • 🆕 Chinese Brand Monitor — Cross-platform brand mention aggregator (Weibo + RedNote + Bilibili + Douban + Xueqiu in one normalized feed, sentiment-tagged, cross-platform deduped — $0.06/mention)
  • Weibo Scraper — You are here (microblogging, hot search, real-time public opinion)
  • Bilibili Scraper — China's video platform: danmaku, comments, Gen-Z creator sentiment
  • RedNote (Xiaohongshu) Scraper — China's Instagram + Pinterest (lifestyle, consumer reviews)
  • RedNote Shop Scraper — RedShop e-commerce (products, vendors, prices)
  • Douban Scraper — Long-form reviews (movies/books/music), group discussions
  • Xueqiu Scraper — Chinese stock-discussion sentiment, cashtag indexing

Together, these cover the five pillars of Chinese consumer signal: microblog opinion, video sentiment, lifestyle reviews, e-commerce, and long-form opinion. Most analysts buy 2-4 of these for cross-platform coverage. Building a cross-platform brand monitoring pipeline? The Chinese Brand Monitor aggregator gives you all 5 platforms in one normalized output — saves 4-6 hours of engineering vs. orchestrating individual scrapers.

Who buys this scraper

Buyer profileUse case
AI / LLM training data teamsReal-time Chinese microblog text for SFT corpora + current-events grounding
Hedge fund / equity research desksBrand mention velocity, hot-search momentum as alt-data on Chinese consumer stocks (POP MART, BYD, Anta, Yum China, BeiGene)
Brand monitoring agenciesReal-time tracking of Western brand mentions, crisis detection on China's public square
Geopolitical / policy analystsMonitor Chinese public discourse, narrative tracking, policy response signal
Academic NLP / sentiment researchersChinese microblog corpus, labeled sentiment data for classifier training
Journalists / investigative teamsSource Chinese public opinion data for reporting on consumer brands, viral events

🆕 Get the replies, not just the posts. Turn on includeComments and every post found also returns its comment thread — one row per comment, tagged mode: post_comment and linked by postId. The post is the prompt; the replies are the sentiment. No login or cookie needed.

Where the volume actually is: a narrow, high-engagement query returns far bigger threads than a broad one — a generic keyword mostly returns posts with zero comments. Posts with no thread are skipped automatically, so you are never billed for an empty round-trip. In search, the newest posts rarely have replies yet: set searchDateTo 2 or more days back (Beijing time) and the same keyword returns older posts, which have had time to collect them.

It multiplies rows and cost — 100 posts × 50 comments is ~5,100 rows instead of 100 — so it is off by default and capped by maxComments.

📈 maxResults now goes to 5,000 per run (was 500), and maxComments to 1,000. Defaults are unchanged, so nothing about your existing runs or bill changes unless you raise them yourself.

There are two different "hot" things on Weibo and confusing them costs you the data:

ModeReturnsHas comment threads?
hot_searchthe trending board — topic words and their rankNo. They are not posts.
hot_timelinethe posts Weibo's own hot feed is pushing right nowYes, and they are enormous.

Measured anonymously on 2026-07-28, the posts this mode returned carried 633,714 / 618,338 / 493,677 reposts and 122,053 / 264,451 comments. That is the difference between this and a keyword search: a broad search returns mostly posts with zero replies, these do not.

So this is the mode to pair with includeComments. One measured run:

hot_timeline (maxResults: 5) → 5 rows
hot_timeline + includeComments (maxComments 20) → 89 rows (5 posts + 84 comments)
hot_timeline + includeComments (maxComments 50) → 114 rows (5 posts + 109 comments)

An 18-23x multiplier from the same 5 posts, depending on maxComments — measured 2026-08-05, and all 5 posts carried a live thread.

{
"mode": "hot_timeline",
"maxResults": 20,
"includeComments": true,
"maxComments": 100,
"sentimentAnalysis": true
}

At $0.035 per row that run bills roughly $4.10 per 5 posts of full threads.

What is Weibo?

Weibo (微博) is China's dominant microblogging platform — think Twitter meets Instagram. With 580M+ monthly active users, it's where Chinese public opinion forms, brands communicate, and news breaks. Government officials, celebrities, and brands all maintain active Weibo accounts. For data buyers, Weibo's hot search ranking is the closest thing China has to a real-time barometer of public attention.

Weibo API alternative

There is no official public Weibo API available for international developers. Weibo's developer API requires a Chinese business license, has severe rate limits, and returns limited data. This Weibo Scraper is an alternative — it extracts trending topics, posts, and comments without any official API access. No Chinese business registration needed.

Use Cases by buyer

WhoWhy they use it
AI / LLM training data teamsReal-time Chinese-language microblog text for SFT, RLHF training, and current-events grounding for Chinese LLMs
Equity research / hedge fundsHot-search velocity + brand mention spikes as alt-data leading indicator on Chinese consumer stocks
Brand monitoring teamsReal-time tracking of brand mentions, viral content, and crisis detection on China's public square
Geopolitical / policy analystsMonitor public discourse on policy, international topics, and narrative trends
PR & communicationsTrack brand mentions and sentiment shifts in real time
Competitive intelligenceTrack Chinese competitor announcements, product launches, and audience reception
Influencer marketingFind and evaluate Weibo KOLs by followers, engagement, verification status
JournalismAccess Chinese public opinion data for investigative reporting
Academic researchPre-built Chinese microblog corpus with engagement metrics for NLP and sociology studies

Scrape Weibo with Python, JavaScript, or no code

You can use the Weibo Scraper directly from the Apify Console (no code), or integrate it into your own scripts with Python or JavaScript.

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("zhorex/weibo-scraper").call(run_input={
"mode": "hot_search",
"maxResults": 50
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('zhorex/weibo-scraper').call({
mode: 'hot_search',
maxResults: 50,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => console.log(item));

Using the raw REST API (Postman / curl)

⚠️ The run endpoint is asynchronous — its response is the run object (IDs + status), NOT your scraped data. If you POST to /acts/.../runs you get back something like { "data": { "status": "READY", "defaultDatasetId": "…" } } with no results in it — that's expected, the run hasn't finished yet. The records land in the run's dataset, not in that response. (The containerUrl link is the live container; once a run finishes it just shows "run has already finished with status SUCCEEDED" — that means success, it is not where the data lives.)

Easiest — one call that waits for the run and returns the records directly:

curl -X POST "https://api.apify.com/v2/acts/zhorex~weibo-scraper/run-sync-get-dataset-items?token=YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mode":"hot_search","maxResults":50}'

The response body is the JSON array of records — no second call needed.

Or async — start the run, then fetch the dataset once it finishes:

# 1) start the run — note the "defaultDatasetId" in the response
curl -X POST "https://api.apify.com/v2/acts/zhorex~weibo-scraper/runs?token=YOUR_API_TOKEN" \
-H "Content-Type: application/json" -d '{"mode":"hot_search","maxResults":50}'
# 2) when the run status is SUCCEEDED, fetch the records from its dataset
curl "https://api.apify.com/v2/datasets/DEFAULT_DATASET_ID/items?token=YOUR_API_TOKEN"

💡 In the Apify Console you can also open any run and click the Output / Storage → Dataset tab to view and download the same data as JSON / CSV / Excel.

Province rollup — where in China the conversation actually is

Turn on geoRollup and every run adds one row per location (Chinese provinces; overseas countries flagged isOverseas), on top of the posts themselves:

fieldwhat it is
province / provinceEn广东 / Guangdong (or 美国 / United States for a post sent from abroad)
isOverseastrue for a foreign country, false for a Chinese province, Hong Kong, Macau or Taiwan, null for a label outside both lists
postCount, sharePcthow many posts came from there, and what share of the located sample
promotedCounthow many of those were promoted posts, not organic
sentimentMeanaverage sentiment for that province, when sentimentAnalysis is on
sampleSize, lowConfidencethe sample the row was computed from, and whether it is thin

It costs no extra scraping. Weibo already attaches a location to each post (发布于 广东) in the same payload the run pays for, so the rollup is computed from bytes you have already fetched — no additional requests, no additional time.

Honest sampling. If fewer than 50 posts in the run carry a location, no province rows are emitted and none are charged. Between 50 and 200 they ship with lowConfidence: true, and every row states its own sampleSize so you can judge it rather than trust it. Measured 2026-08-15 on a 150-post run: 136 posts carried a location across 28 provinces.

Rollup rows bill as ordinary results. Works with search, hot_timeline and user_posts.

Every post row carries isPromoted. Weibo mixes paid promotion into keyword results, and until now those rows arrived indistinguishable from organic opinion. If you are measuring public sentiment, filter them out; if you are tracking competitors' paid social, they are the interesting ones.

Features

ModeWhat it doesCookies needed?
Search PostsFind posts by keyword — returns query-relevant resultsNo
Hot Search / TrendingReal-time trending topics with heat scores and badges (新 / 热 / 爆)No
Hot Search Delta 🆕Scheduled trend monitor — the whole board each run, every topic tagged new / rising / falling / steady / dropped vs the last run (rank velocity, time-on-board, peaks)No
Post CommentsComments + post detail with engagement metricsNo
Post Details 🆕One row per post URL or ID: current text, reposts, comments, likes, author, date, images, video — no comments fetchedNo
User PostsPosts from specific accountsYes — your cookieString
User Profile 🆕One row per account: followers, following, post count, verification, bio, location, creation dateNo
Complaints (黑猫投诉) 🆕Consumer complaints from Black Cat, Sina's complaint platform: the most recently active about the companies you list, or the latest across China — company, complaint text, demand, amount, progressNo
  • No browser needed — Pure HTTP, runs in 256 MB without sentiment (≥512 MB with sentimentAnalysis)
  • No VPN needed — Globally accessible endpoints
  • Zero setup — search, trending, comments, post details, user profiles and Black Cat complaints need no login or cookies
  • Rate-limit handling — Exponential backoff on 418/429 errors
  • Delta mode (deltaMode, off by default) — in search, user_posts and hot_timeline (and per company list in complaints): each run returns only rows that no earlier run with the same deltaStateKey delivered (deduped by ID across runs). In search, a delta run reads only Weibo's first result window. Leave it off to get the full current result set on every run. How it works
  • 🆕 Sentiment scoring (optional) — set sentimentAnalysis: true to tag every post & comment with Chinese sentiment: polarity (positive/neutral/negative) + a −1.0…+1.0 score (SnowNLP model for Chinese text, keyword fallback for English). Built for brand-sentiment tracking & alt-data pipelines. Loads a model, so run with memory ≥512 MB when enabled.
  • 🆕 Auto-localize brand search — search a Latin brand name (e.g. Nike) and the Actor automatically also searches its Chinese name (耐克), and the two names take turns page by page, merged and deduped, so both share maxResults (billed per result, no extra cost). On by default; add custom variants via searchAliases.
  • 🆕 Date window in search (searchDateFrom / searchDateTo) — restrict a keyword search to a period, in Beijing time (+08:00) — the same zone every row's own createdAtIso carries, so the days you ask for are the days you see in the data. Weibo's API has no lower-bound parameter, so the Actor starts at your end date and pages backwards until the posts are older than your start date; anything outside the window is dropped before delivery, so you are not charged for it. The status message reports the oldest and newest post actually delivered and says plainly when the run stopped before reaching your start date. Both fields are optional — leave them empty and nothing about the run changes. Details
  • 🆕 Author profiles (includeAuthorProfiles, off by default) — in search, hot_timeline and user_posts, one user_profile-shaped row per distinct post author (followers, following, post count, verification, bio, location), up to 50 authors per run, each charged as one result ($0.035). Details
  • 🆕 Follow-up comments (followUpComments, off by default) — for a search monitor with deltaMode: each delivered post's thread is re-read once, 1 to 7 days after it was posted, and only comments never delivered before come back, each charged as one result ($0.035). Details

How to Use

Get the current Weibo hot search — the real-time pulse of Chinese internet.

{
"mode": "hot_search",
"maxResults": 50
}

2. Post Comments (no cookies needed)

Extract comments from specific posts. Provide post IDs (mid) or detail URLs.

{
"mode": "post_comments",
"postIds": ["5285773987283226"],
"maxComments": 50
}

3. Search Posts

Search by keyword in Chinese or English. Returns query-relevant results — no cookies needed.

💡 Brand/topic monitoring — autoLocalize does the Chinese↔Latin step for you. Weibo indexes by Chinese text, so 耐克 returns more posts than Nike. With autoLocalize on (the default), searching a common Latin brand name automatically also searches its Chinese name, alternating pages between the two and merging the results, capped at maxResults so it costs no extra. For brands outside the built-in dictionary, add the Chinese term yourself via searchAliases. (Few results for an English keyword usually means the language, not an error.)

{
"mode": "search",
"searchQuery": "人工智能",
"maxResults": 50
}

Brand search with auto-localize — searches Nike and 耐克 (automatic), plus any aliases you add, merged and deduped and capped at maxResults:

{
"mode": "search",
"searchQuery": "Nike",
"searchAliases": ["AJ", "Air Jordan"],
"maxResults": 100
}

🆕 Date window — search a specific period (searchDateFrom / searchDateTo)

Both fields are optional. Leave them empty and the Actor behaves exactly as it always has: same requests, same rows, same price.

{
"mode": "search",
"searchQuery": "人工智能",
"searchDateFrom": "2026-09-01",
"searchDateTo": "2026-09-03",
"maxResults": 500
}
FieldMeaning
searchDateFromEarliest post to return. A plain date is 00:00:00 Beijing time of that day.
searchDateToLatest post to return. A plain date is 23:59:59 Beijing time of that day, so 2026-09-01 → 2026-09-03 covers all three whole days. Leave empty to start from now.

Dates are Beijing time (UTC+08:00) unless you write an offset. That is deliberate: Weibo stamps every post +0800, and the createdAtIso field on every row the Actor delivers carries that offset. If the window were read as UTC, asking for 2026-09-01 → 2026-09-03 would hand you rows dated 2026-09-04 in your own data and silently miss the first 8 hours of 2026-09-01. Reading the window in Beijing time means the calendar days you ask for are exactly the calendar days you can group by.

Accepted forms:

You writeIt means
2026-09-012026-09-01 00:00:00 +08:00 (and 23:59:59 +08:00 for searchDateTo)
2026-09-01T08:00:002026-09-01 08:00:00 +08:00 — no offset means Beijing
2026-09-01T08:00:00Zexactly that UTC instant — an offset you write is always honoured
2026-09-01Zthe whole of 2026-09-01 in UTC

The run's status message prints the window in both zones, so the two never have to be guessed at.

How it works, and what it cannot promise. Weibo's search API has no lower-bound parameter — only an end time (measured 2026-09-23: starttime and timescope are both ignored; endtime is honoured). So the Actor starts at searchDateTo and pages backwards, window by window, until the posts it receives are older than searchDateFrom. Consequences worth knowing before you buy rows:

  • Posts outside your window are dropped before delivery, so you are never charged for them. The same goes for the rare row Weibo returns with an unreadable date.
  • maxResults, your run's charge limit and the run timeout still cap the walk. When one of them stops the run before it reaches searchDateFrom, the status message says so and names the reason — it never implies the period is fully covered.
  • Even a window walked all the way back is a sample of the period, not a complete archive: Weibo caps how far each result window reaches.
  • The status message always reports the oldest and newest post actually delivered, so you can see the real span you paid for.
  • In deltaMode the dates still filter the rows, but the deeper history walk stays off (as it always is in delta mode), so a delta run only checks Weibo's first result window against your dates. The status message says this too.
  • A date the Actor cannot parse, or a searchDateFrom at or after searchDateTo (an empty window), fails the run with an explanatory message and charges nothing — deliberately, so a typo can never quietly bill you for a different period. A searchDateTo in the future simply starts from now.

🆕 Author profiles with the posts (includeAuthorProfiles, off by default — charged per row)

Weibo's search sends no follower count with a post, so authorFollowers is null on search rows. With includeAuthorProfiles: true (in search, hot_timeline and user_posts), after the posts are collected the run also fetches the public profile of each distinct post author, with no login, and adds it as its own row:

  • Same fields as a user_profile row — userId, screenName, followersCount, friendsCount, statusesCount, verified, verifiedType, verifiedReason, description, location, createdAt and the rest — tagged mode: "author_profile". Join it to the posts on authorId = userId.
  • Where a profile came back, it also fills the post's authorFollowers / authorFollowing that Weibo left null (a value Weibo did send is never overwritten).
  • Cost: each profile row is one result, charged like a post ($0.035). 100 posts by 60 different authors come to 160 results ($5.60 instead of $3.50). Each author is profiled once per run.
  • Limits: at most 50 authors per run (the first 50 in post order), one author every 2 s (two requests each), and at most about 3 minutes of profiles per run; never more profiles than your run's charge limit can pay for (posts come first). The rows are delivered once the profiles are in, so a run with this on can take up to ~3 minutes longer. The status message says how many authors were left out and why.
  • Nothing invented: an author whose profile request fails, or whom Weibo does not resolve, gets no row and no charge, and their posts keep authorFollowers: null. A field Weibo did not send is null.
  • In deltaMode only the authors of the run's NEW posts are profiled; a profile row is never added to the delta memory, so an author can be charged again on a later run that brings a new post of theirs. In user_posts, an account whose creator row is already in the run is not fetched again.
{
"mode": "search",
"searchQuery": "耐克",
"maxResults": 50,
"includeAuthorProfiles": true
}

4. User Posts

Returns a user's posts. Requires your cookieString — Weibo serves timelines only to signed-in sessions. Provide numeric user IDs or profile URLs.

{
"mode": "user_posts",
"userIds": ["1642634100"],
"maxResults": 50,
"cookieString": "SUB=...; SUBP=...; (paste your full Cookie request header)"
}

5. Hot Search Delta — scheduled trend monitor (no cookies needed)

Run this on a schedule (hourly or daily) and each run reports what changed on the trending board since the previous run, instead of a flat snapshot. Every topic is tagged new, rising, falling, steady, or dropped, with rank movement, hot-value change, how long it has been trending, and its running peak.

{
"mode": "hot_search_delta",
"deltaStateKey": "default"
}

State persists across runs in a named store, so the first run sets a baseline and every run after it shows the deltas. Use different deltaStateKey values to track independent streams (e.g. hourly vs daily).

maxResults sets how many topics from the top of the board get a row (the default of 100 covers the whole ~50-topic board). The whole board is always compared, so a topic that slips just below your cut is not reported as dropped, and it keeps its firstSeenAt and peaks when it climbs back. dropped means it left the board entirely.

6. User Profiles (no cookies needed) 🆕

One row per Weibo account: screen name, followers, following, total posts, verification (type and reason), bio, gender, location, avatar, cover image, custom domain, account creation date, birthday, company and Weibo's credit rating. No login needed — Weibo serves profiles to the anonymous session this Actor already uses. Paste numeric user IDs or profile URLs (weibo.com/u/<id> or weibo.com/<id>).

{
"mode": "user_profile",
"userIds": ["2803301701", "https://weibo.com/u/1642634100"]
}
  • $0.035 per profile (the same item-scraped result as every other row), charged only for profiles actually delivered.
  • An ID Weibo does not resolve (deleted, banned, or not a user) returns no row and is not charged. Duplicate IDs are fetched and charged once (02803301701 and 2803301701 are the same account to Weibo, so they count as duplicates). maxResults does not apply here: you get one row per ID.
  • Every ID that did not come back is accounted for. The run status names up to 10 per reason; the full list is saved in the run's default key-value store as the record PROFILE_REPORT (notFound, failed, notFetched with notFetchedReason, and retryIds — the IDs worth pasting into a re-run). A run that delivered every ID writes no record.
  • Profiles are delivered in batches of 25 as the run goes, so a run cut off by its timeout keeps what it already delivered. If Weibo refuses or fails 5 IDs in a row, the run renews its anonymous session once, and if that does not help it stops rather than keep hitting Weibo: the rest are listed as not fetched (and not charged), to re-run later.
  • A field Weibo does not send is null, never an invented "", false or 0. location is what the account shows on its profile (self-declared), not an IP location.
  • Follower and following lists are not offered. Weibo answers those endpoints only to a logged-in session (an anonymous request gets an empty {}), so they cannot be delivered reliably without an account. The counts (followersCount, friendsCount) are in every row.
  • Great on a Schedule to track follower growth of a list of accounts: each run returns (and charges) every profile's current counts.

7. Black Cat (黑猫投诉) complaints (no cookies needed) 🆕

Black Cat (黑猫投诉, tousu.sina.com.cn) is Sina's consumer-complaint platform — Sina also runs Weibo. Chinese consumers file a complaint against a named company, say what they want (a refund, an apology, compensation), and the company answers on the record; the site's own counter showed 38.5 million valid complaints on 24 Sep 2026. This mode returns one row per complaint with no login: the company complained about, the complaint's title and text, what the consumer asks for (投诉要求), the problem (投诉问题), the amount in dispute where Black Cat shows it, the complaint's progress, when it was filed, and a link to the complaint page. The complainant's name and avatar are never included.

Two sources:

  • complaintCompanyIds set → each company's complaint list (its 最新投诉 tab), companies read in the order given, with maxResults shared across them. Black Cat ranks a list by recent activity, not by filing date: a complaint filed weeks ago that just got a reply sits next to today's (seen live on 24 Sep 2026).
  • complaintCompanyIds empty → Black Cat's public latest-complaints feed across all companies: what China is complaining about right now. The status message says which source the run used.

Finding a company ID (couid): open any complaint about the company on tousu.sina.com.cn and click the company name next to 投诉对象. Its page URL looks like https://tousu.sina.com.cn/company/view/?couid=2092643773 — paste that URL, or just the number after couid=.

{
"mode": "complaints",
"complaintCompanyIds": [
"https://tousu.sina.com.cn/company/view/?couid=2092643773",
"5650743478"
],
"maxResults": 200
}

Output — the shape of one row. Black Cat's complaint pages forbid reposting their content without permission, so the text values below are placeholders, not a real complaint; the field names, types and label values are exactly what a run returns. A company-list row carries amountInvolvedCny: null; a public-feed row carries the amount Black Cat shows:

{
"mode": "complaints",
"source": "feed",
"complaintId": "17400000000",
"title": "<complaint title as served>",
"summary": "<complaint text as the list serves it>",
"isTruncated": false,
"appeal": "退款",
"issue": "<the problem, as the complainant tagged it>",
"amountInvolvedCny": 99,
"companyName": "<company complained about>",
"companyId": "2092643773",
"statusCode": 6,
"status": "已回复",
"statusEn": "replied",
"createdAt": "2026-09-24T15:55:40+08:00",
"createdAtTimestamp": 1790236540,
"upvoteCount": 0,
"commentsCount": 0,
"shareCount": 0,
"url": "https://tousu.sina.com.cn/complaint/view/17400000000/?sld=<token as served>",
"scrapedAt": "2026-09-24T08:13:41.000000+00:00"
}
  • $0.035 per complaint (the same item-scraped result as every other row), charged only for complaints delivered.
  • Depth: the first 500 complaints of a list — the most recently active. Black Cat shows anonymous visitors 50 pages of 10: page 50 answers, page 51 answers 登录查看更多内容 ("log in to see more") — measured 24 Sep 2026 on a company list and on the public feed, and Black Cat serves 10 per page whatever size is asked for. When a company has more, the run stops there and the status says so; the rest are not reachable without an account. If Black Cat ever serves an empty page before its own count is reached, the run stops there too and says so (endedEarly).
  • summary is what the list serves: the whole complaint, or a ~150-character preview ending in ... — then isTruncated is true. The complaint page (url) has the full text.
  • amountInvolvedCny (涉诉金额) is served by the public feed only; on company-list rows it is null, not a guessed 0. A 0 means Black Cat itself shows 0元.
  • status is Black Cat's own label — 已回复 (replied), 已完成 (completed), 处理中 (in progress), 待分配 (awaiting assignment), 通过审核 (passed review), 已关闭 (closed) — with the English in statusEn and the raw code in statusCode. A code Black Cat has not been seen to label keeps its statusCode with null labels.
  • companyName is the company complained about. companyId is the couid exactly as Black Cat serves it: on the public feed, the company complained about; on a company list Black Cat echoes the ID you asked for on every row. A parent company's list also carries its sub-brands' complaints (京东客服's list holds 京东外卖客服专线 ones), so such a row names the sub-brand in companyName and carries the parent's ID. The status labels such a list with the names its rows carry, e.g. 5650743478 (京东客服 / 京东外卖客服专线).
  • createdAt is Beijing time (+08:00) and createdAtTimestamp the raw Unix seconds. On a company list it is the record's created_at; on the public feed it is the time the complaint page shows as published (发布于). Of 4 complaints seen in both lists on 24 Sep 2026, 3 carried the same time; the fourth was created at 02:34 and published at 15:55 the same day.
  • url is the complaint page exactly as Black Cat serves it: it carries an sld token, and the same page without it answers 页面不存在 (page does not exist). It is never built from the complaint number.
  • An ID Black Cat rejects (code 10002 参数错误) or answers with no complaints — an unknown ID looks exactly like a company with none — gives no row and no charge, and the status names it. Duplicate IDs are read once. Black Cat answers some companies under more than one ID (1003626 and 2092643773 both return 美团's list): a second ID whose page holds only complaints the first already served, with a matching complaint count (within 0.2%), is not read twice, and the status names it (sameListAs). A sub-brand listed next to its parent is still read in full, although many of its complaints are also in the parent's list. A complaint that turns up twice in a run is delivered, and charged, once.
  • Every list not delivered in full is accounted for in the run's default key-value store record COMPLAINTS_REPORT (invalid, empty, failed, depthCapped, endedEarly, sameListAs, notFetched with notFetchedReason, heldBackComplaintIds, perList with the company names each list carried, and retryCompanyIds — the IDs worth re-running). A run with nothing left out writes no record.
  • Polite to the source: at least 0.5 s between requests, and the run stops after 5 requests in a row fail or are refused, listing the companies it did not reach. Rows are delivered, then charged, in batches of 25, so a run cut off by its timeout keeps what it already delivered.
  • Scheduling (e.g. a daily Apify Schedule): with deltaMode on, each run returns only complaints that no earlier run with the same deltaStateKey delivered, tracked by complaint number for each company list (or the feed). Each list is read to its end or to the 500-complaint depth, because complaints already delivered that just got a reply sit above ones never delivered; those already delivered are skipped and not charged. With deltaMode off, each run returns the top of each list again, up to maxResults.
  • sentimentAnalysis scores each complaint's title and summary. includeComments, geoRollup, adFilter and cookieString are Weibo options: set in this mode, the status names them as not applied.
  • No keyword search. Black Cat's search endpoint now answers with a web page rather than data, so this mode reads company lists and the public feed only.

8. Post details (campaign / KOL post tracking) (no cookies needed) 🆕

One row per post you list: the post's current text and engagement — repostsCount, commentsCount, attitudesCount (likes) — plus the author (authorName, authorId, authorVerified, authorAvatar), createdAt / createdAtIso, source, images, videoUrl, isRepost and the quoted post in repostOf, and postUrl. No login needed, and no comments are fetched. The row has the same keys as the post row post_comments returns (mode: "post_detail"), so the two join on postId.

{
"mode": "post_details",
"postIds": [
"https://weibo.com/7087906569/RjB5HwkDj",
"https://m.weibo.cn/detail/5346708137705813",
"5346634918789689"
],
"maxResults": 100
}

Accepted forms in postIds: the numeric post ID (5346708137705813), the short code Weibo puts in post links (RjB5HwkDj), https://weibo.com/<uid>/<id or code>, https://m.weibo.cn/detail/<id> and https://m.weibo.cn/status/<id or code>. The short code is the numeric ID written in base 62, so the Actor resolves it without a request: the first two entries above are the same post, fetched and charged once. Several posts pasted into one entry (url1,url2, a spreadsheet row) count as one entry each. A bare code that reads as a plain word (a header cell such as Marketing) is not taken as a post: it is named as skipped, and a real code of that shape is accepted in its URL form. The status names every post it left out as you pasted it.

Output — the shape of one row (illustrative text; the nulls are what Weibo's post endpoint leaves out, measured 25 Sep 2026):

{
"postId": "5346708137705813",
"mid": "5346708137705813",
"text": "【#示例话题#】帖子的完整正文……",
"isTruncated": false,
"createdAt": "Thu Sep 24 16:20:19 +0800 2026",
"createdAtIso": "2026-09-24T16:20:19+08:00",
"source": "微博视频号",
"repostsCount": 466,
"commentsCount": 330,
"attitudesCount": 493,
"authorName": "示例账号",
"authorId": "7087906569",
"authorDescription": null,
"authorFollowers": null,
"authorFollowing": null,
"authorVerified": true,
"authorVerifiedReason": null,
"authorAvatar": "https://tvax3.sinaimg.cn/crop.0.0.500.500.1024/...",
"images": [],
"videoUrl": "http://f.video.weibocdn.com/...",
"region": null,
"regionEn": null,
"isPromoted": false,
"isRepost": false,
"repostOf": null,
"postUrl": "https://weibo.com/7087906569/5346708137705813",
"scrapedAt": "2026-09-25T00:49:26.349921+00:00",
"mode": "post_detail"
}
  • $0.035 per post (the same item-scraped result as every other row), charged only for posts delivered. maxResults caps the posts per run (the form pre-fills 25 — raise it for a longer list, up to 5,000 per run; split longer lists across runs; the status says how many were left out).
  • Run it on a Schedule to chart each post's engagement over time. Every run returns every listed post again, with its counts as of that run and a scrapedAt timestamp, so the rows of successive runs form a time series per postId. Each run is its own dataset and is charged per post returned.
  • Full text: Weibo's post endpoint cuts long posts too (measured: 150 of 880 characters). With fetchFullText on (the default, no extra charge) the full text is fetched; isTruncated: true remains only when Weibo would not serve it, and the status names those posts.
  • A value Weibo does not send is null, never an invented 0, "" or false. Weibo's post endpoint does not carry the author's follower count or bio (authorFollowers and authorDescription are null; user_profile has them), and many posts carry no location.
  • repostsCount tops out at 1,000,000. Above that Weibo reports exactly 1000000 (it shows 100万+): measured 25 Sep 2026, 2 of 10 hot posts carried exactly 1,000,000 reposts next to exact comment and like counts. Read 1,000,000 as "one million or more".
  • A post Weibo does not show — deleted, hidden, or never existed (Weibo answers 该微博不存在) — gives no row and is not charged. The status names up to 10 per reason; the full list is in the run's default key-value store record POST_DETAILS_REPORT (notFound, failed, notFetched with notFetchedReason, heldBack, previewOnly, unparsed, weiboMessages — what Weibo said — inputs, and retryIds, the IDs worth re-running). A run that delivered every post writes no record.
  • Polite to Weibo: at least 2 s between requests. Rows are delivered, then charged, in batches of 25, so a run cut off by its timeout keeps what it delivered. If Weibo refuses or fails 5 posts in a row, the run renews its anonymous session once, then stops and lists the rest as not fetched (not charged).
  • sentimentAnalysis scores each post's text. includeComments, geoRollup, adFilter, deltaMode and the search date window do not apply here; set them and the status says so. Leave postIds empty and the run takes 3 of Weibo's current hot posts, so you can see the output before pasting your own; those 3 are charged like any other post ($0.105).

⏰ Set up daily monitoring in 2 minutes

Most of this Actor's value is in recurring runs. A single pull is a one-off snapshot — but a daily or hourly schedule turns it into a continuously-updated Chinese brand / public-opinion feed. That's where pay-per-result compounds: instead of paying once for a static dump, you build a living dataset that tracks how the conversation moves week over week.

  1. Run the Actor once with your input — a brand keyword in search mode, or hot_search to capture the trending board — and check the output looks right.
  2. Apify Console → Schedules → Create → pick this Actor and your saved input. (Even faster: open any finished run and click Schedule to reuse its exact input.)
  3. Set a cron expression and save — e.g. 0 8 * * * = daily at 8am, or 0 * * * * = hourly. While you're there, enable the email notification on failed runs so you hear about a hiccup without checking manually.

Each scheduled run delivers its own fresh dataset, so you build a continuously-updated history with zero manual work — perfect for sentiment trend lines, brand-mention velocity, and time-series alt-data.

🔁 How a scheduled run can behave:

  • deltaMode off (the default) — each run returns the current result set, up to maxResults: a full snapshot every time, so a post that was in an earlier run can come back in the next one.
  • deltaMode on (in search, user_posts and hot_timeline) — each run returns only rows that no earlier run with the same deltaStateKey delivered: posts are deduped by post ID across runs, comment rows (includeComments) by their own comment ID, and rows an earlier run delivered are skipped (not delivered, not charged). In search the deeper history walk is off, so a delta run reads only Weibo's first result window. On a narrow keyword or a single account a run can return a handful of rows, or none. Use a distinct deltaStateKey per query/user you track (e.g. nike-weekly).
  • What deltaMode remembers — up to 200,000 IDs per deltaStateKey (in complaints, per company list). A row Weibo returns again counts as recent; when the memory is full, the IDs Weibo has gone longest without returning are forgotten first, and a forgotten row that comes back is delivered and charged again.
  • hot_search_delta — purpose-built for scheduled trend-velocity tracking; each run returns (and bills) the whole trending board, every topic tagged new / rising / falling / steady / dropped versus the last run.

Example — a daily search snapshot (each run returns the query's newest posts, up to maxResults):

{
"mode": "search",
"searchQuery": "Nike",
"maxResults": 200
}

Optional variant — the same search with deltaMode (each run returns only posts no earlier run with this deltaStateKey delivered; one key per query you track):

{
"mode": "search",
"searchQuery": "Nike",
"deltaMode": true,
"deltaStateKey": "nike-daily",
"maxResults": 200
}

🆕 Follow-up comments on earlier posts (followUpComments, off by default — charged per row)

A delta search reads each post minutes after it is published, usually before anyone has replied, and the post then drops out of Weibo's newest results. So a delta monitor delivers the posts but almost never their replies (measured on saved search pages of 26 Sep 2026, 李宁 and 耐克, 291 posts: posts 0-1 h old had 0-0.2 comments each, posts 6-24 h old about 1.1). followUpComments: true (only in search with deltaMode on) comes back for them:

  • The posts each delta run delivers are remembered with their posting time, in the same deltaStateKey record. A later delta run re-reads each one's comment thread once: on the first run at least 24 hours after it was posted and before it is 7 days old, newest posts first.
  • Only comments never delivered under this deltaStateKey come back, as ordinary comment rows (mode: "post_comment", linked by postId) marked isFollowUp: true — their post row came in an earlier run. Up to maxComments per post, reading at most 6 pages of each thread (about 120 new top-level comments). Once the key's delta memory is full (about 200,000 ids; it then forgets its oldest), a re-read returns only comments newer than the newest one already delivered for that post, so a forgotten comment is never delivered — or charged — twice.
  • Cost: each comment row is one result, charged like a post ($0.035). A thread with no new comments adds no row and no charge.
  • Limits: at most 300 threads, 400 comment requests and about 4 minutes of re-reading per run, 2 s apart, stopping 5 minutes before the run's timeout and at your run's charge limit; posts not reached keep their place for the next run. A post whose request failed is tried again on a later run while it is under 7 days old. A post that reaches 7 days without a re-read, or that the 10,000-post queue pushes out, leaves it unread; the status message counts them.
  • Run time: the run's own posts are delivered after the re-reads, so a run with this on can take up to ~4 minutes longer. Keep the schedule interval longer than a run lasts: two runs on the same deltaStateKey at the same time both see the same rows as new.
  • It starts with the posts delivered from the first run with it on (posts delivered before carry no posting time in the record). It does not run with adFilter: "promo_only", which would drop every comment row. Switch it off and the delta memory works exactly as before.
{
"mode": "search",
"searchQuery": "Nike",
"deltaMode": true,
"deltaStateKey": "nike-daily",
"maxResults": 200,
"followUpComments": true
}

🧠 Need this at AI-training-corpus scale?

If you're pulling Weibo's short-form posts, trending-topic chatter, and comment threads to train or fine-tune language models, the Chinese AI Training Corpus Engine assembles all 5 Chinese platforms — Weibo, RedNote, Bilibili, Douban, and Xueqiu — into AI-ready documents in one run: deduplicated, quality-scored, PII-scrubbed, and provenance-stamped for EU AI Act documentation, from $0.025/doc, with rejects and duplicates never charged.

📦 Want the full China feed? — China Monitoring Packages

Weibo is one platform. If you're monitoring a brand across Chinese social media, three pre-configured bundles combine this Actor with the Chinese Brand Monitor (Weibo + RedNote + Bilibili + Douban + Xueqiu in one scheduled call): Brand Starter ($110/mo, daily single-brand watch), Competitive Intel ($440/mo, brand + competitors share-of-voice), Fund Signal Desk (~$635/mo, multi-ticker daily sentiment). Self-serve — clone the preset, attach a Schedule, done. Support by text/issue tracker.

How to Get Cookies (for User Posts)

User posts needs a login cookie:

  1. Open weibo.com in your browser and log in
  2. Open DevTools (F12) → Network and reload the page
  3. Click any weibo.com request
  4. Copy the full Cookie request header (F12 → Network → any weibo.com request → Request Headers) and paste it into cookieString.

Leave cookieString empty for every other mode: a stale cookie replaces the anonymous session those modes use.

The cookie typically lasts several days before expiring.

Output Examples

{
"rank": 1,
"title": "人工智能最新突破",
"hotValue": 2847562,
"labelName": "热",
"isHot": true,
"isPromoted": false,
"url": "https://s.weibo.com/weibo?q=...",
"scrapedAt": "2026-04-10T12:00:00Z"
}

rank is Weibo's own board position. The board also carries a paid slot: that row ships with isPromoted: true and rank: null, and adFilter: "organic_only" removes it (unbilled).

Hot Search Delta record

Each record shows how a topic moved since the previous run (status ∈ new / rising / falling / steady / dropped):

{
"rank": 3,
"title": "某品牌新品发布",
"hotValue": 1820000,
"status": "rising",
"rankDelta": 5,
"hotValueDelta": 640000,
"previousRank": 8,
"firstSeenAt": "2026-06-04T08:00:00+00:00",
"minutesOnBoard": 120,
"peakRank": 3,
"peakHotValue": 1820000,
"isHot": true,
"isNew": false,
"url": "https://s.weibo.com/weibo?q=...",
"snapshotAt": "2026-06-04T10:00:00+00:00"
}

Post

{
"postId": "5285773987283226",
"text": "介绍一下我的老婆!@金莎",
"isTruncated": false,
"createdAt": "Wed Apr 09 12:49:23 +0800 2026",
"createdAtIso": "2026-04-09T12:49:23+08:00",
"repostsCount": 493,
"commentsCount": 4549,
"attitudesCount": 97438,
"authorName": "孙丞潇",
"authorId": "7511222755",
"authorFollowers": 0,
"authorVerified": false,
"images": ["https://wx1.sinaimg.cn/large/..."],
"videoUrl": "",
"isRepost": false,
"repostOf": null,
"postUrl": "https://weibo.com/7511222755/5285773987283226",
"scrapedAt": "2026-04-10T12:00:00Z"
}

Long posts come back in full. Weibo's list endpoints (search, user timelines, hot timeline) send only the first ~140-170 characters of a long post — about half of all search results. With fetchFullText (on by default, no extra charge) each cut-off post is fetched in full: measured 2026-09-18, a 100-row search had 52 truncated rows and 51 came back complete. The post quoted inside a repost (repostOf) and the post rows of post_comments and post_details are expanded the same way. A row (or its repostOf) still carries isTruncated: true only when Weibo would not serve the long text. createdAt keeps Weibo's raw date string; createdAtIso is the same moment in ISO 8601. videoUrl is a signed link that expires about 1 hour after the scrape, so download the video right away if you need it.

repostOf — el post original, sin peticiones extra. Weibo incrusta el post citado entero dentro de la misma respuesta y este Actor lo tiraba: sólo escribía isRepost: true. Medido el 26-ago sobre una búsqueda real de 80 filas, 23 eran reposts (29%). Su text no viene vacío —trae la cadena //@usuario: con el contenido citado dentro— así que esto no rescata filas inútiles: entrega estructurado lo que hoy llegaba como una cadena. Para vigilancia de marca, quién escribió el original y cuánto circuló es justo el dato buscado. null cuando la fila no es un repost.

Comment

{
"commentId": "5285813927600208",
"text": "恭喜恭喜!神仙眷侣,一定要狠狠幸福哦~",
"createdAt": "Thu Apr 09 12:51:31 +0800 2026",
"createdAtIso": "2026-04-09T12:51:31+08:00",
"likeCount": 1268,
"replyCount": 42,
"authorName": "吃瓜罗伯特",
"authorId": "6108685154",
"authorFollowers": 10660,
"authorVerified": true,
"authorRegion": "广东",
"authorRegionEn": "Guangdong",
"postId": "5285773987283226",
"postUrl": "https://weibo.com/detail/5285773987283226",
"scrapedAt": "2026-04-10T12:00:00Z"
}

Profile (user_profile)

{
"userId": "1642634100",
"screenName": "新浪科技",
"description": "新浪科技是中国最有影响力的TMT产业资讯及数码产品服务平台。…",
"gender": "m",
"location": "北京 海淀区",
"followersCount": 23800866,
"friendsCount": 2702,
"statusesCount": 265367,
"verified": true,
"verifiedType": 3,
"verifiedReason": "新浪科技官方微博",
"avatarHd": "https://tvax2.sinaimg.cn/crop.0.0.436.436.1024/...",
"coverImage": "https://ww2.sinaimg.cn/crop.0.0.640.640.640/...",
"domain": "sinatech",
"profileUrl": "https://weibo.com/u/1642634100",
"createdAt": "2009-08-28 21:07:40",
"birthday": "",
"company": "新浪网技术(中国)有限公司",
"sunshineCredit": "阳光信用极好",
"mode": "user_profile",
"scrapedAt": "2026-09-24T07:14:32Z"
}

Sentiment field (optional)

With sentimentAnalysis: true, every post and comment gains a sentiment object:

{
"polarity": "positive",
"score": 0.42,
"method": "snownlp"
}

polarity ∈ positive / neutral / negative · score ∈ −1.0…+1.0 (higher = more positive) · method is snownlp for Chinese text, or keyword for the English-only fallback.

Content is in Chinese

All content is returned in the original Simplified Chinese. Weibo is a Chinese-language platform — posts, comments, trending topics, and user bios are in Chinese.

If you need English translations, pipe the output through a translation API (Google Translate, DeepL, or Claude).

Technical Details

  • No browser: pure HTTP — fast and lightweight, runs in 256 MB without sentiment (≥512 MB with sentimentAnalysis)
  • No authentication except user_posts (your cookie): the other modes read publicly accessible content only
  • Built-in rate limiting: automatic retry with exponential backoff to handle peak-hour throttling
  • Globally accessible: no VPN or proxy required
  • Clean structured JSON output: ready for analysis or downstream pipelines

Pricing

$35 per 1,000 results ($0.035 per result) — pay-per-event

Each scraped item (post, comment, trending topic, user profile, or Black Cat complaint) counts as one result. The two optional add-ons bill the same way: each author profile row (includeAuthorProfiles) and each follow-up comment row (followUpComments) is one result.

You only pay for delivered records — empty or failed results are never charged.

Typical costs (small-scale):

  • Top 50 trending topics snapshot: ~$1.75
  • 100 posts on a brand keyword: ~$3.50
  • 200 comments on a viral post: ~$7.00
  • 50 posts from one account: ~$1.75 (needs your cookieString)
  • 100 user profiles: ~$3.50 (no login)
  • 100 tracked posts, current engagement: ~$3.50 per run (no login)
  • 100 Black Cat complaints about a company: ~$3.50 (no login)

B2B / bulk-scale examples:

  • AI training corpus seed (10,000 posts on a topic): ~$350
  • Daily brand sentiment monitor (500 posts/day for a month): ~$525/month
  • Equity research signal (10 tickers × 200 posts daily): ~$2,100/month
  • Multi-source academic dataset (50,000 posts across 30 keywords): ~$1,750

No separate compute charge: you pay per result, plus Apify's tiny per-run start event.

Limitations

  • User posts needs your own login cookie for the timeline: without cookieString Weibo returns no posts, and the run delivers and charges nothing for that account.
  • Search, hot search, post comments, post details and user profiles work fully without authentication
  • Follower / following lists are not available: Weibo serves them only to logged-in sessions. user_profile returns the counts, not the lists
  • Only public data is accessible — private/locked accounts are not available
  • Weibo may rate-limit requests during peak hours — handled automatically with backoff
  • Very old posts may not be available
  • Black Cat complaints: the first 500 per list, the most recently active. Anonymous visitors see 50 pages of 10 per company (and of the public feed); the run says when a company has more. No keyword search: Black Cat's search endpoint answers with a web page, not data. Company lists do not carry the amount in dispute (amountInvolvedCny is null there; the public feed has it)
  • Date window is search mode only. searchDateFrom / searchDateTo do nothing in user_posts, user_profile, complaints, hot_timeline, hot_search, hot_search_delta, post_comments or post_details; set them there and the run says so in its status message rather than pretending to have filtered
  • Date window is a sample, not an archive. Weibo's search exposes only an end time, and caps how far back each result window reaches. searchDateFrom / searchDateTo bound what is delivered and billed, but a deep period can stop early on maxResults, your charge limit, the run timeout, or Weibo itself — the run's status message always names which, and reports the oldest and newest post actually delivered

FAQ

Is there a Weibo API?

There is no official public Weibo API available for international developers. Weibo's developer platform requires a Chinese business license and imposes strict rate limits. This Weibo Scraper is an alternative — extract trending topics, posts, and comments without any official API access.

How much does it cost to scrape Weibo?

The Weibo Scraper is priced per result: $35 per 1,000 ($0.035 each) (pay-per-event). Each scraped item (post, comment, trending topic, user profile, or Black Cat complaint) counts as one result. You can start with Apify's free plan, which includes $5 of monthly credits.

Can I scrape Weibo in Python?

Yes. Install the Apify Python client (pip install apify-client), then use the ApifyClient to call the zhorex/weibo-scraper actor. See the Python code example above.

This scraper accesses publicly available data through Weibo's public web endpoints; user_posts reads your own signed-in session through the cookie you provide. It does not bypass authentication or access private/locked accounts. Always review your local laws and Weibo's terms of service before scraping.

What does this Weibo scraper cover?

The Weibo Scraper by Zhorex reads Weibo's public web endpoints (user_posts needs your cookie). It supports 9 modes (hot timeline, hot search, hot-search delta, post comments, post details, search, user posts, user profiles, and Black Cat (黑猫投诉) consumer complaints), handles rate limits automatically, and runs without a browser or VPN.

Integrations & data export

The Weibo Scraper integrates with your existing workflow tools:

  • Google Sheets — Send scraped Weibo data directly to a spreadsheet
  • Zapier / Make / n8n — Automate workflows triggered by new Weibo data
  • REST API — Call the actor programmatically and retrieve results via Apify's REST API
  • Webhooks — Get notified when a scraping run finishes and process data in real time
  • Data formats — Download results in JSON, CSV, Excel, XML, or RSS

More scrapers by Zhorex

Chinese Digital Intelligence Suite

  • 🆕 Chinese AI Training Corpus Engine — Weibo + RedNote + Bilibili + Douban + Xueqiu into AI-training-ready documents (MinHash dedup, quality scoring, PII scrub, EU AI Act provenance)
  • 🆕 Chinese Brand Monitor — Cross-platform brand mention aggregator (Weibo + RedNote + Bilibili + Douban + Xueqiu, sentiment + dedup)
  • Bilibili Scraper — China's video platform: danmaku, comments, Gen-Z creator analytics
  • RedNote (Xiaohongshu) Scraper — China's Instagram + Pinterest (lifestyle, consumer reviews)
  • RedNote Shop Scraper — RedShop e-commerce (products, vendors, prices)
  • Douban Scraper — Long-form reviews, ratings, group discussions (movies/books/music)
  • Xueqiu Scraper — Chinese stock-discussion sentiment, cashtag indexing (SH/SZ/HK/US-listed Chinese)

Streaming & Video

Markets & Alt-Data

Price & Availability Monitors

Same scheduling pattern as this actor's monitors — schedule them and every run is tagged with what changed.

B2B Reviews

Other Tools

Support

Having issues? Open an issue on the Actor page.


Reviews ⭐

Reviews, good or bad, help other users decide: leave one on the reviews page.

Found a bug or missing field? Open an issue.