Weibo Scraper — Hot Search, Posts, Comments & Creator Feeds
Pricing
from $35.00 / 1,000 item scrapeds
Weibo Scraper — Hot Search, Posts, Comments & Creator Feeds
Sina Weibo (微博) scraper API, no login: keyword search, trending posts, hot-search board (微博热搜), comments, user profiles and post details, plus Black Cat (黑猫投诉) consumer complaints. Creator timelines (user_posts) need your cookie. Nine modes for social listening and China market research.
Pricing
from $35.00 / 1,000 item scrapeds
Rating
1.0
(1)
Developer
Sami
Maintained by CommunityActor stats
3
Bookmarked
571
Total users
252
Monthly active users
2 days ago
Last modified
Categories
Share
Weibo Scraper - Chinese Social Intelligence
Extract Chinese public opinion, trending topics, and real-time consumer sentiment from Weibo (微博) — China's dominant microblog with 580M+ monthly users producing a dense public-opinion signal. Built for AI training corpora, Chinese consumer equity research alt-data, brand monitoring agencies, and academic NLP teams. No API key, no VPN, and no login except user_posts (your cookie).
💡 Tracking a brand, not just scraping one platform? Pair this with two recurring monitors that watch the whole picture for you: Chinese Brand Monitor — Weibo + RedNote + Bilibili + Douban + Xueqiu mentions in one scheduled, deduped, sentiment-tagged feed — and AI Brand Visibility Monitor — how the AI engines (DeepSeek, Qwen, Kimi, GLM + ChatGPT/Gemini/Claude) answer about your brand.
How to scrape Weibo in 3 easy steps
- Go to the Weibo Scraper page on Apify and click "Try for free"
- Configure your input — choose a mode (
hot_timeline,hot_search,hot_search_delta,post_comments,post_details,search,user_posts,user_profile, orcomplaintsfor Black Cat consumer complaints), enter your keywords, post IDs or URLs, user IDs or company IDs, and set the number of results - Click "Run", wait for the scraper to finish, then download your data in JSON, CSV, or Excel format
No coding required. No API key. Works with Apify's free plan.
📅 Need a specific period, or a deeper search? In
searchmode, setsearchDateFrom/searchDateTo(Beijing time, +08:00) and the run pages back fromsearchDateTothrough that period, up to yourmaxResults; with or without dates, a search keeps paging past Weibo's first result window (about 1,000 posts) toward yourmaxResults(up to 5,000 per run; this deeper walk is off indeltaMode). Need the accounts rather than their posts?user_profilereturns followers, following, post count, verification, bio and location per user ID, no login. Tracking specific posts?post_detailsreturns each post's current text, reposts, comments and likes from its URL, no login. Andcomplaintspulls Black Cat (黑猫投诉) consumer complaints about the companies you list, also no login. → Date window details
Part of the Chinese Digital Intelligence Suite
Built by Zhorex, who maintains a suite of Chinese-platform scrapers — built specifically for AI training data buyers, equity research analysts covering Chinese consumer stocks, and brand monitoring teams:
- 🆕 Chinese Brand Monitor — Cross-platform brand mention aggregator (Weibo + RedNote + Bilibili + Douban + Xueqiu in one normalized feed, sentiment-tagged, cross-platform deduped — $0.06/mention)
- Weibo Scraper — You are here (microblogging, hot search, real-time public opinion)
- Bilibili Scraper — China's video platform: danmaku, comments, Gen-Z creator sentiment
- RedNote (Xiaohongshu) Scraper — China's Instagram + Pinterest (lifestyle, consumer reviews)
- RedNote Shop Scraper — RedShop e-commerce (products, vendors, prices)
- Douban Scraper — Long-form reviews (movies/books/music), group discussions
- Xueqiu Scraper — Chinese stock-discussion sentiment, cashtag indexing
Together, these cover the five pillars of Chinese consumer signal: microblog opinion, video sentiment, lifestyle reviews, e-commerce, and long-form opinion. Most analysts buy 2-4 of these for cross-platform coverage. Building a cross-platform brand monitoring pipeline? The Chinese Brand Monitor aggregator gives you all 5 platforms in one normalized output — saves 4-6 hours of engineering vs. orchestrating individual scrapers.
Who buys this scraper
| Buyer profile | Use case |
|---|---|
| AI / LLM training data teams | Real-time Chinese microblog text for SFT corpora + current-events grounding |
| Hedge fund / equity research desks | Brand mention velocity, hot-search momentum as alt-data on Chinese consumer stocks (POP MART, BYD, Anta, Yum China, BeiGene) |
| Brand monitoring agencies | Real-time tracking of Western brand mentions, crisis detection on China's public square |
| Geopolitical / policy analysts | Monitor Chinese public discourse, narrative tracking, policy response signal |
| Academic NLP / sentiment researchers | Chinese microblog corpus, labeled sentiment data for classifier training |
| Journalists / investigative teams | Source Chinese public opinion data for reporting on consumer brands, viral events |
🆕 Get the replies, not just the posts. Turn on
includeCommentsand every post found also returns its comment thread — one row per comment, taggedmode: post_commentand linked bypostId. The post is the prompt; the replies are the sentiment. No login or cookie needed.Where the volume actually is: a narrow, high-engagement query returns far bigger threads than a broad one — a generic keyword mostly returns posts with zero comments. Posts with no thread are skipped automatically, so you are never billed for an empty round-trip. In
search, the newest posts rarely have replies yet: setsearchDateTo2 or more days back (Beijing time) and the same keyword returns older posts, which have had time to collect them.It multiplies rows and cost — 100 posts × 50 comments is ~5,100 rows instead of 100 — so it is off by default and capped by
maxComments.📈
maxResultsnow goes to 5,000 per run (was 500), andmaxCommentsto 1,000. Defaults are unchanged, so nothing about your existing runs or bill changes unless you raise them yourself.
🆕 hot_timeline — the trending POSTS, not the trending words
There are two different "hot" things on Weibo and confusing them costs you the data:
| Mode | Returns | Has comment threads? |
|---|---|---|
hot_search | the trending board — topic words and their rank | No. They are not posts. |
hot_timeline | the posts Weibo's own hot feed is pushing right now | Yes, and they are enormous. |
Measured anonymously on 2026-07-28, the posts this mode returned carried 633,714 / 618,338 / 493,677 reposts and 122,053 / 264,451 comments. That is the difference between this and a keyword search: a broad search returns mostly posts with zero replies, these do not.
So this is the mode to pair with includeComments. One measured run:
hot_timeline (maxResults: 5) → 5 rowshot_timeline + includeComments (maxComments 20) → 89 rows (5 posts + 84 comments)hot_timeline + includeComments (maxComments 50) → 114 rows (5 posts + 109 comments)
An 18-23x multiplier from the same 5 posts, depending on maxComments — measured
2026-08-05, and all 5 posts carried a live thread.
{"mode": "hot_timeline","maxResults": 20,"includeComments": true,"maxComments": 100,"sentimentAnalysis": true}
At $0.035 per row that run bills roughly $4.10 per 5 posts of full threads.
What is Weibo?
Weibo (微博) is China's dominant microblogging platform — think Twitter meets Instagram. With 580M+ monthly active users, it's where Chinese public opinion forms, brands communicate, and news breaks. Government officials, celebrities, and brands all maintain active Weibo accounts. For data buyers, Weibo's hot search ranking is the closest thing China has to a real-time barometer of public attention.
Weibo API alternative
There is no official public Weibo API available for international developers. Weibo's developer API requires a Chinese business license, has severe rate limits, and returns limited data. This Weibo Scraper is an alternative — it extracts trending topics, posts, and comments without any official API access. No Chinese business registration needed.
Use Cases by buyer
| Who | Why they use it |
|---|---|
| AI / LLM training data teams | Real-time Chinese-language microblog text for SFT, RLHF training, and current-events grounding for Chinese LLMs |
| Equity research / hedge funds | Hot-search velocity + brand mention spikes as alt-data leading indicator on Chinese consumer stocks |
| Brand monitoring teams | Real-time tracking of brand mentions, viral content, and crisis detection on China's public square |
| Geopolitical / policy analysts | Monitor public discourse on policy, international topics, and narrative trends |
| PR & communications | Track brand mentions and sentiment shifts in real time |
| Competitive intelligence | Track Chinese competitor announcements, product launches, and audience reception |
| Influencer marketing | Find and evaluate Weibo KOLs by followers, engagement, verification status |
| Journalism | Access Chinese public opinion data for investigative reporting |
| Academic research | Pre-built Chinese microblog corpus with engagement metrics for NLP and sociology studies |
Scrape Weibo with Python, JavaScript, or no code
You can use the Weibo Scraper directly from the Apify Console (no code), or integrate it into your own scripts with Python or JavaScript.
Python
from apify_client import ApifyClientclient = ApifyClient("YOUR_API_TOKEN")run = client.actor("zhorex/weibo-scraper").call(run_input={"mode": "hot_search","maxResults": 50})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });const run = await client.actor('zhorex/weibo-scraper').call({mode: 'hot_search',maxResults: 50,});const { items } = await client.dataset(run.defaultDatasetId).listItems();items.forEach((item) => console.log(item));
Using the raw REST API (Postman / curl)
⚠️ The run endpoint is asynchronous — its response is the run object (IDs + status), NOT your scraped data. If you
POSTto/acts/.../runsyou get back something like{ "data": { "status": "READY", "defaultDatasetId": "…" } }with no results in it — that's expected, the run hasn't finished yet. The records land in the run's dataset, not in that response. (ThecontainerUrllink is the live container; once a run finishes it just shows "run has already finished with status SUCCEEDED" — that means success, it is not where the data lives.)
Easiest — one call that waits for the run and returns the records directly:
curl -X POST "https://api.apify.com/v2/acts/zhorex~weibo-scraper/run-sync-get-dataset-items?token=YOUR_API_TOKEN" \-H "Content-Type: application/json" \-d '{"mode":"hot_search","maxResults":50}'
The response body is the JSON array of records — no second call needed.
Or async — start the run, then fetch the dataset once it finishes:
# 1) start the run — note the "defaultDatasetId" in the responsecurl -X POST "https://api.apify.com/v2/acts/zhorex~weibo-scraper/runs?token=YOUR_API_TOKEN" \-H "Content-Type: application/json" -d '{"mode":"hot_search","maxResults":50}'# 2) when the run status is SUCCEEDED, fetch the records from its datasetcurl "https://api.apify.com/v2/datasets/DEFAULT_DATASET_ID/items?token=YOUR_API_TOKEN"
💡 In the Apify Console you can also open any run and click the Output / Storage → Dataset tab to view and download the same data as JSON / CSV / Excel.
Province rollup — where in China the conversation actually is
Turn on geoRollup and every run adds one row per location (Chinese provinces; overseas
countries flagged isOverseas), on top of the posts themselves:
| field | what it is |
|---|---|
province / provinceEn | 广东 / Guangdong (or 美国 / United States for a post sent from abroad) |
isOverseas | true for a foreign country, false for a Chinese province, Hong Kong, Macau or Taiwan, null for a label outside both lists |
postCount, sharePct | how many posts came from there, and what share of the located sample |
promotedCount | how many of those were promoted posts, not organic |
sentimentMean | average sentiment for that province, when sentimentAnalysis is on |
sampleSize, lowConfidence | the sample the row was computed from, and whether it is thin |
It costs no extra scraping. Weibo already attaches a location to each post
(发布于 广东) in the same payload the run pays for, so the rollup is computed from bytes
you have already fetched — no additional requests, no additional time.
Honest sampling. If fewer than 50 posts in the run carry a location, no province rows
are emitted and none are charged. Between 50 and 200 they ship with
lowConfidence: true, and every row states its own sampleSize so you can judge it
rather than trust it. Measured 2026-08-15 on a 150-post run: 136 posts carried a
location across 28 provinces.
Rollup rows bill as ordinary results. Works with search, hot_timeline and
user_posts.
Promoted posts are now labelled
Every post row carries isPromoted. Weibo mixes paid promotion into keyword results,
and until now those rows arrived indistinguishable from organic opinion. If you are
measuring public sentiment, filter them out; if you are tracking competitors' paid social,
they are the interesting ones.
Features
| Mode | What it does | Cookies needed? |
|---|---|---|
| Search Posts | Find posts by keyword — returns query-relevant results | No |
| Hot Search / Trending | Real-time trending topics with heat scores and badges (新 / 热 / 爆) | No |
| Hot Search Delta 🆕 | Scheduled trend monitor — the whole board each run, every topic tagged new / rising / falling / steady / dropped vs the last run (rank velocity, time-on-board, peaks) | No |
| Post Comments | Comments + post detail with engagement metrics | No |
| Post Details 🆕 | One row per post URL or ID: current text, reposts, comments, likes, author, date, images, video — no comments fetched | No |
| User Posts | Posts from specific accounts | Yes — your cookieString |
| User Profile 🆕 | One row per account: followers, following, post count, verification, bio, location, creation date | No |
| Complaints (黑猫投诉) 🆕 | Consumer complaints from Black Cat, Sina's complaint platform: the most recently active about the companies you list, or the latest across China — company, complaint text, demand, amount, progress | No |
- No browser needed — Pure HTTP, runs in 256 MB without sentiment (≥512 MB with
sentimentAnalysis) - No VPN needed — Globally accessible endpoints
- Zero setup — search, trending, comments, post details, user profiles and Black Cat complaints need no login or cookies
- Rate-limit handling — Exponential backoff on 418/429 errors
- Delta mode (
deltaMode, off by default) — insearch,user_postsandhot_timeline(and per company list incomplaints): each run returns only rows that no earlier run with the samedeltaStateKeydelivered (deduped by ID across runs). Insearch, a delta run reads only Weibo's first result window. Leave it off to get the full current result set on every run. How it works - 🆕 Sentiment scoring (optional) — set
sentimentAnalysis: trueto tag every post & comment with Chinese sentiment: polarity (positive/neutral/negative) + a −1.0…+1.0 score (SnowNLP model for Chinese text, keyword fallback for English). Built for brand-sentiment tracking & alt-data pipelines. Loads a model, so run with memory ≥512 MB when enabled. - 🆕 Auto-localize brand search — search a Latin brand name (e.g.
Nike) and the Actor automatically also searches its Chinese name (耐克), and the two names take turns page by page, merged and deduped, so both sharemaxResults(billed per result, no extra cost). On by default; add custom variants viasearchAliases. - 🆕 Date window in search (
searchDateFrom/searchDateTo) — restrict a keyword search to a period, in Beijing time (+08:00) — the same zone every row's owncreatedAtIsocarries, so the days you ask for are the days you see in the data. Weibo's API has no lower-bound parameter, so the Actor starts at your end date and pages backwards until the posts are older than your start date; anything outside the window is dropped before delivery, so you are not charged for it. The status message reports the oldest and newest post actually delivered and says plainly when the run stopped before reaching your start date. Both fields are optional — leave them empty and nothing about the run changes. Details - 🆕 Author profiles (
includeAuthorProfiles, off by default) — insearch,hot_timelineanduser_posts, oneuser_profile-shaped row per distinct post author (followers, following, post count, verification, bio, location), up to 50 authors per run, each charged as one result ($0.035). Details - 🆕 Follow-up comments (
followUpComments, off by default) — for asearchmonitor withdeltaMode: each delivered post's thread is re-read once, 1 to 7 days after it was posted, and only comments never delivered before come back, each charged as one result ($0.035). Details
How to Use
1. Trending Topics (no cookies needed)
Get the current Weibo hot search — the real-time pulse of Chinese internet.
{"mode": "hot_search","maxResults": 50}
2. Post Comments (no cookies needed)
Extract comments from specific posts. Provide post IDs (mid) or detail URLs.
{"mode": "post_comments","postIds": ["5285773987283226"],"maxComments": 50}
3. Search Posts
Search by keyword in Chinese or English. Returns query-relevant results — no cookies needed.
💡 Brand/topic monitoring —
autoLocalizedoes the Chinese↔Latin step for you. Weibo indexes by Chinese text, so耐克returns more posts thanNike. WithautoLocalizeon (the default), searching a common Latin brand name automatically also searches its Chinese name, alternating pages between the two and merging the results, capped atmaxResultsso it costs no extra. For brands outside the built-in dictionary, add the Chinese term yourself viasearchAliases. (Few results for an English keyword usually means the language, not an error.)
{"mode": "search","searchQuery": "人工智能","maxResults": 50}
Brand search with auto-localize — searches Nike and 耐克 (automatic), plus any aliases you add, merged and deduped and capped at maxResults:
{"mode": "search","searchQuery": "Nike","searchAliases": ["AJ", "Air Jordan"],"maxResults": 100}
🆕 Date window — search a specific period (searchDateFrom / searchDateTo)
Both fields are optional. Leave them empty and the Actor behaves exactly as it always has: same requests, same rows, same price.
{"mode": "search","searchQuery": "人工智能","searchDateFrom": "2026-09-01","searchDateTo": "2026-09-03","maxResults": 500}
| Field | Meaning |
|---|---|
searchDateFrom | Earliest post to return. A plain date is 00:00:00 Beijing time of that day. |
searchDateTo | Latest post to return. A plain date is 23:59:59 Beijing time of that day, so 2026-09-01 → 2026-09-03 covers all three whole days. Leave empty to start from now. |
Dates are Beijing time (UTC+08:00) unless you write an offset. That is deliberate: Weibo stamps every post +0800, and the createdAtIso field on every row the Actor delivers carries that offset. If the window were read as UTC, asking for 2026-09-01 → 2026-09-03 would hand you rows dated 2026-09-04 in your own data and silently miss the first 8 hours of 2026-09-01. Reading the window in Beijing time means the calendar days you ask for are exactly the calendar days you can group by.
Accepted forms:
| You write | It means |
|---|---|
2026-09-01 | 2026-09-01 00:00:00 +08:00 (and 23:59:59 +08:00 for searchDateTo) |
2026-09-01T08:00:00 | 2026-09-01 08:00:00 +08:00 — no offset means Beijing |
2026-09-01T08:00:00Z | exactly that UTC instant — an offset you write is always honoured |
2026-09-01Z | the whole of 2026-09-01 in UTC |
The run's status message prints the window in both zones, so the two never have to be guessed at.
How it works, and what it cannot promise. Weibo's search API has no lower-bound parameter — only an end time (measured 2026-09-23: starttime and timescope are both ignored; endtime is honoured). So the Actor starts at searchDateTo and pages backwards, window by window, until the posts it receives are older than searchDateFrom. Consequences worth knowing before you buy rows:
- Posts outside your window are dropped before delivery, so you are never charged for them. The same goes for the rare row Weibo returns with an unreadable date.
maxResults, your run's charge limit and the run timeout still cap the walk. When one of them stops the run before it reachessearchDateFrom, the status message says so and names the reason — it never implies the period is fully covered.- Even a window walked all the way back is a sample of the period, not a complete archive: Weibo caps how far each result window reaches.
- The status message always reports the oldest and newest post actually delivered, so you can see the real span you paid for.
- In
deltaModethe dates still filter the rows, but the deeper history walk stays off (as it always is in delta mode), so a delta run only checks Weibo's first result window against your dates. The status message says this too. - A date the Actor cannot parse, or a
searchDateFromat or aftersearchDateTo(an empty window), fails the run with an explanatory message and charges nothing — deliberately, so a typo can never quietly bill you for a different period. AsearchDateToin the future simply starts from now.
🆕 Author profiles with the posts (includeAuthorProfiles, off by default — charged per row)
Weibo's search sends no follower count with a post, so authorFollowers is null on search rows. With includeAuthorProfiles: true (in search, hot_timeline and user_posts), after the posts are collected the run also fetches the public profile of each distinct post author, with no login, and adds it as its own row:
- Same fields as a
user_profilerow —userId,screenName,followersCount,friendsCount,statusesCount,verified,verifiedType,verifiedReason,description,location,createdAtand the rest — taggedmode: "author_profile". Join it to the posts onauthorId=userId. - Where a profile came back, it also fills the post's
authorFollowers/authorFollowingthat Weibo leftnull(a value Weibo did send is never overwritten). - Cost: each profile row is one result, charged like a post ($0.035). 100 posts by 60 different authors come to 160 results ($5.60 instead of $3.50). Each author is profiled once per run.
- Limits: at most 50 authors per run (the first 50 in post order), one author every 2 s (two requests each), and at most about 3 minutes of profiles per run; never more profiles than your run's charge limit can pay for (posts come first). The rows are delivered once the profiles are in, so a run with this on can take up to ~3 minutes longer. The status message says how many authors were left out and why.
- Nothing invented: an author whose profile request fails, or whom Weibo does not resolve, gets no row and no charge, and their posts keep
authorFollowers: null. A field Weibo did not send isnull. - In
deltaModeonly the authors of the run's NEW posts are profiled; a profile row is never added to the delta memory, so an author can be charged again on a later run that brings a new post of theirs. Inuser_posts, an account whose creator row is already in the run is not fetched again.
{"mode": "search","searchQuery": "耐克","maxResults": 50,"includeAuthorProfiles": true}
4. User Posts
Returns a user's posts. Requires your cookieString — Weibo serves timelines only to signed-in sessions. Provide numeric user IDs or profile URLs.
{"mode": "user_posts","userIds": ["1642634100"],"maxResults": 50,"cookieString": "SUB=...; SUBP=...; (paste your full Cookie request header)"}
5. Hot Search Delta — scheduled trend monitor (no cookies needed)
Run this on a schedule (hourly or daily) and each run reports what changed on the trending board since the previous run, instead of a flat snapshot. Every topic is tagged new, rising, falling, steady, or dropped, with rank movement, hot-value change, how long it has been trending, and its running peak.
{"mode": "hot_search_delta","deltaStateKey": "default"}
State persists across runs in a named store, so the first run sets a baseline and every run after it shows the deltas. Use different deltaStateKey values to track independent streams (e.g. hourly vs daily).
maxResults sets how many topics from the top of the board get a row (the default of 100 covers the whole ~50-topic board). The whole board is always compared, so a topic that slips just below your cut is not reported as dropped, and it keeps its firstSeenAt and peaks when it climbs back. dropped means it left the board entirely.
6. User Profiles (no cookies needed) 🆕
One row per Weibo account: screen name, followers, following, total posts, verification (type and reason), bio, gender, location, avatar, cover image, custom domain, account creation date, birthday, company and Weibo's credit rating. No login needed — Weibo serves profiles to the anonymous session this Actor already uses. Paste numeric user IDs or profile URLs (weibo.com/u/<id> or weibo.com/<id>).
{"mode": "user_profile","userIds": ["2803301701", "https://weibo.com/u/1642634100"]}
- $0.035 per profile (the same
item-scrapedresult as every other row), charged only for profiles actually delivered. - An ID Weibo does not resolve (deleted, banned, or not a user) returns no row and is not charged. Duplicate IDs are fetched and charged once (
02803301701and2803301701are the same account to Weibo, so they count as duplicates).maxResultsdoes not apply here: you get one row per ID. - Every ID that did not come back is accounted for. The run status names up to 10 per reason; the full list is saved in the run's default key-value store as the record
PROFILE_REPORT(notFound,failed,notFetchedwithnotFetchedReason, andretryIds— the IDs worth pasting into a re-run). A run that delivered every ID writes no record. - Profiles are delivered in batches of 25 as the run goes, so a run cut off by its timeout keeps what it already delivered. If Weibo refuses or fails 5 IDs in a row, the run renews its anonymous session once, and if that does not help it stops rather than keep hitting Weibo: the rest are listed as not fetched (and not charged), to re-run later.
- A field Weibo does not send is
null, never an invented"",falseor0.locationis what the account shows on its profile (self-declared), not an IP location. - Follower and following lists are not offered. Weibo answers those endpoints only to a logged-in session (an anonymous request gets an empty
{}), so they cannot be delivered reliably without an account. The counts (followersCount,friendsCount) are in every row. - Great on a Schedule to track follower growth of a list of accounts: each run returns (and charges) every profile's current counts.
7. Black Cat (黑猫投诉) complaints (no cookies needed) 🆕
Black Cat (黑猫投诉, tousu.sina.com.cn) is Sina's consumer-complaint platform — Sina also runs Weibo. Chinese consumers file a complaint against a named company, say what they want (a refund, an apology, compensation), and the company answers on the record; the site's own counter showed 38.5 million valid complaints on 24 Sep 2026. This mode returns one row per complaint with no login: the company complained about, the complaint's title and text, what the consumer asks for (投诉要求), the problem (投诉问题), the amount in dispute where Black Cat shows it, the complaint's progress, when it was filed, and a link to the complaint page. The complainant's name and avatar are never included.
Two sources:
complaintCompanyIdsset → each company's complaint list (its 最新投诉 tab), companies read in the order given, withmaxResultsshared across them. Black Cat ranks a list by recent activity, not by filing date: a complaint filed weeks ago that just got a reply sits next to today's (seen live on 24 Sep 2026).complaintCompanyIdsempty → Black Cat's public latest-complaints feed across all companies: what China is complaining about right now. The status message says which source the run used.
Finding a company ID (couid): open any complaint about the company on tousu.sina.com.cn and click the company name next to 投诉对象. Its page URL looks like https://tousu.sina.com.cn/company/view/?couid=2092643773 — paste that URL, or just the number after couid=.
{"mode": "complaints","complaintCompanyIds": ["https://tousu.sina.com.cn/company/view/?couid=2092643773","5650743478"],"maxResults": 200}
Output — the shape of one row. Black Cat's complaint pages forbid reposting their content without permission, so the text values below are placeholders, not a real complaint; the field names, types and label values are exactly what a run returns. A company-list row carries amountInvolvedCny: null; a public-feed row carries the amount Black Cat shows:
{"mode": "complaints","source": "feed","complaintId": "17400000000","title": "<complaint title as served>","summary": "<complaint text as the list serves it>","isTruncated": false,"appeal": "退款","issue": "<the problem, as the complainant tagged it>","amountInvolvedCny": 99,"companyName": "<company complained about>","companyId": "2092643773","statusCode": 6,"status": "已回复","statusEn": "replied","createdAt": "2026-09-24T15:55:40+08:00","createdAtTimestamp": 1790236540,"upvoteCount": 0,"commentsCount": 0,"shareCount": 0,"url": "https://tousu.sina.com.cn/complaint/view/17400000000/?sld=<token as served>","scrapedAt": "2026-09-24T08:13:41.000000+00:00"}
- $0.035 per complaint (the same
item-scrapedresult as every other row), charged only for complaints delivered. - Depth: the first 500 complaints of a list — the most recently active. Black Cat shows anonymous visitors 50 pages of 10: page 50 answers, page 51 answers 登录查看更多内容 ("log in to see more") — measured 24 Sep 2026 on a company list and on the public feed, and Black Cat serves 10 per page whatever size is asked for. When a company has more, the run stops there and the status says so; the rest are not reachable without an account. If Black Cat ever serves an empty page before its own count is reached, the run stops there too and says so (
endedEarly). summaryis what the list serves: the whole complaint, or a ~150-character preview ending in...— thenisTruncatedistrue. The complaint page (url) has the full text.amountInvolvedCny(涉诉金额) is served by the public feed only; on company-list rows it isnull, not a guessed 0. A0means Black Cat itself shows 0元.statusis Black Cat's own label — 已回复 (replied), 已完成 (completed), 处理中 (in progress), 待分配 (awaiting assignment), 通过审核 (passed review), 已关闭 (closed) — with the English instatusEnand the raw code instatusCode. A code Black Cat has not been seen to label keeps itsstatusCodewithnulllabels.companyNameis the company complained about.companyIdis thecouidexactly as Black Cat serves it: on the public feed, the company complained about; on a company list Black Cat echoes the ID you asked for on every row. A parent company's list also carries its sub-brands' complaints (京东客服's list holds 京东外卖客服专线 ones), so such a row names the sub-brand incompanyNameand carries the parent's ID. The status labels such a list with the names its rows carry, e.g.5650743478 (京东客服 / 京东外卖客服专线).createdAtis Beijing time (+08:00) andcreatedAtTimestampthe raw Unix seconds. On a company list it is the record'screated_at; on the public feed it is the time the complaint page shows as published (发布于). Of 4 complaints seen in both lists on 24 Sep 2026, 3 carried the same time; the fourth was created at 02:34 and published at 15:55 the same day.urlis the complaint page exactly as Black Cat serves it: it carries ansldtoken, and the same page without it answers 页面不存在 (page does not exist). It is never built from the complaint number.- An ID Black Cat rejects (code 10002 参数错误) or answers with no complaints — an unknown ID looks exactly like a company with none — gives no row and no charge, and the status names it. Duplicate IDs are read once. Black Cat answers some companies under more than one ID (1003626 and 2092643773 both return 美团's list): a second ID whose page holds only complaints the first already served, with a matching complaint count (within 0.2%), is not read twice, and the status names it (
sameListAs). A sub-brand listed next to its parent is still read in full, although many of its complaints are also in the parent's list. A complaint that turns up twice in a run is delivered, and charged, once. - Every list not delivered in full is accounted for in the run's default key-value store record
COMPLAINTS_REPORT(invalid,empty,failed,depthCapped,endedEarly,sameListAs,notFetchedwithnotFetchedReason,heldBackComplaintIds,perListwith the company names each list carried, andretryCompanyIds— the IDs worth re-running). A run with nothing left out writes no record. - Polite to the source: at least 0.5 s between requests, and the run stops after 5 requests in a row fail or are refused, listing the companies it did not reach. Rows are delivered, then charged, in batches of 25, so a run cut off by its timeout keeps what it already delivered.
- Scheduling (e.g. a daily Apify Schedule): with
deltaModeon, each run returns only complaints that no earlier run with the samedeltaStateKeydelivered, tracked by complaint number for each company list (or the feed). Each list is read to its end or to the 500-complaint depth, because complaints already delivered that just got a reply sit above ones never delivered; those already delivered are skipped and not charged. WithdeltaModeoff, each run returns the top of each list again, up tomaxResults. sentimentAnalysisscores each complaint's title and summary.includeComments,geoRollup,adFilterandcookieStringare Weibo options: set in this mode, the status names them as not applied.- No keyword search. Black Cat's search endpoint now answers with a web page rather than data, so this mode reads company lists and the public feed only.
8. Post details (campaign / KOL post tracking) (no cookies needed) 🆕
One row per post you list: the post's current text and engagement — repostsCount, commentsCount, attitudesCount (likes) — plus the author (authorName, authorId, authorVerified, authorAvatar), createdAt / createdAtIso, source, images, videoUrl, isRepost and the quoted post in repostOf, and postUrl. No login needed, and no comments are fetched. The row has the same keys as the post row post_comments returns (mode: "post_detail"), so the two join on postId.
{"mode": "post_details","postIds": ["https://weibo.com/7087906569/RjB5HwkDj","https://m.weibo.cn/detail/5346708137705813","5346634918789689"],"maxResults": 100}
Accepted forms in postIds: the numeric post ID (5346708137705813), the short code Weibo puts in post links (RjB5HwkDj), https://weibo.com/<uid>/<id or code>, https://m.weibo.cn/detail/<id> and https://m.weibo.cn/status/<id or code>. The short code is the numeric ID written in base 62, so the Actor resolves it without a request: the first two entries above are the same post, fetched and charged once. Several posts pasted into one entry (url1,url2, a spreadsheet row) count as one entry each. A bare code that reads as a plain word (a header cell such as Marketing) is not taken as a post: it is named as skipped, and a real code of that shape is accepted in its URL form. The status names every post it left out as you pasted it.
Output — the shape of one row (illustrative text; the nulls are what Weibo's post endpoint leaves out, measured 25 Sep 2026):
{"postId": "5346708137705813","mid": "5346708137705813","text": "【#示例话题#】帖子的完整正文……","isTruncated": false,"createdAt": "Thu Sep 24 16:20:19 +0800 2026","createdAtIso": "2026-09-24T16:20:19+08:00","source": "微博视频号","repostsCount": 466,"commentsCount": 330,"attitudesCount": 493,"authorName": "示例账号","authorId": "7087906569","authorDescription": null,"authorFollowers": null,"authorFollowing": null,"authorVerified": true,"authorVerifiedReason": null,"authorAvatar": "https://tvax3.sinaimg.cn/crop.0.0.500.500.1024/...","images": [],"videoUrl": "http://f.video.weibocdn.com/...","region": null,"regionEn": null,"isPromoted": false,"isRepost": false,"repostOf": null,"postUrl": "https://weibo.com/7087906569/5346708137705813","scrapedAt": "2026-09-25T00:49:26.349921+00:00","mode": "post_detail"}
- $0.035 per post (the same
item-scrapedresult as every other row), charged only for posts delivered.maxResultscaps the posts per run (the form pre-fills 25 — raise it for a longer list, up to 5,000 per run; split longer lists across runs; the status says how many were left out). - Run it on a Schedule to chart each post's engagement over time. Every run returns every listed post again, with its counts as of that run and a
scrapedAttimestamp, so the rows of successive runs form a time series perpostId. Each run is its own dataset and is charged per post returned. - Full text: Weibo's post endpoint cuts long posts too (measured: 150 of 880 characters). With
fetchFullTexton (the default, no extra charge) the full text is fetched;isTruncated: trueremains only when Weibo would not serve it, and the status names those posts. - A value Weibo does not send is
null, never an invented0,""orfalse. Weibo's post endpoint does not carry the author's follower count or bio (authorFollowersandauthorDescriptionarenull;user_profilehas them), and many posts carry no location. repostsCounttops out at 1,000,000. Above that Weibo reports exactly1000000(it shows 100万+): measured 25 Sep 2026, 2 of 10 hot posts carried exactly 1,000,000 reposts next to exact comment and like counts. Read 1,000,000 as "one million or more".- A post Weibo does not show — deleted, hidden, or never existed (Weibo answers 该微博不存在) — gives no row and is not charged. The status names up to 10 per reason; the full list is in the run's default key-value store record
POST_DETAILS_REPORT(notFound,failed,notFetchedwithnotFetchedReason,heldBack,previewOnly,unparsed,weiboMessages— what Weibo said —inputs, andretryIds, the IDs worth re-running). A run that delivered every post writes no record. - Polite to Weibo: at least 2 s between requests. Rows are delivered, then charged, in batches of 25, so a run cut off by its timeout keeps what it delivered. If Weibo refuses or fails 5 posts in a row, the run renews its anonymous session once, then stops and lists the rest as not fetched (not charged).
sentimentAnalysisscores each post's text.includeComments,geoRollup,adFilter,deltaModeand the search date window do not apply here; set them and the status says so. LeavepostIdsempty and the run takes 3 of Weibo's current hot posts, so you can see the output before pasting your own; those 3 are charged like any other post ($0.105).
⏰ Set up daily monitoring in 2 minutes
Most of this Actor's value is in recurring runs. A single pull is a one-off snapshot — but a daily or hourly schedule turns it into a continuously-updated Chinese brand / public-opinion feed. That's where pay-per-result compounds: instead of paying once for a static dump, you build a living dataset that tracks how the conversation moves week over week.
- Run the Actor once with your input — a brand keyword in
searchmode, orhot_searchto capture the trending board — and check the output looks right. - Apify Console → Schedules → Create → pick this Actor and your saved input. (Even faster: open any finished run and click Schedule to reuse its exact input.)
- Set a cron expression and save — e.g.
0 8 * * *= daily at 8am, or0 * * * *= hourly. While you're there, enable the email notification on failed runs so you hear about a hiccup without checking manually.
Each scheduled run delivers its own fresh dataset, so you build a continuously-updated history with zero manual work — perfect for sentiment trend lines, brand-mention velocity, and time-series alt-data.
🔁 How a scheduled run can behave:
deltaModeoff (the default) — each run returns the current result set, up tomaxResults: a full snapshot every time, so a post that was in an earlier run can come back in the next one.deltaModeon (insearch,user_postsandhot_timeline) — each run returns only rows that no earlier run with the samedeltaStateKeydelivered: posts are deduped by post ID across runs, comment rows (includeComments) by their own comment ID, and rows an earlier run delivered are skipped (not delivered, not charged). Insearchthe deeper history walk is off, so a delta run reads only Weibo's first result window. On a narrow keyword or a single account a run can return a handful of rows, or none. Use a distinctdeltaStateKeyper query/user you track (e.g.nike-weekly).- What
deltaModeremembers — up to 200,000 IDs perdeltaStateKey(incomplaints, per company list). A row Weibo returns again counts as recent; when the memory is full, the IDs Weibo has gone longest without returning are forgotten first, and a forgotten row that comes back is delivered and charged again.hot_search_delta— purpose-built for scheduled trend-velocity tracking; each run returns (and bills) the whole trending board, every topic tagged new / rising / falling / steady / dropped versus the last run.
Example — a daily search snapshot (each run returns the query's newest posts, up to maxResults):
{"mode": "search","searchQuery": "Nike","maxResults": 200}
Optional variant — the same search with deltaMode (each run returns only posts no earlier run with this deltaStateKey delivered; one key per query you track):
{"mode": "search","searchQuery": "Nike","deltaMode": true,"deltaStateKey": "nike-daily","maxResults": 200}
🆕 Follow-up comments on earlier posts (followUpComments, off by default — charged per row)
A delta search reads each post minutes after it is published, usually before anyone has replied, and the post then drops out of Weibo's newest results. So a delta monitor delivers the posts but almost never their replies (measured on saved search pages of 26 Sep 2026, 李宁 and 耐克, 291 posts: posts 0-1 h old had 0-0.2 comments each, posts 6-24 h old about 1.1). followUpComments: true (only in search with deltaMode on) comes back for them:
- The posts each delta run delivers are remembered with their posting time, in the same
deltaStateKeyrecord. A later delta run re-reads each one's comment thread once: on the first run at least 24 hours after it was posted and before it is 7 days old, newest posts first. - Only comments never delivered under this
deltaStateKeycome back, as ordinary comment rows (mode: "post_comment", linked bypostId) markedisFollowUp: true— their post row came in an earlier run. Up tomaxCommentsper post, reading at most 6 pages of each thread (about 120 new top-level comments). Once the key's delta memory is full (about 200,000 ids; it then forgets its oldest), a re-read returns only comments newer than the newest one already delivered for that post, so a forgotten comment is never delivered — or charged — twice. - Cost: each comment row is one result, charged like a post ($0.035). A thread with no new comments adds no row and no charge.
- Limits: at most 300 threads, 400 comment requests and about 4 minutes of re-reading per run, 2 s apart, stopping 5 minutes before the run's timeout and at your run's charge limit; posts not reached keep their place for the next run. A post whose request failed is tried again on a later run while it is under 7 days old. A post that reaches 7 days without a re-read, or that the 10,000-post queue pushes out, leaves it unread; the status message counts them.
- Run time: the run's own posts are delivered after the re-reads, so a run with this on can take up to ~4 minutes longer. Keep the schedule interval longer than a run lasts: two runs on the same
deltaStateKeyat the same time both see the same rows as new. - It starts with the posts delivered from the first run with it on (posts delivered before carry no posting time in the record). It does not run with
adFilter: "promo_only", which would drop every comment row. Switch it off and the delta memory works exactly as before.
{"mode": "search","searchQuery": "Nike","deltaMode": true,"deltaStateKey": "nike-daily","maxResults": 200,"followUpComments": true}
🧠 Need this at AI-training-corpus scale?
If you're pulling Weibo's short-form posts, trending-topic chatter, and comment threads to train or fine-tune language models, the Chinese AI Training Corpus Engine assembles all 5 Chinese platforms — Weibo, RedNote, Bilibili, Douban, and Xueqiu — into AI-ready documents in one run: deduplicated, quality-scored, PII-scrubbed, and provenance-stamped for EU AI Act documentation, from $0.025/doc, with rejects and duplicates never charged.
📦 Want the full China feed? — China Monitoring Packages
Weibo is one platform. If you're monitoring a brand across Chinese social media, three pre-configured bundles combine this Actor with the Chinese Brand Monitor (Weibo + RedNote + Bilibili + Douban + Xueqiu in one scheduled call): Brand Starter ($110/mo, daily single-brand watch), Competitive Intel ($440/mo, brand + competitors share-of-voice), Fund Signal Desk (~$635/mo, multi-ticker daily sentiment). Self-serve — clone the preset, attach a Schedule, done. Support by text/issue tracker.
How to Get Cookies (for User Posts)
User posts needs a login cookie:
- Open weibo.com in your browser and log in
- Open DevTools (F12) → Network and reload the page
- Click any weibo.com request
- Copy the full
Cookierequest header (F12 → Network → any weibo.com request → Request Headers) and paste it intocookieString.
Leave cookieString empty for every other mode: a stale cookie replaces the anonymous session those modes use.
The cookie typically lasts several days before expiring.
Output Examples
Trending Topic
{"rank": 1,"title": "人工智能最新突破","hotValue": 2847562,"labelName": "热","isHot": true,"isPromoted": false,"url": "https://s.weibo.com/weibo?q=...","scrapedAt": "2026-04-10T12:00:00Z"}
rank is Weibo's own board position. The board also carries a paid slot: that row ships with isPromoted: true and rank: null, and adFilter: "organic_only" removes it (unbilled).
Hot Search Delta record
Each record shows how a topic moved since the previous run (status ∈ new / rising / falling / steady / dropped):
{"rank": 3,"title": "某品牌新品发布","hotValue": 1820000,"status": "rising","rankDelta": 5,"hotValueDelta": 640000,"previousRank": 8,"firstSeenAt": "2026-06-04T08:00:00+00:00","minutesOnBoard": 120,"peakRank": 3,"peakHotValue": 1820000,"isHot": true,"isNew": false,"url": "https://s.weibo.com/weibo?q=...","snapshotAt": "2026-06-04T10:00:00+00:00"}
Post
{"postId": "5285773987283226","text": "介绍一下我的老婆!@金莎","isTruncated": false,"createdAt": "Wed Apr 09 12:49:23 +0800 2026","createdAtIso": "2026-04-09T12:49:23+08:00","repostsCount": 493,"commentsCount": 4549,"attitudesCount": 97438,"authorName": "孙丞潇","authorId": "7511222755","authorFollowers": 0,"authorVerified": false,"images": ["https://wx1.sinaimg.cn/large/..."],"videoUrl": "","isRepost": false,"repostOf": null,"postUrl": "https://weibo.com/7511222755/5285773987283226","scrapedAt": "2026-04-10T12:00:00Z"}
Long posts come back in full. Weibo's list endpoints (search, user timelines, hot timeline) send only the first ~140-170 characters of a long post — about half of all search results. With
fetchFullText(on by default, no extra charge) each cut-off post is fetched in full: measured 2026-09-18, a 100-row search had 52 truncated rows and 51 came back complete. The post quoted inside a repost (repostOf) and the post rows ofpost_commentsandpost_detailsare expanded the same way. A row (or itsrepostOf) still carriesisTruncated: trueonly when Weibo would not serve the long text.createdAtkeeps Weibo's raw date string;createdAtIsois the same moment in ISO 8601.videoUrlis a signed link that expires about 1 hour after the scrape, so download the video right away if you need it.
repostOf— el post original, sin peticiones extra. Weibo incrusta el post citado entero dentro de la misma respuesta y este Actor lo tiraba: sólo escribíaisRepost: true. Medido el 26-ago sobre una búsqueda real de 80 filas, 23 eran reposts (29%). Sutextno viene vacío —trae la cadena//@usuario:con el contenido citado dentro— así que esto no rescata filas inútiles: entrega estructurado lo que hoy llegaba como una cadena. Para vigilancia de marca, quién escribió el original y cuánto circuló es justo el dato buscado.nullcuando la fila no es un repost.
Comment
{"commentId": "5285813927600208","text": "恭喜恭喜!神仙眷侣,一定要狠狠幸福哦~","createdAt": "Thu Apr 09 12:51:31 +0800 2026","createdAtIso": "2026-04-09T12:51:31+08:00","likeCount": 1268,"replyCount": 42,"authorName": "吃瓜罗伯特","authorId": "6108685154","authorFollowers": 10660,"authorVerified": true,"authorRegion": "广东","authorRegionEn": "Guangdong","postId": "5285773987283226","postUrl": "https://weibo.com/detail/5285773987283226","scrapedAt": "2026-04-10T12:00:00Z"}
Profile (user_profile)
{"userId": "1642634100","screenName": "新浪科技","description": "新浪科技是中国最有影响力的TMT产业资讯及数码产品服务平台。…","gender": "m","location": "北京 海淀区","followersCount": 23800866,"friendsCount": 2702,"statusesCount": 265367,"verified": true,"verifiedType": 3,"verifiedReason": "新浪科技官方微博","avatarHd": "https://tvax2.sinaimg.cn/crop.0.0.436.436.1024/...","coverImage": "https://ww2.sinaimg.cn/crop.0.0.640.640.640/...","domain": "sinatech","profileUrl": "https://weibo.com/u/1642634100","createdAt": "2009-08-28 21:07:40","birthday": "","company": "新浪网技术(中国)有限公司","sunshineCredit": "阳光信用极好","mode": "user_profile","scrapedAt": "2026-09-24T07:14:32Z"}
Sentiment field (optional)
With sentimentAnalysis: true, every post and comment gains a sentiment object:
{"polarity": "positive","score": 0.42,"method": "snownlp"}
polarity ∈ positive / neutral / negative · score ∈ −1.0…+1.0 (higher = more positive) · method is snownlp for Chinese text, or keyword for the English-only fallback.
Content is in Chinese
All content is returned in the original Simplified Chinese. Weibo is a Chinese-language platform — posts, comments, trending topics, and user bios are in Chinese.
If you need English translations, pipe the output through a translation API (Google Translate, DeepL, or Claude).
Technical Details
- No browser: pure HTTP — fast and lightweight, runs in 256 MB without sentiment (≥512 MB with
sentimentAnalysis) - No authentication except
user_posts(your cookie): the other modes read publicly accessible content only - Built-in rate limiting: automatic retry with exponential backoff to handle peak-hour throttling
- Globally accessible: no VPN or proxy required
- Clean structured JSON output: ready for analysis or downstream pipelines
Pricing
$35 per 1,000 results ($0.035 per result) — pay-per-event
Each scraped item (post, comment, trending topic, user profile, or Black Cat complaint) counts as one result. The two optional add-ons bill the same way: each author profile row (includeAuthorProfiles) and each follow-up comment row (followUpComments) is one result.
You only pay for delivered records — empty or failed results are never charged.
Typical costs (small-scale):
- Top 50 trending topics snapshot: ~$1.75
- 100 posts on a brand keyword: ~$3.50
- 200 comments on a viral post: ~$7.00
- 50 posts from one account: ~$1.75 (needs your cookieString)
- 100 user profiles: ~$3.50 (no login)
- 100 tracked posts, current engagement: ~$3.50 per run (no login)
- 100 Black Cat complaints about a company: ~$3.50 (no login)
B2B / bulk-scale examples:
- AI training corpus seed (10,000 posts on a topic): ~$350
- Daily brand sentiment monitor (500 posts/day for a month): ~$525/month
- Equity research signal (10 tickers × 200 posts daily): ~$2,100/month
- Multi-source academic dataset (50,000 posts across 30 keywords): ~$1,750
No separate compute charge: you pay per result, plus Apify's tiny per-run start event.
Limitations
- User posts needs your own login cookie for the timeline: without
cookieStringWeibo returns no posts, and the run delivers and charges nothing for that account. - Search, hot search, post comments, post details and user profiles work fully without authentication
- Follower / following lists are not available: Weibo serves them only to logged-in sessions.
user_profilereturns the counts, not the lists - Only public data is accessible — private/locked accounts are not available
- Weibo may rate-limit requests during peak hours — handled automatically with backoff
- Very old posts may not be available
- Black Cat complaints: the first 500 per list, the most recently active. Anonymous visitors see 50 pages of 10 per company (and of the public feed); the run says when a company has more. No keyword search: Black Cat's search endpoint answers with a web page, not data. Company lists do not carry the amount in dispute (
amountInvolvedCnyis null there; the public feed has it) - Date window is
searchmode only.searchDateFrom/searchDateTodo nothing inuser_posts,user_profile,complaints,hot_timeline,hot_search,hot_search_delta,post_commentsorpost_details; set them there and the run says so in its status message rather than pretending to have filtered - Date window is a sample, not an archive. Weibo's search exposes only an end time, and caps how far back each result window reaches.
searchDateFrom/searchDateTobound what is delivered and billed, but a deep period can stop early onmaxResults, your charge limit, the run timeout, or Weibo itself — the run's status message always names which, and reports the oldest and newest post actually delivered
FAQ
Is there a Weibo API?
There is no official public Weibo API available for international developers. Weibo's developer platform requires a Chinese business license and imposes strict rate limits. This Weibo Scraper is an alternative — extract trending topics, posts, and comments without any official API access.
How much does it cost to scrape Weibo?
The Weibo Scraper is priced per result: $35 per 1,000 ($0.035 each) (pay-per-event). Each scraped item (post, comment, trending topic, user profile, or Black Cat complaint) counts as one result. You can start with Apify's free plan, which includes $5 of monthly credits.
Can I scrape Weibo in Python?
Yes. Install the Apify Python client (pip install apify-client), then use the ApifyClient to call the zhorex/weibo-scraper actor. See the Python code example above.
Is scraping Weibo legal?
This scraper accesses publicly available data through Weibo's public web endpoints; user_posts reads your own signed-in session through the cookie you provide. It does not bypass authentication or access private/locked accounts. Always review your local laws and Weibo's terms of service before scraping.
What does this Weibo scraper cover?
The Weibo Scraper by Zhorex reads Weibo's public web endpoints (user_posts needs your cookie). It supports 9 modes (hot timeline, hot search, hot-search delta, post comments, post details, search, user posts, user profiles, and Black Cat (黑猫投诉) consumer complaints), handles rate limits automatically, and runs without a browser or VPN.
Integrations & data export
The Weibo Scraper integrates with your existing workflow tools:
- Google Sheets — Send scraped Weibo data directly to a spreadsheet
- Zapier / Make / n8n — Automate workflows triggered by new Weibo data
- REST API — Call the actor programmatically and retrieve results via Apify's REST API
- Webhooks — Get notified when a scraping run finishes and process data in real time
- Data formats — Download results in JSON, CSV, Excel, XML, or RSS
More scrapers by Zhorex
Chinese Digital Intelligence Suite
- 🆕 Chinese AI Training Corpus Engine — Weibo + RedNote + Bilibili + Douban + Xueqiu into AI-training-ready documents (MinHash dedup, quality scoring, PII scrub, EU AI Act provenance)
- 🆕 Chinese Brand Monitor — Cross-platform brand mention aggregator (Weibo + RedNote + Bilibili + Douban + Xueqiu, sentiment + dedup)
- Bilibili Scraper — China's video platform: danmaku, comments, Gen-Z creator analytics
- RedNote (Xiaohongshu) Scraper — China's Instagram + Pinterest (lifestyle, consumer reviews)
- RedNote Shop Scraper — RedShop e-commerce (products, vendors, prices)
- Douban Scraper — Long-form reviews, ratings, group discussions (movies/books/music)
- Xueqiu Scraper — Chinese stock-discussion sentiment, cashtag indexing (SH/SZ/HK/US-listed Chinese)
Streaming & Video
- Twitch Streamer & Channel Analytics — Twitch profiles, live streams, clips, and VODs
- Kick.com Streamer & Channel Analytics — Kick.com profiles, live streams, clips, and categories
- YouTube Shorts Scraper Pro — YouTube Shorts videos, creators, trends
- Letterboxd Scraper — Western film reviews and ratings
Markets & Alt-Data
- TradingView Multi-Market Scraper — Stocks, crypto, forex, indices
- Hyperliquid Pro Scraper — DeFi top traders, vaults, perpetual markets
- Booking.com Reviews Scraper — Hotel reviews and ratings
Price & Availability Monitors
Same scheduling pattern as this actor's monitors — schedule them and every run is tagged with what changed.
- 🆕 Hostelworld Rate & Availability Monitor — forward-dated accommodation rates, promo stack, commission split
- 🆕 Resy Availability & Scarcity Monitor — restaurant slots by date and party size, prime-window scarcity
- 🆕 DE/AT Pharmacy Price & Stock Monitor — OTC prices, RRP breaches, stock status
- 🆕 UK Tyre Price & Fitment Monitor — tyre prices by size and postcode
- 🆕 UK PPE Shelf Monitor — 14K+ SKUs, price moves, new listings and delistings
B2B Reviews
- G2 Reviews Scraper — B2B software reviews and ratings
- Capterra Reviews Scraper — Software product reviews and ratings
Other Tools
- Perplexity AI Scraper — AI-powered search results
- Tech Stack Detector — Detect technologies used by websites
- Telegram Channel Scraper — Public Telegram channel messages
- Phone Number Validator — Validate and format phone numbers
Support
Having issues? Open an issue on the Actor page.
Reviews ⭐
Reviews, good or bad, help other users decide: leave one on the reviews page.
Found a bug or missing field? Open an issue.