Weibo Scraper avatar

Weibo Scraper

Under maintenance

Pricing

from $0.005 / actor start

Go to Apify Store
Weibo Scraper

Weibo Scraper

Under maintenance

Scrape Weibo posts, users and comments.

Pricing

from $0.005 / actor start

Rating

0.0

(0)

Developer

wangyan

wangyan

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Xiaohongshu (RedNote) Scraper — Notes, Creators & Comments

Export Xiaohongshu / RedNote data as JSON, CSV or Excel. Search by keyword, pull full note details, profile creators, and collect comments — without an API key, a Chinese business entity, or an agency retainer.

Xiaohongshu has 300M+ monthly active users and serves 600M+ searches a day. It is where Chinese consumers decide what to buy. It is also almost completely closed to outsiders: there is no public API, the web endpoints are cryptographically signed, and full platform access requires a mainland Chinese company. This Actor gives you the data layer without any of that.


Quick start (3 minutes)

  1. Grab a cookie — open https://www.xiaohongshu.com/ in your browser, log in, press F12 → Console tab, type document.cookie, copy the whole string.
  2. Click "Try for free" at the top of this page (or paste the cookie into the cookie field of a run).
  3. Hit "Start" — results appear in the Dataset tab within ~30 seconds.

Minimal input that works:

{
"searchKeywords": ["护肤"],
"cookie": "a1=...; web_session=...;",
"maxItems": 20
}

What you can do with it

You areYou use it to
A brand entering ChinaMeasure share of voice, find which product claims resonate, track competitors' campaigns
An agencyBuild KOC/KOL shortlists from real engagement numbers instead of media kits
A dropshipper / sourcing operatorSee what is trending in China 3–6 months before it reaches Western marketplaces
A market researcherPull thousands of first-person consumer reviews on any category
A trend / AI teamFeed a genuinely hard-to-obtain Chinese-language dataset into your models

Three ways to use it

1 · Search by keyword

The most common use — scrape notes matching a query.

{
"searchKeywords": ["秋冬护肤", "敏感肌"],
"cookie": "...",
"maxItems": 200,
"sort": "most_liked",
"noteType": "all",
"fetchNoteDetails": true,
"maxCommentsPerNote": 20
}

2 · Deep-scrape specific notes

Already have a shortlist? Paste note URLs — the Actor pulls full body text, tags, all images, IP location, publish time.

{
"noteUrls": [
"https://www.xiaohongshu.com/explore/65f2a1b3000000001203c4d5?xsec_token=ABC123"
],
"cookie": "...",
"maxCommentsPerNote": 50
}

Keep the xsec_token query parameter when copying note URLs — Xiaohongshu refuses the request without it.

3 · Build a KOL shortlist

Pull creator profiles with their recent notes.

{
"creatorUrls": [
"https://www.xiaohongshu.com/user/profile/5ff0e6f0000000000101d4e2"
],
"cookie": "...",
"maxNotesPerCreator": 50
}

Input reference

FieldDefaultNotes
searchKeywords[]Chinese works ×10 better than English — use 护肤, not skincare
noteUrls[]Must include xsec_token= query param
creatorUrls[]Format: https://www.xiaohongshu.com/user/profile/<userId>
cookie""Required in practice. See Cookie walkthrough
maxItems100Hard cap. You only pay for records actually delivered
sortgeneralgeneral / latest / most_liked / most_commented
noteTypeallall / video / image
fetchNoteDetailsfalseOpen each note for full body + tags + IP + publish time
maxCommentsPerNote0Set > 0 to also pull comments (charged separately)
maxNotesPerCreator20Recent notes to collect per creator profile
proxyConfiguration{useApifyProxy: true, apifyProxyGroups: ["RESIDENTIAL"]}CN residential strongly recommended

Xiaohongshu heavily restricts anonymous requests. A real browser cookie unlocks full search.

30-second extraction:

1. Open https://www.xiaohongshu.com/ in Chrome / Edge / Firefox
2. Log in (QR code via Xiaohongshu mobile app)
3. Press F12 → Developer Tools opens
4. Click the "Console" tab
5. Paste this and press Enter:
copy(document.cookie)
6. Your clipboard now holds the cookie string
7. Paste into the "cookie" field of the Actor input

The cookie string looks like:

abRequestId=xxx; a1=xxx; webId=xxx; web_session=xxx; xsecappid=xhs-pc-web; ...

Critical cookies that must be present: a1 and web_session. If either is missing, log in again and re-copy.

Cookies expire. Expect to refresh every 3–7 days under heavy use. The Actor logs a warning when the cookie looks stale.


Xiaohongshu rate-limits aggressively by IP. Without a mainland-China or at least residential proxy, the Actor will often return nothing or get captcha-walled.

Easiest — use Apify's built-in proxy:

"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": ["RESIDENTIAL"],
"apifyProxyCountry": "CN"
}

Bring your own — point to any CN residential pool:

"proxyConfiguration": {
"proxyUrls": ["http://user:pass@cn-proxy.example.com:8080"]
}

Output

Each record is a flat JSON object. Chinese counts like 1.2w / 3.4万 / 2亿 are normalised to integers.

{
"source": "search",
"searchKeyword": "护肤",
"noteId": "65f2a1b3000000001203c4d5",
"type": "normal",
"title": "秋冬护肤 10 件小事",
"description": "...",
"noteUrl": "https://www.xiaohongshu.com/explore/65f2a1b3000000001203c4d5",
"xsecToken": "ABC123...",
"coverUrl": "https://sns-webpic-qc.xhscdn.com/...",
"likedCount": 12000,
"collectedCount": 800,
"commentCount": 36,
"shareCount": 55,
"authorId": "5ff0e6f0000000000101d4e2",
"authorNickname": "护肤小助手",
"authorAvatar": "...",
"authorUrl": "https://www.xiaohongshu.com/user/profile/..."
}

Note detail (with fetchNoteDetails: true)

Adds publishedAt, lastUpdatedAt, ipLocation, imageUrls[], videoUrl, tags[], atUsers[], description (full body).

Creator profile

Adds redId (Xiaohongshu ID), bio, gender, ipLocation, followsCount, fansCount, likesAndCollectsCount, tags[].

Comment

{
"source": "comment",
"noteId": "65f2a1b3000000001203c4d5",
"commentId": "1234...",
"content": "太好用了",
"likeCount": 42,
"subCommentCount": 3,
"createdAt": 1727500000,
"ipLocation": "北京",
"authorId": "...",
"authorNickname": "..."
}

Pricing

Pay-per-event — you are only charged for records actually delivered.

EventPrice (USD)What triggers it
apify-actor-start$0.005Once per run (per GB RAM, min 1)
note-scraped$0.0015Each note delivered to the dataset
creator-scraped$0.004Each creator profile
comment-scraped$0.0004Each comment

Cost estimator

JobSettingApprox cost
Quick keyword scan100 notes, no details$0.16
Deep category sweep1 000 notes + details + 20 comments each$9.50
10-KOL shortlist10 creators × 50 notes each$0.80
Daily brand-mention monitor500 notes/day for 30 days$22.50 / month

Platform margin: Apify takes 20%, you keep 80%. The prices above are what you pay.


Common errors

Error / symptomCauseFix
No notes captured for 'X'Cookie expired / missingRe-copy cookie as above
Only 1-2 notes returned when more existRate-limited by IPEnable CN residential proxy
BrowserType.launch timeoutMemory too lowRaise minMemoryMbytes to 4096
Captcha page screenshots in logsIP blockedSwitch proxy group, wait 10 min
Note URL "refused"Missing xsec_token= in URLCopy URL again from Xiaohongshu web, not mobile share

Why this exists (and why we don't sign anything)

Xiaohongshu signs every API call with x-s / x-t headers derived from obfuscated browser JS that changes without notice. Re-implementing the signature is a losing maintenance race — Xiaohongshu updates it, your Actor breaks, your users churn.

Instead the Actor opens a real Chromium page in the Apify runtime, lets the page's own JS make the real signed calls, and reads JSON off the wire. Xiaohongshu can rotate signatures all they want — the browser catches up for us.


Known limits — be honest with yourself

  • Cookies expire. Users are expected to refresh their own cookie. The Actor does not attempt to re-authenticate.
  • Search depth is finite. Xiaohongshu caps keyword search around a few hundred results per query — no amount of scrolling pulls more.
  • English keywords mostly return nothing. The platform is Chinese-first.
  • Comments are heavily paginated. maxCommentsPerNote above a few hundred will slow substantially.
  • This is a scraper, not a CRM. Data is read-only, best-effort. Fields may be null or "" when the source didn't provide them.

Changelog

  • 0.1.8 · Loosen page.goto wait condition + raise timeout to 90s for cross-border runs
  • 0.1.7 · Pin pydantic<2.12 to avoid upstream crawlee incompatibility
  • 0.1.4 · Add output schema for Store publishing
  • 0.1.1 · Initial public release