Weibo Scraper
Under maintenancePricing
from $0.005 / actor start
Pricing
from $0.005 / actor start
Rating
0.0
(0)
Developer
wangyan
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Xiaohongshu (RedNote) Scraper — Notes, Creators & Comments
Export Xiaohongshu / RedNote data as JSON, CSV or Excel. Search by keyword, pull full note details, profile creators, and collect comments — without an API key, a Chinese business entity, or an agency retainer.
Xiaohongshu has 300M+ monthly active users and serves 600M+ searches a day. It is where Chinese consumers decide what to buy. It is also almost completely closed to outsiders: there is no public API, the web endpoints are cryptographically signed, and full platform access requires a mainland Chinese company. This Actor gives you the data layer without any of that.
Quick start (3 minutes)
- Grab a cookie — open https://www.xiaohongshu.com/ in your browser, log in, press
F12→ Console tab, typedocument.cookie, copy the whole string. - Click "Try for free" at the top of this page (or paste the cookie into the
cookiefield of a run). - Hit "Start" — results appear in the Dataset tab within ~30 seconds.
Minimal input that works:
{"searchKeywords": ["护肤"],"cookie": "a1=...; web_session=...;","maxItems": 20}
What you can do with it
| You are | You use it to |
|---|---|
| A brand entering China | Measure share of voice, find which product claims resonate, track competitors' campaigns |
| An agency | Build KOC/KOL shortlists from real engagement numbers instead of media kits |
| A dropshipper / sourcing operator | See what is trending in China 3–6 months before it reaches Western marketplaces |
| A market researcher | Pull thousands of first-person consumer reviews on any category |
| A trend / AI team | Feed a genuinely hard-to-obtain Chinese-language dataset into your models |
Three ways to use it
1 · Search by keyword
The most common use — scrape notes matching a query.
{"searchKeywords": ["秋冬护肤", "敏感肌"],"cookie": "...","maxItems": 200,"sort": "most_liked","noteType": "all","fetchNoteDetails": true,"maxCommentsPerNote": 20}
2 · Deep-scrape specific notes
Already have a shortlist? Paste note URLs — the Actor pulls full body text, tags, all images, IP location, publish time.
{"noteUrls": ["https://www.xiaohongshu.com/explore/65f2a1b3000000001203c4d5?xsec_token=ABC123"],"cookie": "...","maxCommentsPerNote": 50}
Keep the xsec_token query parameter when copying note URLs — Xiaohongshu refuses the
request without it.
3 · Build a KOL shortlist
Pull creator profiles with their recent notes.
{"creatorUrls": ["https://www.xiaohongshu.com/user/profile/5ff0e6f0000000000101d4e2"],"cookie": "...","maxNotesPerCreator": 50}
Input reference
| Field | Default | Notes |
|---|---|---|
searchKeywords | [] | Chinese works ×10 better than English — use 护肤, not skincare |
noteUrls | [] | Must include xsec_token= query param |
creatorUrls | [] | Format: https://www.xiaohongshu.com/user/profile/<userId> |
cookie | "" | Required in practice. See Cookie walkthrough |
maxItems | 100 | Hard cap. You only pay for records actually delivered |
sort | general | general / latest / most_liked / most_commented |
noteType | all | all / video / image |
fetchNoteDetails | false | Open each note for full body + tags + IP + publish time |
maxCommentsPerNote | 0 | Set > 0 to also pull comments (charged separately) |
maxNotesPerCreator | 20 | Recent notes to collect per creator profile |
proxyConfiguration | {useApifyProxy: true, apifyProxyGroups: ["RESIDENTIAL"]} | CN residential strongly recommended |
Cookie walkthrough
Xiaohongshu heavily restricts anonymous requests. A real browser cookie unlocks full search.
30-second extraction:
1. Open https://www.xiaohongshu.com/ in Chrome / Edge / Firefox2. Log in (QR code via Xiaohongshu mobile app)3. Press F12 → Developer Tools opens4. Click the "Console" tab5. Paste this and press Enter:copy(document.cookie)6. Your clipboard now holds the cookie string7. Paste into the "cookie" field of the Actor input
The cookie string looks like:
abRequestId=xxx; a1=xxx; webId=xxx; web_session=xxx; xsecappid=xhs-pc-web; ...
Critical cookies that must be present: a1 and web_session. If either is missing, log
in again and re-copy.
Cookies expire. Expect to refresh every 3–7 days under heavy use. The Actor logs a warning when the cookie looks stale.
CN proxy — strongly recommended
Xiaohongshu rate-limits aggressively by IP. Without a mainland-China or at least residential proxy, the Actor will often return nothing or get captcha-walled.
Easiest — use Apify's built-in proxy:
"proxyConfiguration": {"useApifyProxy": true,"apifyProxyGroups": ["RESIDENTIAL"],"apifyProxyCountry": "CN"}
Bring your own — point to any CN residential pool:
"proxyConfiguration": {"proxyUrls": ["http://user:pass@cn-proxy.example.com:8080"]}
Output
Each record is a flat JSON object. Chinese counts like 1.2w / 3.4万 / 2亿 are normalised
to integers.
Note (from search)
{"source": "search","searchKeyword": "护肤","noteId": "65f2a1b3000000001203c4d5","type": "normal","title": "秋冬护肤 10 件小事","description": "...","noteUrl": "https://www.xiaohongshu.com/explore/65f2a1b3000000001203c4d5","xsecToken": "ABC123...","coverUrl": "https://sns-webpic-qc.xhscdn.com/...","likedCount": 12000,"collectedCount": 800,"commentCount": 36,"shareCount": 55,"authorId": "5ff0e6f0000000000101d4e2","authorNickname": "护肤小助手","authorAvatar": "...","authorUrl": "https://www.xiaohongshu.com/user/profile/..."}
Note detail (with fetchNoteDetails: true)
Adds publishedAt, lastUpdatedAt, ipLocation, imageUrls[], videoUrl, tags[],
atUsers[], description (full body).
Creator profile
Adds redId (Xiaohongshu ID), bio, gender, ipLocation, followsCount, fansCount,
likesAndCollectsCount, tags[].
Comment
{"source": "comment","noteId": "65f2a1b3000000001203c4d5","commentId": "1234...","content": "太好用了","likeCount": 42,"subCommentCount": 3,"createdAt": 1727500000,"ipLocation": "北京","authorId": "...","authorNickname": "..."}
Pricing
Pay-per-event — you are only charged for records actually delivered.
| Event | Price (USD) | What triggers it |
|---|---|---|
apify-actor-start | $0.005 | Once per run (per GB RAM, min 1) |
note-scraped | $0.0015 | Each note delivered to the dataset |
creator-scraped | $0.004 | Each creator profile |
comment-scraped | $0.0004 | Each comment |
Cost estimator
| Job | Setting | Approx cost |
|---|---|---|
| Quick keyword scan | 100 notes, no details | $0.16 |
| Deep category sweep | 1 000 notes + details + 20 comments each | $9.50 |
| 10-KOL shortlist | 10 creators × 50 notes each | $0.80 |
| Daily brand-mention monitor | 500 notes/day for 30 days | $22.50 / month |
Platform margin: Apify takes 20%, you keep 80%. The prices above are what you pay.
Common errors
| Error / symptom | Cause | Fix |
|---|---|---|
No notes captured for 'X' | Cookie expired / missing | Re-copy cookie as above |
| Only 1-2 notes returned when more exist | Rate-limited by IP | Enable CN residential proxy |
BrowserType.launch timeout | Memory too low | Raise minMemoryMbytes to 4096 |
| Captcha page screenshots in logs | IP blocked | Switch proxy group, wait 10 min |
| Note URL "refused" | Missing xsec_token= in URL | Copy URL again from Xiaohongshu web, not mobile share |
Why this exists (and why we don't sign anything)
Xiaohongshu signs every API call with x-s / x-t headers derived from obfuscated browser
JS that changes without notice. Re-implementing the signature is a losing maintenance race —
Xiaohongshu updates it, your Actor breaks, your users churn.
Instead the Actor opens a real Chromium page in the Apify runtime, lets the page's own JS make the real signed calls, and reads JSON off the wire. Xiaohongshu can rotate signatures all they want — the browser catches up for us.
Known limits — be honest with yourself
- Cookies expire. Users are expected to refresh their own cookie. The Actor does not attempt to re-authenticate.
- Search depth is finite. Xiaohongshu caps keyword search around a few hundred results per query — no amount of scrolling pulls more.
- English keywords mostly return nothing. The platform is Chinese-first.
- Comments are heavily paginated.
maxCommentsPerNoteabove a few hundred will slow substantially. - This is a scraper, not a CRM. Data is read-only, best-effort. Fields may be
nullor""when the source didn't provide them.
Changelog
- 0.1.8 · Loosen
page.gotowait condition + raise timeout to 90s for cross-border runs - 0.1.7 · Pin
pydantic<2.12to avoid upstream crawlee incompatibility - 0.1.4 · Add output schema for Store publishing
- 0.1.1 · Initial public release