# Xiaohongshu (RedNote) Scraper — Notes, Creators, Comments (`edgy_dock/xiaohongshu-rednote-scraper`) Actor

Search Xiaohongshu / RedNote by keyword and export notes, creator profiles and comments as structured JSON, CSV or Excel. No signing keys, no API access needed.

- **URL**: https://apify.com/edgy\_dock/xiaohongshu-rednote-scraper.md
- **Developed by:** [wangyan](https://apify.com/edgy_dock) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.005 / actor start

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Xiaohongshu (RedNote) Scraper — Notes, Creators & Comments

Export Xiaohongshu / RedNote data as JSON, CSV or Excel. Search by keyword, pull full note
details, profile creators, and collect comments — without an API key, a Chinese business
entity, or an agency retainer.

Xiaohongshu has **300M+ monthly active users** and serves **600M+ searches a day**. It is where
Chinese consumers decide what to buy. It is also almost completely closed to outsiders: there
is no public API, the web endpoints are cryptographically signed, and full platform access
requires a mainland Chinese company. This Actor gives you the data layer without any of that.

***

### Quick start (3 minutes)

1. **Grab a cookie** — open <https://www.xiaohongshu.com/> in your browser, log in, press
   `F12` → Console tab, type `document.cookie`, copy the whole string.
2. **Click "Try for free"** at the top of this page (or paste the cookie into the `cookie`
   field of a run).
3. **Hit "Start"** — results appear in the Dataset tab within ~30 seconds.

Minimal input that works:

```json
{
  "searchKeywords": ["护肤"],
  "cookie": "a1=...; web_session=...;",
  "maxItems": 20
}
```

***

### What you can do with it

| You are | You use it to |
|---|---|
| **A brand entering China** | Measure share of voice, find which product claims resonate, track competitors' campaigns |
| **An agency** | Build KOC/KOL shortlists from real engagement numbers instead of media kits |
| **A dropshipper / sourcing operator** | See what is trending in China 3–6 months before it reaches Western marketplaces |
| **A market researcher** | Pull thousands of first-person consumer reviews on any category |
| **A trend / AI team** | Feed a genuinely hard-to-obtain Chinese-language dataset into your models |

***

### Three ways to use it

#### 1 · Search by keyword

The most common use — scrape notes matching a query.

```json
{
  "searchKeywords": ["秋冬护肤", "敏感肌"],
  "cookie": "...",
  "maxItems": 200,
  "sort": "most_liked",
  "noteType": "all",
  "fetchNoteDetails": true,
  "maxCommentsPerNote": 20
}
```

#### 2 · Deep-scrape specific notes

Already have a shortlist? Paste note URLs — the Actor pulls full body text, tags, all images,
IP location, publish time.

```json
{
  "noteUrls": [
    "https://www.xiaohongshu.com/explore/65f2a1b3000000001203c4d5?xsec_token=ABC123"
  ],
  "cookie": "...",
  "maxCommentsPerNote": 50
}
```

**Keep the `xsec_token` query parameter** when copying note URLs — Xiaohongshu refuses the
request without it.

#### 3 · Build a KOL shortlist

Pull creator profiles with their recent notes.

```json
{
  "creatorUrls": [
    "https://www.xiaohongshu.com/user/profile/5ff0e6f0000000000101d4e2"
  ],
  "cookie": "...",
  "maxNotesPerCreator": 50
}
```

***

### Input reference

| Field | Default | Notes |
|---|---|---|
| `searchKeywords` | `[]` | Chinese works ×10 better than English — use `护肤`, not `skincare` |
| `noteUrls` | `[]` | Must include `xsec_token=` query param |
| `creatorUrls` | `[]` | Format: `https://www.xiaohongshu.com/user/profile/<userId>` |
| `cookie` | `""` | **Required in practice.** See [Cookie walkthrough](#cookie-walkthrough) |
| `maxItems` | `100` | Hard cap. You only pay for records actually delivered |
| `sort` | `general` | `general` / `latest` / `most_liked` / `most_commented` |
| `noteType` | `all` | `all` / `video` / `image` |
| `fetchNoteDetails` | `false` | Open each note for full body + tags + IP + publish time |
| `maxCommentsPerNote` | `0` | Set > 0 to also pull comments (charged separately) |
| `maxNotesPerCreator` | `20` | Recent notes to collect per creator profile |
| `proxyConfiguration` | `{useApifyProxy: true, apifyProxyGroups: ["RESIDENTIAL"]}` | **CN residential strongly recommended** |

***

### Cookie walkthrough

Xiaohongshu heavily restricts anonymous requests. A real browser cookie unlocks full search.

**30-second extraction:**

```
1. Open https://www.xiaohongshu.com/ in Chrome / Edge / Firefox
2. Log in (QR code via Xiaohongshu mobile app)
3. Press F12 → Developer Tools opens
4. Click the "Console" tab
5. Paste this and press Enter:
      copy(document.cookie)
6. Your clipboard now holds the cookie string
7. Paste into the "cookie" field of the Actor input
```

The cookie string looks like:

```
abRequestId=xxx; a1=xxx; webId=xxx; web_session=xxx; xsecappid=xhs-pc-web; ...
```

**Critical cookies that must be present:** `a1` and `web_session`. If either is missing, log
in again and re-copy.

**Cookies expire.** Expect to refresh every 3–7 days under heavy use. The Actor logs a
warning when the cookie looks stale.

***

### CN proxy — strongly recommended

Xiaohongshu rate-limits aggressively by IP. Without a mainland-China or at least residential
proxy, the Actor will often return nothing or get captcha-walled.

**Easiest** — use Apify's built-in proxy:

```json
"proxyConfiguration": {
  "useApifyProxy": true,
  "apifyProxyGroups": ["RESIDENTIAL"],
  "apifyProxyCountry": "CN"
}
```

**Bring your own** — point to any CN residential pool:

```json
"proxyConfiguration": {
  "proxyUrls": ["http://user:pass@cn-proxy.example.com:8080"]
}
```

***

### Output

Each record is a flat JSON object. Chinese counts like `1.2w` / `3.4万` / `2亿` are normalised
to integers.

#### Note (from search)

```json
{
  "source": "search",
  "searchKeyword": "护肤",
  "noteId": "65f2a1b3000000001203c4d5",
  "type": "normal",
  "title": "秋冬护肤 10 件小事",
  "description": "...",
  "noteUrl": "https://www.xiaohongshu.com/explore/65f2a1b3000000001203c4d5",
  "xsecToken": "ABC123...",
  "coverUrl": "https://sns-webpic-qc.xhscdn.com/...",
  "likedCount": 12000,
  "collectedCount": 800,
  "commentCount": 36,
  "shareCount": 55,
  "authorId": "5ff0e6f0000000000101d4e2",
  "authorNickname": "护肤小助手",
  "authorAvatar": "...",
  "authorUrl": "https://www.xiaohongshu.com/user/profile/..."
}
```

#### Note detail (with `fetchNoteDetails: true`)

Adds `publishedAt`, `lastUpdatedAt`, `ipLocation`, `imageUrls[]`, `videoUrl`, `tags[]`,
`atUsers[]`, `description` (full body).

#### Creator profile

Adds `redId` (Xiaohongshu ID), `bio`, `gender`, `ipLocation`, `followsCount`, `fansCount`,
`likesAndCollectsCount`, `tags[]`.

#### Comment

```json
{
  "source": "comment",
  "noteId": "65f2a1b3000000001203c4d5",
  "commentId": "1234...",
  "content": "太好用了",
  "likeCount": 42,
  "subCommentCount": 3,
  "createdAt": 1727500000,
  "ipLocation": "北京",
  "authorId": "...",
  "authorNickname": "..."
}
```

***

### Pricing

Pay-per-event — you are only charged for records actually delivered.

| Event | Price (USD) | What triggers it |
|---|---|---|
| `apify-actor-start` | $0.005 | Once per run (per GB RAM, min 1) |
| `note-scraped` | $0.0015 | Each note delivered to the dataset |
| `creator-scraped` | $0.004 | Each creator profile |
| `comment-scraped` | $0.0004 | Each comment |

#### Cost estimator

| Job | Setting | Approx cost |
|---|---|---|
| Quick keyword scan | 100 notes, no details | **$0.16** |
| Deep category sweep | 1 000 notes + details + 20 comments each | **$9.50** |
| 10-KOL shortlist | 10 creators × 50 notes each | **$0.80** |
| Daily brand-mention monitor | 500 notes/day for 30 days | **$22.50 / month** |

Platform margin: Apify takes 20%, you keep 80%. The prices above are what you pay.

***

### Common errors

| Error / symptom | Cause | Fix |
|---|---|---|
| `No notes captured for 'X'` | Cookie expired / missing | Re-copy cookie as above |
| Only 1-2 notes returned when more exist | Rate-limited by IP | Enable CN residential proxy |
| `BrowserType.launch` timeout | Memory too low | Raise `minMemoryMbytes` to 4096 |
| Captcha page screenshots in logs | IP blocked | Switch proxy group, wait 10 min |
| Note URL "refused" | Missing `xsec_token=` in URL | Copy URL again from Xiaohongshu web, not mobile share |

***

### Why this exists (and why we don't sign anything)

Xiaohongshu signs every API call with `x-s` / `x-t` headers derived from obfuscated browser
JS that changes without notice. Re-implementing the signature is a losing maintenance race —
Xiaohongshu updates it, your Actor breaks, your users churn.

Instead the Actor opens a real Chromium page in the Apify runtime, lets the page's own JS
make the real signed calls, and reads JSON off the wire. Xiaohongshu can rotate signatures
all they want — the browser catches up for us.

***

### Known limits — be honest with yourself

- **Cookies expire.** Users are expected to refresh their own cookie. The Actor does not
  attempt to re-authenticate.
- **Search depth is finite.** Xiaohongshu caps keyword search around a few hundred results
  per query — no amount of scrolling pulls more.
- **English keywords mostly return nothing.** The platform is Chinese-first.
- **Comments are heavily paginated.** `maxCommentsPerNote` above a few hundred will slow
  substantially.
- **This is a scraper, not a CRM.** Data is read-only, best-effort. Fields may be `null` or
  `""` when the source didn't provide them.

***

### Changelog

- **0.1.8** · Loosen `page.goto` wait condition + raise timeout to 90s for cross-border runs
- **0.1.7** · Pin `pydantic<2.12` to avoid upstream crawlee incompatibility
- **0.1.4** · Add output schema for Store publishing
- **0.1.1** · Initial public release

# Actor input Schema

## `searchKeywords` (type: `array`):

Keywords to search on Xiaohongshu. Chinese keywords return far more results than English ones — try 护肤 instead of skincare.

## `noteUrls` (type: `array`):

Individual note URLs to scrape in full detail. Keep the xsec\_token query parameter if the URL has one — without it Xiaohongshu refuses the request.

## `creatorUrls` (type: `array`):

Creator profile URLs, e.g. https://www.xiaohongshu.com/user/profile/5ff0e6f0000000000101d4e2

## `cookie` (type: `string`):

Required for reliable results. Open xiaohongshu.com while logged in, press F12, go to Console, type document.cookie and paste the result here. See the README for a 30-second walkthrough.

## `maxItems` (type: `integer`):

Hard cap on records written to the dataset. You are only charged for records actually delivered.

## `sort` (type: `string`):

How search results are ordered.

## `noteType` (type: `string`):

Filter search results by media type.

## `fetchNoteDetails` (type: `boolean`):

Open every search hit to collect full body text, tags, all images, IP location and publish time. Slower and uses more compute, but returns much richer records.

## `maxCommentsPerNote` (type: `integer`):

Set above 0 to also collect comments. Comments are charged separately and count towards Max items.

## `maxNotesPerCreator` (type: `integer`):

How many recent notes to collect from each creator profile.

## `proxyConfiguration` (type: `object`):

Xiaohongshu rate-limits aggressively by IP. Residential proxies are strongly recommended for runs above a few hundred items.

## Actor input object example

```json
{
  "searchKeywords": [
    "护肤",
    "露营装备"
  ],
  "maxItems": 100,
  "sort": "general",
  "noteType": "all",
  "fetchNoteDetails": false,
  "maxCommentsPerNote": 0,
  "maxNotesPerCreator": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `datasetUrl` (type: `string`):

Browse the full scraped dataset in Apify Console.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchKeywords": [
        "护肤"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("edgy_dock/xiaohongshu-rednote-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchKeywords": ["护肤"] }

# Run the Actor and wait for it to finish
run = client.actor("edgy_dock/xiaohongshu-rednote-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchKeywords": [
    "护肤"
  ]
}' |
apify call edgy_dock/xiaohongshu-rednote-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,edgy_dock/xiaohongshu-rednote-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5RSHbRUasRzo2fEiL/builds/TrwKnLRoGjxxaqfpc/openapi.json
