# Xiaohongshu (Rednote) Note scraper (Keyword search) (`lexac94/xiaohongshu-scraper`) Actor

Use your own browser cookies to search xiaohongshu.com for keywords with a logged-in browser session and extracts note title, description, author, engagement counts, images, tags and publish date.

- **URL**: https://apify.com/lexac94/xiaohongshu-scraper.md
- **Developed by:** [Lexa N](https://apify.com/lexac94) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 note discovereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Xiaohongshu Search Scraper do?

**Xiaohongshu Search Scraper** searches [xiaohongshu.com](https://www.xiaohongshu.com) (RED / RedNote / 小红书) for the **keywords you give it** and returns one clean record per note: title, full description, author, likes, collects, comments, shares, tags, publish and edit times, the poster's **IP region**, every image with dimensions and live-photo video, and every video stream variant. It runs a real logged-in browser session, so it uses the same sort, note-type, time-window and **location** filters a person sees on the site, including **同城 (same city)**, which lets you restrict search to one country or city when combined with a proxy there.

It is built for **monitoring, not one-off dumps**: each run remembers what it has already collected, queues what it found but hasn't opened yet, and caps how many notes it opens, so you can schedule it hourly and let it work through a long keyword list over days without re-scraping or overloading one account.

### Why use Xiaohongshu Search Scraper?

- **Local-market listening.** The `ip_location` field is Xiaohongshu's own stamp of where the poster was. Combined with the 同城 filter and a residential proxy, you can collect what people in Singapore, Malaysia, Hong Kong or any other region are actually saying, instead of wading through mainland results.
- **Brand and competitor tracking.** Run 20 to 100 keywords on a schedule and get only new notes each time.
- **Trend research.** Sort by newest, restrict to the past day or week, and see what is rising.
- **Complete media metadata.** Full-resolution image URLs, live-photo streams, and all video variants with codec, resolution and bitrate.
- **Cheap per note.** Pay-per-event pricing well below comparable scrapers; see the pricing section.

### How to use Xiaohongshu Search Scraper

1. Log in to xiaohongshu.com in Chrome with the account you want to use. A dedicated account is strongly recommended.
2. Export its cookies. The **Cookie-Editor** extension (Export → JSON) is the easiest; Netscape `cookies.txt` and a raw `Cookie:` header string also work.
3. Paste the export into the **Session cookies** field. It is stored encrypted, and the actor keeps it in its own private store after the first run, so later runs and schedules can leave the field empty.
4. Enter your **keywords**. Chinese and English both work.
5. Set **Max note opens per run** (default 40) and click **Start**.
6. For an ongoing monitor, create a **Schedule** (hourly is a good default) and set **Keywords per run** so a long list rotates across runs.

### Input

| Field | Description | Default |
|---|---|---|
| `keywords` | Search terms, one search each | required |
| `keywords_per_run` | Rotate through a long list N keywords per run (0 = all every run) | 0 |
| `sort` | relevant / latest / most liked / most commented / most collected | relevant |
| `sweep_sorts` | Run every sort order per keyword and merge, for more unique notes | false |
| `note_type` | all / video / image | all |
| `time_window` | all / past day / past week / past 6 months | all |
| `location` | anywhere / same city (同城) / nearby (附近), relative to the session's IP | anywhere |
| `region_keywords` | Words that flag a note as about your region (`region_marker`) | Singapore set |
| `max_notes_per_keyword` | Stop scrolling a search once this many new notes are found | 40 |
| `max_note_opens_per_run` | Hard cap on note pages opened per run | 40 |
| `max_run_minutes` | Stop cleanly after this many minutes; keep below your schedule interval | 45 |
| `scrape_details` | Open each note for full fields. Off = search-card data only | true |
| `delay_min_ms` / `delay_max_ms` | Random pause between page actions | 2000 / 5000 |
| `cookies` | Logged-in session cookies (secret); needed once | |
| `proxy_configuration` | Apify Proxy. Use RESIDENTIAL with a country code to make 同城 mean that country | off |
| `block_media` | Skip downloading images and video; URLs are still captured | true |
| `skip_seen` | Do not re-open notes scraped in earlier runs | true |
| `state_store_name` | Named store holding seen IDs, the pending queue and the session | xiaohongshu-scraper-state |

#### Filtering to one region

Set `proxy_configuration` to Apify Proxy, group **RESIDENTIAL**, country **SG** (or your target), and `location` to **same city**. Xiaohongshu's 同城 filter is relative to the browsing session's location, so search results come back local before a single note is opened. Every record also carries `ip_location` (the platform's own poster-region stamp) and `region_marker` (true when the text matches your **Region keywords**; defaults cover Singapore), so you can filter downstream as well.

### Output

One dataset item per note. Download as JSON, CSV, Excel or HTML, or read it through the API.

```json
{
  "note_id": "66f1a2b3c4d5e6f7a8b9c0d1",
  "url": "https://www.xiaohongshu.com/explore/66f1a2b3c4d5e6f7a8b9c0d1",
  "keyword": "咖啡店 推荐",
  "found_with": {"sort": "general", "note_type": "all", "time_window": "all", "location": "same_city"},
  "title": "…",
  "description": "…full text…",
  "note_type": "normal",
  "author_id": "5f9aab180000000001008d92",
  "author_name": "…",
  "author_avatar": "https://sns-avatar-…",
  "author_url": "https://www.xiaohongshu.com/user/profile/5f9aab18…",
  "likes": 175, "collects": 40, "comments": 31, "shares": 6,
  "tags": ["…"],
  "at_users": [],
  "ip_location": "新加坡",
  "region_marker": true,
  "published_at": "2026-09-01T04:49:26+00:00",
  "last_edited_at": "2026-09-01T05:10:02+00:00",
  "image_count": 4,
  "image_urls": ["https://sns-webpic-…"],
  "images": [{"url": "…", "width": 1440, "height": 1920, "scenes": {"WB_DFT": "…", "WB_PRV": "…"}, "live_photo": false, "live_photo_url": null}],
  "video_url": null,
  "video": null,
  "scraped_at": "2026-09-28T04:40:00+00:00",
  "source": "state"
}
```

For video notes, `video` holds duration, thumbnail and every stream variant with codec, resolution, bitrate, audio codec and master/backup URLs. Media URLs are signed by Xiaohongshu and expire after a while; download at scrape time if you need the files.

#### Data fields

| Field | Description |
|---|---|
| `note_id`, `url` | Note identifier and canonical URL |
| `keyword`, `found_with` | The search and filters that discovered the note |
| `title`, `description`, `note_type` | Content; `note_type` is `normal` (images) or `video` |
| `author_id`, `author_name`, `author_avatar`, `author_url` | Poster |
| `likes`, `collects`, `comments`, `shares` | Engagement counts at scrape time |
| `tags`, `at_users` | Hashtags and @-mentions |
| `ip_location` | Poster's region as stamped by Xiaohongshu |
| `region_marker` | True when text or IP region matches your region keywords |
| `published_at`, `last_edited_at` | ISO timestamps, UTC |
| `image_count`, `image_urls`, `images` | Gallery, with dimensions and live-photo video |
| `video_url`, `video` | Playable URL and all stream variants |
| `scraped_at`, `source` | When it was collected; `state` (page data) or `dom` (fallback) |

### How much does it cost to scrape Xiaohongshu?

Pay-per-event. You are charged only for what you get:

| Event | Price |
|---|---|
| Actor start | $0.01 per run (per GB of memory, minimum one) |
| Note detail | $0.003 per full note record |
| Note discovered | $0.0005 per search-card stub (search-only mode) |

A scheduled hourly run that opens 40 notes costs about **$0.13 per run**, or roughly **$3 a day for ~1,000 notes** including compute. Runs that stop early because the account was challenged or the spending limit was reached are charged only for the notes they delivered. Set **Maximum total charge** on the run to cap spend.

### Tips

- **One account, one hourly run.** Xiaohongshu flags sessions that appear from many places in quick succession. Keep runs at least an hour apart per account, keep delays at 2 seconds or more, and do not switch proxy countries between runs on the same account.
- **Turn on `sweep_sorts`** for a fresh keyword to find more unique notes; turn it off for a monitor that only needs what's new.
- **Use `time_window: 1w` for monitors.** Discovery becomes fast and cheap once the backlog is done.
- **If a run ends with "Blocked"** the account was challenged. Open Xiaohongshu in your own browser, complete the verification, re-export cookies, and paste them once more. The pending queue is preserved.
- **`debug_raw`** saves screenshots of the filter panel and any error text to the run's key-value store, which is the fastest way to see what happened.

### FAQ and disclaimer

**Why does it need my cookies?** Xiaohongshu shows search results and note details only to logged-in users, and its API is signed. A real browser session is the only stable way to read the page. Your cookies are stored as a secret and never appear in logs or datasets.

**Can it collect comments or follower counts?** Not in this version. Comments would roughly double the cost per note; follower counts need a profile visit per author. Both are on the list if there is demand.

**Is this legal?** Scraping publicly visible content for research is generally lawful, but Xiaohongshu's terms restrict automated access and notes may contain personal information. You are responsible for how you use the output and for complying with applicable law and the platform's terms.

**Known limitations.** The filter panel is driven by clicking the site's UI, so a redesign can break sort/time/location filters until the actor is updated; the run logs a warning and falls back to default ordering. Media URLs expire. Relevance-sorted results for brand terms drift toward mainland China after the first ~20 notes; use 同城, `time_window`, region-anchored keywords, or the `ip_location` field.

Report problems in the **Issues** tab. Custom fields, other regions or a managed schedule are available on request.

# Actor input Schema

## `site` (type: `string`):

Which front end your account uses. Accounts registered with a mainland China number use xiaohongshu.com (Chinese UI). Accounts registered overseas are routed to rednote.com (international UI). Export cookies from the same site you select here.

## `keywords` (type: `array`):

Search terms, one search each. Chinese and English both work, e.g. 咖啡店 推荐, 装修 灵感, skincare routine, lululemon. Add a place name to anchor results to a region.

## `keywords_per_run` (type: `integer`):

For long keyword lists: search only this many keywords per run, cycling through the list across runs (position remembered in the state store). 0 = search every keyword every run.

## `sort` (type: `string`):

Search result ordering. Ignored when 'Sweep all sort orders' is on.

## `sweep_sorts` (type: `boolean`):

Run each keyword under every sort order and merge the results. Finds more unique notes per keyword at the cost of a few extra search page loads (no extra note opens).

## `note_type` (type: `string`):

Restrict results to one note format.

## `time_window` (type: `string`):

Only notes published within this window.

## `location` (type: `string`):

XHS's 位置距离 filter. 'Same city' (同城) and 'Nearby' (附近) are relative to the browsing session's location, which on the web is its IP. Combine with a residential proxy in your target country to restrict search to local notes.

## `region_keywords` (type: `array`):

Words that mark a note as being about your target region. Each record gets region\_marker = true when its title, text, tags or IP region contain any of them. Defaults are for Singapore; replace for another market.

## `max_notes_per_keyword` (type: `integer`):

Stop scrolling search once this many NEW (unseen) notes are found for a keyword. A single sort order usually exhausts at 200–300.

## `max_note_opens_per_run` (type: `integer`):

Hard cap on note detail pages opened in one run. This is the account-exposure dial: ~40 per run on an hourly schedule is a reasonable ceiling for one account. Notes discovered but not opened stay pending for the next run.

## `max_run_minutes` (type: `integer`):

Stop searching and opening notes once the run has used this many minutes, then save state and exit cleanly. Keep it below the schedule interval so runs never overlap (two sessions on one account at once is the pattern to avoid). 0 = no budget.

## `scrape_details` (type: `boolean`):

If off, only search-card data (id, title, author, likes) is saved and no note pages are opened. Much lighter on the account.

## `delay_min_ms` (type: `integer`):

Random pause lower bound between page actions. Keep it human-like; the account is the bottleneck, not the machine.

## `delay_max_ms` (type: `integer`):

Random pause upper bound between page actions.

## `cookies` (type: `string`):

Logged-in cookies exported from the site selected above (xiaohongshu.com or rednote.com). Accepts Cookie-Editor JSON export, Netscape cookies.txt, or a raw Cookie header string. Paste once: the actor stores them in its private state store and later runs can leave this empty.

## `proxy_configuration` (type: `object`):

Off by default. Switch on Apify Proxy with the RESIDENTIAL group and your target country code so the session's location matches the account and the 'Same city' filter means that country. Also the fix if datacenter IPs start hitting captchas.

## `block_media` (type: `boolean`):

Do not download images, video or fonts; URLs are still captured. Cuts proxy traffic by most of its volume. Turn off only for debugging screenshots.

## `skip_seen` (type: `boolean`):

Uses the state store below to avoid re-opening notes already scraped. Turn off to re-scrape everything (e.g. to refresh engagement counts).

## `state_store_name` (type: `string`):

Named key-value store that holds seen note IDs and the pending queue. Use a different name per project to keep their histories separate.

## `debug_raw` (type: `boolean`):

Saves a screenshot after each filter click (filter-\*.png) and the raw page state for video notes missing a stream URL (raw-note-<id>.json) to the run key-value store. Off in production.

## `headless` (type: `boolean`):

Turn off only for local debugging so you can watch the browser.

## Actor input object example

```json
{
  "site": "xiaohongshu",
  "keywords": [
    "咖啡店 推荐",
    "skincare routine"
  ],
  "keywords_per_run": 0,
  "sort": "general",
  "sweep_sorts": false,
  "note_type": "all",
  "time_window": "all",
  "location": "all",
  "region_keywords": [
    "新加坡",
    "🇸🇬",
    "HDB",
    "BTO",
    "组屋",
    "Singapore",
    "SG",
    "狮城",
    "坡县",
    "坡岛",
    "condo"
  ],
  "max_notes_per_keyword": 40,
  "max_note_opens_per_run": 40,
  "max_run_minutes": 45,
  "scrape_details": true,
  "delay_min_ms": 2000,
  "delay_max_ms": 5000,
  "proxy_configuration": {
    "useApifyProxy": false
  },
  "block_media": true,
  "skip_seen": true,
  "state_store_name": "xiaohongshu-scraper-state",
  "debug_raw": false,
  "headless": true
}
```

# Actor output Schema

## `notes` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "咖啡店 推荐",
        "skincare routine"
    ],
    "proxy_configuration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("lexac94/xiaohongshu-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "咖啡店 推荐",
        "skincare routine",
    ],
    "proxy_configuration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("lexac94/xiaohongshu-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "咖啡店 推荐",
    "skincare routine"
  ],
  "proxy_configuration": {
    "useApifyProxy": false
  }
}' |
apify call lexac94/xiaohongshu-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lexac94/xiaohongshu-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1NeWByyxoodYIvbfM/builds/VdLWaVCykDzxSh9qr/openapi.json
