# Xiaohongshu RedNote Scraper (`axiomworks/xiaohongshu-note-scraper`) Actor

Scrape public Xiaohongshu (RedNote) notes from note URLs or 24-character note IDs: title, description, hashtags, publish date, author, like, collect, comment and share counts, image URLs and video URL. No keyword or user search: you supply the notes. No login needed.

- **URL**: https://apify.com/axiomworks/xiaohongshu-note-scraper.md
- **Developed by:** [Axiom Works](https://apify.com/axiomworks) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.79 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Xiaohongshu RedNote Scraper

### What does Xiaohongshu RedNote Scraper do?

Xiaohongshu RedNote Scraper collects details from public Xiaohongshu (RedNote) note pages when you provide note URLs or note IDs. It reads the note data embedded in each public page and returns the post text, author identity, engagement counts, tags, image links, and available video details. You can also draw a sample of notes from Xiaohongshu's public sitemap. No login or cookies are required for pages that are publicly accessible.

This Actor is designed for collecting a defined list of notes, checking public post details, or exploring a small sample from the site's sitemap. The sitemap is a fixed list chosen by Xiaohongshu; it is not a keyword search or a chronological feed. You can run the Actor manually, schedule it, or call it through the Apify API and integrations. It does not fetch comments or profile pages.

The source can change without notice. Some notes are removed, restricted, or require a token. When a page cannot be read after three attempts, the dataset includes an `unavailable` row for its ID so you can identify the gap. Those rows are not charged as results. The Actor sends requests at a modest pace and stops at the item limit you set.

### What data can you get?

Each successful row describes one note. Fields the page does not provide are null. Counts are parsed from the values shown on the page, so a rounded display such as `1.2万` becomes an approximate number. Turn on `includeRawCounts` when you also need the original display strings. Image and video links are https media URLs that may expire; store or use them promptly if your workflow needs them. The Actor returns media URLs without downloading the media.

| Field | What it means |
| --- | --- |
| `id`, `noteId` | Stable 24-character note ID; useful for joins and deduplication. |
| `sourceUrl`, `url` | Requested page URL and canonical public note URL. |
| `type` | Note type, usually `normal` or `video`. |
| `title`, `description` | Public title and description. The description has #topic# markers and sticker codes removed; topics are in `hashtags`. |
| `hashtags` | Tag names attached to the note. |
| `publishedAt`, `lastUpdatedAt` | Publication and update times as ISO 8601 strings, when present. |
| `ipLocation` | Public location label, when the page provides one. |
| `authorId`, `authorNickname` | Public author ID and nickname. |
| `authorAvatarUrl`, `authorProfileUrl` | Avatar link and constructed profile URL; the profile itself is not scraped. |
| `likeCount`, `collectCount` | Approximate likes and saves parsed from display text. |
| `commentCount`, `shareCount` | Approximate comment and share counts. Comments themselves are not collected. |
| `imageUrls` | https URLs for the note's images. |
| `videoUrl`, `videoDurationSec` | Video URL and duration when available for a video note. |
| `rawCounts` | Original displayed count strings, only with `includeRawCounts`. |
| `scrapedAt` | Time this Actor read the note. |
| `error` | `unavailable` for a page that could not be read. |

The `sourceUrl` is the exact normalized page or endpoint requested. A URL supplied in `/discovery/item/` form is normalized to `/explore/` because the latter is the public page used by this scraper. If an input URL includes `xsec_token`, that query parameter is retained in `sourceUrl`. Duplicate IDs are emitted once, even when your input contains different URLs for the same note.

### How to use

1. Open the Actor's Input tab and paste one or more public note URLs or IDs into **Note URLs or IDs**. You can copy an `/explore/` URL, a `/discovery/item/` URL, another Xiaohongshu URL containing a 24-character note ID, or the bare ID.
2. Set **Maximum items** to the number of unique note records you want. The limit includes unavailable rows, which helps keep a requested list aligned with the output. For example, a list of ten valid unique notes with `maxItems: 7` produces seven rows.
3. Leave the proxy setting as provided unless you need a different Apify Proxy configuration. The Actor uses the proxy only when that setting is supplied.
4. Click **Start**. When the run finishes, inspect the Dataset tab. Export the dataset as JSON, CSV, or another supported format, or use the API link in the Output tab.

For sitemap discovery, clear the URL list and enable **Discover from sitemap**. That mode reads the site's public note sitemap and works through entries in its published order. It is useful as a bulk sample, but it cannot select a topic, author, or date range. Start with a small `maxItems` to check whether the sample fits your needs. If you need a specific set of notes, provide their URLs instead.

The Actor validates input before requesting a page. A missing source, malformed URL, URL on another host, ID of the wrong length, or item limit outside the supported range causes a clear error. This avoids a run that silently returns nothing. A note that exists but cannot be loaded is different: it produces an `unavailable` row after bounded retries.

### Input

| Field | Type | Use |
| --- | --- | --- |
| `noteUrls` | array of strings | Public note URLs or bare IDs. The prefilled examples are ready for a quick test. |
| `discoverFromSitemap` | boolean | Default false. Discover a fixed public sample when `noteUrls` is empty. |
| `maxItems` | integer | Stop after this many unique records, from 1 to 10,000. Default 20. |
| `includeRawCounts` | boolean | Default false. Include original text for the four engagement counts. |
| `proxyConfiguration` | object | Apify Proxy settings for outbound requests. |

Example input for two known notes:

```json
{
  "noteUrls": [
    "68c8d6d7000000001d01bdbd",
    "https://www.xiaohongshu.com/discovery/item/68c90c25000000000e00c789"
  ],
  "maxItems": 2,
  "includeRawCounts": true,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

The optional `xsec_token` query parameter is preserved when you provide it. Other query parameters are dropped during normalization because they are not needed to identify the note. The Actor never asks for account cookies or passwords.

For a one-off run, begin with the prefilled input and confirm that the current public pages remain accessible. The prefilled list includes several examples and a three-item limit. To collect your own list, replace those examples with your own note URLs. A note URL with a 24-character hexadecimal ID anywhere in its path can be normalized; the host must still be Xiaohongshu. The Actor intentionally rejects arbitrary external URLs, which keeps requests within its declared target hosts.

### Output

The following is a real JSON record from a local sitemap run. Keys the page did not provide (such as `videoUrl`, `ipLocation`, `rawCounts`, `error`) are present with a null value. Text and media links are returned as seen during that run, so later runs may show updated counts or fresh signed URLs.

```json
{
  "id": "6895c3670000000004005f5d",
  "sourceUrl": "https://www.xiaohongshu.com/explore/6895c3670000000004005f5d",
  "noteId": "6895c3670000000004005f5d",
  "url": "https://www.xiaohongshu.com/explore/6895c3670000000004005f5d",
  "type": "normal",
  "title": "总算知道怎么挑精华！速抄作业走上护肤捷径",
  "description": "本人肤质：冬混干夏混油，敏皮上岸\n护肤这么些年我是既要又要\n用过的精华没有一百也有五十瓶\n今天就总结一波值回票价的精华使用感\n大家可以对症下药～\n-\n娇韵诗第九代黄金双萃【进阶抗老】\n赫莲娜绿宝瓶精华【急救维稳】\n海蓝之谜鎏金精华【紧致饱满】\n雅诗兰黛小棕瓶精华【夜间抗老】",
  "hashtags": [],
  "publishedAt": "2025-10-10T03:18:16.000Z",
  "lastUpdatedAt": "2025-12-31T05:06:28.000Z",
  "authorId": "5dea493c00000000010039a2",
  "authorNickname": "Ong菜大王",
  "authorAvatarUrl": "https://sns-avatar-qc.xhscdn.com/avatar/1040g2jo31hggvighjs705nfa94u08ed2imjrj4o",
  "authorProfileUrl": "https://www.xiaohongshu.com/user/profile/5dea493c00000000010039a2",
  "likeCount": 290,
  "collectCount": 198,
  "commentCount": 9,
  "imageUrls": [
    "http://sns-webpic-qc.xhscdn.com/202609300450/ed89499765e78785913902f59b6e0e0e/1040g2sg31qocjum200eg5nfa94u08ed2qt64898!nd_dft_wlteh_jpg_3",
    "http://sns-webpic-qc.xhscdn.com/202609300450/02cb67c8d407cc4b8bc53cfde983641d/1040g00831omtk55r0g1g5nfa94u08ed207g8rq0!nd_dft_wlteh_jpg_3",
    "http://sns-webpic-qc.xhscdn.com/202609300450/eff99a8e87047722333054ce66b4f985/1040g00831omtk55r0g105nfa94u08ed2l39apig!nd_dft_wlteh_jpg_3",
    "http://sns-webpic-qc.xhscdn.com/202609300450/ed624e3080a9f027ed59708be891211c/1040g00831omtk55r0g0g5nfa94u08ed2pklmfdg!nd_dft_wlteh_jpg_3",
    "http://sns-webpic-qc.xhscdn.com/202609300450/42dcef7db705642194f35ee50ae9eed9/1040g2sg31qocjum200e05nfa94u08ed24ckune8!nd_dft_wlteh_jpg_3"
  ],
  "scrapedAt": "2026-09-29T20:50:52.953Z"
}
```

A failed page is represented by a row with `id`, `sourceUrl`, `noteId`, `url`, `scrapedAt`, and `error: "unavailable"`; the other fields are null. It is still a dataset item, so you can compare the output against your input list. The Actor charges the `result` event only when a note is successfully scraped. There is no enrichment or other extra charge.

The output schema lists every field the Actor can produce, while the overview view places the most useful fields in the table. Date strings are in UTC. Counts are integers when a displayed count could be parsed, but a popular note may have a rounded count. Optional media and author fields depend on what Xiaohongshu embeds in the public page at the time of the request.

### How much does it cost?

This Actor uses per-result charging: one `result` event for each successfully scraped note. An unavailable row does not trigger that event. Check the Actor's **Pricing** tab for the current per-result rate; prices can change. Set `maxItems` before starting a large run so the output volume stays within your planned budget. A sitemap sample can generate many requests, so test a small limit first.

An input list with repeated IDs does not increase the result count because the Actor deduplicates by note ID. Failed notes are not charged. For reliable cost estimates, run a short representative batch and inspect its result count, duration, and proxy usage. The Actor does not download image or video files, which keeps the transfer focused on note HTML and sitemap XML.

### Use with the API

You can call the Actor from a scheduled Apify run, a workflow integration, or code. Supply your own Apify API token through the client environment. The examples show an input list and a limit; adapt them to your own public notes. The API result is a run, whose default dataset contains the rows described above.

Python with the Apify client:

```python
from apify_client import ApifyClient
import os

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("axiomworks/xiaohongshu-note-scraper").call(run_input={
    "noteUrls": ["68c8d6d7000000001d01bdbd"],
    "maxItems": 1,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["id"], item.get("title"))
```

JavaScript with the Apify client:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('axiomworks/xiaohongshu-note-scraper').call({
  noteUrls: ['68c8d6d7000000001d01bdbd'],
  maxItems: 1,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => process.stdout.write(`${item.id}: ${item.title ?? item.error}\n`));
```

cURL with a synchronous dataset response:

```bash
curl -X POST \
  'https://api.apify.com/v2/acts/axiomworks~xiaohongshu-note-scraper/run-sync-get-dataset-items' \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"noteUrls":["68c8d6d7000000001d01bdbd"],"maxItems":1}'
```

Keep your API token in your environment or a secret manager. The Actor input needs note URLs or discovery settings, not an account password. You can also connect the default dataset to supported Apify integrations for storage and processing. Since image links may expire, an integration that needs media should process those links soon after the run.

### Use with AI agents (MCP)

Apify's MCP server can expose this Actor as a tool to an AI agent. Configure an MCP client with the server URL below. Use the Actor's input schema to give the agent clear, bounded requests. Your agent can then inspect the dataset rather than trying to infer post details from a link alone.

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=axiomworks/xiaohongshu-note-scraper"
    }
  }
}
```

Example prompts: “Collect the public title, author, hashtags, and like count for these three Xiaohongshu note URLs.” “Read ten notes from the public sitemap and return only their titles and publication dates.” “Check which of these note IDs are unavailable and list the IDs.” Specify a small item limit when exploring, and remember that the sitemap cannot search by keyword.

The Actor's dataset schema tells the agent which fields exist (absent values are null). If your agent needs exact display strings for engagement counts, ask it to set `includeRawCounts` to true. The source record contains both a stable `id` and a `sourceUrl`, making it easier for downstream steps to cite the public page behind a result.

### FAQ

**Can this Actor search Xiaohongshu by keyword?** No. It reads specific public notes or a fixed sample from the public sitemap. The sitemap is not a search index with filters. If you already have note URLs from another workflow, paste those URLs into `noteUrls`.

**Why did a note return `unavailable`?** The public page may have been removed, restricted, redirected, or temporarily blocked. The Actor retries up to three times with backoff and a fresh proxy session when configured. It does not log in, solve CAPTCHAs, or request signed private APIs. The row includes the note ID so you can retry it in a later run if appropriate.

**Are counts exact?** They reflect displayed page values. Large values may be rounded by Xiaohongshu, and the integer parser turns units such as `万` and `千` into numbers. Use `includeRawCounts` to keep the displayed form beside the parsed estimate.

**Does it scrape comments or user profiles?** No. It returns a comment count when available, and it builds a profile URL from the public author ID. It does not open profile pages or request comments. Those data are outside the scope of this Actor.

**Are images and videos downloaded?** No. The Actor returns links embedded in the public note. They can be signed and short lived. Video fields are null unless a video stream is available in the page data. Image notes have a null video URL.

**What happens when I supply the same note twice?** The Actor normalizes each input to its note ID and emits the ID once. This applies across `/explore/`, `/discovery/item/`, and bare-ID forms. The first supplied URL determines any retained `xsec_token` for that ID.

### Is it legal to scrape Xiaohongshu?

The Actor reads publicly visible note pages without logging in. Xiaohongshu's robots.txt disallows crawling, and its terms may restrict automated use. Assess those rules and the law that applies to your use case before collecting or reusing data. This Actor is an independent tool and is not affiliated with Xiaohongshu or RedNote.

Public posts can include personal data, including author IDs and nicknames. Handle it lawfully, respect privacy rights, and apply relevant data protection rules, including China's PIPL where it applies. Only collect and retain data you need, and avoid using the output for harassment, profiling, or unauthorized republication. Access to a public page can change, and a successful run is not permission to republish its content.

### Feedback

If a public note page changes format, share a sample URL, the input settings, and the relevant run log through the Actor's Issues or feedback channel. Do not post tokens, cookies, proxy passwords, or private data. A report that identifies the affected note ID and which fields are missing is especially useful. You can also report a sitemap change or a case where a displayed count is parsed incorrectly.

The Actor is maintained around public note details. Requests for search results, comments, or login-only profile information require a different scope and may not be possible without access this Actor intentionally avoids. For this Actor, the most useful feedback is a reproducible public page and the expected field value visible on that page.

# Actor input Schema

## `noteUrls` (type: `array`):

Public Xiaohongshu note URLs or 24-character note IDs to scrape. Include xsec\_token in a URL when needed.

## `discoverFromSitemap` (type: `boolean`):

Read public note sitemap entries when no note URLs are supplied. This is a fixed sample, not keyword search.

## `maxItems` (type: `integer`):

Maximum number of note records to return, including unavailable-note rows, 1-10000. Default 20.

## `includeRawCounts` (type: `boolean`):

Include engagement counts as displayed on Xiaohongshu alongside parsed numbers.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings for requesting public note pages. Choose a proxy group if your access requires it.

## Actor input object example

```json
{
  "noteUrls": [
    "https://www.xiaohongshu.com/explore/68c8d6d7000000001d01bdbd",
    "https://www.xiaohongshu.com/explore/69217182000000001e03bc69",
    "https://www.xiaohongshu.com/explore/68dbcfca0000000004010eba"
  ],
  "discoverFromSitemap": false,
  "maxItems": 3,
  "includeRawCounts": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "noteUrls": [
        "https://www.xiaohongshu.com/explore/68c8d6d7000000001d01bdbd",
        "https://www.xiaohongshu.com/explore/69217182000000001e03bc69",
        "https://www.xiaohongshu.com/explore/68dbcfca0000000004010eba"
    ],
    "discoverFromSitemap": false,
    "maxItems": 3,
    "includeRawCounts": false,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiomworks/xiaohongshu-note-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "noteUrls": [
        "https://www.xiaohongshu.com/explore/68c8d6d7000000001d01bdbd",
        "https://www.xiaohongshu.com/explore/69217182000000001e03bc69",
        "https://www.xiaohongshu.com/explore/68dbcfca0000000004010eba",
    ],
    "discoverFromSitemap": False,
    "maxItems": 3,
    "includeRawCounts": False,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("axiomworks/xiaohongshu-note-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "noteUrls": [
    "https://www.xiaohongshu.com/explore/68c8d6d7000000001d01bdbd",
    "https://www.xiaohongshu.com/explore/69217182000000001e03bc69",
    "https://www.xiaohongshu.com/explore/68dbcfca0000000004010eba"
  ],
  "discoverFromSitemap": false,
  "maxItems": 3,
  "includeRawCounts": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call axiomworks/xiaohongshu-note-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiomworks/xiaohongshu-note-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/g5kaNbTUqurmdtPmY/builds/Nde8VzDIrfhSCNut2/openapi.json
