RedNote (Xiaohongshu) Comments Scraper + Replies avatar

RedNote (Xiaohongshu) Comments Scraper + Replies

Pricing

from $4.50 / 1,000 results

Go to Apify Store
RedNote (Xiaohongshu) Comments Scraper + Replies

RedNote (Xiaohongshu) Comments Scraper + Replies

Extract RedNote/Xiaohongshu comments and nested replies with authors, likes, locations, timestamps, and structured export-ready data.

Pricing

from $4.50 / 1,000 results

Rating

0.0

(0)

Developer

Julio Gonzalez

Julio Gonzalez

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

18 days ago

Last modified

Share

Extract public RedNote / Xiaohongshu comments and nested replies into clean, structured data.

Use the results for market research, brand monitoring, audience analysis, creator research, customer-language discovery, AI workflows, spreadsheets, dashboards, or your own applications.

Unofficial tool. Not affiliated with, endorsed by, or operated by RedNote / Xiaohongshu.

What you get

Each result can include:

  • comment or reply text
  • author name and author ID when available
  • author avatar URL when available
  • like count
  • reply count
  • timestamp
  • location label when supplied by the source
  • comment and parent-comment IDs
  • original post URL

Results are saved to an Apify dataset and can be exported to JSON, CSV, Excel, or accessed through the Apify API.

Common use cases

  • analyze what customers are saying about products or brands
  • collect recurring questions, complaints, and pain points
  • monitor competitors and public discussion
  • research creators and their audiences
  • identify frequently discussed topics
  • collect comment datasets for AI or NLP analysis
  • feed RedNote discussion data into dashboards or automated workflows

Input

Provide one or more public RedNote / Xiaohongshu post URLs.

You can choose:

  • Maximum comments per post — default 100
  • Include replies — optionally collect nested replies
  • Maximum replies per comment — default 20
  • Comment orderLatest or Hot / most liked

Up to 50 post URLs can be submitted in one run.

Output

Each dataset item represents a normalized top-level comment or nested reply.

Fields can include:

  • item type (comment or reply)
  • comment ID and parent comment ID
  • comment/reply text
  • author ID, name, and avatar URL
  • like count
  • timestamp
  • location
  • reply count
  • source post URL

The public Actor returns normalized data only and does not expose the upstream TikHub credential or raw TikHub response objects.

TikHub configuration

The Actor uses TikHub's Xiaohongshu App V2 API as its upstream provider. Configure this secret environment variable in the Actor settings:

TIKHUB_API_KEY

TIKHUB_TOKEN is accepted as an operator fallback.

For full Xiaohongshu URLs, the Actor extracts and sends note_id, which TikHub documents as preferred. For supported short/share links where the note ID cannot be extracted locally, it uses TikHub's share_text parameter.

Pagination and replies

Top-level comments follow TikHub's documented pagination contract:

  • first request: empty cursor, index 0, pageArea UNFOLDED
  • next request: reuse the previous response's cursor, index, and pageArea

Second-level replies follow TikHub's documented sub-comment pagination contract:

  • first request: empty cursor, index 1
  • next request: reuse cursor and index from the previous response

Replies are fetched only when the top-level source metadata indicates that replies exist. This avoids blindly issuing a paid sub-comment request for every comment.

If a later pagination request is rejected after earlier pages succeeded, the Actor keeps the already returned dataset results and stops that pagination branch cleanly.

Upstream safety controls

The Actor implements conservative TikHub request handling:

  • one TikHub request at a time per Actor run
  • minimum 500 ms interval between upstream requests by default
  • maximum three attempts only for transient network/5xx failures
  • exponential backoff with jitter
  • HTTP/API 401, 403, and 429 immediately halt further TikHub requests in that run
  • HTTP 400 and other request-specific client errors are not hammered with retries
  • hard upstream request cap per run (default 2,000)
  • pagination loop protection
  • duplicate dataset protection
  • Apify spending-limit detection stops further upstream work as soon as the SDK reports the run limit has been reached
  • clearly unrelated URLs are rejected before making a TikHub request

Operator-only environment variables:

  • TIKHUB_MIN_INTERVAL_MS — default 500; minimum enforced 250
  • TIKHUB_MAX_REQUESTS_PER_RUN — default 2000

API routes used

  • GET /api/v1/xiaohongshu/app_v2/get_note_comments
  • GET /api/v1/xiaohongshu/app_v2/get_note_sub_comments

Verification

Run from inside the Actor folder:

$python3.14 self_test.py

Expected output:

Social Note Comments Scraper v1.0 final self-test passed (22 checks).

The release self-test covers exact first-request parameters, top-level pagination handoff, sub-comment pagination, preservation of first-page results after a later-page HTTP 400, nested replies, duplicate protection, spending-limit stop behavior, authentication/rate-limit circuit breakers, retry caps, request caps, serialization of concurrent calls, input validation, and Actor schemas.

Responsible use

Use this Actor only for lawful purposes and public data, subject to applicable privacy, intellectual-property, and platform requirements. Do not use it to redistribute protected content without authorization or to expose upstream credentials/direct API access.