# Changelog of RedNote (Xiaohongshu) Scraper — Trending Notes & Creators (`logiover/rednote-xiaohongshu-scraper`) Actor

- **URL**: https://apify.com/logiover/rednote-xiaohongshu-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/logiover/rednote-xiaohongshu-scraper.md

## Changelog

### 2026-09-23

- Fleet-wide quality audit. Verified end to end against live data and re-checked the input schema, the output columns and the run configuration.
- Output verified on a live run: 30 columns returned, 80.0% of cells populated.
- Run reliability reviewed: 100.0% of public runs succeeded in the last 30 days.
- Input schema, output schema and pricing configuration reviewed.
- Noted that 6 column(s) came back empty in this sample (`description`, `collectedCount`, `commentCount`, `shareCount`, `ipLocation`, `publishTime`); these are under review.

### 2026-09-01

- Fleet-wide health check. Verified against this Actor's real run history: 30-day success rate, output row counts, per-field fill rates, and peak memory against the configured memory limit.
- Reviewed for the failure patterns that have cost this fleet runs — unguarded proxy setup, retry loops that can outlast the run's own time budget, and full-page HTML parsing that can exhaust a small container.
- No change to input, output fields or scraping logic.

### 2026-08-11

- Maintenance release: refreshed the build and dependencies.
- Re-verified live execution, non-empty structured output and dataset field/type integrity.
- Reviewed reliability (retries, pagination) and output quality as part of a full-fleet QA pass.

### 2026-08-01

- Completed the August 2026 full health check: verified empty/programmatic default, Console UI default, and two source-informed alternative inputs on Apify.
- Confirmed successful live execution, non-empty structured output, dataset-field/type integrity, and logical sample quality within the 5-minute quality window.
- Replaced the field-level output schema with a canonical `results` dataset link so run results open correctly in the Apify Console.
- Declared 30 dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder.
- Declared 30 nullable dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder or brittle to sparse modes.

### 2026-07-24 — v2.1

- **Guaranteed 500+ notes per run.** Widened the keyless discovery pool from 12 to **67 live-probed `channel_type` verticals** (all confirmed to return an SSR `__INITIAL_STATE__` feed with no cookie), and confirmed the anonymous feed **refreshes on every request** (0% overlap over 6 straight passes of one channel). A single full round-robin sweep now yields ~1,400 unique notes; the multi-cycle sweep re-loops until `maxResults` or a dry-stop.
- **Smarter multi-cycle sweep + dry-stop.** The trending mode keeps re-sweeping all channels; it stops only when `maxResults` is reached, the ~4-minute time budget runs out (pushes what it has, exits success), or the last 3 *completed* sweeps together add fewer than 5 new unique notes (pool drained).
- **Higher default volume.** `maxResults` default & prefill raised 300 → **500**.
- **More channels selectable** in `channel` mode (48-option dropdown: Beauty, Parenting, Tech, Car, Pets, Sports, Music, Photography, Finance, Anime, Astrology, Jewelry, and more).
- Output field keys unchanged (fully backward compatible).

### 2026-07-24

- **Major rewrite → v2.0: fully keyless engine.** Removed the cookie/login requirement entirely. The scraper now fetches RedNote's server-rendered `/explore` HTML and parses the inline `window.__INITIAL_STATE__` → `feed.feeds[]`. No cookie, no login, no API key.
- **New "trending" discovery engine.** Sweeps all 12 discovery channels (Recommend, Fashion, Food, Cosmetics, Movie & TV, Career, Love, Household, Gaming, Travel, Fitness, Video). Each channel returns fresh notes on every request and channels never overlap, so the scraper loops and dedupes by note ID to accumulate hundreds of unique trending notes per run.
- **Two modes** via a dropdown: `trending` (sweep all channels, default) and `channel` (one specific channel).
- **Empty input now works.** `{}` runs a trending sweep with `maxResults` = 300. Nothing is required.
- **Numbers as numbers.** `likedCount` is now a real integer (Chinese "1.4万" / "10万+" are parsed to 14000 / 100000); the original string is kept in `likedCountText`.
- **Richer output**: added `noteUrl`, `coverImage`, `coverImagePreview`, `coverWidth`, `coverHeight`, `channel`, `channelName`, flat creator fields (`userId`, `nickname`, `avatar`, `userUrl`, `creatorXsecToken`). All legacy field keys (`noteId`, `type`, `title`, `likedCount`, `cover`, `images`, `url`, `xsecToken`, `author.*`, `scrapedAt`) are preserved.
- **Lighter & faster**: switched from Playwright/Chromium + crawlee to pure `got-scraping` over HTTP (512 MB, no browser). Residential proxy with a fresh IP per retry.
- **Graceful time budget**: stops after ~4 minutes, pushes everything collected, exits success.
- Honest scope: full note body text, comments, hashtags, search and profile stats require a logged-in cookie and are out of scope for this keyless actor.

### 2026-06-05

- Reliability fix: results were no longer dropped by strict output validation.
- Stability & performance hardening; fresh rebuild.
