# Changelog of Douban Pro Scraper — Reviews, Discussions & Subject Data (`zhorex/douban-scraper`) Actor

- **URL**: https://apify.com/zhorex/douban-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/zhorex/douban-scraper.md

## Changelog

All notable changes to the Douban Scraper Actor.

### 2026-09-30: author and date on movie reviews

- Every movie review row came out with `authorUsername`, `authorUrl`, `publishedAt` and `publishedAtIso` null: the mobile list, the only list movies have, shows no author link and no date.
- `authorUrl` is now built from the numeric user id in the avatar file the mobile list does show (`icon/u1000152-23.jpg` -> `https://www.douban.com/people/1000152/`). On the desktop lists in the test fixtures, every numeric author link (10 of 10) matches its avatar's id. A default avatar (`user_normal.jpg`) names nobody, so that row keeps `authorUrl` null. The same applies to book reviews read from the mobile list (`listSource: "mobile"`).
- With `fetchFullReviewBody` on (the default), the movie body already comes from Douban's review JSON. Where that JSON carries them, the row now also takes `authorUsername` (`user.name`), the author's own profile link (`user.url`) and the date (`create_time`, Beijing time, into `publishedAt` / `publishedAtIso`). Anything missing stays null. The field names are read defensively and have not yet been checked on a live run.
- No new request, no price, default or prefill change; charges and requests are identical offline. Offline tests: `python -m pytest tests/test_movie_review_author.py`.

### 2026-09-26: deltaMode no longer re-bills the rows at the top of a list

- Fixed: the delta memory kept rows in the order they were first delivered and was cut to the last 20,000, but a row Douban listed again kept its old place. `subject_reviews` and `subject_comments` re-read the same most-popular-first list every run, so the rows at its top — delivered on the first run and listed on every run since — were the first forgotten once a stream passed 20,000 rows, and the next run delivered and charged them again. A row listed again now counts as recent; only the rows Douban has gone longest without listing are forgotten.
- Fixed: when a run could not read the delta memory (an Apify API error), it saved that run's few rows over the whole memory, and the next run delivered and charged everything again (offline: 80 remembered, then 5, then 75 charged again). It now re-reads the memory when saving and adds to it, or leaves it untouched if it still cannot be read. A failed read of the key an earlier build used for a repeated target no longer discards the stream's own memory.
- The memory holds 200,000 rows per stream (about 192,000 for short comments; was 20,000), under about 5 MB. Memories saved by earlier builds load as they are.
- No price, default or prefill change; runs without `deltaMode` are unchanged (checked offline against the previous commit).

### 2026-09-24: new `group_search` mode

#### Changed

- `group_topic`: Douban now answers topic pages with its login wall (4 of 4 measured on 24 Sep 2026, also through Apify). Such runs used to end with only "no billable items"; the status now names the topics Douban did not serve (not charged) and points to `group_search`, which works without login. The mode's title in the input form says so too.
- One-off runs of the existing modes no longer end their status with a "turn on deltaMode + a daily Schedule" tip. Nothing else about those runs changed (same requests, rows, charges; checked against the saved pre-change output). `deltaMode` itself works and is documented as before.

#### Added

- **`group_search` mode.** Put a keyword in **Search query**. The run returns one row for each Douban Group discussion whose **title contains that exact keyword**, newest first, up to `maxResults`. Matching is case-insensitive for Latin letters.
- Douban's own group search is fuzzy: 李宁 also lists titles that only contain 李 and 宁 apart. Those looser matches are dropped, are not delivered and are not charged. The run status says how many were dropped.
- Row fields: `topicId`, `title`, `url` (clean topic link, without the tracking parameter), `groupId`, `groupName`, `groupUrl`, `replyCount`, `postedAt` (Beijing time, `+08:00`), `keyword`, `mode`, `type` (`group_search_result`) and `scrapedAt`. If the listing does not show a value, the field is `null`, never a made-up `0`.
- One run reads at most 10 search pages (about 500 listed topics), because Douban sends a connection that keeps paging through its search to an anti-bot check (measured on 24 Sep 2026 after about 15 requests in an hour). It stops earlier when it has `maxResults` mentions, when the list ends, after 3 pages in a row without a title mention, when a page repeats topics already read, or at the first page Douban refuses (login wall or anti-bot check). The status says which of these ended the run.
- If Douban's page layout changes so that listed rows cannot be read, the run says so instead of reporting "no results". Rows it cannot read are counted in the status and never charged.
- When a run stops early after delivering rows, the status says that a re-run starts again from page 1 and would bill those rows again.
- Billing: each delivered row is billed with the existing `subject-search-result` event ($0.005), the same event and price as `subject_search` results. No new events and no price changes.

#### Unchanged

- The other four modes and all input defaults and prefills are the same as before. An offline test checks that each of them produces the same rows, charges and status messages as the previous build.
