Scrape long-form reviews of movies, books and music, short book and music comments with star ratings, and group discussion titles from Douban (豆瓣), China's reviews and interest community. Plus subject and group search. For Chinese-LLM training corpora, sentiment analysis and NLP. API, no login.
Every movie review row came out with authorUsername, authorUrl, publishedAt and publishedAtIso null: the mobile list, the only list movies have, shows no author link and no date.
authorUrl is now built from the numeric user id in the avatar file the mobile list does show (icon/u1000152-23.jpg -> https://www.douban.com/people/1000152/). On the desktop lists in the test fixtures, every numeric author link (10 of 10) matches its avatar's id. A default avatar (user_normal.jpg) names nobody, so that row keeps authorUrl null. The same applies to book reviews read from the mobile list (listSource: "mobile").
With fetchFullReviewBody on (the default), the movie body already comes from Douban's review JSON. Where that JSON carries them, the row now also takes authorUsername (user.name), the author's own profile link (user.url) and the date (create_time, Beijing time, into publishedAt / publishedAtIso). Anything missing stays null. The field names are read defensively and have not yet been checked on a live run.
No new request, no price, default or prefill change; charges and requests are identical offline. Offline tests: python -m pytest tests/test_movie_review_author.py.
2026-09-26: deltaMode no longer re-bills the rows at the top of a list
Fixed: the delta memory kept rows in the order they were first delivered and was cut to the last 20,000, but a row Douban listed again kept its old place. subject_reviews and subject_comments re-read the same most-popular-first list every run, so the rows at its top — delivered on the first run and listed on every run since — were the first forgotten once a stream passed 20,000 rows, and the next run delivered and charged them again. A row listed again now counts as recent; only the rows Douban has gone longest without listing are forgotten.
Fixed: when a run could not read the delta memory (an Apify API error), it saved that run's few rows over the whole memory, and the next run delivered and charged everything again (offline: 80 remembered, then 5, then 75 charged again). It now re-reads the memory when saving and adds to it, or leaves it untouched if it still cannot be read. A failed read of the key an earlier build used for a repeated target no longer discards the stream's own memory.
The memory holds 200,000 rows per stream (about 192,000 for short comments; was 20,000), under about 5 MB. Memories saved by earlier builds load as they are.
No price, default or prefill change; runs without deltaMode are unchanged (checked offline against the previous commit).
2026-09-24: new group_search mode
Changed
group_topic: Douban now answers topic pages with its login wall (4 of 4 measured on 24 Sep 2026, also through Apify). Such runs used to end with only "no billable items"; the status now names the topics Douban did not serve (not charged) and points to group_search, which works without login. The mode's title in the input form says so too.
One-off runs of the existing modes no longer end their status with a "turn on deltaMode + a daily Schedule" tip. Nothing else about those runs changed (same requests, rows, charges; checked against the saved pre-change output). deltaMode itself works and is documented as before.
Added
group_search mode. Put a keyword in Search query. The run returns one row for each Douban Group discussion whose title contains that exact keyword, newest first, up to maxResults. Matching is case-insensitive for Latin letters.
Douban's own group search is fuzzy: 李宁 also lists titles that only contain 李 and 宁 apart. Those looser matches are dropped, are not delivered and are not charged. The run status says how many were dropped.
Row fields: topicId, title, url (clean topic link, without the tracking parameter), groupId, groupName, groupUrl, replyCount, postedAt (Beijing time, +08:00), keyword, mode, type (group_search_result) and scrapedAt. If the listing does not show a value, the field is null, never a made-up 0.
One run reads at most 10 search pages (about 500 listed topics), because Douban sends a connection that keeps paging through its search to an anti-bot check (measured on 24 Sep 2026 after about 15 requests in an hour). It stops earlier when it has maxResults mentions, when the list ends, after 3 pages in a row without a title mention, when a page repeats topics already read, or at the first page Douban refuses (login wall or anti-bot check). The status says which of these ended the run.
If Douban's page layout changes so that listed rows cannot be read, the run says so instead of reporting "no results". Rows it cannot read are counted in the status and never charged.
When a run stops early after delivering rows, the status says that a re-run starts again from page 1 and would bill those rows again.
Billing: each delivered row is billed with the existing subject-search-result event ($0.005), the same event and price as subject_search results. No new events and no price changes.
Unchanged
The other four modes and all input defaults and prefills are the same as before. An offline test checks that each of them produces the same rows, charges and status messages as the previous build.