Douban Pro Scraper — Reviews, Discussions & Subject Data
Pricing
from $30.00 / 1,000 review scrapeds
Douban Pro Scraper — Reviews, Discussions & Subject Data
Scrape long-form reviews of movies, books and music, short book and music comments with star ratings, and group discussion titles from Douban (豆瓣), China's reviews and interest community. Plus subject and group search. For Chinese-LLM training corpora, sentiment analysis and NLP. API, no login.
Pricing
from $30.00 / 1,000 review scrapeds
Rating
0.0
(0)
Developer
Sami
Maintained by CommunityActor stats
0
Bookmarked
249
Total users
129
Monthly active users
3 days ago
Last modified
Categories
Share
Douban Scraper — Reviews, Comments & Group Discussions
Extract long-form reviews, ratings, comments, and group discussions from Douban (豆瓣) — China's leading reviews + interest community. Movies, books, and music. No API key, no browser, no VPN. Built for Chinese AI training corpora and consumer research.
How to scrape Douban in 3 easy steps
- Go to the Douban Scraper page on Apify and click "Try for free"
- Configure your input — choose a mode (
subject_reviews,subject_comments,group_topic,subject_search, orgroup_search), enter your Douban URLs or query, and set the number of results - Click "Run", wait for the scraper to finish, then download your data in JSON, CSV, or Excel format
No coding required. No API key. Works with Apify's free plan.
Part of the Chinese Digital Intelligence Suite
Built by Zhorex, who maintains a suite of Chinese-platform scrapers:
- 🆕 Chinese Brand Monitor — Cross-platform brand mention aggregator (Weibo + RedNote + Bilibili + Douban + Xueqiu in one normalized feed, sentiment + dedup, $0.06/mention)
- Weibo Scraper — China's Twitter (microblogging, hot search, public opinion)
- RedNote (Xiaohongshu) Scraper — China's Instagram + Pinterest
- Xueqiu Scraper — Chinese stock-discussion sentiment, cashtag indexing
- RedNote Shop Scraper — RedShop e-commerce data
- Douban Scraper — You are here (reviews, ratings, group discussions)
Together, these cover the pillars of Chinese digital intelligence: microblogging, video, lifestyle, e-commerce, finance, and long-form reviews. For cross-platform brand monitoring (the most common multi-platform use case), the Chinese Brand Monitor aggregator orchestrates all 5 in one call with normalized output — saves 4-6 hours of engineering vs. wiring up individual scrapers.
What is Douban?
Douban (豆瓣) is China's reviews and interest-community platform — Goodreads + Letterboxd + Rate Your Music + niche-Reddit fused into one site, with 200M+ monthly users. It's where Chinese readers, cinephiles, music fans, and hobby communities post the longest-form opinion content on the Chinese internet. Movies, books, music, TV shows, and tens of thousands of user-run discussion groups.
For anyone building a Chinese-language LLM, sentiment classifier, or consumer research dataset, Douban is a rich source of opinion-heavy long-form Chinese text.
Modes
| Mode | What it does | Records |
|---|---|---|
subject_reviews | Long-form reviews (500-5,000+ Chinese chars each) for a movie/book/music album | One per review |
subject_comments | Short comments + star ratings under a subject's discussion page | One per comment |
subject_search | Search Douban for movies / books / music by keyword | One per result |
group_topic ⚠️ Beta | Pull a discussion thread + its replies from a Douban Group | One per topic (with nested replies) |
group_search | Find Douban Group discussions whose title contains your keyword (brand, product, person), newest first | One per matching topic |
v1.0 Known Limitations (read before buying)
- Movie comments require browser rendering. Douban serves movie short-comments through a JS-only mobile widget that v1.0 cannot extract —
subject_commentsfor movies returns a diagnostic record explaining the limitation. Usesubject_reviewsfor movie data instead — long-form movie reviews are richer for AI training anyway. Books and music short-comments work normally. - Movie review list bodies are excerpt-only by default. Mobile movie list pages don't show the author's name, the publication date or the full body — only review IDs, titles, ratings and the author's avatar.
authorUrlis built from the numeric user id in the avatar file name (https://www.douban.com/people/<id>/), and staysnullfor authors with Douban's default avatar. SetfetchFullReviewBody: true(default) to fetch each review's detail page and fill in the full markdown body; when that detail names the author and the date, the row also getsauthorUsername,publishedAtandpublishedAtIso, andauthorUrlbecomes the author's own profile link. Where it does not, those fields staynull— never guessed. - Book / music comment coverage varies by subject. Popular subjects (rating count > 10K — e.g. 三体, OK Computer) reliably serve inline short-comments; some less-popular books have begun moving comment lists to AJAX-only rendering and will return 0 records. If
subject_commentsreturns 0 records for a URL, fall back tosubject_reviews(which works on all subjects). - Movie search returns Douban's tag-matched feed. For precise targeting of a specific film, use
subject_reviewswith the movie's subject URL directly. - Book search caps at ~10 discovery results per query. Douban's book suggestion endpoint doesn't paginate. For bulk book review extraction, supply multiple subject URLs to
subject_reviewsmode. group_topicmode is Beta. On 24 Sep 2026, every topic page we opened logged out (4 of 4, from a home connection, not yet re-checked through Apify's proxies) showed Douban's login wall instead of the topic, so expect many topics to fail; others return 404 (deleted). When a topic fails, the run logs a warning and continues — you are not charged for failed topics. To find discussions about a keyword without opening topic pages, usegroup_search.- Residential proxies are strongly recommended (default in input). Datacenter IPs degrade movie-mode and may trigger generic anti-bot challenges.
Use Cases
| Who | Why |
|---|---|
| AI / LLM training data buyers | A rich source of Chinese long-form opinion text — key for Chinese-language model fine-tuning |
| Sentiment analysis researchers | Star-rating-labelled Chinese review text, ideal for supervised sentiment classifiers |
| Brand monitoring teams | Find Chinese consumer reviews mentioning your product, competitor films, or book titles |
| Cultural trend analysts | Track which films / books / albums are gaining traction in Chinese-speaking markets |
| Academic NLP researchers | Pre-built corpus of opinion text with engagement metrics — citable in cross-cultural studies |
| Localization / translation teams | Real Chinese phrasing patterns for entertainment vocabulary |
Scrape Douban with Python, JavaScript, or no code
You can use the Douban Scraper directly from the Apify Console (no code), or integrate it into your scripts.
Python
from apify_client import ApifyClientclient = ApifyClient("YOUR_API_TOKEN")run = client.actor("zhorex/douban-scraper").call(run_input={"mode": "subject_reviews","subjectUrls": ["https://book.douban.com/subject/1084336/"],"maxResults": 50,})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });const run = await client.actor('zhorex/douban-scraper').call({mode: 'subject_reviews',subjectUrls: ['https://book.douban.com/subject/1084336/'],maxResults: 50,});const { items } = await client.dataset(run.defaultDatasetId).listItems();items.forEach((item) => console.log(item));
Input examples
1. Subject reviews (long-form)
Pull long-form reviews for one or more movies, books, or music albums. Provide subject URLs or numeric subject IDs.
{"mode": "subject_reviews","subjectUrls": ["https://movie.douban.com/subject/1292052/","https://book.douban.com/subject/1084336/","https://music.douban.com/subject/1419463/"],"maxResults": 100,"fetchFullReviewBody": true}
2. Subject comments + ratings
Short comments + star ratings for a book or music album. Movie comments are not supported in v1.0 — use subject_reviews mode for movies.
{"mode": "subject_comments","subjectUrls": ["https://book.douban.com/subject/1084336/"],"maxResults": 200}
3. Subject search
Search Douban for movies, books, or music by keyword. Returns Douban's discovery feed for that query.
💡 Search the Chinese title/name for far more results. Douban indexes subjects by their Chinese title (e.g.
三体, not "The Three-Body Problem"), so the native term returns far more than the Latin/English one. Few results usually means the query language, not an error.
{"mode": "subject_search","searchQuery": "三体","searchType": "all","maxResults": 30}
4. Group topic (Beta)
Pull one or more group discussion threads with embedded replies, when Douban serves the topic page. Logged out, it often shows its login wall instead (see the Beta note above); those topics return no row and are not charged.
{"mode": "group_topic","topicUrls": ["https://www.douban.com/group/topic/319929381/"],"maxRepliesPerTopic": 100}
5. Group search — who is talking about a keyword in Douban Groups
Douban Groups (豆瓣小组) are where Chinese consumers argue about brands, products and celebrities. group_search searches them for your keyword and returns one row per discussion whose title contains the exact keyword, newest first. Put the keyword in Search query.
{"mode": "group_search","searchQuery": "李宁","maxResults": 50}
What to know before you run it:
- Only real title mentions are returned. Douban's group search is fuzzy: searching 李宁 also lists titles that merely contain 李 and 宁 apart (宁艺卓, 李庚希) or do not contain the word at all. In a measured search only 11 of the 48 topics on the first page had 李宁 in the title. Every other row is dropped, never delivered and never charged, and the run's status says how many were dropped. Matching is an exact substring, case-insensitive for Latin letters (
nikematches Nike and NIKE), so 李宁 also matches 李宁玉. - How deep it goes. Douban lists up to 3,000 results (60 pages of ~50) per keyword, but it does not let one connection page through them for long. In our tests on 24 Sep 2026, pages 1 and 2 opened and page 60 showed the login wall; after about 15 requests to Douban in an hour, the same connection was sent to Douban's anti-bot check for the next search pages. So one run reads at most 10 search pages (about 500 listed topics), a few seconds apart. It stops earlier when it has
maxResultstitle mentions, when the list ends, after 3 pages in a row without a single title mention, or at the first page Douban refuses (login wall or anti-bot check). The status says which of these ended the run. In our test, 2 pages gave 25 title mentions of 李宁, so the defaultmaxResultsof 50 takes about 4 pages. - Titles only, no topic text. Rows carry what the search listing shows: title, group, reply count and time. They do not include the topic's body, author or replies: Douban showed its login wall instead of the topic page on every topic we opened logged out (4 of 4, 24 Sep 2026).
postedAtis when the topic was created, in Beijing time (+08:00). Across 97 measured results, the listed time went up exactly as the topic id went up, even for threads with 200+ replies, so it is not the time of the last reply.- Group names can be cut short. The search listing shortens long group names with
....groupIdandgroupUrlalways identify the group exactly. - Search for the keyword the way Douban users write it: 李宁, not "Li-Ning". Latin brand names such as nike work as written.
Output examples
Review record
{"type": "review","reviewId": "1000104","subjectId": "1084336","subjectName": "小王子","subjectType": "book","title": "长大就笨了","content": "(Full review body in markdown — Chinese long-form text)","rating": 5,"ratingLabel": "力荐","authorUsername": "示例用户A","authorUrl": "https://www.douban.com/people/000000001/","authorAvatarUrl": "https://img3.doubanio.com/icon/u000000001-1.jpg","publishedAt": "2005-04-06 11:51:52","publishedAtIso": "2005-04-06T03:51:52Z","stats": { "replyCount": 444 },"reviewUrl": "https://book.douban.com/review/1000104/","scrapedAt": "2026-05-13T01:39:22Z"}
Comment record
{"type": "comment","commentId": "10287387","subjectId": "1084336","subjectName": "小王子","subjectType": "book","content": "十几岁的时候渴慕着小王子,一天之间可以看四十四次日落。","rating": 5,"ratingLabel": "力荐","authorUsername": "示例用户B","authorUrl": "https://www.douban.com/people/000000002/","publishedAt": "2007-02-08 11:16:40","publishedAtIso": "2007-02-08T03:16:40Z","stats": { "votesCount": 9232 },"scrapedAt": "2026-05-13T01:39:22Z"}
Subject (search result)
{"type": "subject","subjectId": "2567698","subjectName": "三体","subjectType": "book","year": "2008","author": "刘慈欣","rating": null,"cover": "https://img1.doubanio.com/view/subject/s/public/s2768378.jpg","subjectUrl": "https://book.douban.com/subject/2567698/","scrapedAt": "2026-05-13T01:39:22Z"}
Group topic record (Beta)
{"type": "group_topic","topicId": "319929381","groupName": "(Group name)","title": "(Discussion title)","content": "(Topic body in markdown — Chinese long-form text)","authorUsername": "(Author handle)","publishedAt": "2026-04-01 10:20:30","publishedAtIso": "2026-04-01T02:20:30Z","stats": { "repliesCaptured": 50 }, // how many this run captured,// capped by maxRepliesPerTopic — not// the thread's total reply count"replies": [{"replyId": "...","authorUsername": "...","content": "...","publishedAt": "...","votesCount": 12}],"topicUrl": "https://www.douban.com/group/topic/319929381/","scrapedAt": "2026-05-13T01:39:22Z"}
Group search record
{"type": "group_search_result","topicId": "500559514","title": "你们觉得李宁特步名气大还是ask韩潮名气大","url": "https://www.douban.com/group/topic/500559514/","groupId": "757619","groupName": "长江国际售楼部","groupUrl": "https://www.douban.com/group/757619/","replyCount": 52,"postedAt": "2026-09-22T11:40:05+08:00","keyword": "李宁","mode": "group_search","scrapedAt": "2026-09-24T17:19:43Z"}
replyCount is the count the search listing shows (52回复). If the listing shows none, it is null, never a made-up 0. url is the clean topic link, without Douban's tracking parameter.
Pricing
Pay per result — no monthly fee, no minimum.
| Event | Price | When charged |
|---|---|---|
review-scraped | $0.030 | Per long-form review record extracted |
comment-scraped | $0.005 | Per short comment record extracted |
group-topic-scraped | $0.030 | Per group topic Douban served (with embedded replies); login-walled topics are not charged |
subject-search-result | $0.005 | Per search result row — subject_search results and group_search title mentions |
Concrete cost examples:
- 100 long-form reviews of one popular movie's reviews page: $3.00
- 1,000 short comments across multiple books: $5.00
- 200 search results to seed a crawl: $1.00
- 50 group topics whose title mentions a brand (
group_search): $0.25. Fuzzy matches Douban lists but whose title does not contain the keyword are free, because they are never delivered.
Diagnostic / log records (e.g. movie comment limitation notices) are NEVER charged.
Content is in Chinese
All content is returned in the original Simplified Chinese. Douban is a Chinese-language platform — reviews, comments, group discussions, and user names are in Chinese.
If you need English translations, pipe the output through a translation API (Google Translate, DeepL, or Claude).
Technical Details
- No browser — pure HTTP, runs in 512MB RAM
- No login required — works against publicly accessible content only
- Built-in rate limiting — exponential backoff on 429 / 503
- Globally accessible — residential proxy recommended (default in input)
- UTF-8 throughout — Chinese text round-trips cleanly
- Markdown review bodies —
<p>,<a>,<strong>etc. converted to lightweight markdown for downstream LLM ingestion deltaModememory — 200,000 rows per stream (about 192,000 for short comments) (mode + targets, or yourdeltaStateKey). A row Douban lists again counts as recent; when the memory is full, the rows Douban has gone longest without listing are forgotten first, and one listed again after that is delivered and charged again.
FAQ
Is there a Douban API?
Douban's official developer API has been deprecated for several years. There is no working public Douban API for international developers. This Douban Scraper extracts reviews, comments, ratings, and group discussions from publicly accessible web endpoints.
Do I need a Douban login or cookies?
No. The Actor never logs in and takes no cookies. subject_reviews, subject_comments, subject_search and group_search read pages Douban shows to logged-out visitors. group_topic is the exception to know about: on 24 Sep 2026, every topic page we opened logged out (4 of 4, from a home connection, not yet re-checked through Apify's proxies) showed Douban's login wall instead of the topic. Such topics are skipped and not charged. Login-walled content (private groups, blocked users) is not in scope.
Why are movie comments not supported?
Douban serves movie short-comments through a JavaScript-only widget on the mobile site that requires headless browser execution. v1.0 returns long-form movie reviews instead, which contain richer opinion text and are the primary value for AI training data. Books and music short-comments work normally.
Can I scrape Douban in Python?
Yes. Install the Apify Python client (pip install apify-client), then call the zhorex/douban-scraper actor. See the Python code example above.
How much does it cost to scrape Douban?
Each record type has its own price (see the Pricing table). A typical research run extracting 100 movie reviews costs about $3. There is no monthly fee or minimum spend — pay only for what you extract. Diagnostic records (e.g. movie-comment-mode limitation notices) are never charged.
Is scraping Douban legal?
This scraper accesses publicly available content through Douban's public web endpoints. It does not bypass authentication and does not access private/locked content. Always review your local laws and Douban's terms of service before scraping.
What if a group topic URL fails?
Group topics are marked Beta in v1.0. Topic pages can answer with Douban's login wall (4 of 4 opened logged out on 24 Sep 2026), and some fail for other reasons (private group, moderated topic, deleted post). When a topic fails, the run logs a warning and continues with the next URL — you are not charged for failed topics.
What does this Douban scraper cover?
The Douban Scraper by Zhorex covers reviews, comments, group discussions, and search across movies, books, and music. Built specifically for Chinese AI training data buyers and sentiment research teams. Part of the Chinese Digital Intelligence Suite (Weibo, Bilibili, RedNote, Douban).
Integrations & data export
The Douban Scraper integrates with your existing workflow:
- Google Sheets — Send scraped reviews + ratings directly to a spreadsheet
- Zapier / Make / n8n — Automate workflows triggered by new Douban records
- REST API — Call the actor programmatically and retrieve data via Apify's REST API
- Webhooks — Get notified when a run finishes
- Data formats — Download as JSON, CSV, Excel, XML, or RSS
More scrapers by Zhorex
Chinese Digital Intelligence Suite
- 🆕 Chinese Brand Monitor — Cross-platform brand mention aggregator (all 5 platforms, sentiment + dedup)
- Weibo Scraper — China's Twitter (microblog, hot search)
- RedNote (Xiaohongshu) Scraper — China's Instagram + Pinterest
- Xueqiu Scraper — Chinese stock-discussion sentiment, cashtag indexing
- RedNote Shop Scraper — RedShop e-commerce
Reviews & ratings (cross-vertical)
- Letterboxd Scraper — Western film reviews and ratings
- G2 Reviews Scraper — B2B software reviews
- Capterra Reviews Scraper — Software product reviews
- Booking.com Reviews Scraper — Hotel reviews
Streaming & video
- Twitch Scraper — Twitch profiles, live streams, clips, VODs
- Kick Scraper — Kick.com profiles, streams, clips
- YouTube Shorts Scraper Pro — Shorts metadata + analytics
Markets & alt-data
- TradingView Scraper — Stocks, crypto, forex, indices
Other tools
- Perplexity AI Scraper — AI-powered search results
- Telegram Channel Scraper — Public Telegram channel messages
- Tech Stack Detector — Detect technologies used by websites
- LinkedIn Company Enrichment — Enrich company records
- Domain Authority Checker — Domain SEO metrics
- Phone Number Validator — Validate and format phone numbers
Support
Found a bug or want a new field? Open an issue on the Actor's Issues page.
💡 Used this Actor? Please leave a star rating, good or bad — it helps other users judge this tool.