Douban Pro Scraper — Reviews, Discussions & Subject Data avatar

Douban Pro Scraper — Reviews, Discussions & Subject Data

Pricing

from $30.00 / 1,000 review scrapeds

Go to Apify Store
Douban Pro Scraper — Reviews, Discussions & Subject Data

Douban Pro Scraper — Reviews, Discussions & Subject Data

Scrape long-form reviews of movies, books and music, short book and music comments with star ratings, and group discussion titles from Douban (豆瓣), China's reviews and interest community. Plus subject and group search. For Chinese-LLM training corpora, sentiment analysis and NLP. API, no login.

Pricing

from $30.00 / 1,000 review scrapeds

Rating

0.0

(0)

Developer

Sami

Sami

Maintained by Community

Actor stats

0

Bookmarked

249

Total users

129

Monthly active users

3 days ago

Last modified

Share

Douban Scraper — Reviews, Comments & Group Discussions

Extract long-form reviews, ratings, comments, and group discussions from Douban (豆瓣) — China's leading reviews + interest community. Movies, books, and music. No API key, no browser, no VPN. Built for Chinese AI training corpora and consumer research.

How to scrape Douban in 3 easy steps

  1. Go to the Douban Scraper page on Apify and click "Try for free"
  2. Configure your input — choose a mode (subject_reviews, subject_comments, group_topic, subject_search, or group_search), enter your Douban URLs or query, and set the number of results
  3. Click "Run", wait for the scraper to finish, then download your data in JSON, CSV, or Excel format

No coding required. No API key. Works with Apify's free plan.

Part of the Chinese Digital Intelligence Suite

Built by Zhorex, who maintains a suite of Chinese-platform scrapers:

  • 🆕 Chinese Brand Monitor — Cross-platform brand mention aggregator (Weibo + RedNote + Bilibili + Douban + Xueqiu in one normalized feed, sentiment + dedup, $0.06/mention)
  • Weibo Scraper — China's Twitter (microblogging, hot search, public opinion)
  • RedNote (Xiaohongshu) Scraper — China's Instagram + Pinterest
  • Xueqiu Scraper — Chinese stock-discussion sentiment, cashtag indexing
  • RedNote Shop Scraper — RedShop e-commerce data
  • Douban Scraper — You are here (reviews, ratings, group discussions)

Together, these cover the pillars of Chinese digital intelligence: microblogging, video, lifestyle, e-commerce, finance, and long-form reviews. For cross-platform brand monitoring (the most common multi-platform use case), the Chinese Brand Monitor aggregator orchestrates all 5 in one call with normalized output — saves 4-6 hours of engineering vs. wiring up individual scrapers.

What is Douban?

Douban (豆瓣) is China's reviews and interest-community platform — Goodreads + Letterboxd + Rate Your Music + niche-Reddit fused into one site, with 200M+ monthly users. It's where Chinese readers, cinephiles, music fans, and hobby communities post the longest-form opinion content on the Chinese internet. Movies, books, music, TV shows, and tens of thousands of user-run discussion groups.

For anyone building a Chinese-language LLM, sentiment classifier, or consumer research dataset, Douban is a rich source of opinion-heavy long-form Chinese text.

Modes

ModeWhat it doesRecords
subject_reviewsLong-form reviews (500-5,000+ Chinese chars each) for a movie/book/music albumOne per review
subject_commentsShort comments + star ratings under a subject's discussion pageOne per comment
subject_searchSearch Douban for movies / books / music by keywordOne per result
group_topic ⚠️ BetaPull a discussion thread + its replies from a Douban GroupOne per topic (with nested replies)
group_searchFind Douban Group discussions whose title contains your keyword (brand, product, person), newest firstOne per matching topic

v1.0 Known Limitations (read before buying)

  • Movie comments require browser rendering. Douban serves movie short-comments through a JS-only mobile widget that v1.0 cannot extract — subject_comments for movies returns a diagnostic record explaining the limitation. Use subject_reviews for movie data instead — long-form movie reviews are richer for AI training anyway. Books and music short-comments work normally.
  • Movie review list bodies are excerpt-only by default. Mobile movie list pages don't show the author's name, the publication date or the full body — only review IDs, titles, ratings and the author's avatar. authorUrl is built from the numeric user id in the avatar file name (https://www.douban.com/people/<id>/), and stays null for authors with Douban's default avatar. Set fetchFullReviewBody: true (default) to fetch each review's detail page and fill in the full markdown body; when that detail names the author and the date, the row also gets authorUsername, publishedAt and publishedAtIso, and authorUrl becomes the author's own profile link. Where it does not, those fields stay null — never guessed.
  • Book / music comment coverage varies by subject. Popular subjects (rating count > 10K — e.g. 三体, OK Computer) reliably serve inline short-comments; some less-popular books have begun moving comment lists to AJAX-only rendering and will return 0 records. If subject_comments returns 0 records for a URL, fall back to subject_reviews (which works on all subjects).
  • Movie search returns Douban's tag-matched feed. For precise targeting of a specific film, use subject_reviews with the movie's subject URL directly.
  • Book search caps at ~10 discovery results per query. Douban's book suggestion endpoint doesn't paginate. For bulk book review extraction, supply multiple subject URLs to subject_reviews mode.
  • group_topic mode is Beta. On 24 Sep 2026, every topic page we opened logged out (4 of 4, from a home connection, not yet re-checked through Apify's proxies) showed Douban's login wall instead of the topic, so expect many topics to fail; others return 404 (deleted). When a topic fails, the run logs a warning and continues — you are not charged for failed topics. To find discussions about a keyword without opening topic pages, use group_search.
  • Residential proxies are strongly recommended (default in input). Datacenter IPs degrade movie-mode and may trigger generic anti-bot challenges.

Use Cases

WhoWhy
AI / LLM training data buyersA rich source of Chinese long-form opinion text — key for Chinese-language model fine-tuning
Sentiment analysis researchersStar-rating-labelled Chinese review text, ideal for supervised sentiment classifiers
Brand monitoring teamsFind Chinese consumer reviews mentioning your product, competitor films, or book titles
Cultural trend analystsTrack which films / books / albums are gaining traction in Chinese-speaking markets
Academic NLP researchersPre-built corpus of opinion text with engagement metrics — citable in cross-cultural studies
Localization / translation teamsReal Chinese phrasing patterns for entertainment vocabulary

Scrape Douban with Python, JavaScript, or no code

You can use the Douban Scraper directly from the Apify Console (no code), or integrate it into your scripts.

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("zhorex/douban-scraper").call(run_input={
"mode": "subject_reviews",
"subjectUrls": ["https://book.douban.com/subject/1084336/"],
"maxResults": 50,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('zhorex/douban-scraper').call({
mode: 'subject_reviews',
subjectUrls: ['https://book.douban.com/subject/1084336/'],
maxResults: 50,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => console.log(item));

Input examples

1. Subject reviews (long-form)

Pull long-form reviews for one or more movies, books, or music albums. Provide subject URLs or numeric subject IDs.

{
"mode": "subject_reviews",
"subjectUrls": [
"https://movie.douban.com/subject/1292052/",
"https://book.douban.com/subject/1084336/",
"https://music.douban.com/subject/1419463/"
],
"maxResults": 100,
"fetchFullReviewBody": true
}

2. Subject comments + ratings

Short comments + star ratings for a book or music album. Movie comments are not supported in v1.0 — use subject_reviews mode for movies.

{
"mode": "subject_comments",
"subjectUrls": ["https://book.douban.com/subject/1084336/"],
"maxResults": 200
}

Search Douban for movies, books, or music by keyword. Returns Douban's discovery feed for that query.

💡 Search the Chinese title/name for far more results. Douban indexes subjects by their Chinese title (e.g. 三体, not "The Three-Body Problem"), so the native term returns far more than the Latin/English one. Few results usually means the query language, not an error.

{
"mode": "subject_search",
"searchQuery": "三体",
"searchType": "all",
"maxResults": 30
}

4. Group topic (Beta)

Pull one or more group discussion threads with embedded replies, when Douban serves the topic page. Logged out, it often shows its login wall instead (see the Beta note above); those topics return no row and are not charged.

{
"mode": "group_topic",
"topicUrls": ["https://www.douban.com/group/topic/319929381/"],
"maxRepliesPerTopic": 100
}

5. Group search — who is talking about a keyword in Douban Groups

Douban Groups (豆瓣小组) are where Chinese consumers argue about brands, products and celebrities. group_search searches them for your keyword and returns one row per discussion whose title contains the exact keyword, newest first. Put the keyword in Search query.

{
"mode": "group_search",
"searchQuery": "李宁",
"maxResults": 50
}

What to know before you run it:

  • Only real title mentions are returned. Douban's group search is fuzzy: searching 李宁 also lists titles that merely contain 李 and 宁 apart (宁艺卓, 李庚希) or do not contain the word at all. In a measured search only 11 of the 48 topics on the first page had 李宁 in the title. Every other row is dropped, never delivered and never charged, and the run's status says how many were dropped. Matching is an exact substring, case-insensitive for Latin letters (nike matches Nike and NIKE), so 李宁 also matches 李宁玉.
  • How deep it goes. Douban lists up to 3,000 results (60 pages of ~50) per keyword, but it does not let one connection page through them for long. In our tests on 24 Sep 2026, pages 1 and 2 opened and page 60 showed the login wall; after about 15 requests to Douban in an hour, the same connection was sent to Douban's anti-bot check for the next search pages. So one run reads at most 10 search pages (about 500 listed topics), a few seconds apart. It stops earlier when it has maxResults title mentions, when the list ends, after 3 pages in a row without a single title mention, or at the first page Douban refuses (login wall or anti-bot check). The status says which of these ended the run. In our test, 2 pages gave 25 title mentions of 李宁, so the default maxResults of 50 takes about 4 pages.
  • Titles only, no topic text. Rows carry what the search listing shows: title, group, reply count and time. They do not include the topic's body, author or replies: Douban showed its login wall instead of the topic page on every topic we opened logged out (4 of 4, 24 Sep 2026).
  • postedAt is when the topic was created, in Beijing time (+08:00). Across 97 measured results, the listed time went up exactly as the topic id went up, even for threads with 200+ replies, so it is not the time of the last reply.
  • Group names can be cut short. The search listing shortens long group names with .... groupId and groupUrl always identify the group exactly.
  • Search for the keyword the way Douban users write it: 李宁, not "Li-Ning". Latin brand names such as nike work as written.

Output examples

Review record

{
"type": "review",
"reviewId": "1000104",
"subjectId": "1084336",
"subjectName": "小王子",
"subjectType": "book",
"title": "长大就笨了",
"content": "(Full review body in markdown — Chinese long-form text)",
"rating": 5,
"ratingLabel": "力荐",
"authorUsername": "示例用户A",
"authorUrl": "https://www.douban.com/people/000000001/",
"authorAvatarUrl": "https://img3.doubanio.com/icon/u000000001-1.jpg",
"publishedAt": "2005-04-06 11:51:52",
"publishedAtIso": "2005-04-06T03:51:52Z",
"stats": { "replyCount": 444 },
"reviewUrl": "https://book.douban.com/review/1000104/",
"scrapedAt": "2026-05-13T01:39:22Z"
}

Comment record

{
"type": "comment",
"commentId": "10287387",
"subjectId": "1084336",
"subjectName": "小王子",
"subjectType": "book",
"content": "十几岁的时候渴慕着小王子,一天之间可以看四十四次日落。",
"rating": 5,
"ratingLabel": "力荐",
"authorUsername": "示例用户B",
"authorUrl": "https://www.douban.com/people/000000002/",
"publishedAt": "2007-02-08 11:16:40",
"publishedAtIso": "2007-02-08T03:16:40Z",
"stats": { "votesCount": 9232 },
"scrapedAt": "2026-05-13T01:39:22Z"
}

Subject (search result)

{
"type": "subject",
"subjectId": "2567698",
"subjectName": "三体",
"subjectType": "book",
"year": "2008",
"author": "刘慈欣",
"rating": null,
"cover": "https://img1.doubanio.com/view/subject/s/public/s2768378.jpg",
"subjectUrl": "https://book.douban.com/subject/2567698/",
"scrapedAt": "2026-05-13T01:39:22Z"
}

Group topic record (Beta)

{
"type": "group_topic",
"topicId": "319929381",
"groupName": "(Group name)",
"title": "(Discussion title)",
"content": "(Topic body in markdown — Chinese long-form text)",
"authorUsername": "(Author handle)",
"publishedAt": "2026-04-01 10:20:30",
"publishedAtIso": "2026-04-01T02:20:30Z",
"stats": { "repliesCaptured": 50 }, // how many this run captured,
// capped by maxRepliesPerTopic — not
// the thread's total reply count
"replies": [
{
"replyId": "...",
"authorUsername": "...",
"content": "...",
"publishedAt": "...",
"votesCount": 12
}
],
"topicUrl": "https://www.douban.com/group/topic/319929381/",
"scrapedAt": "2026-05-13T01:39:22Z"
}

Group search record

{
"type": "group_search_result",
"topicId": "500559514",
"title": "你们觉得李宁特步名气大还是ask韩潮名气大",
"url": "https://www.douban.com/group/topic/500559514/",
"groupId": "757619",
"groupName": "长江国际售楼部",
"groupUrl": "https://www.douban.com/group/757619/",
"replyCount": 52,
"postedAt": "2026-09-22T11:40:05+08:00",
"keyword": "李宁",
"mode": "group_search",
"scrapedAt": "2026-09-24T17:19:43Z"
}

replyCount is the count the search listing shows (52回复). If the listing shows none, it is null, never a made-up 0. url is the clean topic link, without Douban's tracking parameter.

Pricing

Pay per result — no monthly fee, no minimum.

EventPriceWhen charged
review-scraped$0.030Per long-form review record extracted
comment-scraped$0.005Per short comment record extracted
group-topic-scraped$0.030Per group topic Douban served (with embedded replies); login-walled topics are not charged
subject-search-result$0.005Per search result row — subject_search results and group_search title mentions

Concrete cost examples:

  • 100 long-form reviews of one popular movie's reviews page: $3.00
  • 1,000 short comments across multiple books: $5.00
  • 200 search results to seed a crawl: $1.00
  • 50 group topics whose title mentions a brand (group_search): $0.25. Fuzzy matches Douban lists but whose title does not contain the keyword are free, because they are never delivered.

Diagnostic / log records (e.g. movie comment limitation notices) are NEVER charged.

Content is in Chinese

All content is returned in the original Simplified Chinese. Douban is a Chinese-language platform — reviews, comments, group discussions, and user names are in Chinese.

If you need English translations, pipe the output through a translation API (Google Translate, DeepL, or Claude).

Technical Details

  • No browser — pure HTTP, runs in 512MB RAM
  • No login required — works against publicly accessible content only
  • Built-in rate limiting — exponential backoff on 429 / 503
  • Globally accessible — residential proxy recommended (default in input)
  • UTF-8 throughout — Chinese text round-trips cleanly
  • Markdown review bodies — <p>, <a>, <strong> etc. converted to lightweight markdown for downstream LLM ingestion
  • deltaMode memory — 200,000 rows per stream (about 192,000 for short comments) (mode + targets, or your deltaStateKey). A row Douban lists again counts as recent; when the memory is full, the rows Douban has gone longest without listing are forgotten first, and one listed again after that is delivered and charged again.

FAQ

Is there a Douban API?

Douban's official developer API has been deprecated for several years. There is no working public Douban API for international developers. This Douban Scraper extracts reviews, comments, ratings, and group discussions from publicly accessible web endpoints.

Do I need a Douban login or cookies?

No. The Actor never logs in and takes no cookies. subject_reviews, subject_comments, subject_search and group_search read pages Douban shows to logged-out visitors. group_topic is the exception to know about: on 24 Sep 2026, every topic page we opened logged out (4 of 4, from a home connection, not yet re-checked through Apify's proxies) showed Douban's login wall instead of the topic. Such topics are skipped and not charged. Login-walled content (private groups, blocked users) is not in scope.

Why are movie comments not supported?

Douban serves movie short-comments through a JavaScript-only widget on the mobile site that requires headless browser execution. v1.0 returns long-form movie reviews instead, which contain richer opinion text and are the primary value for AI training data. Books and music short-comments work normally.

Can I scrape Douban in Python?

Yes. Install the Apify Python client (pip install apify-client), then call the zhorex/douban-scraper actor. See the Python code example above.

How much does it cost to scrape Douban?

Each record type has its own price (see the Pricing table). A typical research run extracting 100 movie reviews costs about $3. There is no monthly fee or minimum spend — pay only for what you extract. Diagnostic records (e.g. movie-comment-mode limitation notices) are never charged.

This scraper accesses publicly available content through Douban's public web endpoints. It does not bypass authentication and does not access private/locked content. Always review your local laws and Douban's terms of service before scraping.

What if a group topic URL fails?

Group topics are marked Beta in v1.0. Topic pages can answer with Douban's login wall (4 of 4 opened logged out on 24 Sep 2026), and some fail for other reasons (private group, moderated topic, deleted post). When a topic fails, the run logs a warning and continues with the next URL — you are not charged for failed topics.

What does this Douban scraper cover?

The Douban Scraper by Zhorex covers reviews, comments, group discussions, and search across movies, books, and music. Built specifically for Chinese AI training data buyers and sentiment research teams. Part of the Chinese Digital Intelligence Suite (Weibo, Bilibili, RedNote, Douban).

Integrations & data export

The Douban Scraper integrates with your existing workflow:

  • Google Sheets — Send scraped reviews + ratings directly to a spreadsheet
  • Zapier / Make / n8n — Automate workflows triggered by new Douban records
  • REST API — Call the actor programmatically and retrieve data via Apify's REST API
  • Webhooks — Get notified when a run finishes
  • Data formats — Download as JSON, CSV, Excel, XML, or RSS

More scrapers by Zhorex

Chinese Digital Intelligence Suite

Reviews & ratings (cross-vertical)

Streaming & video

Markets & alt-data

Other tools

Support

Found a bug or want a new field? Open an issue on the Actor's Issues page.


💡 Used this Actor? Please leave a star rating, good or bad — it helps other users judge this tool.