# Douban Scraper: Movies, TV, Books, Ratings & Reviews (`sinodata/douban-scraper`) Actor

Scrape Douban (豆瓣) without login: movie, TV and book details with ratings, Top 250 and other charts, search, short reviews with star ratings, and full reviews.

- **URL**: https://apify.com/sinodata/douban-scraper.md
- **Developed by:** [SinoData](https://apify.com/sinodata) (community)
- **Categories:** Social media, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Douban Scraper: Movies, TV, Books, Ratings & Reviews (豆瓣)

Extract public data from **Douban** (豆瓣), China's largest community for rating and reviewing movies, TV shows and books, without logging in. No Chinese phone number, no cookies from you, no browser.

### What you can scrape

| Source | What you get |
|---|---|
| **Search** | Movies, TV shows or books for any keyword (Chinese or English) |
| **Douban URLs** | Full details: Douban score and number of ratings, genres, countries, languages, directors, cast, runtime, release dates, episodes, intro, cover. For books: authors, translators, publisher, publish date, pages, price |
| **Charts** | Top 250 movies, Top 250 books, in cinemas now, popular movies, weekly best, popular TV (Chinese, US, Korean, Japanese) |
| **Short reviews (短评)** | The most popular short reviews with the reviewer's stars (1 to 5), helpful votes, date and reviewer location |
| **Full reviews (影评/书评)** | Long-form reviews with title, full text, stars and "useful" count |

### Use cases

- **Film and publishing market research**: see how Chinese audiences rate a movie, show or book, and why.
- **Box office and release tracking**: monitor what's in cinemas and trending on Douban.
- **Sentiment analysis**: collect rated short reviews (text plus 1 to 5 stars) as labelled training data.
- **Recommendation and catalogue enrichment**: add Douban scores and metadata to your own catalogue.

### How to use

1. Add **search keywords**, **Douban URLs**, or pick one or more **charts**.
2. Turn on **Short reviews** and/or **Full reviews** if you want them.
3. Run, then download the results as JSON, CSV or Excel, or use the API.

Accepted inputs: `https://movie.douban.com/subject/1292052/`, `https://book.douban.com/subject/1007305/`, `https://m.douban.com/movie/subject/35588177/`, or a bare id like `1292052` (treated as movie or TV).

#### Example input

```json
{
  "charts": ["movie_top250"],
  "maxSubjectsPerSource": 50,
  "includeShortReviews": true,
  "maxShortReviewsPerSubject": 20
}
```

### Output

Every record has a `type` field: `subject` (a movie, TV show or book), `shortReview` or `review`. Use the views in the dataset tab, or filter by `type` when exporting.

#### Title

```json
{
  "type": "subject",
  "subjectType": "movie",
  "id": "1292052",
  "url": "https://movie.douban.com/subject/1292052/",
  "title": "肖申克的救赎",
  "originalTitle": "The Shawshank Redemption",
  "year": "1994",
  "rating": 9.7,
  "ratingCount": 3346721,
  "genres": ["剧情", "犯罪"],
  "countries": ["美国"],
  "directors": ["弗兰克·德拉邦特"],
  "actors": ["蒂姆·罗宾斯", "摩根·弗里曼"]
}
```

#### Short review

```json
{
  "type": "shortReview",
  "subjectTitle": "肖申克的救赎",
  "text": "...",
  "stars": 4,
  "votes": 5225,
  "createdAt": "2008-01-13T17:53:08.000Z",
  "userName": "..."
}
```

### Pricing

Pay only for results:

| Event | Price |
|---|---|
| Title (movie, TV show or book) | $0.003 |
| Short review | $0.001 |
| Full review | $0.002 |

The run stops cleanly when it reaches the maximum cost you set.

### Limits to know

- **Douban limits requests per IP address.** Without a proxy, the scraper keeps to about one request every 4.5 seconds and pauses for about 4 minutes if it still hits the limit, so big runs are slow. Turn on **Apify Proxy** in the Advanced section to switch IP instead and run much faster.
- Logged-out visitors see about the **top 180 short reviews** and the **top 20 full reviews** per title, sorted by popularity. Newest-first sorting needs a logged-in account and is not offered.
- Only public data visible to a logged-out visitor is collected.
- Respect Douban's terms and applicable law, and don't use personal data (names, reviews) for spam or harassment.

### Feedback

Found a bug or need a field that's missing? Open an issue on the Issues tab and it will usually be fixed within a day or two.

# Actor input Schema

## `searchQueries` (type: `array`):

Search Douban for movies, TV shows or books, e.g. 流浪地球, 三体, Oppenheimer.

## `searchType` (type: `string`):

Movies & TV, or books.

## `subjectUrls` (type: `array`):

Movie, TV or book pages (movie.douban.com/subject/<id>, book.douban.com/subject/<id>) or numeric ids (treated as movie/TV).

## `charts` (type: `array`):

Douban charts to scrape: movie\_top250, movie\_showing (in cinemas), movie\_hot\_gaia, movie\_weekly\_best, tv\_hot, tv\_domestic, tv\_american, tv\_korean, tv\_japanese, book\_top250.

## `maxSubjectsPerSource` (type: `integer`):

How many movies, shows or books to save for each search keyword and each chart.

## `fetchFullDetails` (type: `boolean`):

Load each title's page for the full record (cast, genres, countries, intro, rating count). Off = faster, with only the basics from search and charts.

## `includeShortReviews` (type: `boolean`):

Save the most popular short reviews with star rating and votes. Douban shows logged-out visitors about the top 180 per title.

## `maxShortReviewsPerSubject` (type: `integer`):

Short reviews to save for each title (Douban's logged-out limit is about 180).

## `includeReviews` (type: `boolean`):

Save long-form reviews. Douban shows logged-out visitors about the top 20 per title.

## `maxReviewsPerSubject` (type: `integer`):

Long-form reviews to save for each title.

## `fetchReviewFullText` (type: `boolean`):

Load the complete text of each long-form review, not just the abstract (one extra request per review).

## `proxyConfiguration` (type: `object`):

Douban limits requests per IP. Without a proxy the scraper pauses when it hits the limit; with Apify Proxy it switches IP instead, which is much faster for big runs.

## Actor input object example

```json
{
  "searchQueries": [
    "流浪地球"
  ],
  "searchType": "movie",
  "maxSubjectsPerSource": 20,
  "fetchFullDetails": true,
  "includeShortReviews": false,
  "maxShortReviewsPerSubject": 50,
  "includeReviews": false,
  "maxReviewsPerSubject": 10,
  "fetchReviewFullText": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `titles` (type: `string`):

No description

## `shortReviews` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "流浪地球"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("sinodata/douban-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["流浪地球"] }

# Run the Actor and wait for it to finish
run = client.actor("sinodata/douban-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "流浪地球"
  ]
}' |
apify call sinodata/douban-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sinodata/douban-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/YhyPsnYbWhEE0GjWE/builds/bqbofjQ9iZjkKwvEd/openapi.json
