# Bilibili Scraper: Video Search, Comments & Danmaku (B站) (`datagleaner/bilibili-scraper`) Actor

Bilibili scraper for video search by keyword, full video details, comments with IP location and danmaku (bullet comments). Views, likes, coins, tags and uploader per video. Plain HTTP, no browser, no login. $8 per 1,000 videos, $2 per 1,000 comments.

- **URL**: https://apify.com/datagleaner/bilibili-scraper.md
- **Developed by:** [Data Gleaner](https://apify.com/datagleaner) (community)
- **Categories:** Social media, Videos, AI
- **Stats:** 2 total users, 1 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bilibili Scraper: videos, comments and danmaku (B站)

Bilibili scraper for video search by keyword, video stats, comments with IP location and danmaku (bullet comments), exported to JSON, CSV or Excel. **$8 per 1,000 videos, against about $20 elsewhere**, with comment IP location, raw danmaku with timestamps, and no login or browser.

It works on **Bilibili (哔哩哔哩, bilibili.com)**, China's largest video platform. Search videos by keyword, get full details (views, likes, coins, favorites, tags, uploader) for specific videos by link or BV id, and scrape comments, replies and danmaku. Plain HTTP against Bilibili's own web API, with the request signing handled for you. No cookies needed.

### What you can do with it

- **Creator and brand research:** find who is making videos about your product or niche, and how well they perform (views, likes, coins, favorites, shares).
- **China social listening:** read what Chinese viewers actually say in comments, including the city or province they post from.
- **Trend tracking:** search by newest or by most views, and watch which topics and tags take off.
- **LLM and research datasets:** titles, descriptions, tags and comment text in Chinese, with timestamps, ready for analysis.

### Input

| Field | Meaning | Default |
|---|---|---|
| `searchKeywords` | Keywords to search (Chinese works best) | none |
| `searchOrder` | `totalrank` relevance, `click` most views, `pubdate` newest, `dm` most danmaku, `stow` most favorites, `scores` most comments | `totalrank` |
| `durationFilter` | Search only: `any`, `under10`, `10to30`, `30to60`, `over60` (minutes) | `any` |
| `publishedAfter`, `publishedBefore` | Search only: keep videos published in a range. Accepts `2026-01-31` (a whole day in China time), an ISO datetime, or a relative age such as `7 days` for a rolling window on scheduled runs | none |
| `trendingCategory` | `popular` also returns Bilibili's site-wide trending feed (综合热门, a few hundred videos), no keyword needed, with coins and shares included (`tags` stay empty unless `enrichSearchResults` is on). Works alone or with keywords and URLs | `none` |
| `maxItems` | Max videos per keyword or trending list (search ends near 1,000) | 30 |
| `enrichSearchResults` | Fetch each search hit's full record (adds coins, shares, exact category, full tags) at 2 extra requests per video | off |
| `videoUrls` | Video links, BV ids (`BV1xx411c7mD`) or av ids (`av170001`) | none |
| `includeComments` | Also scrape comments for every video | off |
| `maxCommentsPerVideo` | Top-level comments per video | 20 |
| `commentSort` | `hot` or `newest` | `hot` |
| `includeReplies`, `maxRepliesPerComment` | Also scrape replies under each comment | off, 10 |
| `includeDanmaku` | Also scrape danmaku (bullet comments, 弹幕) for every video | off |
| `maxDanmakuPerVideo` | Danmaku per video, across all its parts | 1000 |
| `requestDelaySecs` | Pause between requests, with random jitter | 1.5 |
| `sessionCookie` | Optional `SESSDATA` cookie of your own Bilibili account, for full comment depth (see Limits) | none |
| `proxyConfiguration` | Optional proxy | none |

Provide at least one of `searchKeywords`, `videoUrls` or `trendingCategory`. With no input at all, the Actor runs a built-in 5-video example for `python 教程`.

Search and comments for 50 videos:

```json
{
  "searchKeywords": ["python 教程"],
  "searchOrder": "click",
  "maxItems": 50,
  "includeComments": true,
  "maxCommentsPerVideo": 40
}
```

Trending videos, no keyword:

```json
{ "trendingCategory": "popular", "maxItems": 50 }
```

Videos from the last 7 days that run 10 to 30 minutes:

```json
{ "searchKeywords": ["露营装备"], "publishedAfter": "7 days", "durationFilter": "10to30" }
```

Specific videos with full details:

```json
{ "videoUrls": ["https://www.bilibili.com/video/BV1xx411c7mD", "BV1rpWjevEip"] }
```

### Output

One dataset, three item types, told apart by `type`. The **Videos**, **Comments** and **Danmaku** views in the Console show each as a table.

Video:

```json
{
  "type": "video",
  "bvid": "BV1rpWjevEip",
  "aid": 113006243481679,
  "url": "https://www.bilibili.com/video/BV1rpWjevEip",
  "title": "【全748集】目前B站最全最细的Python零基础全套教程",
  "description": "本套教程从零开始讲解……",
  "tags": ["程序员", "Python教程", "Python入门零基础"],
  "category": "计算机技术",
  "durationSec": 143894,
  "publishedAt": "2024-08-22T14:59:18Z",
  "stats": { "views": 20516496, "danmaku": 132072, "comments": 356871, "likes": 563224, "coins": null, "favorites": 866293, "shares": null },
  "owner": { "mid": 3546597933714079, "name": "Python官方课程", "face": "https://i1.hdslb.com/bfs/face/6822e4fdc8f8.jpg" },
  "thumbnailUrl": "https://i2.hdslb.com/bfs/archive/a979056b1a32.jpg",
  "publishedTimestamp": 1724338758,
  "source": "search",
  "sourceQuery": "python 教程",
  "scrapedAt": "2026-10-08T12:00:00Z"
}
```

Search results do not carry `coins`, `shares` or the exact category, so those are `null` (or the search category name) unless you turn on `enrichSearchResults` or pass the video through `videoUrls`.

Comment:

```json
{
  "type": "comment",
  "rpid": 123456789,
  "videoBvid": "BV1xx411c7mD",
  "videoAid": 2,
  "url": "https://www.bilibili.com/video/BV1xx411c7mD#reply123456789",
  "text": "经典老视频",
  "likes": 88,
  "createdAt": "2023-11-14T22:13:20Z",
  "replyCount": 3,
  "parentRpid": null,
  "author": { "mid": 4242, "name": "测试用户", "level": 5 },
  "ipLocation": "上海",
  "createdTimestamp": 1699999999,
  "scrapedAt": "2026-10-08T12:00:00Z"
}
```

`parentRpid` is set on replies. `ipLocation` is the province or country Bilibili shows next to the comment, when present.

Danmaku (one item per bullet comment):

```json
{
  "type": "danmaku",
  "bvid": "BV1xx411c7mD",
  "cid": 62131,
  "part": 1,
  "videoTitle": "字幕君交流场所",
  "text": "经典老视频",
  "timeInVideoSeconds": 12.345,
  "mode": 1,
  "fontSize": 25,
  "color": "#ffffff",
  "sentAt": "2009-09-09T01:10:00Z",
  "senderHash": "a1b2c3d4",
  "danmakuId": "1001",
  "scrapedAt": "2026-10-08T12:00:00Z"
}
```

`mode` is the display type (1 scrolling, 4 bottom, 5 top, 6 reverse, 7 positioned, 8 code). `senderHash` is Bilibili's hashed sender id, not a user id.

#### Output fields

| Field | Rows | Meaning |
|---|---|---|
| `type` | all | `video`, `comment` or `danmaku` |
| `bvid`, `aid`, `url` | video | BV id, numeric av id, link |
| `title`, `description`, `tags`, `category` | video | Text fields; category is Bilibili's 分区 |
| `durationSec` | video | Length in seconds |
| `publishedAt`, `publishedTimestamp` | video | Publish time, ISO UTC and Unix seconds |
| `stats` | video | `views`, `danmaku`, `comments`, `likes`, `coins`, `favorites`, `shares` |
| `owner` | video | Uploader `mid`, `name`, `face` |
| `thumbnailUrl` | video | Cover image |
| `source`, `sourceQuery` | video | Which input produced it: `search`, `trending` or `videoUrl`, and the keyword, category or URL |
| `rpid`, `videoBvid`, `videoAid`, `parentRpid` | comment | Comment id, its video, and the parent comment for replies |
| `text`, `likes`, `replyCount` | comment | Content and counters |
| `createdAt`, `createdTimestamp` | comment | Comment time, ISO UTC and Unix seconds |
| `author` | comment | `mid`, `name`, account `level` (0-6) |
| `ipLocation` | comment | Commenter's IP location (IP属地) |
| `cid`, `part`, `videoTitle`, `timeInVideoSeconds`, `mode`, `fontSize`, `color`, `sentAt`, `senderHash`, `danmakuId` | danmaku | Raw bullet comment with its playback time |
| `scrapedAt` | all | When the Actor collected the item, ISO UTC |

### How much does it cost to scrape Bilibili? (pay per event)

| Event | Price |
|---|---|
| `video` | **$8 per 1,000 videos** ($0.008 each) |
| `comment` | **$2 per 1,000 comments** ($0.002 each, replies included) |
| `danmaku` | **$0.50 per 1,000 danmaku** ($0.0005 each) |

You pay only for items actually delivered. The Actor respects the maximum cost you set for a run and stops cleanly when it is reached.

Example: 100 videos with 40 comments each (logged-in comment depth) is 100 x $0.008 + 4,000 x $0.002 = **$8.80**. Without a `sessionCookie`, expect about 4 comments per video, so the same run costs about $1.60.

Example with danmaku: 20 videos with up to 1,000 danmaku each is 20 x $0.008 + 20,000 x $0.0005 = **$10.16**.

### Use with Python

```python
## pip install apify-client
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("datagleaner/bilibili-scraper").call(
    run_input={"searchKeywords": ["python 教程"], "searchOrder": "click", "maxItems": 5}
)
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["title"], item["stats"]["views"], item["url"])
```

Five videos cost $0.04.

### Use with JavaScript / Node.js

```js
// npm install apify-client
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('datagleaner/bilibili-scraper').call({
    searchKeywords: ['python 教程'],
    searchOrder: 'click',
    maxItems: 5,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) console.log(item.title, item.stats.views, item.url);
```

### Use it from n8n, Make, Zapier or an AI agent

Actor ID: `datagleaner/bilibili-scraper`

Minimal input:

```
{
  "searchKeywords": ["python 教程"],
  "maxItems": 5
}
```

Each tool below runs this Actor with your own Apify API token.

- **n8n:** add the **Apify** node (`@apify/n8n-nodes-apify`). On n8n Cloud you install it from the community node registry. Choose **Run an Actor and get dataset**, set Actor to `datagleaner/bilibili-scraper` and paste the input above.

- **Make:** use the Apify app's **Run an Actor** module, then **Get Dataset Items** to read the results. **Watch Actor Runs** can trigger a scenario when a run finishes.

- **Zapier:** use the Apify action **Run Actor**, then the search **Fetch dataset items**. The trigger **Finished Actor run** starts a Zap when a run ends.

- **AI agents (MCP):** connect to `https://mcp.apify.com?tools=datagleaner/bilibili-scraper`. In Claude Code:

  ```
  claude mcp add --transport http apify "https://mcp.apify.com?tools=datagleaner/bilibili-scraper"
  ```

  Then run `/mcp` to sign in to Apify in your browser, or send the header `Authorization: Bearer YOUR_APIFY_TOKEN`. The agent gets this Actor as a ready tool (`datagleaner--bilibili-scraper`); without preloading it can still find it with the `search-actors` tool and run it with `call-actor`. Clients that run local MCP servers (Claude Desktop's `mcpServers` config) can use Apify's package `@apify/actors-mcp-server`, run with `npx -y` and `APIFY_TOKEN` set. Then ask in plain words, for example:

  > Find the 20 most-viewed Bilibili videos for "露营装备", with 10 hot comments each, and summarize which brands viewers mention and which provinces they comment from.

- **LangChain (Python):**

```python
## pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/bilibili-scraper")
result = tool.invoke({"run_input": json.loads('{"searchKeywords": ["python 教程"], "maxItems": 5}')})
```

### Limits

- Public data only. Nothing that needs a Bilibili login (follower lists, private videos, watch history) is available.
- **Comment depth without login is shallow.** Bilibili serves anonymous visitors only about 3 to 4 curated top comments per video. Replies under a comment page normally. For full comment lists, paste the `SESSDATA` cookie of a Bilibili account you own into `sessionCookie` (use a throwaway account; this path has not been verified by us yet).
- Search returns at most about 1,000 results per keyword. Bilibili's comment list is also finite for very large threads.
- Bilibili rate-limits aggressively, sometimes with a silent empty answer instead of an error (the Actor detects this). The Actor paces requests, renews its session and retries with backoff, but a very large run from a datacenter IP can still be blocked. If that happens, lower the speed (raise `requestDelaySecs`) or add a proxy.
- If Bilibili blocks a run completely (no data collected), it ends without charge and with a status message asking you to retry in a few minutes or turn on `proxyConfiguration`. If the block comes mid-run, everything already collected is kept.
- Videos that were removed or are region-locked are skipped with a log line.
- **Not included: a user's video list.** Bilibili's uploader-page endpoint is behind stricter risk control and failed in testing, so it is left out rather than shipped unreliable.
- Short `b23.tv` links are not resolved; use the full URL or the BV id.

### FAQ

**Is there a Bilibili API for Python?** Bilibili has no documented public API for this data. This Actor calls the same web endpoints the website uses, handles the request signing, and returns structured JSON you can fetch from Python or Node.js with `apify-client` (see above).

**Can I use it as a Bilibili comment scraper without an account?** Yes: turn on `includeComments` and give keywords or video links. Each comment has its text, likes, time, reply count and the commenter's IP location. Without an account Bilibili shows only about 3 to 4 top comments per video; for full lists see `sessionCookie` under Limits.

**How do I scrape Bilibili videos by keyword?** Put the keywords in `searchKeywords`, choose `searchOrder` (relevance, most views, newest, most danmaku, most favorites, most comments) and set `maxItems`. Search results lack `coins` and `shares`; turn on `enrichSearchResults` or pass the videos through `videoUrls` to get them.

**How many danmaku do I get per video?** The public XML returns the most recent danmaku pool, typically up to a few thousand per part, not the full history of a very popular video. Multi-part videos are read part by part until `maxDanmakuPerVideo` is reached.

**Can I download the videos themselves?** No. It collects metadata, comments and danmaku, not video files.

### Responsible use

Only public data is collected, at a polite pace. You are responsible for using it lawfully: respect Bilibili's terms of service, copyright on video content, and privacy law (comment authors are real people; do not use the data to profile or harass them). Do not use this Actor for spam or to circumvent access controls.

### 中文简介

抓取哔哩哔哩（Bilibili）公开数据：按关键词搜索视频、按 BV 号或链接获取完整视频信息（标题、简介、标签、分区、时长、播放/弹幕/评论/点赞/投币/收藏/分享数、UP 主）、抓取评论（含 IP 属地）与回复、弹幕（公开 XML，通常为最近的数千条）。无需浏览器、无需登录，已内置 wbi 签名与限流重试。按事件计费：视频 $8 / 1000 条，评论 $2 / 1000 条，弹幕 $0.5 / 1000 条。暂不支持 UP 主视频列表。

# Actor input Schema

## `searchKeywords` (type: `array`):

Keywords to search for videos (Chinese works best, e.g. `python 教程`). Each keyword returns up to Max videos per keyword.

## `searchOrder` (type: `string`):

How Bilibili ranks search results.

## `durationFilter` (type: `string`):

Only return search results of this length (Bilibili's own duration filter).

## `publishedAfter` (type: `string`):

Only return search results published on or after this date: `2026-01-31` (a whole day in China time, as on Bilibili), an ISO datetime, or a relative age such as `7 days` for a rolling window on scheduled runs.

## `publishedBefore` (type: `string`):

Only return search results published on or before this date (same formats as Published after).

## `maxItems` (type: `integer`):

Search caps out around 1,000 results per keyword (50 pages of 20).

## `trendingCategory` (type: `string`):

Also return what is trending on Bilibili now, no keyword needed (the site-wide trending feed, 综合热门, a few hundred videos). Returns up to Max videos per keyword, with full stats (coins and shares included). Works alone or with keywords and URLs.

## `enrichSearchResults` (type: `boolean`):

Search results lack coins, shares and an accurate category and tag list. Turn on to fetch each video's full record (2 extra requests per video, slower).

## `videoUrls` (type: `array`):

Bilibili video links, BV ids (`BV1xx411c7mD`) or av ids (`av170001`). Returns full details for each.

## `userIds` (type: `array`):

Bilibili creators as a UID (`546195`) or a space link (`https://space.bilibili.com/546195`). Returns each creator's latest uploads, newest first, as video rows (billed as videos). Bilibili guards this list closely: if it refuses from the platform's IPs, the run says so, charges nothing for it and carries on with the other sources.

## `maxVideosPerUser` (type: `integer`):

Latest uploads returned per creator in Creator UIDs.

## `includeComments` (type: `boolean`):

Also scrape comments for every video returned.

## `maxCommentsPerVideo` (type: `integer`):

Top-level comments per video.

## `commentSort` (type: `string`):

Order of top-level comments.

## `includeReplies` (type: `boolean`):

Also scrape replies under each comment (one extra request per commented thread).

## `maxRepliesPerComment` (type: `integer`):

Replies fetched under each comment when Include replies is on.

## `includeDanmaku` (type: `boolean`):

Also scrape danmaku (弹幕) for every video returned: one extra request per video part. The public XML holds the most recent danmaku pool, typically up to a few thousand per part.

## `maxDanmakuPerVideo` (type: `integer`):

Total danmaku per video across all its parts.

## `requestDelaySecs` (type: `number`):

Politeness pacing with random jitter. Lower is faster but triggers Bilibili's risk control sooner.

## `proxyConfiguration` (type: `object`):

Optional. Default is no proxy. Use Apify Proxy (a China or Hong Kong residential group works best) if Bilibili blocks your run.

## `sessionCookie` (type: `string`):

Without login Bilibili serves only about 3 curated top comments per video. To get full comment lists, paste the SESSDATA cookie of a Bilibili account you own. Use a throwaway account: heavy scraping while logged in can get it restricted.

## Actor input object example

```json
{
  "searchKeywords": [
    "python 教程"
  ],
  "searchOrder": "totalrank",
  "durationFilter": "any",
  "maxItems": 30,
  "trendingCategory": "none",
  "enrichSearchResults": false,
  "videoUrls": [],
  "userIds": [],
  "maxVideosPerUser": 30,
  "includeComments": false,
  "maxCommentsPerVideo": 20,
  "commentSort": "hot",
  "includeReplies": false,
  "maxRepliesPerComment": 10,
  "includeDanmaku": false,
  "maxDanmakuPerVideo": 1000,
  "requestDelaySecs": 1.5
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchKeywords": [
        "python 教程"
    ],
    "videoUrls": [],
    "userIds": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagleaner/bilibili-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchKeywords": ["python 教程"],
    "videoUrls": [],
    "userIds": [],
}

# Run the Actor and wait for it to finish
run = client.actor("datagleaner/bilibili-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchKeywords": [
    "python 教程"
  ],
  "videoUrls": [],
  "userIds": []
}' |
apify call datagleaner/bilibili-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagleaner/bilibili-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yGkffe5MU9idBd0Fu/builds/U0qWtsy2OcMmSsp4Q/openapi.json
