# Bilibili Scraper: Videos, Search, Uploaders & Comments (`changefeeds/bilibili-scraper`) Actor

Search Bilibili, export video details and stats, an uploader's videos and profile, or a video's comments — via the public api.bilibili.com web API. No Bilibili login required.

- **URL**: https://apify.com/changefeeds/bilibili-scraper.md
- **Developed by:** [Changefeeds Tools](https://apify.com/changefeeds) (community)
- **Categories:** Social media, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 item returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bilibili Scraper: Videos, Search, Uploaders & Comments

A Bilibili scraper and Bilibili API client in one actor. Search Bilibili by
keyword, pull full stats for specific videos, export an uploader's video list
and profile, or grab a video's comments — all through the same public
`api.bilibili.com` web API a logged-out browser on bilibili.com uses. No
Bilibili account, cookie, or login required.

### Modes

#### `search` — videos matching a keyword

```json
{ "mode": "search", "keywords": ["minecraft"], "maxItems": 20, "order": "totalrank" }
```

Each keyword is searched separately and paginated until `maxItems` videos are
collected (or results run out). `order` is one of `totalrank` (comprehensive
ranking, default), `click` (most viewed), `pubdate` (newest), `dm` (most
danmaku) or `stow` (most favorited).

#### `videos` — details for specific videos

```json
{ "mode": "videos", "urls": ["BV1m5aY69E8D", "av117342986110880", "https://b23.tv/BV1m5aY69E8D"], "includeTags": true }
```

Accepts BV ids (`BV1xxxxxxxxx`), av ids (`av123456` or a bare number), full
`bilibili.com/video/...` URLs, and `b23.tv` short links (the redirect is
followed for you).

#### `user` — an uploader's videos, profile and follower count

```json
{ "mode": "user", "urls": ["686127", "https://space.bilibili.com/686127"], "maxItems": 50 }
```

Accepts a bare numeric `mid`, a `space.bilibili.com/<mid>` URL, or a `b23.tv`
short link. Emits the uploader's recent videos plus one profile row with
name, bio, level, follower/following counts and avatar.

#### `comments` — a video's comments

```json
{ "mode": "comments", "urls": ["BV1m5aY69E8D"], "maxComments": 100 }
```

Same video-reference formats as `videos` mode. See **Limits** below: in
practice this returns Bilibili's own small "hot comments" preview, not a full
comment thread — that's a limit of the logged-out API itself, not a bug.

### Sample output

A `video` row (live run, 2026-09-30):

```json
{
  "type": "video",
  "bvid": "BV1m5aY69E8D",
  "aid": 117342986110880,
  "url": "https://www.bilibili.com/video/BV1m5aY69E8D",
  "title": "十四年的等待，我的世界终于迎来全新第四维度：筛界 Minecraft Live2026",
  "description": "",
  "owner_name": "籽岷",
  "owner_mid": 686127,
  "published_at": "2026-09-27T12:29:08.000Z",
  "duration_seconds": 531,
  "views": 1194280,
  "likes": 71174,
  "coins": null,
  "favorites": 13774,
  "shares": null,
  "danmaku": 2267,
  "reply_count": 4684,
  "tags": ["筛界", "籽岷", "方块制片局", "新维度", "像素风", "沙盒游戏", "新版本", "我的世界", "LIVE", "MINECRAFT"],
  "cover_url": "https://i0.hdslb.com/bfs/archive/39e6cfdf143c830aaf3685bb0c2b15e6d51fb030.jpg",
  "category": "单机游戏",
  "scraped_at": "2026-09-30T04:27:15.499Z"
}
```

`coins` and `shares` are only available from `videos` mode (the full
`/view` endpoint); `search` and `user` mode results don't carry them, so
they're `null` there.

A `user` row:

```json
{
  "type": "user",
  "mid": 686127,
  "name": "籽岷",
  "sign": "十四年的等待...",
  "level": 6,
  "followers": 5454479,
  "following": 855,
  "videos_count": 30,
  "face_url": "https://i0.hdslb.com/bfs/face/7efb679569b2faeff38fa08f6f992fa1ada5e948.webp",
  "scraped_at": "2026-09-30T04:28:08.401Z"
}
```

A `comment` row:

```json
{
  "type": "comment",
  "video": "BV1m5aY69E8D",
  "rpid": 318812679792,
  "author_name": "空悟明理",
  "author_mid": 1759281120,
  "text": "岷叔这种新内容介绍真不用写这么公式的稿子...",
  "likes": 521,
  "reply_count": 7,
  "created_at": "2026-09-27T12:36:39.000Z"
}
```

A failed input gets a free `status` row instead, e.g.:

```json
{ "type": "status", "input": "686127", "mode": "user", "status": "risk_control", "error": "HTTP 412: request was banned (risk control)", "checked_at": "2026-09-30T04:28:08.401Z" }
```

`status` is one of:

- `not_found` — the keyword/id resolved to nothing.
- `invalid` — the input entry isn't a recognised BV id, av id, mid or URL.
- `risk_control` — Bilibili's anti-scraping check rejected the request (HTTP
  412, or a JSON `code` of -352/-412). See Limits.
- `error` — any other HTTP/network failure.

The key-value store record `OUTPUT` holds a per-run summary (counts, HTTP
stats, `stopped_reason`, `charged_events`). Every run leaves at least one
dataset row, even one where every input fails, so a quiet run is never
mistaken for a broken one.

### Pricing

Pay per event, nothing else: **$0.002 per item returned** ($2 per 1,000
video/uploader/comment rows). Failed inputs (`not_found`, `invalid`,
`risk_control`, `error`) are never charged. If you set a maximum total charge
for the run, the actor stops cleanly once it's reached and says so in
`OUTPUT.stopped_reason` (`"max_total_charge_reached"`).

### Limits, stated plainly

- **Public data only, no login.** This reads exactly what `api.bilibili.com`
  serves a logged-out browser. No private messages, no member-only content,
  no video/audio downloads.
- **Bilibili's risk control is real and IP-dependent.** The uploader
  video-listing endpoint (`/x/space/wbi/arc/search`) and the profile endpoint
  (`/x/space/wbi/acc/info`) can reject requests from some IPs (including some
  datacenter ranges) with HTTP 412 or a `-352` code, even with correct WBI
  signing. When that happens the actor reports a `risk_control` status row
  instead of crashing, and — for `user` mode — still emits whatever it could
  get (e.g. the follower count from `/x/relation/stat`, which is far more
  reliable, even when the profile or video list is blocked).
- **Comments are capped by Bilibili itself, not by this actor.** The
  logged-out `/x/v2/reply/main` endpoint returns only its small "hot
  comments" set (a handful of rows) for both its "hot" and "recent" modes and
  reports `is_end: true` immediately — full-thread pagination needs a logged
  in session, which this actor deliberately does not use. `maxComments`
  caps what you get; it can't make Bilibili return more than it's willing to.
- **View/like/coin counts are a snapshot.** They change between your run and
  the next one, same as on the site.
- **It stays polite:** at most 2 requests in flight, ~300 ms between
  requests, HTTP 429/5xx retried with backoff (Retry-After honoured, capped
  at 60 s, up to 3 retries), a realistic desktop Chrome User-Agent and
  Referer `https://www.bilibili.com` (what the site's own web client sends —
  this is not evasion, it's what the public API expects of a browser
  client).

### FAQ

**Do I need a Bilibili account or cookie?** No — this uses only the public,
logged-out `api.bilibili.com` endpoints.

**Can it download videos?** No, never. Only metadata (title, stats, tags,
cover URL).

**Why did I get a `risk_control` status instead of data?** Bilibili's
anti-scraping layer rejected that specific request from this run's IP. Retry
later, or lower request volume — the actor already backs off and won't burn
your run.

**Can I search in English?** Yes, `keywords` are passed straight to
Bilibili's own search; results follow whatever that returns for the term.

**How is this a "bilibili api" client?** Every mode is a thin export of a few
Bilibili web-API endpoints (search, view, space listing, comments) — this
actor exists so you don't have to reimplement WBI request signing yourself.

### Troubleshooting: Bilibili v\_voucher / -352 / 412 risk control

If you're calling Bilibili's logged-out web endpoints from a server and occasionally (or always) get a response with `code: 0` but no real payload — just a `v_voucher` field where your data should be — you've hit Bilibili's **risk control** system (风控), not a bug. The same family of errors shows up as HTTP 412 or a `-352` error code.

**Why it happens**

Bilibili doesn't publish an official third-party API. Community libraries reverse-engineer the same endpoints the website and apps use. Bilibili's risk-control layer inspects request volume, IP reputation, and whether the client presents cookies a real browser session would have. When suspicious, it doesn't return an error — it returns a `200 OK` with `code: 0` and a `v_voucher` token instead of your data. That token is meant for a captcha flow that a scraper can't complete.

**Free fixes**

- **Send visitor cookies before making API calls.** Bilibili's frontend sets `buvid3`, `buvid4` and `b_nut` tokens via `/x/frontend/finger/spi` before calling data endpoints:

```python
import requests

session = requests.Session()
resp = session.get("https://api.bilibili.com/x/frontend/finger/spi")
data = resp.json()["data"]
session.cookies.set("buvid3", data["b_3"])
session.cookies.set("buvid4", data["b_4"])
```

- **Set a real `Referer` and browser-like `User-Agent`.** A bare Python default user agent with no referer is an easy tell.
- **Throttle and randomize request timing.** Risk control weighs request velocity heavily; add delay and jitter between calls.
- **Rotate sessions across proxy IPs for volume.** A single IP making many requests gets flagged eventually.

Bilibili tightens its rules periodically, so whatever works today may need adjusting later.

**This actor handles it**

This actor manages visitor-token handshake, rotating residential proxy, and retries. Failed or risk-controlled inputs are not charged.

### Local development

```bash
pnpm --filter @mmnm/bilibili test        # unit tests, no network
pnpm --filter @mmnm/bilibili build
```

`node src/main.ts` runs the actor locally with Apify's local storage
(`./storage`); `ACTOR_TEST_PAY_PER_EVENT=true ACTOR_MAX_TOTAL_CHARGE_USD=1`
exercises the charging path with the SDK's $1 local test price.

# Actor input Schema

## `mode` (type: `string`):

What to scrape. "search": videos matching keywords. "videos": details for specific videos. "user": an uploader's videos + profile + follower count. "comments": top-level comments on a video.

## `keywords` (type: `array`):

Search terms. Each keyword is searched separately, newest-relevance-ranked results first (or by `order`), up to `maxItems` videos per keyword.

## `urls` (type: `array`):

For "videos"/"comments": BV ids (BV1xxxxxxxxx), av ids (av123456), full bilibili.com/video/... URLs, or b23.tv short links. For "user": a numeric mid, a space.bilibili.com/<mid> URL, or a b23.tv short link.

## `maxItems` (type: `integer`):

Max videos returned per keyword (search) or per uploader (user). 1 to 2,000.

## `order` (type: `string`):

Result ordering for search mode.

## `maxComments` (type: `integer`):

Max comments returned per video. 1 to 2,000. In practice Bilibili's logged-out comment API returns only its top handful of comments per video (see README Limits).

## `includeTags` (type: `boolean`):

Also fetch and include each video's tag list. One extra request per video in "videos" mode; free in "search" mode (tags are already in the search result).

## `proxyConfiguration` (type: `object`):

Bilibili's risk control often returns empty results or -352 errors to cloud-server IPs, so requests go through Apify Proxy by default (residential). Proxy traffic is included in the price. Turn it off only if you run the Actor from your own network.

## Actor input object example

```json
{
  "mode": "search",
  "keywords": [
    "minecraft"
  ],
  "maxItems": 20,
  "order": "totalrank",
  "maxComments": 100,
  "includeTags": false,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "minecraft"
    ],
    "maxItems": 20,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("changefeeds/bilibili-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["minecraft"],
    "maxItems": 20,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("changefeeds/bilibili-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "minecraft"
  ],
  "maxItems": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call changefeeds/bilibili-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,changefeeds/bilibili-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/NT6GyIieZ2E2APWcv/builds/IDkxxp8lEARGlzn0W/openapi.json
