Bilibili Transcript
Pricing
from $0.3483 / transcript
Bilibili Transcript
Bilibili Transcript processes one public video part, keeping the selected part's speech and metadata together for multilingual research and media analysis. Results include detected language, ordered timestamps, and optional translation into 133 languages. Transcript pricing opens at $0.3483.
Pricing
from $0.3483 / transcript
Rating
0.0
(0)
Developer
AgentX
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
3 days ago
Last modified
Categories
Share
Bilibili Transcript - Timestamped Bilibili Speech-to-Text API
Bilibili Transcript converts one public Bilibili video part into detected-language text, ordered timestamped segments, source metadata, and an optional translated transcript. It processes audible speech from the media; it does not promise to return Bilibili caption tracks or existing subtitle files.
- Supply one public Bilibili video-part URL;
?p=selects the part and the Actor never expands an entire multipart work. - Short links and redirects are resolved before the platform gate confirms the Bilibili extractor.
- One successful part produces one 22-field Dataset record containing source context and generated speech segments.
- The output excludes Bilibili subtitles, danmaku, comments, OCR, and the source media file.
Run Bilibili Transcript · Open the API page
For the first run, choose a public part with clear dialogue and verify that the selected p value matches the intended episode before automating more items.
Why Choose This API
A narrow contract for one known Bilibili video
Use this Actor after identifying an exact Bilibili BV page and, for multipart media, the desired part. The contract stays narrow: one URL plus one optional translation target. Platform confirmation uses resolved extractor information, so string appearance alone is not treated as proof of support.
The result combines source context and spoken content in one record:
| Capability | What is returned |
|---|---|
| Audio transcription | Detected language, joined text, and ordered start / end / text ranges for the selected part |
| BV-part context | Title, description, uploader identifiers, duration, publish time, thumbnail, categories, and tags when available |
| Source counters | Current views, likes, shares, comments, and other supported nullable metrics |
| Optional translation | A complete target-language version aligned to the original segment intervals |
This is not a creator-space crawler, bangumi or series expander, live recorder, subtitle/danmaku exporter, speaker-labeling tool, summarizer, or downloader. When Bilibili omits a source value, the Dataset preserves its nullable, zero, or empty form rather than inferring it.
Recognition operates on the chosen part's audio. Mandarin dialects, mixed Chinese and English, names, background music, overlapping voices, and specialist vocabulary can affect accuracy. Check consequential wording against the selected part.
Quick Start Guide
Run in Apify Console
- Open Bilibili Transcript on Apify.
- Open the required Bilibili part and copy the URL including its
?p=value when present. - Paste it into Video URL and optionally choose one Translate language.
- Start the run and inspect the one Dataset record after transcription completes.
The current schema pre-fills a BV URL selecting part 1:
{"video_url": "https://www.bilibili.com/video/BV1o84y1V7EY/?p=1","translate": "spanish"}
The integration and response examples reuse that URL for a consistent scenario. A source can later be deleted, region-limited, made login-only, or reorganized into different parts. Replace an unavailable example with another public Bilibili part containing audible speech.
What success means
Success requires playable media for the requested part, a usable audio stream, detected speech, and a published result. Reading BV metadata without speech output does not satisfy the transcript contract.
Input Parameters
Input configuration
| Input | Type | Required | Description |
|---|---|---|---|
video_url | string | Yes | One public Bilibili video URL to download and transcribe. |
translate | string | No | One target language from the schema's 133 selectable values. Leave empty to skip translation. |
The input processes one resolved media item. Creator spaces, searches, collections, series expansion, private or membership-only pages, uploads, cookies, credentials, and URL arrays are excluded. Playlist processing is disabled.
Correct BV syntax is only the first condition. The selected part must still be publicly accessible through the current extractor and expose audio; regional policy, deletion, authentication, or platform response changes can prevent preparation.
After source speech succeeds, translation can process segments concurrently and restores their original order. A single unrecoverable segment causes the whole optional translation to be omitted without its event charge.
Output Data Schema
One 22-field Dataset item
The selected part maps to one Dataset item with 22 fields:
| Group | Fields | Notes |
|---|---|---|
| Processing | processor, processed_at | Actor URL and processing timestamp |
| Identity | platform, title, description, thumbnail, published_at | Source-provided video identity |
| Author | author, author_id, author_url | Channel or uploader context when available |
| Media | duration, audio_title, audio_artist | Source-reported values; duration can be 0 when metadata is absent |
| Engagement | view_count, like_count, shares_count, dislike_count, comment_count | Nullable source metrics |
| Labels | categories, tags | Source classifications when exposed |
| Speech | transcript, translation | Timestamped source transcript and optional translated version |
Abbreviated schema illustration for the part-1-and-Spanish example:
{"processor": "https://apify.com/agentx/bilibili-transcript?fpr=aiagentapi","processed_at": "2026-07-21T13:50:00+00:00","platform": "BiliBili","title": "English Speech - All About Me","author": "嘀嗒英语Kidstalk","duration": 20,"view_count": 120128,"like_count": 1769,"categories": [],"tags": ["英文", "演讲", "英语口语", "儿童英语"],"transcript": {"language": "English","text": "Everyone, nice to meet you. Let me introduce myself.","segments": [{"start": "00:00:00.830","end": "00:00:10.660","text": "Everyone, nice to meet you. Let me introduce myself."}]},"translation": {"language": "Spanish","text": "Hola a todos, encantada de conocerlos. Permítanme presentarme.","segments": [{"start": "00:00:00.830","end": "00:00:10.660","text": "Hola a todos, encantada de conocerlos. Permítanme presentarme."}]}}
Displayed counts and text are illustrative snapshots. Actual metadata can change, and the Dataset does not add video/audio files, Bilibili subtitle tracks, danmaku, SRT/VTT, speaker labels, word timing, confidence, summaries, or legal verification.
Integration Examples
REST API
The synchronous endpoint returns Dataset items directly:
curl -L "https://api.apify.com/v2/actors/agentx~bilibili-transcript/run-sync-get-dataset-items" \-H "Authorization: Bearer $APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"video_url": "https://www.bilibili.com/video/BV1o84y1V7EY/?p=1","translate": "spanish"}'
For long videos or workflows that must not hold one HTTP connection open, start an asynchronous run and retrieve its Dataset afterward. Apify documents that synchronous Dataset responses can time out after 300 seconds while the Actor run continues.
Python client
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("agentx/bilibili-transcript").call(run_input={"video_url": "https://www.bilibili.com/video/BV1o84y1V7EY/?p=1","translate": "spanish",})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["transcript"]["text"])
MCP for AI clients
Configure the Apify MCP server with the Actor-scoped tool URL:
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=agentx/bilibili-transcript","headers": {"Authorization": "Bearer YOUR_APIFY_TOKEN"}}}}
Call agentx/bilibili-transcript with the same input fields. See the Apify MCP documentation and the generated Actor API page.
Pricing & Cost Calculator
Bilibili Transcript uses event pricing. The repository metadata is the current build authority; any public Store page remains a separate deployed snapshot and can lag these values.
| Event | Current price |
|---|---|
| Actor start | $0.001 per charged start event; the 8192 MB run configuration charges eight start events |
| Actor usage | $0.00001 per usage unit; total depends on runtime resources |
| Transcript - Free | $0.38700 |
| Transcript - Bronze | $0.37410 |
| Transcript - Silver | $0.36120 |
| Transcript - Gold, Platinum, Diamond | $0.34830 |
| Translation - Free | $0.15 |
| Translation - Bronze | $0.145 |
| Translation - Silver | $0.14 |
| Translation - Gold, Platinum, Diamond | $0.135 |
At Free tier, the fixed subtotal for a transcript on 8192 MB is $0.387 + (8 × $0.001) = $0.395, plus usage. A completed translation makes it $0.387 + $0.15 + $0.008 = $0.545. Start and runtime usage may accrue even when later work fails.
Check the live pricing page before scheduling a large workload. Estimate with representative media because speech density, duration, network behavior, and translation length affect runtime.
Use Cases & Applications
Search and retrieval over one video
Index the joined transcript for part-level search, or index individual segments so a match can point back to an approximate interval. Embeddings and answers are downstream responsibilities.
Editorial and research review
Researchers, bilingual editors, education teams, and analysts can keep spoken claims beside the BV title, uploader, publish time, and available engagement. Preserve the source and part selector for reproducible review.
Accessibility drafts and content repurposing
Segment ranges can seed bilingual notes, caption editing, summaries, or search, but they remain machine-generated drafts rather than platform subtitles or certified accessibility output.
Multilingual review
One target language yields a parallel segment-aligned reading layer. Separate runs are required for additional target languages.
Automation boundaries
Discovery should first produce authorized BV links and part choices; this Actor then handles selected parts individually. A technically successful run does not grant publication or redistribution rights.
FAQ
Does the Actor return Bilibili subtitles or danmaku?
No. The implemented path generates speech text from audio. Platform subtitles, danmaku, comments, and pixels rendered into the video are not input to that process.
Can it process a multi-part Bilibili video?
Use a URL that selects one specific part, such as ?p=1. The run processes one resolved media item; it does not expand an entire multi-part series.
Can it process an entire creator page or series?
No. The public input accepts one video URL, and playlist and multi-part expansion are disabled. Use a discovery Actor first and submit individual video URLs in separate runs.
What happens if the video has no speech?
The run fails without a transcript Dataset item. Silent clips, music-only media, missing audio, or speech that cannot be detected do not satisfy the successful output contract.
Why are some metadata fields null?
Bilibili or the extraction response did not expose them for that video. The schema intentionally permits nullable values instead of fabricating counts, dates, audio attribution, or author details.
Are timestamps word-accurate?
No. They are segment time ranges formatted as HH:MM:SS.mmm. They are useful for navigation and downstream processing but are not guaranteed word-level alignments.
How are long videos handled?
The Actor runs with 8192 MB and 2 CPUs. A part longer than 3600 seconds is prepared as sequential 900-second core blocks with 15-second edge context, then merged by owned time ranges. Long parts still cost more wall time and usage.
Is translation always charged when requested?
No. Translation is charged only when the complete translated structure is produced. If any segment ultimately fails, the output keeps translation empty and the translation event is not charged.
SEO Keywords & Search Terms
Bilibili video to text, BV transcript with timestamps, transcribe one Bilibili part, and Bilibili speech for RAG describe this exact single-part workflow.
It is not a subtitle or danmaku downloader, creator-space scraper, multipart batch API, media downloader, summarizer, OCR tool, or guaranteed-verbatim service.
Use agentx/bilibili-transcript in API clients. Evaluate current Input, Output, Pricing, Reviews, and Issues only after deployment, because online Store state is not the authority for these local edits.
Trust & Certifications
Repository evidence proves bounded interfaces: two closed input fields, exact part-oriented example data, 22 Dataset/view/display keys, platform confirmation, and complete-or-absent translation. It does not prove that a volatile Bilibili URL will remain accessible.
No rating, user count, monthly activity, or bookmark figure is copied from the Store because those values depend on publication state and age quickly.
No independent accuracy certificate, SLA, legal-transcript designation, privacy certification, or uptime guarantee is claimed. Bilibili controls source access and machine-generated speech requires review.
Legal & Compliance
Process only content you are authorized to use. Public accessibility does not remove copyright, privacy, contractual, publicity, data-protection, or platform-policy obligations. Review Bilibili's Terms of Service and applicable law for your use case.
The Actor returns source metadata and generated speech text. It does not grant a license to the video, verify ownership, determine fair use, provide legal advice, or certify the accuracy of quotations. Avoid submitting private credentials or confidential URLs because they are not part of the supported input contract.
If a source owner deletes or restricts a video, future runs can fail even when an older run succeeded. Retain the source URL and processing time when provenance matters.
Related Tools
Related AgentX Actors
- TikTok Transcript handles one public TikTok video.
- YouTube Transcript handles one public YouTube video.
- Video Transcript accepts uploaded media and public URLs from many platforms.
Enrich with broader discovery
- All Video Scraper provides broader video metadata and media workflows.
- All Jobs Scraper supplies cross-platform job data for content or labor-market research.
- All Shopping Scraper supplies cross-market product data for commerce research.
Choose a discovery Actor when you do not yet have a specific video URL. Keep transcription runs one video at a time.
Support & Community
When reporting a problem, include the public video URL, run ID, expected result, observed error, and whether translation was enabled. Do not post private tokens or credentials.
The local schemas, code behavior, pricing metadata, and cited Bilibili terms were reviewed on July 21, 2026. Example availability and online Store state remain volatile.