Wikimedia Commons TimedText Subtitles Scraper avatar

Wikimedia Commons TimedText Subtitles Scraper

Pricing

from $3.40 / 1,000 result items

Go to Apify Store
Wikimedia Commons TimedText Subtitles Scraper

Wikimedia Commons TimedText Subtitles Scraper

Extract TimedText subtitles and closed captions from Wikimedia Commons. Pulls page metadata and subtitle content including page ID, title, revision content, cue index, start and end times, and cue text. Use it to build subtitle datasets for media accessibility research or video indexing.

Pricing

from $3.40 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

ParseForge

Wikimedia Commons TimedText Subtitles Scraper

Scrape TimedText subtitles and closed captions from Wikimedia Commons, up to a million per run. Every subtitle cue comes with its start and end timestamps, cue text, and page metadata. No login or API key. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons hosts thousands of subtitle files in its TimedText namespace, but there is no direct way to bulk download them for analysis or reuse. This Actor reads the TimedText category via the MediaWiki API and returns each subtitle cue as a structured row, including timestamps and text.

Who uses itWhat they scrape Wikimedia Commons for
Video archivistsWhich subtitles are available for a given language or topic on Commons.
Accessibility researchersHow closed captions are structured across different media files.
Language learnersExtract subtitle text for vocabulary study from educational videos.
Data journalistsAnalyze the metadata and timing patterns of public domain subtitles.

What it does

This Actor collects TimedText subtitle pages from Wikimedia Commons by category and optional search term, and returns each subtitle cue as a flat row with start time, end time, and cue text.

  • ๐Ÿ“‚ Category-based scraping: Start from any TimedText category (default 'TimedText') and get all pages.
  • ๐Ÿ” Search filter: Narrow results to pages matching a keyword within the category.
  • โฑ๏ธ Cue-level output: Each subtitle cue is a separate row with start_time, end_time, and cue_text.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with Wikimedia Commons data

๐Ÿ“‚ Archive subtitle collections.

A digital librarian runs the Actor on the 'TimedText' category to download all available subtitle files for preservation.

๐Ÿ” Research caption patterns.

An accessibility researcher uses the search filter to find subtitles for a specific language and analyzes timing and text structure.

๐Ÿ“š Build language learning datasets.

A language teacher extracts subtitle text from educational videos on Commons to create vocabulary exercises.

๐Ÿ“Š Analyze subtitle metadata.

A data journalist scrapes all TimedText pages to study the distribution of languages and topics in public domain subtitles.

Why choose this scraper

What you get
Structured cue dataEach subtitle cue is a separate row with start_time, end_time, and cue_text, ready for analysis.
No authenticationThe MediaWiki API is public; no login or API key required.
Bulk exportExport to CSV, JSON, Excel, or XML for offline use.
ScalableScrape up to 1,000,000 items per run (paid plan).

What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

{
"pageid": 143229778,
"title": "TimedText:Sandbox.webm.en.srt",
"cueIndex": 1,
"startTime": "00:00:20,000",
"endTime": "00:00:24,400",
"cueText": "should be shown from second 20 to second 24",
"category": "TimedText",
"scrapedAt": "2026-09-24T10:59:03.382Z"
}

Every value above comes from a real run. A field a record does not have comes back as null.

Configure the run

Drive the Actor from a Wikimedia Commons category name and an optional search term. The category filter selects which TimedText pages to scrape, and the search term further narrows results to pages whose title contains that term. The Input tab lists every parameter.

A first run with the defaults:

{
"maxItems": 10,
"category": "TimedText"
}

A larger pull:

{
"maxItems": 200,
"category": "TimedText"
}

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account.
  2. Set your inputs and any filters, then click Start.
  3. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$undefined

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Ensure the category name is correct and exists on Commons. Also check that the search term (if used) matches page titles in that category.

The run fails with an API error.

Wikimedia Commons may be temporarily unavailable. Try again later. If the issue persists, check the category name for special characters.

I only see 10 items even though I set maxItems higher.

Free accounts are limited to 10 items. Upgrade to a paid plan to scrape more.

The output has duplicate rows.

Each subtitle cue is a separate row. If a page has multiple cues, you will see multiple rows for that page. This is expected.

Can I scrape all TimedText pages at once?

Yes, set maxItems to a high number (up to 1,000,000) and leave the category as 'TimedText' with no search term.

FAQ

QuestionAnswer
What is the TimedText namespace on Wikimedia Commons?It is a special namespace that stores subtitle and closed caption files for media files hosted on Commons. Each page contains cues with timestamps and text.
Do I need an API key or login?No. The MediaWiki API used by this Actor is publicly accessible without authentication.
Can I scrape subtitles for a specific video?This Actor scrapes by category and search term, not by video file. To find subtitles for a specific video, search for its filename in the category.
What data does each row contain?Each row includes pageid, title, cueIndex, startTime, endTime, cueText, category, scrapedAt, and error.
How many items can I scrape?Free users are limited to 10 items (preview). Paid users can set maxItems up to 1,000,000.
What export formats are supported?CSV, JSON, Excel, and XML.
Does this scrape subtitles from Wikipedia?No, it scrapes from Wikimedia Commons, which hosts media files and their subtitles. Wikipedia uses a different system.
Can I filter by language?Not directly, but you can use the search term to filter page titles that include language codes (e.g., 'en' for English).
Is the data up to date?Yes, the Actor queries the live MediaWiki API each run, so you get the current state of the category.
What if the category name is misspelled?The API will return an empty result set. Double-check the category name on Commons.

Browse the full ParseForge collection for more scrapers.

๐Ÿ†˜ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

โš ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.