# Wikimedia Commons Subcategory Scraper (`parseforge/wikimedia-commons-subcategory-scraper`) Actor

Scrapes Wikimedia Commons subcategory names and IDs from a starting category. Returns each subcategory as a flat row with page ID, namespace, and title.

- **URL**: https://apify.com/parseforge/wikimedia-commons-subcategory-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Business, Other, Education
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.62 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Wikimedia Commons Subcategory Scraper

**Scrape Wikimedia Commons subcategory trees from any category, up to a million per run.** Every subcategory comes with its page ID, namespace, and full title. No login or API key. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons holds millions of freely licensed media files, but they are organized in a deep hierarchy of categories. Finding every subcategory under a topic by hand is slow and error-prone. This scraper reads the public Wikimedia API directly, starting from one category and walking down the tree, so you get a complete list of subcategories in one dataset.

| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Mapping the full category structure of a Commons topic for a GLAM project. |
| Data curators | Building a taxonomy of freely licensed media categories for a search tool. |
| Researchers | Analyzing how a subject area is categorized and organized on Commons. |
| Content managers | Finding all relevant subcategories to source images for a website or publication. |

### What it does

This Actor collects Wikimedia Commons subcategory names and IDs from a starting category and returns each one as a flat row.

- 🌳 **Full tree traversal:** starts at one category and recursively fetches all descendant subcategories.
- ⚙️ **Configurable depth:** set a maximum number of items to control the size of your dataset.
- 📄 **Flat row output:** each subcategory is one row with its page ID, namespace, and title, ready for spreadsheets.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Wikimedia Commons data

**🗂️ Map a media category tree.**

A digital archivist starts from 'Category:Paintings' and collects every subcategory to understand how the collection is organized before a migration.

**🔍 Build a searchable taxonomy.**

A developer scrapes all subcategories under 'Category:Animals' and uses the list to power a faceted search on a stock media website.

**📊 Analyze category growth.**

A researcher runs the scraper monthly on 'Category:COVID-19' and tracks how the number of subcategories changes over time.

**🖼️ Source images by topic.**

A content manager extracts all subcategories of 'Category:Historical photographs of cities' to find relevant image sets for a blog series.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key needed** | Uses the public Wikimedia API with no registration or authentication. |
| **Recursive discovery** | Automatically walks down the category tree so you do not miss nested subcategories. |
| **Fixed schema** | Every row has the same fields: page ID, namespace, and full category title. |
| **Flexible export** | Save results as CSV, JSON, Excel, or XML for any downstream tool. |

### What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "pageid": 703561,
 "ns": 14,
 "title": "Category:Wikimedia Commons",
 "url": "https://commons.wikimedia.org/wiki/Category%3AWikimedia_Commons",
 "scrapedAt": "2026-09-04T01:16:46.381Z"
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor from a single starting category title, and set a maximum item count to limit how many subcategories are collected. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "startCategory": "Category:Commons",
 "maxItems": 10
}
```

A larger pull:

```json
{
 "startCategory": "Category:Commons",
 "maxItems": 200
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [Wikimedia Commons Subcategory Scraper](https://apify.com/parseforge/wikimedia-commons-subcategory-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/wikimedia-commons-subcategory-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check that your start category title is spelled exactly as it appears on Commons, including the 'Category:' prefix and any spaces or punctuation. Also verify the category contains subcategories.

**The run stopped before collecting all subcategories.**

You likely hit the Max Items limit you set. Increase the number in the input to allow more subcategories to be collected.

**I see duplicate subcategories in my dataset.**

A subcategory can appear under multiple parent categories. The scraper returns every occurrence it finds during traversal, so duplicates are expected and reflect the real Commons structure.

**The scraper is taking a long time to finish.**

Large category trees with many thousands of subcategories take time because the API is queried page by page. Reduce the Max Items limit to get a smaller, faster sample.

### FAQ

| Question | Answer |
|---|---|
| What is a Wikimedia Commons subcategory? | A subcategory is a child category nested inside a parent category. For example, 'Category:Dogs' is a subcategory of 'Category:Animals'. This scraper collects the titles and IDs of those child categories. |
| Do I need a Wikimedia account or API key? | No. The scraper uses the public Wikimedia API, which requires no login, no token, and no registration. |
| How deep does the scraper go into the category tree? | It starts at the category you provide and recursively fetches all descendant subcategories until it reaches the maximum item count you set. |
| Can I scrape media files, not subcategories? | No, this Actor is designed to collect subcategory metadata only. To scrape actual media file URLs or metadata, you would need a different scraper. |
| What does the output look like? | Each row is a flat object with the subcategory's page ID, namespace (always 14 for categories), and full title including the 'Category:' prefix. |
| Is there a limit on how many subcategories I can scrape? | Free users are limited to 10 items for a preview. Paid users can set a maximum up to 1,000,000 items per run. |
| Can I start from any category? | Yes, you provide the full title with the 'Category:' prefix, such as 'Category:Commons' or 'Category:Paintings from Italy'. |
| What export formats are supported? | You can export your dataset to CSV, JSON, Excel, or XML from the Apify platform. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

### Pricing

This Actor uses **pay-per-result** pricing: **$0.004 per result** collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `startCategory` (type: `string`):

The full title of the Wikimedia Commons category to start from, including the 'Category:' prefix. For example, 'Category:Commons'.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## Actor input object example

```json
{
  "startCategory": "Category:Commons",
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startCategory": "Category:Commons",
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/wikimedia-commons-subcategory-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startCategory": "Category:Commons",
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/wikimedia-commons-subcategory-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startCategory": "Category:Commons",
  "maxItems": 10
}' |
apify call parseforge/wikimedia-commons-subcategory-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/wikimedia-commons-subcategory-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SnSNnV3htenuy4lIg/builds/tT8ukKhKQHjTFQkif/openapi.json
