# Wikimedia Commons Category Files Scraper (`parseforge/wikimedia-commons-category-files-scraper`) Actor

- **URL**: https://apify.com/parseforge/wikimedia-commons-category-files-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.62 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

![ParseForge Banner](https://github.com/ParseForge/apify-assets/blob/main/banner.jpg?raw=true)

## 🖼️ Wikimedia Commons Category Files Scraper

> 🚀 **Export every file in a Wikimedia Commons category, with image URLs and license, in seconds.**

This Actor reads the official MediaWiki API and returns every file in a Wikimedia Commons category with its direct image URL, license, author, dimensions, MIME type, uploader, and dates. No login, no API key, no HTML scraping.

Wikimedia Commons holds over 100 million freely licensed media files. Point this Actor at any category and get a clean, structured dataset you can filter, join, and reuse (with attribution).

| For | Use it to |
|---|---|
| Researchers & educators | Collect openly licensed images by topic |
| Designers & media teams | Source public-domain and CC media with author + license |
| Data teams | Build image datasets with rich metadata |

### 📋 What it does

- Lists all files in a Wikimedia Commons category via the MediaWiki API.
- Returns each file with its direct download URL and metadata.
- Paginates automatically up to your `maxItems`.

> 💡 **Why it matters:** every row carries the license and author, so you can reuse the media correctly.

### 📊 Output

| Field | Description |
|---|---|
| 🖼️ `imageUrl` | Direct file URL on upload.wikimedia.org |
| 📕 `title` | File title |
| 🆔 `pageId` | Commons page id |
| 🔗 `descriptionUrl` | File description page |
| 🗂️ `mime` | MIME type (image/jpeg, etc.) |
| 📐 `width` / `height` | Pixel dimensions |
| 💾 `sizeBytes` | File size in bytes |
| 👤 `uploader` | Uploading user |
| 🕒 `uploadTimestamp` | Upload time |
| 📄 `license` | License short name (e.g. CC BY-SA 4.0) |
| ✍️ `artist` | Author / creator |
| 🏷️ `credit` | Credit line |
| 📅 `dateOriginal` | Original date of the work |

Sample record:

```json
{
  "imageUrl": "https://upload.wikimedia.org/wikipedia/commons/2/21/example.jpg",
  "title": "File:Example.jpg",
  "mime": "image/jpeg",
  "width": 4342,
  "height": 1995,
  "uploader": "Podzemnik",
  "license": "CC BY-SA 4.0",
  "artist": "Michal Klajban"
}
```

### 🚀 How to use

1. [Create a free account w/ $5 credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Enter a `category` (without the `Category:` prefix) and `maxItems`.
3. Run it and download the dataset as JSON, CSV, Excel, or XML.

### ❓ FAQ

**Do I need an API key?** No. The Actor uses the public MediaWiki API.

**Where is the category name from?** Any category page on commons.wikimedia.org, e.g. "Featured pictures on Wikimedia Commons".

**Can I reuse the images?** Yes, subject to each file's license (see the `license` and `artist` fields). Always attribute per the license.

**How fresh is the data?** Every run reads the MediaWiki API live.

### 🔗 Recommended Actors

- [OurAirports Global Airport Database Scraper](https://apify.com/parseforge/ourairports-scraper)
- [DOAJ Journals & Subject Classification Scraper](https://apify.com/parseforge/doaj-subject-classification-scraper)

> 💡 **Pro Tip:** browse the complete [ParseForge collection](https://apify.com/parseforge) for more data Actors.

***

*This Actor is not affiliated with the Wikimedia Foundation. It reads publicly available data from the MediaWiki API. Respect each file's license and Wikimedia's terms.*

# Actor input Schema

## `category` (type: `string`):

Wikimedia Commons category name (without the 'Category:' prefix). Every file in the category is returned with its metadata.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## Actor input object example

```json
{
  "category": "Featured pictures on Wikimedia Commons",
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped files.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "category": "Featured pictures on Wikimedia Commons",
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/wikimedia-commons-category-files-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "category": "Featured pictures on Wikimedia Commons",
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/wikimedia-commons-category-files-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "category": "Featured pictures on Wikimedia Commons",
  "maxItems": 10
}' |
apify call parseforge/wikimedia-commons-category-files-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/wikimedia-commons-category-files-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6UdpDm6XsreIDoYTs/builds/jjdOLJo1SMNisGLZJ/openapi.json
