# Wikimedia Commons Files Scraper (`parseforge/wikimedia-commons-files-scraper`) Actor

Retrieve file metadata from Wikimedia Commons filtered by date range. Pulls file name, upload timestamp, uploader, comment, direct URL, dimensions, size, MIME type, and media type. Ideal for tracking recent uploads and building media archives.

- **URL**: https://apify.com/parseforge/wikimedia-commons-files-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Other, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.68 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)](https://apify.com/parseforge?fpr=vmoqkp)

### Wikimedia Commons Files Scraper

**Scrape Wikimedia Commons file metadata by date range, up to a million files per run.** Every file comes with its upload timestamp, author, comment, direct URL, dimensions, MIME type, and page ID. No login or API key. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons hosts over 100 million freely licensed media files, but the official API returns only 500 results per request and forces you to page through them manually. This Actor reads the public allimages endpoint directly, walks the full date range you give it, and returns every matching file as one flat row. You get the file URL, uploader, dimensions, MIME type, and description link for each upload, ready for bulk download or analysis.

| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Building a complete inventory of Commons uploads for a given period |
| Researchers | Studying upload patterns, contributor activity, or media type distribution over time |
| Dataset builders | Assembling a corpus of freely licensed images for machine learning or analysis |
| Content curators | Finding recently uploaded media on a topic before it appears in search indexes |
| GLAM professionals | Tracking institutional uploads and verifying metadata completeness |

### What it does

This Actor collects Wikimedia Commons file records by upload date range and returns each file as a flat row with its URL, uploader, timestamp, dimensions, MIME type, and page metadata.

- 📅 **Date range driven:** set a start and end timestamp in ISO 8601 format and the Actor walks every upload in between.
- 🔗 **Direct file URLs:** each row includes the original file URL so you can download the media without extra API calls.
- 📐 **Dimensions and MIME type:** width, height, media type, and bit depth are returned for every file.
- 👤 **Uploader attribution:** the username of the uploader and their upload comment are included in each row.
- 📄 **Page metadata:** title, namespace, page ID, and description URL let you jump straight to the Commons file page.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Wikimedia Commons data

**📦 Bulk download a date range.**

A dataset builder sets a start and end date, runs the Actor, and uses the returned file URLs to download every Commons upload from that period into a training corpus.

**📊 Analyze upload trends.**

A researcher scrapes a year of uploads and groups the output by MIME type and uploader to measure how Commons media composition has shifted.

**🏛️ Audit institutional collections.**

A GLAM professional filters a date range matching a museum's upload campaign and verifies that every file has the expected dimensions, author, and description link.

**🔎 Find recent media on a topic.**

A content curator scrapes the last week of uploads, filters the output locally for relevant keywords in file names, and shortlists candidates for reuse.

**🧾 Build a file inventory.**

An archivist runs the Actor over a multi-year range and stores the flat rows as a searchable index of Commons files with direct URLs and page IDs.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | The public allimages endpoint needs no authentication, registration, or OAuth flow |
| **Bulk by default** | The Actor pages through the API automatically, so a million-file range is one run, not a manual loop |
| **Fixed schema** | Every file returns the same fields, so your CSV or JSON output is ready for analysis without cleanup |
| **Date precision** | ISO 8601 timestamps with second-level granularity let you target exact upload windows |
| **Free preview** | Free users can pull up to 10 files to verify the schema before paying for a full run |

### What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "imageUrl": "https://upload.wikimedia.org/wikipedia/commons/3/3e/48th_CMS_Conducts_Routine_Maintenance_%286744734%29.jpg?utm_source=commons.wikimedia.org&utm_campaign=imageinfo&utm_content=original",
 "name": "48th_CMS_Conducts_Routine_Maintenance_(6744734).jpg",
 "title": "File:48th CMS Conducts Routine Maintenance (6744734).jpg",
 "ns": 6,
 "timestamp": "2025-01-01T00:00:01Z",
 "user": "OptimusPrimeBot",
 "comment": "#Spacemedia - Upload of https://d34w7g4gy10iej.cloudfront.net/photos/2107/6744734.jpg via [[:Commons:Spacemedia]]",
 "url": "https://upload.wikimedia.org/wikipedia/commons/3/3e/48th_CMS_Conducts_Routine_Maintenance_%286744734%29.jpg?utm_source=commons.wikimedia.org&utm_campaign=imageinfo&utm_content=original",
 "descriptionurl": "https://commons.wikimedia.org/wiki/File:48th_CMS_Conducts_Routine_Maintenance_(6744734).jpg",
 "descriptionshorturl": "https://commons.wikimedia.org/w/index.php?curid=157392694",
 "size": 1102117,
 "width": 2638,
 "height": 1755,
 "scrapedAt": "2026-09-24T04:46:04.546Z"
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor from a start and end date in ISO 8601 format, and set a maximum item count to cap the run. The date range is inclusive, so files uploaded exactly at the boundary timestamps are included. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "startDate": "2025-01-01T00:00:00Z",
 "endDate": "2025-01-31T23:59:59Z",
 "maxItems": 10
}
```

A larger pull:

```json
{
 "startDate": "2025-01-01T00:00:00Z",
 "endDate": "2025-01-31T23:59:59Z",
 "maxItems": 200
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Set your inputs and any filters, then click **Start**.
3. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check that your startDate is earlier than your endDate and that both are valid ISO 8601 timestamps. Also confirm that files were uploaded to Commons during that window. A very narrow range with no uploads will return an empty dataset.

**Why did the run stop before reaching my maxItems?**

The Actor stops when it has read every file in your date range. If the range contains fewer uploads than your maxItems value, the run ends early with a complete dataset.

**Why are some file URLs returning a 404 when I download them?**

Files on Commons can be renamed or deleted after the Actor reads them. The metadata reflects the state at scrape time. Re-run the Actor to get the current URLs.

**Why is my free run limited to 10 items?**

Free Apify accounts include a preview limit of 10 items per run for this Actor. Upgrade to a paid plan to set maxItems up to 1,000,000.

**Why do some rows have empty comment or user fields?**

Some uploads, especially early ones or those imported by bots, have no upload comment or a system user. Empty strings in these fields are expected and not an error.

### FAQ

| Question | Answer |
|---|---|
| Do I need a Wikimedia account or API key? | No. The Actor uses the public allimages endpoint, which requires no authentication. You can run it immediately with a date range. |
| What date format should I use? | ISO 8601 with a time component, for example 2025-01-01T00:00:00Z for the start and 2025-01-31T23:59:59Z for the end. The range is inclusive. |
| How many files can I scrape in one run? | Free users are limited to 10 files as a preview. Paid users can set maxItems up to 1,000,000 files per run. |
| Does the output include the actual image or metadata? | The output includes the direct file URL for each upload, so you can download the media yourself. The Actor returns metadata, not the binary file content. |
| Can I filter by file type or category? | The Actor filters by upload date range only. You can filter the output locally by MIME type, media type, or file name after the run completes. |
| What is the difference between this and the Wikimedia Commons API? | This Actor wraps the same public API but handles pagination, rate limits, and retries for you, and returns a flat dataset instead of nested JSON. |
| Does it scrape deleted or hidden files? | No. The allimages endpoint returns only files that are currently visible on Commons. Deleted or suppressed uploads are not included. |
| Can I scrape a specific user's uploads? | Not directly with this Actor. It filters by date range. To get a specific user's uploads, run a date range and filter the output by the uploader field. |
| What export formats are supported? | You can export the dataset to CSV, JSON, Excel, or XML from the Apify platform after the run finishes. |
| Is the data returned in a stable order? | Yes. The Actor sorts by upload timestamp, so files are returned oldest to newest within your date range. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `startDate` (type: `string`):

Start of the date range in ISO 8601 format (e.g. 2025-01-01T00:00:00Z). Files uploaded on or after this timestamp will be included.

## `endDate` (type: `string`):

End of the date range in ISO 8601 format (e.g. 2025-01-31T23:59:59Z). Files uploaded on or before this timestamp will be included.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## Actor input object example

```json
{
  "startDate": "2025-01-01T00:00:00Z",
  "endDate": "2025-01-31T23:59:59Z",
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startDate": "2025-01-01T00:00:00Z",
    "endDate": "2025-01-31T23:59:59Z",
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/wikimedia-commons-files-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startDate": "2025-01-01T00:00:00Z",
    "endDate": "2025-01-31T23:59:59Z",
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/wikimedia-commons-files-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startDate": "2025-01-01T00:00:00Z",
  "endDate": "2025-01-31T23:59:59Z",
  "maxItems": 10
}' |
apify call parseforge/wikimedia-commons-files-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/wikimedia-commons-files-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZrZt9ESqDxLUypweq/builds/RaHCvdFc7jXqy9hav/openapi.json
