# Wikimedia Commons Revision Diff Scraper (`parseforge/wikimedia-commons-revision-diff-scraper`) Actor

Track every change on Wikimedia Commons by pulling recent revision diffs. Extracts revision IDs, page titles, usernames, edit summaries, size deltas, and bot flags. Ideal for monitoring uploads, deletions, and metadata updates across the world's largest free media repository.

- **URL**: https://apify.com/parseforge/wikimedia-commons-revision-diff-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Other, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.85 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)](https://apify.com/parseforge?fpr=vmoqkp)

### Wikimedia Commons Revision Diff Scraper

**Scrape revision diffs from Wikimedia Commons by change type, log action, or search term, up to a million per run.** Each diff comes with its old and new revision IDs, size delta, editor, and edit summary. No API key required. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons tracks every file upload, page edit, and deletion through its public revision history, but the web interface shows only one change at a time. This Actor reads the recent changes API directly, filters by edit, new, or log actions like uploads and deletions, and returns each revision pair in one flat row. You can also limit results to pages whose title contains a specific search term.

| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Monitor which files are being uploaded or deleted from the Commons each day. |
| Content moderators | Track recent edits and log actions to spot vandalism or policy violations. |
| Researchers | Analyze editing patterns and contributor activity across the Commons file repository. |
| Data journalists | Collect revision metadata to report on changes to publicly hosted media files. |

### What it does

This Actor collects revision diffs from Wikimedia Commons by change type, log action, or title search, and returns each one as a flat row with old and new revision IDs, size change, editor, and edit summary.

- 📋 **Change type filter:** fetch only edits, new page creations, or logged actions like uploads and deletions.
- 📁 **Log type filter:** when fetching logs, narrow results to upload, delete, move, protect, block, or user rights actions.
- 🔍 **Title search:** limit results to pages whose title contains a case-sensitive search term you provide.
- 📊 **Flat row output:** each revision pair returns old\_revid, new\_revid, size delta, editor, edit summary, and page title.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Wikimedia Commons data

**📸 Monitor new file uploads.**

A digital archivist runs the Actor with change type set to log and log type set to upload to collect every new file added to the Commons in the last 24 hours.

**🗑️ Track file deletions.**

A content moderator filters by log type delete to review which files were removed and by whom, supporting copyright compliance checks.

**✏️ Audit page edits.**

A researcher sets change type to edit and a search term for a specific file name to see every revision made to that file's description page.

**📈 Analyze contributor activity.**

A data journalist collects all recent changes without filters to study which editors are most active and what kinds of changes they make.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | The Wikimedia Commons API is public and requires no authentication or registration. |
| **Bulk diffs** | Collect up to a million revision pairs in a single run instead of clicking through pages one by one. |
| **Fixed schema** | Every row has the same fields: revision IDs, size change, editor, comment, and page title. |
| **Flexible export** | Save results as CSV, JSON, Excel, or XML for use in any analysis tool. |

### What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "type": "new",
 "ns": 14,
 "title": "Category:Centro de detención Estadio Nacional de Chile",
 "pageid": 200137569,
 "revid": 1280029664,
 "old_revid": 0,
 "rcid": 3475809139,
 "user": "Nicolescribe",
 "temp": false,
 "bot": false,
 "minor": false,
 "oldlen": 0,
 "newlen": 0,
 "comment": "[[Special:MyLanguage/COM:AES|←]]Created blank page",
 "diffUrl": "https://commons.wikimedia.org/w/index.php?title=Category%3ACentro%20de%20detenci%C3%B3n%20Estadio%20Nacional%20de%20Chile&diff=1280029664&oldid=0"
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor by change type, log action, and a title search term, alone or together, and filters run as each revision is read so only matches reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "maxItems": 10
}
```

A larger pull:

```json
{
 "maxItems": 200
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Set your inputs and any filters, then click **Start**.
3. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check your search term. It is case-sensitive, so a mismatch in capitalization will return zero results. Also verify that your change type and log type filters are not too restrictive. Try leaving all filters blank to confirm the API is reachable.

**Why does the Actor stop after only 10 items?**

Free Apify users are limited to a 10-item preview. Upgrade to a paid plan to increase the max items limit up to 1,000,000.

**Why are some revisions missing an editor name?**

Edits made by logged-out users show an IP address in the user field. If the field is blank, the revision may have been performed by a system account or the data was suppressed for privacy reasons.

**Why do I see log actions that do not match my log type filter?**

The log type filter only applies when change type is set to log. If change type is blank or set to edit or new, the log type filter is ignored and you will see all matching revisions regardless of log action.

### FAQ

| Question | Answer |
|---|---|
| Do I need a Wikimedia account or API key to use this Actor? | No. The Wikimedia Commons API is publicly accessible and does not require authentication. You can run the Actor immediately with no registration. |
| What is the difference between change type edit, new, and log? | Edit returns revisions to existing pages. New returns page creations. Log returns logged actions like file uploads, deletions, moves, and user rights changes. You can filter by one type or leave it blank to get all three. |
| Can I get the actual diff content, not the metadata? | This Actor returns diff metadata: old and new revision IDs, size change in bytes, editor, and edit summary. To fetch the full diff text, you would need to call the compare API separately using the revision IDs this Actor provides. |
| How do I filter results to a specific file or page? | Use the search term input. The Actor will return only revisions where the page title contains your term. The search is case-sensitive, so match the exact capitalization of the page name. |
| What does the size delta field represent? | It is the difference in bytes between the old and new revision. A positive number means content was added, a negative number means content was removed. |
| Can I scrape more than 500 revisions at once? | Yes. The Actor paginates through the API automatically. Paid users can set max items up to 1,000,000. Free users are limited to a 10-item preview. |
| Does this Actor handle the hCaptcha on the main Wikimedia Commons site? | The Actor calls the API directly, not the web interface. The API does not present a CAPTCHA, so no solving is required. |
| What export formats are supported? | You can export your dataset as CSV, JSON, Excel, or XML directly from the Apify platform. |
| Can I schedule this Actor to run daily? | Yes. Apify supports scheduled runs. You can set the Actor to collect recent changes every hour, day, or week and receive notifications on completion. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `changeType` (type: `string`):

Filter by type of change. Choose edit for page edits, new for page creations, log for logged actions like uploads or deletions, or leave blank for all.

## `logType` (type: `string`):

When change type is log, filter by specific log action. Options include upload for new files, delete for deletions, move for page moves, and others.

## `searchTerm` (type: `string`):

Limit results to pages whose title contains this text. Case-sensitive. Leave empty to fetch all recent changes.

## Actor input object example

```json
{
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/wikimedia-commons-revision-diff-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("parseforge/wikimedia-commons-revision-diff-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call parseforge/wikimedia-commons-revision-diff-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/wikimedia-commons-revision-diff-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KvLWaAGnIdfBGGrxF/builds/zYOgbVpMStgE1T91O/openapi.json
