Wikimedia Commons Revision Diff Scraper avatar

Wikimedia Commons Revision Diff Scraper

Pricing

from $0.85 / 1,000 result items

Go to Apify Store
Wikimedia Commons Revision Diff Scraper

Wikimedia Commons Revision Diff Scraper

Track every change on Wikimedia Commons by pulling recent revision diffs. Extracts revision IDs, page titles, usernames, edit summaries, size deltas, and bot flags. Ideal for monitoring uploads, deletions, and metadata updates across the world's largest free media repository.

Pricing

from $0.85 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

ParseForge

Wikimedia Commons Revision Diff Scraper

Scrape revision diffs from Wikimedia Commons by change type, log action, or search term, up to a million per run. Each diff comes with its old and new revision IDs, size delta, editor, and edit summary. No API key required. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons tracks every file upload, page edit, and deletion through its public revision history, but the web interface shows only one change at a time. This Actor reads the recent changes API directly, filters by edit, new, or log actions like uploads and deletions, and returns each revision pair in one flat row. You can also limit results to pages whose title contains a specific search term.

Who uses itWhat they scrape Wikimedia Commons for
Digital archivistsMonitor which files are being uploaded or deleted from the Commons each day.
Content moderatorsTrack recent edits and log actions to spot vandalism or policy violations.
ResearchersAnalyze editing patterns and contributor activity across the Commons file repository.
Data journalistsCollect revision metadata to report on changes to publicly hosted media files.

What it does

This Actor collects revision diffs from Wikimedia Commons by change type, log action, or title search, and returns each one as a flat row with old and new revision IDs, size change, editor, and edit summary.

  • 📋 Change type filter: fetch only edits, new page creations, or logged actions like uploads and deletions.
  • 📁 Log type filter: when fetching logs, narrow results to upload, delete, move, protect, block, or user rights actions.
  • 🔍 Title search: limit results to pages whose title contains a case-sensitive search term you provide.
  • 📊 Flat row output: each revision pair returns old_revid, new_revid, size delta, editor, edit summary, and page title.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with Wikimedia Commons data

📸 Monitor new file uploads.

A digital archivist runs the Actor with change type set to log and log type set to upload to collect every new file added to the Commons in the last 24 hours.

🗑️ Track file deletions.

A content moderator filters by log type delete to review which files were removed and by whom, supporting copyright compliance checks.

✏️ Audit page edits.

A researcher sets change type to edit and a search term for a specific file name to see every revision made to that file's description page.

📈 Analyze contributor activity.

A data journalist collects all recent changes without filters to study which editors are most active and what kinds of changes they make.

Why choose this scraper

What you get
No API keyThe Wikimedia Commons API is public and requires no authentication or registration.
Bulk diffsCollect up to a million revision pairs in a single run instead of clicking through pages one by one.
Fixed schemaEvery row has the same fields: revision IDs, size change, editor, comment, and page title.
Flexible exportSave results as CSV, JSON, Excel, or XML for use in any analysis tool.

What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

{
"type": "new",
"ns": 14,
"title": "Category:Centro de detención Estadio Nacional de Chile",
"pageid": 200137569,
"revid": 1280029664,
"old_revid": 0,
"rcid": 3475809139,
"user": "Nicolescribe",
"temp": false,
"bot": false,
"minor": false,
"oldlen": 0,
"newlen": 0,
"comment": "[[Special:MyLanguage/COM:AES|←]]Created blank page",
"diffUrl": "https://commons.wikimedia.org/w/index.php?title=Category%3ACentro%20de%20detenci%C3%B3n%20Estadio%20Nacional%20de%20Chile&diff=1280029664&oldid=0"
}

Every value above comes from a real run. A field a record does not have comes back as null.

Configure the run

Drive the Actor by change type, log action, and a title search term, alone or together, and filters run as each revision is read so only matches reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

{
"maxItems": 10
}

A larger pull:

{
"maxItems": 200
}

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Set your inputs and any filters, then click Start.
  3. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$undefined

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check your search term. It is case-sensitive, so a mismatch in capitalization will return zero results. Also verify that your change type and log type filters are not too restrictive. Try leaving all filters blank to confirm the API is reachable.

Why does the Actor stop after only 10 items?

Free Apify users are limited to a 10-item preview. Upgrade to a paid plan to increase the max items limit up to 1,000,000.

Why are some revisions missing an editor name?

Edits made by logged-out users show an IP address in the user field. If the field is blank, the revision may have been performed by a system account or the data was suppressed for privacy reasons.

Why do I see log actions that do not match my log type filter?

The log type filter only applies when change type is set to log. If change type is blank or set to edit or new, the log type filter is ignored and you will see all matching revisions regardless of log action.

FAQ

QuestionAnswer
Do I need a Wikimedia account or API key to use this Actor?No. The Wikimedia Commons API is publicly accessible and does not require authentication. You can run the Actor immediately with no registration.
What is the difference between change type edit, new, and log?Edit returns revisions to existing pages. New returns page creations. Log returns logged actions like file uploads, deletions, moves, and user rights changes. You can filter by one type or leave it blank to get all three.
Can I get the actual diff content, not the metadata?This Actor returns diff metadata: old and new revision IDs, size change in bytes, editor, and edit summary. To fetch the full diff text, you would need to call the compare API separately using the revision IDs this Actor provides.
How do I filter results to a specific file or page?Use the search term input. The Actor will return only revisions where the page title contains your term. The search is case-sensitive, so match the exact capitalization of the page name.
What does the size delta field represent?It is the difference in bytes between the old and new revision. A positive number means content was added, a negative number means content was removed.
Can I scrape more than 500 revisions at once?Yes. The Actor paginates through the API automatically. Paid users can set max items up to 1,000,000. Free users are limited to a 10-item preview.
Does this Actor handle the hCaptcha on the main Wikimedia Commons site?The Actor calls the API directly, not the web interface. The API does not present a CAPTCHA, so no solving is required.
What export formats are supported?You can export your dataset as CSV, JSON, Excel, or XML directly from the Apify platform.
Can I schedule this Actor to run daily?Yes. Apify supports scheduled runs. You can set the Actor to collect recent changes every hour, day, or week and receive notifications on completion.

Browse the full ParseForge collection for more scrapers.

🆘 Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.