Wikimedia Commons Revision Diff Scraper
Pricing
from $0.85 / 1,000 result items
Wikimedia Commons Revision Diff Scraper
Track every change on Wikimedia Commons by pulling recent revision diffs. Extracts revision IDs, page titles, usernames, edit summaries, size deltas, and bot flags. Ideal for monitoring uploads, deletions, and metadata updates across the world's largest free media repository.
Pricing
from $0.85 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Wikimedia Commons Revision Diff Scraper
Scrape revision diffs from Wikimedia Commons by change type, log action, or search term, up to a million per run. Each diff comes with its old and new revision IDs, size delta, editor, and edit summary. No API key required. Export to CSV, JSON, Excel, or XML.
Wikimedia Commons tracks every file upload, page edit, and deletion through its public revision history, but the web interface shows only one change at a time. This Actor reads the recent changes API directly, filters by edit, new, or log actions like uploads and deletions, and returns each revision pair in one flat row. You can also limit results to pages whose title contains a specific search term.
| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Monitor which files are being uploaded or deleted from the Commons each day. |
| Content moderators | Track recent edits and log actions to spot vandalism or policy violations. |
| Researchers | Analyze editing patterns and contributor activity across the Commons file repository. |
| Data journalists | Collect revision metadata to report on changes to publicly hosted media files. |
What it does
This Actor collects revision diffs from Wikimedia Commons by change type, log action, or title search, and returns each one as a flat row with old and new revision IDs, size change, editor, and edit summary.
- 📋 Change type filter: fetch only edits, new page creations, or logged actions like uploads and deletions.
- 📁 Log type filter: when fetching logs, narrow results to upload, delete, move, protect, block, or user rights actions.
- 🔍 Title search: limit results to pages whose title contains a case-sensitive search term you provide.
- 📊 Flat row output: each revision pair returns old_revid, new_revid, size delta, editor, edit summary, and page title.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with Wikimedia Commons data
📸 Monitor new file uploads.
A digital archivist runs the Actor with change type set to log and log type set to upload to collect every new file added to the Commons in the last 24 hours.
🗑️ Track file deletions.
A content moderator filters by log type delete to review which files were removed and by whom, supporting copyright compliance checks.
✏️ Audit page edits.
A researcher sets change type to edit and a search term for a specific file name to see every revision made to that file's description page.
📈 Analyze contributor activity.
A data journalist collects all recent changes without filters to study which editors are most active and what kinds of changes they make.
Why choose this scraper
| What you get | |
|---|---|
| No API key | The Wikimedia Commons API is public and requires no authentication or registration. |
| Bulk diffs | Collect up to a million revision pairs in a single run instead of clicking through pages one by one. |
| Fixed schema | Every row has the same fields: revision IDs, size change, editor, comment, and page title. |
| Flexible export | Save results as CSV, JSON, Excel, or XML for use in any analysis tool. |
What a Wikimedia Commons record looks like
Every record returns as one flat JSON row. Here is a real one from a run:
{"type": "new","ns": 14,"title": "Category:Centro de detención Estadio Nacional de Chile","pageid": 200137569,"revid": 1280029664,"old_revid": 0,"rcid": 3475809139,"user": "Nicolescribe","temp": false,"bot": false,"minor": false,"oldlen": 0,"newlen": 0,"comment": "[[Special:MyLanguage/COM:AES|←]]Created blank page","diffUrl": "https://commons.wikimedia.org/w/index.php?title=Category%3ACentro%20de%20detenci%C3%B3n%20Estadio%20Nacional%20de%20Chile&diff=1280029664&oldid=0"}
Every value above comes from a real run. A field a record does not have comes back as null.
Configure the run
Drive the Actor by change type, log action, and a title search term, alone or together, and filters run as each revision is read so only matches reach your dataset. The Input tab lists every parameter.
A first run with the defaults:
{"maxItems": 10}
A larger pull:
{"maxItems": 200}
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.
Run it
- Create a free Apify account with $5 in credit.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
$undefined
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check your search term. It is case-sensitive, so a mismatch in capitalization will return zero results. Also verify that your change type and log type filters are not too restrictive. Try leaving all filters blank to confirm the API is reachable.
Why does the Actor stop after only 10 items?
Free Apify users are limited to a 10-item preview. Upgrade to a paid plan to increase the max items limit up to 1,000,000.
Why are some revisions missing an editor name?
Edits made by logged-out users show an IP address in the user field. If the field is blank, the revision may have been performed by a system account or the data was suppressed for privacy reasons.
Why do I see log actions that do not match my log type filter?
The log type filter only applies when change type is set to log. If change type is blank or set to edit or new, the log type filter is ignored and you will see all matching revisions regardless of log action.
FAQ
| Question | Answer |
|---|---|
| Do I need a Wikimedia account or API key to use this Actor? | No. The Wikimedia Commons API is publicly accessible and does not require authentication. You can run the Actor immediately with no registration. |
| What is the difference between change type edit, new, and log? | Edit returns revisions to existing pages. New returns page creations. Log returns logged actions like file uploads, deletions, moves, and user rights changes. You can filter by one type or leave it blank to get all three. |
| Can I get the actual diff content, not the metadata? | This Actor returns diff metadata: old and new revision IDs, size change in bytes, editor, and edit summary. To fetch the full diff text, you would need to call the compare API separately using the revision IDs this Actor provides. |
| How do I filter results to a specific file or page? | Use the search term input. The Actor will return only revisions where the page title contains your term. The search is case-sensitive, so match the exact capitalization of the page name. |
| What does the size delta field represent? | It is the difference in bytes between the old and new revision. A positive number means content was added, a negative number means content was removed. |
| Can I scrape more than 500 revisions at once? | Yes. The Actor paginates through the API automatically. Paid users can set max items up to 1,000,000. Free users are limited to a 10-item preview. |
| Does this Actor handle the hCaptcha on the main Wikimedia Commons site? | The Actor calls the API directly, not the web interface. The API does not present a CAPTCHA, so no solving is required. |
| What export formats are supported? | You can export your dataset as CSV, JSON, Excel, or XML directly from the Apify platform. |
| Can I schedule this Actor to run daily? | Yes. Apify supports scheduled runs. You can set the Actor to collect recent changes every hour, day, or week and receive notifications on completion. |
Related actors
Browse the full ParseForge collection for more scrapers.
🆘 Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
⚠️ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
