Wikimedia Commons Files Scraper
Pricing
from $2.68 / 1,000 result items
Wikimedia Commons Files Scraper
Retrieve file metadata from Wikimedia Commons filtered by date range. Pulls file name, upload timestamp, uploader, comment, direct URL, dimensions, size, MIME type, and media type. Ideal for tracking recent uploads and building media archives.
Pricing
from $2.68 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Wikimedia Commons Files Scraper
Scrape Wikimedia Commons file metadata by date range, up to a million files per run. Every file comes with its upload timestamp, author, comment, direct URL, dimensions, MIME type, and page ID. No login or API key. Export to CSV, JSON, Excel, or XML.
Wikimedia Commons hosts over 100 million freely licensed media files, but the official API returns only 500 results per request and forces you to page through them manually. This Actor reads the public allimages endpoint directly, walks the full date range you give it, and returns every matching file as one flat row. You get the file URL, uploader, dimensions, MIME type, and description link for each upload, ready for bulk download or analysis.
| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Building a complete inventory of Commons uploads for a given period |
| Researchers | Studying upload patterns, contributor activity, or media type distribution over time |
| Dataset builders | Assembling a corpus of freely licensed images for machine learning or analysis |
| Content curators | Finding recently uploaded media on a topic before it appears in search indexes |
| GLAM professionals | Tracking institutional uploads and verifying metadata completeness |
What it does
This Actor collects Wikimedia Commons file records by upload date range and returns each file as a flat row with its URL, uploader, timestamp, dimensions, MIME type, and page metadata.
- ๐ Date range driven: set a start and end timestamp in ISO 8601 format and the Actor walks every upload in between.
- ๐ Direct file URLs: each row includes the original file URL so you can download the media without extra API calls.
- ๐ Dimensions and MIME type: width, height, media type, and bit depth are returned for every file.
- ๐ค Uploader attribution: the username of the uploader and their upload comment are included in each row.
- ๐ Page metadata: title, namespace, page ID, and description URL let you jump straight to the Commons file page.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with Wikimedia Commons data
๐ฆ Bulk download a date range.
A dataset builder sets a start and end date, runs the Actor, and uses the returned file URLs to download every Commons upload from that period into a training corpus.
๐ Analyze upload trends.
A researcher scrapes a year of uploads and groups the output by MIME type and uploader to measure how Commons media composition has shifted.
๐๏ธ Audit institutional collections.
A GLAM professional filters a date range matching a museum's upload campaign and verifies that every file has the expected dimensions, author, and description link.
๐ Find recent media on a topic.
A content curator scrapes the last week of uploads, filters the output locally for relevant keywords in file names, and shortlists candidates for reuse.
๐งพ Build a file inventory.
An archivist runs the Actor over a multi-year range and stores the flat rows as a searchable index of Commons files with direct URLs and page IDs.
Why choose this scraper
| What you get | |
|---|---|
| No API key | The public allimages endpoint needs no authentication, registration, or OAuth flow |
| Bulk by default | The Actor pages through the API automatically, so a million-file range is one run, not a manual loop |
| Fixed schema | Every file returns the same fields, so your CSV or JSON output is ready for analysis without cleanup |
| Date precision | ISO 8601 timestamps with second-level granularity let you target exact upload windows |
| Free preview | Free users can pull up to 10 files to verify the schema before paying for a full run |
What a Wikimedia Commons record looks like
Every record returns as one flat JSON row. Here is a real one from a run:
{"imageUrl": "https://upload.wikimedia.org/wikipedia/commons/3/3e/48th_CMS_Conducts_Routine_Maintenance_%286744734%29.jpg?utm_source=commons.wikimedia.org&utm_campaign=imageinfo&utm_content=original","name": "48th_CMS_Conducts_Routine_Maintenance_(6744734).jpg","title": "File:48th CMS Conducts Routine Maintenance (6744734).jpg","ns": 6,"timestamp": "2025-01-01T00:00:01Z","user": "OptimusPrimeBot","comment": "#Spacemedia - Upload of https://d34w7g4gy10iej.cloudfront.net/photos/2107/6744734.jpg via [[:Commons:Spacemedia]]","url": "https://upload.wikimedia.org/wikipedia/commons/3/3e/48th_CMS_Conducts_Routine_Maintenance_%286744734%29.jpg?utm_source=commons.wikimedia.org&utm_campaign=imageinfo&utm_content=original","descriptionurl": "https://commons.wikimedia.org/wiki/File:48th_CMS_Conducts_Routine_Maintenance_(6744734).jpg","descriptionshorturl": "https://commons.wikimedia.org/w/index.php?curid=157392694","size": 1102117,"width": 2638,"height": 1755,"scrapedAt": "2026-09-24T04:46:04.546Z"}
Every value above comes from a real run. A field a record does not have comes back as null.
Configure the run
Drive the Actor from a start and end date in ISO 8601 format, and set a maximum item count to cap the run. The date range is inclusive, so files uploaded exactly at the boundary timestamps are included. The Input tab lists every parameter.
A first run with the defaults:
{"startDate": "2025-01-01T00:00:00Z","endDate": "2025-01-31T23:59:59Z","maxItems": 10}
A larger pull:
{"startDate": "2025-01-01T00:00:00Z","endDate": "2025-01-31T23:59:59Z","maxItems": 200}
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.
Run it
- Create a free Apify account with $5 in credit.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
$undefined
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check that your startDate is earlier than your endDate and that both are valid ISO 8601 timestamps. Also confirm that files were uploaded to Commons during that window. A very narrow range with no uploads will return an empty dataset.
Why did the run stop before reaching my maxItems?
The Actor stops when it has read every file in your date range. If the range contains fewer uploads than your maxItems value, the run ends early with a complete dataset.
Why are some file URLs returning a 404 when I download them?
Files on Commons can be renamed or deleted after the Actor reads them. The metadata reflects the state at scrape time. Re-run the Actor to get the current URLs.
Why is my free run limited to 10 items?
Free Apify accounts include a preview limit of 10 items per run for this Actor. Upgrade to a paid plan to set maxItems up to 1,000,000.
Why do some rows have empty comment or user fields?
Some uploads, especially early ones or those imported by bots, have no upload comment or a system user. Empty strings in these fields are expected and not an error.
FAQ
| Question | Answer |
|---|---|
| Do I need a Wikimedia account or API key? | No. The Actor uses the public allimages endpoint, which requires no authentication. You can run it immediately with a date range. |
| What date format should I use? | ISO 8601 with a time component, for example 2025-01-01T00:00:00Z for the start and 2025-01-31T23:59:59Z for the end. The range is inclusive. |
| How many files can I scrape in one run? | Free users are limited to 10 files as a preview. Paid users can set maxItems up to 1,000,000 files per run. |
| Does the output include the actual image or metadata? | The output includes the direct file URL for each upload, so you can download the media yourself. The Actor returns metadata, not the binary file content. |
| Can I filter by file type or category? | The Actor filters by upload date range only. You can filter the output locally by MIME type, media type, or file name after the run completes. |
| What is the difference between this and the Wikimedia Commons API? | This Actor wraps the same public API but handles pagination, rate limits, and retries for you, and returns a flat dataset instead of nested JSON. |
| Does it scrape deleted or hidden files? | No. The allimages endpoint returns only files that are currently visible on Commons. Deleted or suppressed uploads are not included. |
| Can I scrape a specific user's uploads? | Not directly with this Actor. It filters by date range. To get a specific user's uploads, run a date range and filter the output by the uploader field. |
| What export formats are supported? | You can export the dataset to CSV, JSON, Excel, or XML from the Apify platform after the run finishes. |
| Is the data returned in a stable order? | Yes. The Actor sorts by upload timestamp, so files are returned oldest to newest within your date range. |
Related actors
Browse the full ParseForge collection for more scrapers.
๐ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
โ ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
