Wikimedia Commons Image Metadata Scraper
Pricing
from $2.68 / 1,000 result items
Wikimedia Commons Image Metadata Scraper
Extract detailed image metadata from Wikimedia Commons using its official API. Pulls fields like title, page ID, object name, description, categories, artist, credit, license, and date for any file. Ideal for researchers cataloging open media or building searchable image databases.
Pricing
from $2.68 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Wikimedia Commons Image Metadata Scraper
Scrape image metadata from Wikimedia Commons by search query or category, up to a million files per run. Every record returns the title, author, license, categories, description, and EXIF dates. No API key required. Export to CSV, JSON, Excel, or XML.
Wikimedia Commons hosts millions of freely licensed images, but downloading their metadata one file at a time is slow. This Actor reads the public Wikimedia Commons API directly, so you can pull structured metadata for thousands of images from a single search term or a whole category. It returns each image as one flat row with its title, author, license, description, and more.
| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Building a catalog of freely licensed images with full attribution metadata. |
| Content creators | Finding images they can legally use and getting the required credit line in bulk. |
| Researchers | Analyzing which topics have the most contributions on Wikimedia Commons. |
| SEO specialists | Sourcing images with verifiable license metadata for web content. |
What it does
This Actor collects image metadata from Wikimedia Commons by search query or category and returns each file as a flat row with title, author, license, categories, and EXIF dates.
- 🔍 Search query mode: Feed a keyword or a structured Wikidata statement like 'haswbstatement:P180=Q5' to find matching images.
- 📁 Category mode: Point the Actor at a category such as 'Featured pictures on Wikimedia Commons' and pull every image inside it.
- 📄 Flat row output: Each image becomes one row with its title, author, license short name, categories, description, and original date.
- ⚙️ Max items control: Set a ceiling from 1 to 1,000,000 files so you control the size of the run.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with Wikimedia Commons data
📸 Build a reusable image library.
A content team scrapes a category like 'Featured pictures' to get a CSV of high-quality images with verified license and author fields for their CMS.
🔎 Audit image usage rights.
A legal reviewer runs a search for a brand term, exports the metadata, and checks the LicenseShortName and UsageTerms columns to confirm every image is safe to use.
📊 Study contribution patterns.
A researcher scrapes a topic category, then groups the results by Artist and DateTimeOriginal to see who contributes when.
🏷️ Enrich a dataset with Wikidata links.
A data analyst uses a structured Wikidata query as the search input, then joins the resulting image titles to an external knowledge graph.
Why choose this scraper
| What you get | |
|---|---|
| No API key | The Wikimedia Commons API is public and requires no registration or token. |
| License metadata | Every row includes the license short name, usage terms, and whether attribution is required. |
| Structured queries | Use Wikidata statements to find images by the entities they depict, not text. |
| Bulk export | Save results to CSV, JSON, Excel, or XML for use in any downstream tool. |
What a Wikimedia Commons record looks like
Every record returns as one flat JSON row. Here is a real one from a run:
{"pageid": 6580719,"ns": 6,"title": "File:Torre Belém April 2009-4a.jpg","imagerepository": "local","DateTime": "2013-01-29 23:58:46","ObjectName": "Torre Belém April 2009-4a","CommonsMetadataExtension": 1.2,"Assessments": "featured|valued|potd","ImageDescription": "The Tower of Belém, Lisbon, Portugal. View from Northeast.","DateTimeOriginal": "2009-04","Credit": "Own work","Artist": "Alvesgaspar","LicenseShortName": "CC BY-SA 3.0","UsageTerms": "Creative Commons Attribution-Share Alike 3.0","AttributionRequired": true}
Every value above comes from a real run. A field a record does not have comes back as null.
Configure the run
Drive the Actor from a search query or a category name, and set a max items limit so only the number of records you need reaches your dataset. The Input tab lists every parameter.
A first run with the defaults:
{"searchQuery": "haswbstatement:P180=Q5","maxItems": 10}
A larger pull:
{"searchQuery": "haswbstatement:P180=Q5","maxItems": 200}
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.
Run it
- Create a free Apify account with $5 in credit.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
$undefined
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check that your search query or category name is spelled exactly as it appears on Wikimedia Commons. Also confirm that the max items value is set to at least 1.
The run stopped before reaching my max items limit.
The category or search query may have fewer images than the limit you set. Try a broader search term or a larger category.
Some metadata fields are empty in my output.
Not every image on Wikimedia Commons has all metadata fields filled in. Fields like Artist, Credit, or DateTimeOriginal are optional and may be blank for some files.
I got an error about the input schema.
Make sure you provided either a search query or a category. If both are empty, the Actor has nothing to scrape. Fill one of them and retry.
FAQ
| Question | Answer |
|---|---|
| Do I need a Wikimedia account or API key to use this Actor? | No. The Actor calls the public Wikimedia Commons API, which does not require authentication or an API key. |
| What metadata fields does the Actor return? | It returns the page ID, title, image repository, date and time, object name, Commons metadata extension fields, categories, assessments, image description, original date and time, credit, artist, permission, author count, license short name, usage terms, attribution requirement, copyright status, and restrictions. |
| Can I scrape images by a Wikidata statement instead of a text search? | Yes. You can enter a structured query like 'haswbstatement:P180=Q5' in the search query field to find images that depict a specific Wikidata entity. |
| How many image metadata records can I get in one run? | Free users are limited to 10 items as a preview. Paid users can set the max items field up to 1,000,000. |
| Does this Actor download the actual image files? | No. It scrapes only the metadata. You get the title, description, license, and other fields, but not the image binary. |
| Can I filter by license type? | The Actor returns the license short name for every image. You can filter the output dataset after the run by that column to keep only the licenses you want. |
| What export formats are supported? | You can export your results to CSV, JSON, Excel, or XML from the Apify dataset tab. |
| How do I scrape a whole Wikimedia Commons category? | Enter the exact category name, such as 'Featured pictures on Wikimedia Commons', in the category input field and leave the search query empty. |
| Is the Actor rate-limited by Wikimedia? | The Actor respects the public API. If you request a very large number of items, the run may take longer, but it is designed to work within standard usage limits. |
Related actors
Browse the full ParseForge collection for more scrapers.
🆘 Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
⚠️ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
