Wikimedia Commons Media Scraper
Pricing
from $2.85 / 1,000 results
Wikimedia Commons Media Scraper
Pull media file metadata from Wikimedia Commons by search query, category, or exact File titles. Each record carries the full image URL, thumbnail, dimensions, MIME type, byte size, license, author, and categories. Handy for media libraries, attribution, and open content research.
Pricing
from $2.85 / 1,000 results
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
1
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share

๐ผ Wikimedia Commons Media Scraper
๐ Turn Wikimedia Commons into clean media records in one run. Pull file metadata by search query, by category, or by exact File titles straight from the official MediaWiki API.
Wikimedia Commons hosts over 100 million freely licensed images, audio, and video files. This Actor reads its public MediaWiki API and returns one tidy record per media file, with the full file URL first so previews work instantly. Point it at a search term, a category, or a list of File: titles and get back the metadata that matters for reuse and attribution.
Coverage is whatever Commons exposes through prop=imageinfo: the original file URL and a thumbnail, dimensions and byte size, MIME type, uploader, plus the license, author, credit, description, and category list pulled from the file's extmetadata.
| ๐ฏ Target Audience | ๐ก Primary Use Cases |
|---|---|
| Content and media teams | Source freely licensed images with attribution |
| Wiki and dataset builders | Seed a library from a category or search |
| Researchers and archivists | Catalog open media with license and author data |
| App and bot developers | Feed image lookups without scraping HTML |
๐ What the Wikimedia Commons Media Scraper does
This Actor calls the public Wikimedia Commons MediaWiki API and returns one clean record per media file for the mode you choose:
- Search โ run a full text file search and collect matching files with full metadata.
- Category โ list every file in a category, with or without the
Category:prefix. - Titles โ fetch an exact list of
File:pages you already know.
Every record leads with the image URL, carries license and author fields parsed out of the HTML extmetadata, and includes a scrapedAt timestamp. Files that do not exist are reported as error records, not silently dropped.
๐ Data fields
Each record includes: artist, categories, credit, dateOriginal, description, descriptionUrl, height, imageUrl, licenseShortName, licenseUrl, mime, pageId, scrapedAt, size, thumbUrl, title, uploader, usageTerms, width. All 19 field names come from a real production run, so what you see here is what lands in your dataset.
๐ How to use
- Create a free Apify account using this sign-up link.
- Open the Wikimedia Commons Media Scraper.
- Pick a
mode(search, category, or titles) and fill in the matching field. - Set
maxItemsto the number of records you want. - Click Start and grab your results when the run finishes.
๐ Recommended Actors
- Have I Been Pwned Breaches Catalog Scraper
- Brawlify Brawl Stars Database Scraper
- More reference and open data Actors in the ParseForge collection
๐ก Pro Tip: browse the complete ParseForge collection.
โ ๏ธ Disclaimer: independent tool, not affiliated with Wikimedia Commons or the Wikimedia Foundation. Only publicly available data is collected.
๐ Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.