Wikimedia Commons Media Scraper avatar

Wikimedia Commons Media Scraper

Pricing

from $2.85 / 1,000 results

Go to Apify Store
Wikimedia Commons Media Scraper

Wikimedia Commons Media Scraper

Pull media file metadata from Wikimedia Commons by search query, category, or exact File titles. Each record carries the full image URL, thumbnail, dimensions, MIME type, byte size, license, author, and categories. Handy for media libraries, attribution, and open content research.

Pricing

from $2.85 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

1

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

ParseForge Banner

๐Ÿ–ผ Wikimedia Commons Media Scraper

๐Ÿš€ Turn Wikimedia Commons into clean media records in one run. Pull file metadata by search query, by category, or by exact File titles straight from the official MediaWiki API.

Wikimedia Commons hosts over 100 million freely licensed images, audio, and video files. This Actor reads its public MediaWiki API and returns one tidy record per media file, with the full file URL first so previews work instantly. Point it at a search term, a category, or a list of File: titles and get back the metadata that matters for reuse and attribution.

Coverage is whatever Commons exposes through prop=imageinfo: the original file URL and a thumbnail, dimensions and byte size, MIME type, uploader, plus the license, author, credit, description, and category list pulled from the file's extmetadata.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Content and media teamsSource freely licensed images with attribution
Wiki and dataset buildersSeed a library from a category or search
Researchers and archivistsCatalog open media with license and author data
App and bot developersFeed image lookups without scraping HTML

๐Ÿ“‹ What the Wikimedia Commons Media Scraper does

This Actor calls the public Wikimedia Commons MediaWiki API and returns one clean record per media file for the mode you choose:

  • Search โ€” run a full text file search and collect matching files with full metadata.
  • Category โ€” list every file in a category, with or without the Category: prefix.
  • Titles โ€” fetch an exact list of File: pages you already know.

Every record leads with the image URL, carries license and author fields parsed out of the HTML extmetadata, and includes a scrapedAt timestamp. Files that do not exist are reported as error records, not silently dropped.

๐Ÿ“Š Data fields

Each record includes: artist, categories, credit, dateOriginal, description, descriptionUrl, height, imageUrl, licenseShortName, licenseUrl, mime, pageId, scrapedAt, size, thumbUrl, title, uploader, usageTerms, width. All 19 field names come from a real production run, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. Create a free Apify account using this sign-up link.
  2. Open the Wikimedia Commons Media Scraper.
  3. Pick a mode (search, category, or titles) and fill in the matching field.
  4. Set maxItems to the number of records you want.
  5. Click Start and grab your results when the run finishes.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection.

โš ๏ธ Disclaimer: independent tool, not affiliated with Wikimedia Commons or the Wikimedia Foundation. Only publicly available data is collected.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.