# Wikimedia Commons Image Metadata Scraper (`parseforge/wikimedia-commons-image-scraper`) Actor

Extract detailed image metadata from Wikimedia Commons using its official API. Pulls fields like title, page ID, object name, description, categories, artist, credit, license, and date for any file. Ideal for researchers cataloging open media or building searchable image databases.

- **URL**: https://apify.com/parseforge/wikimedia-commons-image-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Other, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.68 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)](https://apify.com/parseforge?fpr=vmoqkp)

### Wikimedia Commons Image Metadata Scraper

**Scrape image metadata from Wikimedia Commons by search query or category, up to a million files per run.** Every record returns the title, author, license, categories, description, and EXIF dates. No API key required. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons hosts millions of freely licensed images, but downloading their metadata one file at a time is slow. This Actor reads the public Wikimedia Commons API directly, so you can pull structured metadata for thousands of images from a single search term or a whole category. It returns each image as one flat row with its title, author, license, description, and more.

| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Building a catalog of freely licensed images with full attribution metadata. |
| Content creators | Finding images they can legally use and getting the required credit line in bulk. |
| Researchers | Analyzing which topics have the most contributions on Wikimedia Commons. |
| SEO specialists | Sourcing images with verifiable license metadata for web content. |

### What it does

This Actor collects image metadata from Wikimedia Commons by search query or category and returns each file as a flat row with title, author, license, categories, and EXIF dates.

- 🔍 **Search query mode:** Feed a keyword or a structured Wikidata statement like 'haswbstatement:P180=Q5' to find matching images.
- 📁 **Category mode:** Point the Actor at a category such as 'Featured pictures on Wikimedia Commons' and pull every image inside it.
- 📄 **Flat row output:** Each image becomes one row with its title, author, license short name, categories, description, and original date.
- ⚙️ **Max items control:** Set a ceiling from 1 to 1,000,000 files so you control the size of the run.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Wikimedia Commons data

**📸 Build a reusable image library.**

A content team scrapes a category like 'Featured pictures' to get a CSV of high-quality images with verified license and author fields for their CMS.

**🔎 Audit image usage rights.**

A legal reviewer runs a search for a brand term, exports the metadata, and checks the LicenseShortName and UsageTerms columns to confirm every image is safe to use.

**📊 Study contribution patterns.**

A researcher scrapes a topic category, then groups the results by Artist and DateTimeOriginal to see who contributes when.

**🏷️ Enrich a dataset with Wikidata links.**

A data analyst uses a structured Wikidata query as the search input, then joins the resulting image titles to an external knowledge graph.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | The Wikimedia Commons API is public and requires no registration or token. |
| **License metadata** | Every row includes the license short name, usage terms, and whether attribution is required. |
| **Structured queries** | Use Wikidata statements to find images by the entities they depict, not text. |
| **Bulk export** | Save results to CSV, JSON, Excel, or XML for use in any downstream tool. |

### What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "pageid": 6580719,
 "ns": 6,
 "title": "File:Torre Belém April 2009-4a.jpg",
 "imagerepository": "local",
 "DateTime": "2013-01-29 23:58:46",
 "ObjectName": "Torre Belém April 2009-4a",
 "CommonsMetadataExtension": 1.2,
 "Assessments": "featured|valued|potd",
 "ImageDescription": "The Tower of Belém, Lisbon, Portugal. View from Northeast.",
 "DateTimeOriginal": "2009-04",
 "Credit": "Own work",
 "Artist": "Alvesgaspar",
 "LicenseShortName": "CC BY-SA 3.0",
 "UsageTerms": "Creative Commons Attribution-Share Alike 3.0",
 "AttributionRequired": true
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor from a search query or a category name, and set a max items limit so only the number of records you need reaches your dataset. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "searchQuery": "haswbstatement:P180=Q5",
 "maxItems": 10
}
```

A larger pull:

```json
{
 "searchQuery": "haswbstatement:P180=Q5",
 "maxItems": 200
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Set your inputs and any filters, then click **Start**.
3. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check that your search query or category name is spelled exactly as it appears on Wikimedia Commons. Also confirm that the max items value is set to at least 1.

**The run stopped before reaching my max items limit.**

The category or search query may have fewer images than the limit you set. Try a broader search term or a larger category.

**Some metadata fields are empty in my output.**

Not every image on Wikimedia Commons has all metadata fields filled in. Fields like Artist, Credit, or DateTimeOriginal are optional and may be blank for some files.

**I got an error about the input schema.**

Make sure you provided either a search query or a category. If both are empty, the Actor has nothing to scrape. Fill one of them and retry.

### FAQ

| Question | Answer |
|---|---|
| Do I need a Wikimedia account or API key to use this Actor? | No. The Actor calls the public Wikimedia Commons API, which does not require authentication or an API key. |
| What metadata fields does the Actor return? | It returns the page ID, title, image repository, date and time, object name, Commons metadata extension fields, categories, assessments, image description, original date and time, credit, artist, permission, author count, license short name, usage terms, attribution requirement, copyright status, and restrictions. |
| Can I scrape images by a Wikidata statement instead of a text search? | Yes. You can enter a structured query like 'haswbstatement:P180=Q5' in the search query field to find images that depict a specific Wikidata entity. |
| How many image metadata records can I get in one run? | Free users are limited to 10 items as a preview. Paid users can set the max items field up to 1,000,000. |
| Does this Actor download the actual image files? | No. It scrapes only the metadata. You get the title, description, license, and other fields, but not the image binary. |
| Can I filter by license type? | The Actor returns the license short name for every image. You can filter the output dataset after the run by that column to keep only the licenses you want. |
| What export formats are supported? | You can export your results to CSV, JSON, Excel, or XML from the Apify dataset tab. |
| How do I scrape a whole Wikimedia Commons category? | Enter the exact category name, such as 'Featured pictures on Wikimedia Commons', in the category input field and leave the search query empty. |
| Is the Actor rate-limited by Wikimedia? | The Actor respects the public API. If you request a very large number of items, the run may take longer, but it is designed to work within standard usage limits. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `searchQuery` (type: `string`):

A search term or structured query to find images on Wikimedia Commons. For example, a keyword or a Wikidata statement like 'haswbstatement:P180=Q5'.

## `category` (type: `string`):

A specific Wikimedia Commons category to scrape images from, such as 'Featured pictures on Wikimedia Commons'. Leave empty to use the search query instead.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## Actor input object example

```json
{
  "searchQuery": "haswbstatement:P180=Q5",
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "haswbstatement:P180=Q5",
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/wikimedia-commons-image-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "haswbstatement:P180=Q5",
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/wikimedia-commons-image-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "haswbstatement:P180=Q5",
  "maxItems": 10
}' |
apify call parseforge/wikimedia-commons-image-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/wikimedia-commons-image-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WPYt2NrNGDgAEE3i7/builds/aZierxh3ewpflIb2X/openapi.json
