# Wikimedia Commons Page Creation Log Scraper (`parseforge/wikimedia-commons-creation-log-scraper`) Actor

Pull the latest page creation log entries from Wikimedia Commons using the official API. Retrieve log IDs, page titles, page IDs, usernames, timestamps, action types, and parsed comments. Monitor new file and article uploads for moderation, archival research, or contribution tracking.

- **URL**: https://apify.com/parseforge/wikimedia-commons-creation-log-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Other, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.43 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)](https://apify.com/parseforge?fpr=vmoqkp)

### Wikimedia Commons Page Creation Log Scraper

**Scrape the Wikimedia Commons page creation log via the official API, filtered by user, title, or action, up to a million entries per run.** Each row returns the log ID, page title, creator, timestamp, and comment. No login or API key required. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons does not offer a built-in export for its page creation history. This Actor reads the public creation log through the MediaWiki API, letting you pull every new page event without writing a query. Filter by username, page title, or log action so only the entries you need land in your dataset.

| Who uses it | What they scrape Wikimedia Commons for |
|---|---|
| Digital archivists | Monitor which new media pages are being created on Commons in real time. |
| Wikimedia researchers | Study upload patterns and contributor activity across different namespaces. |
| Content moderators | Audit recent page creations by specific users or within a given title prefix. |
| GLAM professionals | Track batch uploads from a partner institution by filtering on a known username. |

### What it does

This Actor collects page creation log entries from Wikimedia Commons and returns each event as a flat row with the log ID, namespace, title, page ID, action, user, timestamp, and comment.

- 📋 **Flat row output:** every log entry is one row, ready for spreadsheets or databases.
- 🔍 **Three filters:** narrow results by username, page title, or log action before the data is written.
- ⚡ **API-first design:** reads directly from the live Wikimedia Commons API, no scraping of HTML pages.
- 📦 **Bulk exports:** download up to a million entries as CSV, JSON, Excel, or XML.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Wikimedia Commons data

**📈 Monitor new uploads by a known user.**

A GLAM institution runs the Actor daily with their Commons username to verify that a scheduled batch upload completed and to log every created page.

**🔎 Audit page creations in a specific namespace.**

A wiki administrator filters by title prefix (e.g., 'File:') to review all new file pages created over the weekend and spot potential copyright issues.

**📊 Analyze contributor activity over time.**

A researcher pulls a full month of creation log data, groups it by user, and charts upload frequency to understand community health.

**🗂️ Build an external index of Commons pages.**

A digital archivist runs the Actor weekly, exports the log as JSON, and feeds it into a custom catalog that mirrors Commons page creation events.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | The Wikimedia Commons API is public; you do not need to register an app or manage credentials. |
| **Fixed schema** | Every run returns the same fields: log ID, namespace, title, page ID, action, user, timestamp, and comment. |
| **Large volume** | Paid users can pull up to 1,000,000 log entries in a single run. |
| **Filter early** | Username, title, and action filters are sent to the API, reducing noise before storage. |

### What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "logid": 408142545,
 "ns": 6,
 "title": "File:SKQS 2938.svg",
 "pageid": 68722651,
 "logpage": 68722651,
 "type": "upload",
 "action": "overwrite",
 "user": "PacmanD",
 "timestamp": "2026-09-25T00:57:59Z",
 "comment": "調整畫框為字型字身框，輪廓不變",
 "scrapedAt": "2026-09-25T00:58:03.270Z"
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor with optional filters for username, page title, and log action, and set a maximum number of items. Filters are applied as the API is queried so only matching entries reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "maxItems": 10
}
```

A larger pull:

```json
{
 "maxItems": 200
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Set your inputs and any filters, then click **Start**.
3. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check your filters. A username that does not exist, a title that has never been created, or an invalid log action will return an empty dataset. Try running with all filters empty and a small maxItems value first to confirm the API is reachable.

**The Actor returns fewer items than my maxItems setting.**

This is normal when there are not enough log entries matching your filters. The Actor stops when the API has no more results to return. Increase maxItems only if you need a larger window of recent history.

**I see an error about an unrecognized log action.**

The MediaWiki API expects specific values for the leaction parameter. Try 'create' or 'create/create'. If you are unsure, leave the field empty to retrieve all creation-related actions without filtering.

**The run timed out.**

Large maxItems values can take time. If you are a paid user, the platform allows long-running tasks. For very large pulls, consider splitting the work into multiple runs with different title prefixes or usernames.

### FAQ

| Question | Answer |
|---|---|
| Do I need a Wikimedia account or API key to use this Actor? | No. The Wikimedia Commons API is publicly accessible. You do not need to register an application or provide any credentials. |
| What exactly is in the page creation log? | It is a record of every new page created on Wikimedia Commons, including file pages, category pages, and gallery pages. Each entry shows who created the page, when, and the edit summary they left. |
| Can I filter by date range? | The current input schema does not expose start and end date filters. The Actor retrieves the most recent entries up to your maxItems limit. If you need a specific date window, you can filter the exported dataset by the timestamp field. |
| What does the 'log action' filter do? | It lets you narrow results to a specific creation-related action. For example, you can set it to 'create' to get only direct page creations, or leave it empty to include all creation-related log events. |
| How many entries can I get in one run? | Free users are limited to 10 items as a preview. Paid users can set maxItems up to 1,000,000 and pull the full volume in a single run. |
| Which export formats are supported? | You can download your dataset as CSV, JSON, Excel, or XML from the Apify platform after the run finishes. |
| Does this Actor scrape the HTML version of the log? | No. It calls the official Wikimedia Commons API endpoint directly and parses the structured JSON response, which is faster and more reliable than screen scraping. |
| Can I get the full wikitext of each created page? | This Actor returns only the log entry metadata. If you need the page content, you would need a separate Actor that fetches page revisions by page ID. |
| Is this Actor compliant with Wikimedia's terms of use? | Yes. It uses the public API without authentication and respects the site's rate limits. You should still review Wikimedia's terms if you plan to redistribute the data. |
| What happens if I set maxItems higher than the available log entries? | The Actor will return all available entries that match your filters and then stop. You will not be charged for empty results beyond the actual data retrieved. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `leaction` (type: `string`):

Filter by a specific log action. Leave empty to retrieve all creation-related actions. Example values include 'create' or 'create/create'.

## `leuser` (type: `string`):

Filter log entries by the performing user. Leave empty to include all users.

## `letitle` (type: `string`):

Filter log entries by the target page title. Leave empty to include all pages.

## Actor input object example

```json
{
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/wikimedia-commons-creation-log-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("parseforge/wikimedia-commons-creation-log-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call parseforge/wikimedia-commons-creation-log-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/wikimedia-commons-creation-log-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KtphwanGNo9a4sfXI/builds/3hv7UPoJAPy1QeJjM/openapi.json
