Wikimedia Commons Page Creation Log Scraper avatar

Wikimedia Commons Page Creation Log Scraper

Pricing

from $0.43 / 1,000 result items

Go to Apify Store
Wikimedia Commons Page Creation Log Scraper

Wikimedia Commons Page Creation Log Scraper

Pull the latest page creation log entries from Wikimedia Commons using the official API. Retrieve log IDs, page titles, page IDs, usernames, timestamps, action types, and parsed comments. Monitor new file and article uploads for moderation, archival research, or contribution tracking.

Pricing

from $0.43 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

ParseForge

Wikimedia Commons Page Creation Log Scraper

Scrape the Wikimedia Commons page creation log via the official API, filtered by user, title, or action, up to a million entries per run. Each row returns the log ID, page title, creator, timestamp, and comment. No login or API key required. Export to CSV, JSON, Excel, or XML.

Wikimedia Commons does not offer a built-in export for its page creation history. This Actor reads the public creation log through the MediaWiki API, letting you pull every new page event without writing a query. Filter by username, page title, or log action so only the entries you need land in your dataset.

Who uses itWhat they scrape Wikimedia Commons for
Digital archivistsMonitor which new media pages are being created on Commons in real time.
Wikimedia researchersStudy upload patterns and contributor activity across different namespaces.
Content moderatorsAudit recent page creations by specific users or within a given title prefix.
GLAM professionalsTrack batch uploads from a partner institution by filtering on a known username.

What it does

This Actor collects page creation log entries from Wikimedia Commons and returns each event as a flat row with the log ID, namespace, title, page ID, action, user, timestamp, and comment.

  • 📋 Flat row output: every log entry is one row, ready for spreadsheets or databases.
  • 🔍 Three filters: narrow results by username, page title, or log action before the data is written.
  • ⚡ API-first design: reads directly from the live Wikimedia Commons API, no scraping of HTML pages.
  • 📦 Bulk exports: download up to a million entries as CSV, JSON, Excel, or XML.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with Wikimedia Commons data

📈 Monitor new uploads by a known user.

A GLAM institution runs the Actor daily with their Commons username to verify that a scheduled batch upload completed and to log every created page.

🔎 Audit page creations in a specific namespace.

A wiki administrator filters by title prefix (e.g., 'File:') to review all new file pages created over the weekend and spot potential copyright issues.

📊 Analyze contributor activity over time.

A researcher pulls a full month of creation log data, groups it by user, and charts upload frequency to understand community health.

🗂️ Build an external index of Commons pages.

A digital archivist runs the Actor weekly, exports the log as JSON, and feeds it into a custom catalog that mirrors Commons page creation events.

Why choose this scraper

What you get
No API keyThe Wikimedia Commons API is public; you do not need to register an app or manage credentials.
Fixed schemaEvery run returns the same fields: log ID, namespace, title, page ID, action, user, timestamp, and comment.
Large volumePaid users can pull up to 1,000,000 log entries in a single run.
Filter earlyUsername, title, and action filters are sent to the API, reducing noise before storage.

What a Wikimedia Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

{
"logid": 408142545,
"ns": 6,
"title": "File:SKQS 2938.svg",
"pageid": 68722651,
"logpage": 68722651,
"type": "upload",
"action": "overwrite",
"user": "PacmanD",
"timestamp": "2026-09-25T00:57:59Z",
"comment": "調整畫框為字型字身框,輪廓不變",
"scrapedAt": "2026-09-25T00:58:03.270Z"
}

Every value above comes from a real run. A field a record does not have comes back as null.

Configure the run

Drive the Actor with optional filters for username, page title, and log action, and set a maximum number of items. Filters are applied as the API is queried so only matching entries reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

{
"maxItems": 10
}

A larger pull:

{
"maxItems": 200
}

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Set your inputs and any filters, then click Start.
  3. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to Wikimedia Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$undefined

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check your filters. A username that does not exist, a title that has never been created, or an invalid log action will return an empty dataset. Try running with all filters empty and a small maxItems value first to confirm the API is reachable.

The Actor returns fewer items than my maxItems setting.

This is normal when there are not enough log entries matching your filters. The Actor stops when the API has no more results to return. Increase maxItems only if you need a larger window of recent history.

I see an error about an unrecognized log action.

The MediaWiki API expects specific values for the leaction parameter. Try 'create' or 'create/create'. If you are unsure, leave the field empty to retrieve all creation-related actions without filtering.

The run timed out.

Large maxItems values can take time. If you are a paid user, the platform allows long-running tasks. For very large pulls, consider splitting the work into multiple runs with different title prefixes or usernames.

FAQ

QuestionAnswer
Do I need a Wikimedia account or API key to use this Actor?No. The Wikimedia Commons API is publicly accessible. You do not need to register an application or provide any credentials.
What exactly is in the page creation log?It is a record of every new page created on Wikimedia Commons, including file pages, category pages, and gallery pages. Each entry shows who created the page, when, and the edit summary they left.
Can I filter by date range?The current input schema does not expose start and end date filters. The Actor retrieves the most recent entries up to your maxItems limit. If you need a specific date window, you can filter the exported dataset by the timestamp field.
What does the 'log action' filter do?It lets you narrow results to a specific creation-related action. For example, you can set it to 'create' to get only direct page creations, or leave it empty to include all creation-related log events.
How many entries can I get in one run?Free users are limited to 10 items as a preview. Paid users can set maxItems up to 1,000,000 and pull the full volume in a single run.
Which export formats are supported?You can download your dataset as CSV, JSON, Excel, or XML from the Apify platform after the run finishes.
Does this Actor scrape the HTML version of the log?No. It calls the official Wikimedia Commons API endpoint directly and parses the structured JSON response, which is faster and more reliable than screen scraping.
Can I get the full wikitext of each created page?This Actor returns only the log entry metadata. If you need the page content, you would need a separate Actor that fetches page revisions by page ID.
Is this Actor compliant with Wikimedia's terms of use?Yes. It uses the public API without authentication and respects the site's rate limits. You should still review Wikimedia's terms if you plan to redistribute the data.
What happens if I set maxItems higher than the available log entries?The Actor will return all available entries that match your filters and then stop. You will not be charged for empty results beyond the actual data retrieved.

Browse the full ParseForge collection for more scrapers.

🆘 Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.