Wikimedia Commons WikiProject Scraper
Pricing
from $3.62 / 1,000 results
Wikimedia Commons WikiProject Scraper
Scrape Wikimedia Commons WikiProject pages: name, status, shortcut, participants, sections, categories, edit history. Export to CSV, JSON, Excel or XML.
Pricing
from $3.62 / 1,000 results
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share

ποΈ Wikimedia Commons WikiProject Scraper
π Export Wikimedia Commons WikiProject pages in seconds. 20 fields per project, one row per page, up to 1,000,000 rows per run.
Wikimedia Commons hosts hundreds of WikiProjects, the volunteer groups that curate media around a topic such as insects, aviation or chemistry. Their pages hold rosters, goals, section structure and edit history, but the site offers no export for them. This Actor reads each WikiProject page through the public Commons API and returns it as one flat row with 20 fields: image, name, status, shortcut, description, participants, sections, categories, creation and last-edit dates. Export to CSV, JSON, Excel, or XML.
| π― Target Audience | π‘ Primary Use Cases |
|---|---|
| Wikimedia community organizers | Audit which WikiProjects are active and who is on their rosters |
| Digital archivists and GLAM staff | Map thematic coverage and find projects to partner with |
| Open-knowledge researchers | Study volunteer coordination across projects over time |
| Data journalists | Track when projects were founded and when they went quiet |
π What the Wikimedia Commons WikiProject Scraper does
π‘ Why it matters: WikiProject rosters and status are scattered across wikitext, subpages and hidden categories. This Actor normalizes all of it into one table.
- π Direct URLs: paste any
Commons:WikiProject ...page URLs; each becomes one row. Titles are matched case-insensitively, soWikiProject_insectsstill resolves toWikiProject Insects. - π Topic search: give search terms such as
birdsorheraldryand the Actor finds matching WikiProjects through the Commons search API, up to your max items. - π Full listing: leave both empty and it lists WikiProjects alphabetically.
- π₯ Rosters from anywhere: participants are read from
ParticipantsorMemberssections and from/Membersor/Participantssubpages. - π’ Status detection: pages tagged historical or inactive come back as
Inactive, everything else asActive. - π Edit history: creation date and creator, last edit date and editor, page size in bytes.
- π No browser, no login: it uses the open MediaWiki API with a polite delay between calls; a proxy is available but off by default.
π¬ Full Demo (π§ Coming soon)
π Output
| Field | Description |
|---|---|
πΌ imageUrl | First image on the page as a direct upload.wikimedia.org URL, or N/A |
π projectName | WikiProject name without the Commons: prefix |
π pageTitle | Full page title on Wikimedia Commons |
π url | Canonical page URL |
π pageId | MediaWiki page id |
π’ status | Active, or Inactive when the page is marked historical or inactive |
β‘ shortcut | COM: shortcuts declared on the page, or Not Disclosed |
π description | First paragraph of the page as plain text, or Not Disclosed |
π₯ participantCount | Number of participants found |
π§© sectionCount | Number of section headings |
π contentLength | Page size in bytes |
π createdAt | Timestamp of the first revision (ISO 8601) |
βοΈ createdBy | User who created the page |
π lastEditedAt | Timestamp of the latest revision (ISO 8601) |
βοΈ lastEditedBy | User of the latest revision |
π sections | Section headings in page order |
π participants | User names listed as participants or members |
π· categories | Categories of the page, hidden ones included |
π scrapedAt | When the row was collected |
β error | null on success; a page that could not be read produces a row with only this field |
Three real records from a run:
[{"imageUrl": "https://upload.wikimedia.org/wikipedia/commons/6/67/System-file-manager-brown.svg","projectName": "WikiProject Insects","pageTitle": "Commons:WikiProject Insects","url": "https://commons.wikimedia.org/wiki/Commons:WikiProject_Insects","pageId": 495944,"status": "Inactive","shortcut": "Not Disclosed","description": "This WikiProject aims primarily to document in photograph, sound and video all species of Insecta.","participantCount": 5,"sectionCount": 10,"contentLength": 4923,"createdAt": "2006-01-04T21:54:08Z","createdBy": "TeunSpaans","lastEditedAt": "2026-07-16T20:15:46Z","lastEditedBy": "TenshiBot","sections": ["Article titles and common names","Categories","Status","Quality","Resources","Requests","Sister WikiProjects","Participants","Sample articles","Cleanup"],"participants": ["TeunSpaans","Svdmolen","Kulac","Keith Edkins","Giancarlodessi"],"categories": ["Inactive Commons pages"],"scrapedAt": "2026-09-08T01:11:55.043Z","error": null},{"imageUrl": "https://upload.wikimedia.org/wikipedia/commons/2/2e/Gnome-applications-science.svg","projectName": "WikiProject Chemistry","pageTitle": "Commons:WikiProject Chemistry","url": "https://commons.wikimedia.org/wiki/Commons:WikiProject_Chemistry","pageId": 1473569,"status": "Active","shortcut": "COM:CHEM","description": "This WikiProject is the Commons branch of the WikiProject Chemistry and other language Wikipedia workgroups (see interwiki links in the left column).","participantCount": 0,"sectionCount": 2,"contentLength": 2554,"createdAt": "2006-12-18T17:24:45Z","createdBy": "Benjah-bmm27","lastEditedAt": "2023-11-03T16:05:50Z","lastEditedBy": "Ameisenigel","sections": ["To-do list","Quality assurance"],"participants": [],"categories": ["Commons WikiProjects","Chemistry","WikiProject Chemistry"],"scrapedAt": "2026-09-08T01:11:56.458Z","error": null},{"imageUrl": "https://upload.wikimedia.org/wikipedia/commons/3/37/People_icon.svg","projectName": "WikiProject Birds","pageTitle": "Commons:WikiProject Birds","url": "https://commons.wikimedia.org/wiki/Commons:WikiProject_Birds","pageId": 1787159,"status": "Active","shortcut": "Not Disclosed","description": "This WikiProject descends from commons:WikiProject Tree of Life. It aims primarily to document in photograph, sound and video all species of birds.","participantCount": 16,"sectionCount": 13,"contentLength": 5717,"createdAt": "2007-03-14T21:41:49Z","createdBy": "Tony Wills","lastEditedAt": "2024-11-07T12:29:20Z","lastEditedBy": "Manojk","sections": ["Overview","New to Commons?","Fill the gaps!","Categories","Status","Quality","Resources","Requests","Identification","Sister WikiProjects","Participants","Sample articles","Cleanup"],"participants": ["Tony Wills","Wsiegmund","Anniolek","Mindaugas Urbonas","Dysmorodrepanis","Tigershrike","MeegsC","Shyamal","Dger","innotata","Kersti Nebelsiek","Llywelyn2000","Jcfidy","Sharadapte","ΰ€Έΰ₯ΰ€¬ΰ₯ΰ€§ ΰ€ΰ₯ΰ€²ΰ€ΰ€°ΰ₯ΰ€£ΰ₯","Manojk"],"categories": ["Commons WikiProjects","WikiProject Tree of Life"],"scrapedAt": "2026-09-08T01:11:58.988Z","error": null}]
A value the page does not provide comes back as Not Disclosed (shortcut, description) or N/A (image). Lists that are genuinely empty, such as a project without a roster, come back as [].
β¨ Why choose this Actor
| What you get | |
|---|---|
| API-backed, not screen-scraped | Reads wikitext, revisions and categories through the MediaWiki API, so layout changes on the site do not break it. |
| One page, one row | Every WikiProject becomes exactly one row with the same 20 keys, ready for a spreadsheet or a join on pageId. |
| Rosters that are actually complete | Reads `{{User |
| Three ways in | Exact URLs, topic search, or an alphabetical listing when you give no input at all. |
| Polite by design | Serial requests with a delay and a descriptive User-Agent, as Wikimedia's API etiquette asks. |
π How it compares to alternatives
| Approach | What you get | What you give up |
|---|---|---|
| This Actor | Flat rows with rosters, status, dates and sections; search-based discovery; CSV/JSON/Excel/XML export | Media files themselves are not downloaded |
| Manual browsing | Full context of each page | Hours of copy-paste for a handful of projects, no dates or counts |
| Raw MediaWiki API calls | Everything, if you script it | Four requests per page, wikitext parsing, template and subpage handling written by you |
| Generic Wikipedia scrapers | Article text | No WikiProject roster or status logic, no Commons namespace support |
π How to use
- Create a free Apify account; it comes with $5 of monthly credit.
- Open the Wikimedia Commons WikiProject Scraper.
- Keep the five example URLs or paste your own WikiProject URLs, or type search terms such as
birds. - Set Max items (free plans return up to 10 rows as a preview) and click Start.
- Open the Dataset tab and export as CSV, Excel, JSON, or XML, or fetch it through the Apify API.
Leave both URLs and search terms empty to list WikiProjects alphabetically. Tick Include subpages if you also want pages such as Commons:WikiProject Aviation/Members as their own rows.
πΌ Business use cases
π€ GLAM partnership scouting
A museum searches paintings, architecture and photography, filters the rows to status = Active, and contacts the projects with the largest participantCount about a batch upload.
π Community health reporting
A Wikimedia affiliate runs the full listing monthly and charts lastEditedAt and participantCount per project to spot WikiProjects that need attention.
πΊοΈ Coverage mapping
An archive compares the categories and sections of every WikiProject against its own collection taxonomy to decide where its media will be found and curated.
π§βπ» Contributor outreach
A program coordinator merges the participants lists of related projects to build an invite list for an edit-a-thon, with createdBy and lastEditedBy as first contacts.
π Automating Wikimedia Commons WikiProject Scraper
- Make / Zapier: trigger a run on a schedule and append new rows to a sheet or CRM.
- Slack: post a message when a tracked project flips from
ActivetoInactive. - Airbyte: load the dataset into a warehouse and join on
pageIdacross runs. - GitHub Actions: run the Actor in CI and commit the CSV to a data repository.
- Google Drive: export to Excel and drop the file into a shared folder after every run.
π Beyond business use cases
- Research: longitudinal studies of volunteer coordination on Commons using
createdAt,lastEditedAtand roster sizes. - Personal: find the right WikiProject to join for your hobby and see who is active there.
- Non-profit: document which topics have organized curation and which are gaps.
- Experimentation: feed
descriptionandsectionsinto a classifier to cluster projects by scope.
π€ Ask an AI assistant about this scraper
Give an AI agent live access to the Actor through the Model Context Protocol:
$claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/wikimedia-commons-wikiproject-scraper"
Then ask in plain language, for example "list the active bird-related WikiProjects on Commons with their participant counts".
β Frequently Asked Questions
β What is a Wikimedia Commons WikiProject?
A group of contributors who organize media curation around a topic. Each project lives at Commons:WikiProject <Topic> and typically lists goals, participants, tasks and categories.
β Do I need a Wikimedia account or API key?
No. The Actor reads the public MediaWiki API without authentication.
β How do I get a specific project?
Paste its URL as a start URL. Capitalization after the first letter is forgiven: WikiProject_insects resolves to WikiProject Insects.
β How do search terms work?
Each term runs a Commons search restricted to WikiProject titles. Matching projects are scraped in relevance order until Max items is reached. Search terms take precedence over start URLs.
β Why is participants empty for some projects?
Some projects, for example WikiProject Chemistry, have no roster on the page or on a /Participants or /Members subpage. The row is still returned with participantCount: 0.
β What does status: Inactive mean?
The page carries a historical or inactive tag or sits in the "Inactive Commons pages" category. Everything else is reported as Active.
β Where does imageUrl come from?
The first [[File:...]] in the page wikitext, or the first image the page renders, turned into a direct upload.wikimedia.org URL. Pages without images return N/A.
β Are subpages included?
Not by default: one row is one WikiProject. Turn on Include subpages to also return pages such as Commons:WikiProject Aviation/Members from search or the listing.
β Does it download media files?
No. It extracts metadata and text about the project page only.
β Is there a rate limit?
The Actor makes three to four API calls per page in series with a short pause, in line with Wikimedia's API etiquette. Use Max items to bound the run.
β Can I run it on a free plan?
Yes. Free plans return up to 10 rows per run as a preview; paid plans return up to 1,000,000.
β What happens if a page does not exist?
That page produces a row containing only an error field with the reason; the other pages are unaffected.
π Integrate with any app
Run the Actor from the Apify API or the JavaScript and Python clients, schedule it in the Apify Console, and connect the dataset to Make, Zapier, Airbyte, Google Sheets, Slack or any webhook target.
π Recommended Actors
- Wikimedia Commons Category Hierarchy Scraper: walk a category tree and its subcategories.
- Wikimedia Commons Media Scraper: file metadata, licenses and download URLs.
- Wikimedia Commons Users Scraper: profiles and edit counts for the participants you find here.
- Wikimedia Commons Page History Scraper: full revision history of any page.
π‘ Pro Tip: browse the complete ParseForge collection for more data sources.
π Need Help? Open our contact form with your run ID, your input, and what you expected.
β οΈ Disclaimer: This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Wikimedia Foundation. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of use and applicable data-protection laws, including GDPR, CCPA, and PIPL.