MediaWiki Extractor - Wikipedia Text, Sections & Infobox
Pricing
from $1.00 / 1,000 article extracteds
MediaWiki Extractor - Wikipedia Text, Sections & Infobox
Extracts plain text, table of contents (sections) and structured infobox data from Wikipedia pages in any language, using the official MediaWiki API. Great for LLM knowledge bases and research pipelines.
Pricing
from $1.00 / 1,000 article extracteds
Rating
0.0
(0)
Developer
Cuantic Data
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
MediaWiki Content Extractor — Wikipedia Text, Sections & Infobox
Extracts plain text, a table of contents (sections) and structured infobox data from Wikipedia pages, in the language you ask for, via the official MediaWiki API.
What it does
- Plain-text article body (
extract), with no wiki markup. - Sections (table of contents) with their nesting level.
- Structured infobox as
{ key, value }pairs, when the article has one and its template name starts with "Infobox" (see limitation below). - Supports any Wikipedia language edition (
lang, e.g.en,es,pt). - If a page doesn't exist, it doesn't fail the whole run: it's reported as
found: falsein its own item.
Who it's for
- Knowledge bases for AI agents / RAG.
- Research that needs structured data from many Wikipedia pages quickly (instead of copy-pasting by hand).
Input
| Field | Type | Required | Description |
|---|---|---|---|
titles | array of string | Yes | Exact page titles, maximum 50 per run. |
lang | string | No (default en) | Wikipedia language code. |
{ "titles": ["Argentina", "Chile"], "lang": "es" }
Output
{"title": "Argentina","resolvedTitle": "Argentina","lang": "es","found": true,"extract": "Argentina, oficialmente...","sections": [{ "level": 2, "title": "History" }, { "level": 2, "title": "Geography" }],"infobox": [{ "key": "capital", "value": "Buenos Aires" }],"sourceUrl": "https://es.wikipedia.org/wiki/Argentina"}
Pricing
Pay-per-event, provisional. Starting reference: USD 1.50 per 1,000 pages
extracted (03-plan.md §2).
How to call it
curl "https://api.apify.com/v2/acts/cuantic-data~mediawiki-content-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"titles": ["Argentina", "Chile"], "lang": "es"}'
Terms and limits
See ./TERMS.md: the public, official MediaWiki API, with an identifiable User-Agent and sequential (not parallel) requests, as Wikimedia's policy requires.
Known limitation: the infobox field is only populated when the page's
template name starts with "Infobox" (works in English and many other
wikis). On wikis that use a translated template name (e.g. "Ficha de
persona" in Spanish), infobox comes back empty — plain text and sections
still work the same in any language.
FAQ
Why is the infobox empty when the article does have one? It's the limitation above: the template name doesn't start with "Infobox" in that language/wiki. The rest of the data (text, sections) still works.
Does it work with non-Wikipedia wikis?
This version targets {lang}.wikipedia.org. Other MediaWiki wikis are
planned for a future version.
Found a bug?
Email cuanticwindows@gmail.com — we reply within 72 hours.
See build/README.md for how to run tests and publish this Actor.