MediaWiki Extractor - Wikipedia Text, Sections & Infobox avatar

MediaWiki Extractor - Wikipedia Text, Sections & Infobox

Pricing

from $1.00 / 1,000 article extracteds

Go to Apify Store
MediaWiki Extractor - Wikipedia Text, Sections & Infobox

MediaWiki Extractor - Wikipedia Text, Sections & Infobox

Extracts plain text, table of contents (sections) and structured infobox data from Wikipedia pages in any language, using the official MediaWiki API. Great for LLM knowledge bases and research pipelines.

Pricing

from $1.00 / 1,000 article extracteds

Rating

0.0

(0)

Developer

Cuantic Data

Cuantic Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

MediaWiki Content Extractor — Wikipedia Text, Sections & Infobox

Extracts plain text, a table of contents (sections) and structured infobox data from Wikipedia pages, in the language you ask for, via the official MediaWiki API.

What it does

  • Plain-text article body (extract), with no wiki markup.
  • Sections (table of contents) with their nesting level.
  • Structured infobox as { key, value } pairs, when the article has one and its template name starts with "Infobox" (see limitation below).
  • Supports any Wikipedia language edition (lang, e.g. en, es, pt).
  • If a page doesn't exist, it doesn't fail the whole run: it's reported as found: false in its own item.

Who it's for

  • Knowledge bases for AI agents / RAG.
  • Research that needs structured data from many Wikipedia pages quickly (instead of copy-pasting by hand).

Input

FieldTypeRequiredDescription
titlesarray of stringYesExact page titles, maximum 50 per run.
langstringNo (default en)Wikipedia language code.
{ "titles": ["Argentina", "Chile"], "lang": "es" }

Output

{
"title": "Argentina",
"resolvedTitle": "Argentina",
"lang": "es",
"found": true,
"extract": "Argentina, oficialmente...",
"sections": [{ "level": 2, "title": "History" }, { "level": 2, "title": "Geography" }],
"infobox": [{ "key": "capital", "value": "Buenos Aires" }],
"sourceUrl": "https://es.wikipedia.org/wiki/Argentina"
}

Pricing

Pay-per-event, provisional. Starting reference: USD 1.50 per 1,000 pages extracted (03-plan.md §2).

How to call it

curl "https://api.apify.com/v2/acts/cuantic-data~mediawiki-content-extractor/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"titles": ["Argentina", "Chile"], "lang": "es"}'

Terms and limits

See ./TERMS.md: the public, official MediaWiki API, with an identifiable User-Agent and sequential (not parallel) requests, as Wikimedia's policy requires.

Known limitation: the infobox field is only populated when the page's template name starts with "Infobox" (works in English and many other wikis). On wikis that use a translated template name (e.g. "Ficha de persona" in Spanish), infobox comes back empty — plain text and sections still work the same in any language.

FAQ

Why is the infobox empty when the article does have one? It's the limitation above: the template name doesn't start with "Infobox" in that language/wiki. The rest of the data (text, sections) still works.

Does it work with non-Wikipedia wikis? This version targets {lang}.wikipedia.org. Other MediaWiki wikis are planned for a future version.

Found a bug? Email cuanticwindows@gmail.com — we reply within 72 hours.


See build/README.md for how to run tests and publish this Actor.