Wikipedia Scraper - Full Article Text and Search avatar

Wikipedia Scraper - Full Article Text and Search

Pricing

$1.00 / 1,000 page returneds

Go to Apify Store
Wikipedia Scraper - Full Article Text and Search

Wikipedia Scraper - Full Article Text and Search

Search Wikipedia by keyword, or fetch exact article titles. MediaWiki gives you the intro paragraph, not the whole article. Full text here is one checkbox. Each page gives plain-text extracts, thumbnails, categories, page IDs and URLs. 50 titles a batch, any language. $1.00 per 1,000 pages.

Pricing

$1.00 / 1,000 page returneds

Rating

5.0

(1)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

0

Monthly active users

10 hours ago

Last modified

Share

Wikipedia Scraper: search results or full article text, in any language edition

It runs one of two jobs. Give it a keyword and you get matching articles with a snippet, word count and link. Give it exact article titles and you get the article text, the lead image and the categories instead.

The two modes return different fields, and that catches people out: search rows have a snippet but no article text, and title rows have the text but no search ranking. Pick the mode that matches what you actually need.

InputA search query, or a list of exact article titles
OutputOne row per article
Ceiling5,000 pages per run
Account neededNone, and no API key
Price$1.00 per 1,000 pages, flat on every plan

📚 What Wikipedia Scraper does

In search mode it runs Wikipedia's own search and pages through the results until it reaches your limit. Each row carries the page id, the title, a plain-text snippet with the highlighting stripped, the article's word count and size, and when it was last edited.

In page-titles mode it fetches the articles you named. By default you get the intro paragraphs; tick Fetch full article text and you get the whole article as plain text instead. Those rows also carry the lead image and the visible category names.

Pick any language edition with language. The code sets the host, so fr reads the French Wikipedia and returns French articles.

📥 What you give it

{
"searchQuery": "machine learning",
"language": "en",
"maxItems": 50
}
FieldDefaultWhat it is
searchQuerybox starts at machine learningKeywords to search for. Leave it empty to use pageTitles instead.
pageTitles[], box starts at Apify, Web scrapingExact article titles. Use this or a search query.
fullTextfalseWhole article instead of the intro. Only applies in page-titles mode, and is ignored in search mode.
languageenThe language edition code: en, fr, de, es, ja and so on.
maxItems50How many pages to return, up to 5,000. In search mode it pages until it gets there.
notionConnectornoneOptional. Write every page into your own Notion as well as the dataset.
notionParentIdnoneOptional. The Notion data source ID to write into.
proxyConfigurationoffOptional, and off by default because a normal run does not need it.

Both boxes start filled in, and a filled searchQuery wins. If you want titles mode, clear the search box first, or the sample titles are ignored.

Titles have to be exact, including capitalisation, exactly as they appear in the article URL.

📤 What you get back

A real search row:

{
"ok": true,
"mode": "search",
"pageid": 233488,
"title": "Machine learning",
"url": "https://en.wikipedia.org/?curid=233488",
"snippet": "Machine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn",
"wordcount": 16471,
"size": 145190,
"timestamp": "2026-09-13T21:33:50Z"
}

And a real page-titles row, with the extract and category list cut short:

{
"ok": true,
"mode": "page",
"fullText": false,
"pageid": 11188,
"title": "French Revolution",
"url": "https://en.wikipedia.org/wiki/French_Revolution",
"extract": "The French Revolution was a period of political and societal change in France that began with the Estates General of 1789 ...",
"thumbnail": "https://upload.wikimedia.org/wikipedia/commons/thumb/5/57/Anonymous_-_Prise_de_la_Bastille.jpg/500px-...",
"categories": ["1789 conflicts", "Atlantic Revolutions", "Civil wars in France", "..."]
}
FieldWhat it is
modesearch or page. It tells you which set of fields to expect.
pageidWikipedia's own numeric id. Stable across renames, so use it as your key.
snippetSearch mode only. The matched text with the markup stripped. There is no extract on a search row.
extractPage mode only. Plain text, intro or whole article depending on fullText. null when the API returned none.
thumbnailPage mode only. The lead image, up to 400 px, or null when the article has no image.
categoriesPage mode only. The visible category names. Hidden maintenance categories are left out.
wordcount, sizeSearch mode only. Article length in words and bytes, useful for ranking what to read.
timestampSearch mode only. When the article was last edited.

🧾 Reading the output

Two kinds of row can land in your dataset.

RowHow to spot itCharged
A pageok: true and a pageidyes
A diagnosticok: false and an errorCodeno
CodeWhat it means
BAD_INPUTNeither a search query nor any page titles were given.
NOT_FOUNDThat exact title does not exist on that language edition. One row per missing title, in the order you listed them.
NO_RESULTSThe search ran and matched nothing, or no title in the list resolved.
RATE_LIMITEDWikipedia asked for a slower pace. Re-run with a smaller maxItems.
SERVER_ERRORWikipedia answered with a server error. Usually passes on its own.
BLOCKEDThe request was not served that time. Run it again.
NETWORKWikipedia could not be reached. A misspelled language code also lands here, because the host does not exist.

Rows arrive at the end, not as they are found. Everything is fetched first and written afterwards, so a large full-text run shows nothing in the dataset until it is done.

The default table view has no ok column, so a NOT_FOUND row looks like a blank line. Switch to All fields or export as JSON.

▶️ How to run it

  1. Open Wikipedia Scraper and click Try for free.
  2. For a search, type into Search query and leave Page titles alone.
  3. For articles, clear Search query, then put exact titles into Page titles.
  4. Set Wikipedia language and Max items, tick Fetch full article text if you want the whole article, then click Start.
  5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost?

$1.00 per 1,000 pages, which is $0.001 each. Flat on every Apify plan, no volume tiers.

You pay per page row delivered. Duplicate page ids are dropped before they reach you, diagnostic rows are not charged, and a run that finds nothing costs you nothing.

💡 What people use it for

  • Building a plain-text corpus on a subject, with fullText on and a list of titles.
  • Grabbing the intro paragraph and lead image for a glossary or an internal wiki.
  • Ranking what to read on a topic, using search mode and sorting on wordcount.
  • Pulling the same article across language editions to compare how each one covers it.

🚧 What it does not do

  • Search rows have no article text. If you need the text, take the titles from a search run and feed them into a second run in page-titles mode.
  • A long title list thins out. Asking for many titles at once, the API stops returning extracts partway through the batch, so some rows come back with extract: null and no categories while still counting as pages. Keep a batch under about 20 titles and they come back complete.
  • No infoboxes, references, tables or wikitext. Plain text only.
  • No revision history, no edit diffs, no talk pages.
  • Titles are exact. A near miss returns NOT_FOUND rather than the closest article.
  • No images beyond the lead thumbnail, and nothing is downloaded.
  • The run finishes even when every request failed, so read ok rather than the run status.
  • Hidden maintenance categories are left out, on purpose. You get the ones a reader sees.

🧭 Which research scraper do you need?

If you wantUse
Wikipedia search results or article textThis one
Preprints with abstracts and PDF linksarXiv Scraper
Published papers with DOIs and citation countsOpenAlex Scraper
Crossref metadata for a DOICrossref Scraper
Developer articles and their tagsDEV.to Scraper

❓ Questions people ask

Do I need an API key? No. Wikipedia's API is open and nothing here needs an account.

Why is extract empty on a search row? Because search mode never returns one. Take the titles and run them in page-titles mode.

Can I get a non-English Wikipedia? Yes. Set language to the edition code and both modes follow it.

What happens if I fill in both boxes? Search mode runs and the titles are ignored. Clear the search box to use titles.

Can I get the whole article? Yes, tick Fetch full article text. It only affects page-titles mode.

Is scraping Wikipedia legal? Wikipedia's content is published under a free licence and its API is open to the public. Check the licence terms for how you attribute reuse. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID, the query or titles you used, and the language code. The errorCode on the diagnostic row usually names the problem on its own.