Wikipedia Scraper avatar

Wikipedia Scraper

Pricing

from $0.60 / 1,000 article extracteds

Go to Apify Store
Wikipedia Scraper

Wikipedia Scraper

Search and extract Wikipedia articles — titles, summaries, full content, categories, and images. Uses the free MediaWiki API.

Pricing

from $0.60 / 1,000 article extracteds

Rating

0.0

(0)

Developer

Automation Lab

Automation Lab

Maintained by Community

Actor stats

1

Bookmarked

48

Total users

5

Monthly active users

10 days ago

Last modified

Categories

Share

Search Wikipedia by keyword or fetch known pages by exact title. Export summaries, optional full plaintext, categories, Wikidata IDs, page descriptions, URLs, and input provenance from any of Wikipedia's 300+ language editions.

What does Wikipedia Scraper do?

Wikipedia Scraper uses the official MediaWiki API to search by topic, fetch exact article titles, or combine both workflows. Every result includes the introductory extract, categories, Wikidata ID and short page description when available, page metadata, and clear provenance. Enable includeFullContent to also receive the full article as plain text in fullContent, ready for RAG and knowledge-base datasets.

Keyword results follow Wikipedia's relevance ranking and include searchQuery and searchRank. Exact-title results include requestedTitle and automatically resolve redirects. Mixed inputs are deduplicated by Wikipedia page ID, so an overlapping page is emitted and charged only once while matchedSearchQueries and requestedTitles retain all matching inputs.

Who is it for?

  • 🎓 Academic researchers — extracting structured knowledge from Wikipedia articles at scale
  • 🤖 NLP engineers — building training datasets from Wikipedia text and metadata
  • 📊 Data analysts — collecting factual data and statistics from Wikipedia pages
  • 💻 App developers — enriching applications with Wikipedia content and summaries
  • 📝 Content creators — gathering reference material and structured facts for writing

Why scrape Wikipedia?

Wikipedia is the world's largest free encyclopedia with over 60 million articles across 300+ languages. It's a primary source for:

  • Knowledge base construction — build reference datasets for AI training, chatbots, or research databases
  • LLM and RAG pipelines — feed clean, structured article text into retrieval-augmented generation systems, fine-tuning datasets, or AI agent knowledge bases
  • Content enrichment — add Wikipedia summaries to product catalogs, educational platforms, or content management systems
  • Research and analysis — analyze article coverage, word counts, and edit patterns across topics
  • Multilingual data — gather information in any language Wikipedia supports
  • SEO and content strategy — understand topic coverage and find content gaps

How much does it cost to scrape Wikipedia?

Wikipedia Scraper uses pay-per-event pricing:

EventPrice
Run started$0.001
Article extracted$0.001 per article

Example costs:

  • 10 articles on "machine learning": ~$0.011
  • 100 articles on "history": ~$0.101
  • 500 articles across 5 keywords: ~$0.506

Platform costs are minimal — a typical run uses under $0.001 in compute. Wikipedia's API is fast and does not require proxies.

Input parameters

ParameterTypeDescriptionDefault
searchQueriesstring[]Keywords to search on Wikipedia. Each keyword runs a separate ranked search.Optional*
articleTitlesstring[]Exact Wikipedia article titles to fetch directly; redirects are resolved.Optional*
languagestringWikipedia language code (e.g., en, de, fr, es, ja, zh)"en"
maxResultsPerSearchintegerMaximum articles per keyword (1–500); does not limit exact titles50
includeFullContentbooleanAdd the full article plaintext in fullContent; extract remains the introductory summaryfalse

*Provide at least one non-blank value in searchQueries or articleTitles.

Mixed input example

{
"searchQueries": ["history of computing"],
"articleTitles": ["Ada Lovelace", "Alan Turing"],
"language": "en",
"maxResultsPerSearch": 20,
"includeFullContent": true
}

For exact pages only, omit searchQueries and provide articleTitles.

Output example

Each article is returned as a JSON object:

{
"pageId": 1164,
"title": "Artificial intelligence",
"extract": "Artificial intelligence (AI) is the capability of computational systems to perform tasks typically associated with human intelligence...",
"fullContent": "Artificial intelligence (AI), in its broadest sense, is intelligence exhibited by machines...",
"url": "https://en.wikipedia.org/wiki/Artificial_intelligence",
"wordCount": 26473,
"size": 266568,
"lastEdited": "2025-01-15T12:00:00Z",
"thumbnail": "https://upload.wikimedia.org/wikipedia/commons/thumb/...",
"categories": ["Artificial intelligence", "Computational fields of study"],
"wikidataId": "Q11660",
"description": "intelligence demonstrated by machines",
"searchQuery": "artificial intelligence",
"searchRank": 1,
"matchedSearchQueries": ["artificial intelligence"],
"scrapedAt": "2025-01-15T12:05:00.000Z"
}

Output fields

FieldTypeDescription
pageIdnumberWikipedia internal page identifier
titlestringArticle title
extractstringIntroductory summary (plain text, no HTML)
fullContentstringFull article plaintext when includeFullContent is enabled; empty if Wikipedia has no full extract
urlstringDirect link to the Wikipedia article
wordCountnumberTotal word count of the article
sizenumberArticle size in bytes
lastEditedstringISO timestamp of the last edit
thumbnailstringURL to article thumbnail image (if available)
categoriesstring[]Wikipedia category labels associated with the page
wikidataIdstringLinked Wikidata entity ID, or an empty string when unavailable
descriptionstringShort page description, or an empty string when unavailable
searchQuerystringFirst keyword that matched this page (search results only)
searchRanknumberOne-based rank for the first matching keyword (search results only)
requestedTitlestringFirst exact-title input that resolved to this page (exact results only)
matchedSearchQueriesstring[]All keyword inputs that matched this deduplicated page
requestedTitlesstring[]All exact-title inputs that resolved to this deduplicated page
scrapedAtstringISO timestamp when the data was extracted

wordCount comes from Wikipedia's search response and is 0 for exact-title-only results. All other base fields retain the same types in both modes.

Supported languages

Wikipedia Scraper supports all 300+ Wikipedia language editions. Use the standard language code:

CodeLanguageArticles
enEnglish6.9M+
deGerman2.9M+
frFrench2.6M+
esSpanish2.0M+
jaJapanese1.4M+
ruRussian1.9M+
zhChinese1.4M+
ptPortuguese1.1M+
itItalian1.8M+
arArabic1.2M+

Any valid Wikipedia language code works — see the full list.

How to scrape Wikipedia articles

  1. Open Wikipedia Scraper on Apify.
  2. Enter search keywords in searchQueries, known page names in articleTitles, or both.
  3. Set the language code (e.g., en, de, fr) for the Wikipedia edition you want.
  4. Adjust maxResultsPerSearch to control how many articles each keyword returns (default: 50).
  5. Optionally enable includeFullContent to add each full article in the fullContent field.
  6. Click Start and wait for the scrape to finish.
  7. Download the deduplicated articles as JSON, CSV, or Excel from the Dataset tab.

API usage

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("automation-lab/wikipedia-scraper").call(run_input={
"searchQueries": ["climate change", "renewable energy"],
"language": "en",
"maxResultsPerSearch": 20,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(f"{item['title']} — {item['wordCount']} words")
print(f" {item['url']}")
print(f" {item['extract'][:200]}...")

Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('automation-lab/wikipedia-scraper').call({
searchQueries: ['climate change', 'renewable energy'],
language: 'en',
maxResultsPerSearch: 20,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach(item => {
console.log(`${item.title} — ${item.wordCount} words`);
console.log(` ${item.url}`);
});

REST API

curl -X POST "https://api.apify.com/v2/acts/automation-lab/wikipedia-scraper/runs?token=YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"searchQueries": ["artificial intelligence"],
"language": "en",
"maxResultsPerSearch": 10
}'

Integrations

Connect Wikipedia Scraper to hundreds of apps using built-in integrations:

  • Google Sheets — export article data to spreadsheets
  • Slack / Microsoft Teams — get notifications when scraping completes
  • Zapier / Make — trigger workflows with scraped Wikipedia data
  • Amazon S3 / Google Cloud Storage — store large datasets in cloud storage
  • Webhook — send results to your own API endpoint

Tips and best practices

  1. Use specific keywords — more specific searches return more relevant results. "Quantum entanglement" is better than "quantum".
  2. Batch keywords efficiently — combine related keywords in one run to save on startup costs.
  3. Language parameter — set the language code to search non-English Wikipedias. Results, summaries, and URLs will all be in the selected language.
  4. Word count filtering — use the wordCount field to filter out stub articles (typically < 500 words).
  5. Rate limits — Wikipedia's API is generous but has rate limits. The scraper handles pagination and batching automatically.
  6. Choose the content depth — the extract field always remains the short introductory summary. Enable includeFullContent when your workflow needs full article plaintext in fullContent; this opt-in mode makes one additional batched Wikipedia API request.
  7. Metadata scope — the Actor returns categories, Wikidata IDs, and short descriptions when Wikipedia provides them. It does not return citations, external links, coordinates, HTML, or category pagination beyond the API's per-request limit.
  8. Exact titles and redirects — use the page title, not a full URL. Wikipedia redirects are resolved automatically; missing titles are logged and skipped without discarding valid pages.
  9. Max 500 results per keyword — this is a Wikipedia API limit. For broader coverage, use multiple related keywords.

Data source, privacy, and support

  • The Actor sends your search keywords, exact titles, language, and ordinary request metadata to Wikimedia's public MediaWiki API. It does not use AI, paid APIs, proxies, source accounts, or cookies.
  • Results can include public biographical information when you request pages about people. Collect only data you need and follow applicable privacy and data-protection rules.
  • The Actor stores no external cache or private copy. Run inputs, logs, and datasets remain in Apify storage under your account's retention and deletion settings.
  • Wikipedia and Wikimedia are named only to identify the public data source. This Actor is not affiliated with or endorsed by the Wikimedia Foundation.
  • For help, use the Actor's Issues tab on Apify and include a sanitized input and run ID.

Legality

Scraping publicly available data is generally legal according to the US Court of Appeals ruling (HiQ Labs v. LinkedIn). This actor only accesses publicly available information and does not require authentication. Always review and comply with the target website's Terms of Service before scraping. For personal data, ensure compliance with GDPR, CCPA, and other applicable privacy regulations.

FAQ

Q: Does this scraper get the full article text? A: Yes, as an opt-in feature. Set includeFullContent to true and read the full article plaintext from fullContent. The existing extract field still contains only the introductory section. If Wikipedia returns no full extract for an article, fullContent is an empty string and the base article record is retained.

Q: How fast is it? A: Very fast. Wikipedia's API is highly optimized. A typical run extracting 50 articles completes in under 5 seconds.

Q: Does it need proxies? A: No. Wikipedia's API is open and does not block automated requests. The scraper identifies itself with a proper User-Agent header.

Q: Can I search in multiple languages at once? A: Each run uses one language. To search multiple languages, run the scraper once per language.

Use with Claude AI (MCP)

This actor is available as a tool in Claude AI through the Model Context Protocol (MCP). Add it to Claude Desktop, Cursor, Windsurf, or any MCP-compatible client.

Setup for Claude Code

$claude mcp add --transport http apify "https://mcp.apify.com?tools=automation-lab/wikipedia-scraper"

Setup for Claude Desktop, Cursor, or VS Code

Add this to your MCP config file:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=automation-lab/wikipedia-scraper"
}
}
}

Example prompts

  • "Search Wikipedia for articles about quantum computing and give me the summaries"
  • "Fetch Wikipedia articles on these 5 historical events and compare their word counts"
  • "Look up Wikipedia articles on machine learning in both English and German and extract the introductions"

Learn more in the Apify MCP documentation.

The extract is truncated or too short. The extract field intentionally contains only the article's introductory section. Set includeFullContent to true to receive the full article plaintext in fullContent while preserving the concise introductory extract.

I'm getting irrelevant results for my search query. Wikipedia's search API ranks by relevance, which may include loosely related articles. Use more specific keywords (e.g., "quantum entanglement" instead of "quantum") and reduce maxResultsPerSearch to get only the top matches. If you already know the page, put its title in articleTitles to fetch it directly.

Other research and news scrapers on Apify