Wikipedia Article Scraper avatar

Wikipedia Article Scraper

Pricing

from $3.00 / 1,000 results

Go to Apify Store
Wikipedia Article Scraper

Wikipedia Article Scraper

Scrape Wikipedia articles by search keyword or exact title. Returns summaries, full article text, categories, and links. Supports 300+ languages.

Pricing

from $3.00 / 1,000 results

Rating

0.0

(0)

Developer

cloud9

cloud9

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

19 days ago

Last modified

Categories

Share

Scrape Wikipedia articles by search keyword or exact title. Returns summaries, full article text, categories, and links. Supports 300+ languages.

Use cases

  • Search Wikipedia and get structured results for a RAG pipeline
  • Build a glossary or entity list for a domain
  • Enrich records with canonical article URLs
  • Measure article length and freshness as a coverage signal
  • Bootstrap reference content for an app

Input

ParameterTypeRequiredDefaultDescription
modestringNo"search"search: search articles by keyword | article: get full article by title | summary: get short summary by title Allowed: search, article, summary.
querystringNo"artificial intelligence"Keyword to search, or exact Wikipedia article title
languagestringNo"en"Wikipedia language code (en, ja, de, fr, es, zh, etc.)
maxResultsintegerNo10Maximum number of search results (only for search mode)

Example input

{
"mode": "search",
"query": "artificial intelligence",
"language": "en",
"maxResults": 10
}

Output

The exact fields depend on the mode you run. This is real output from an actual run of this Actor:

{
"title": "Artificial intelligence",
"snippet": "Artificial intelligence (AI) is the capability of computational systems to perform tasks typically associated with human intelligence, such as learnin…",
"pageid": 1164,
"url": "https://en.wikipedia.org/wiki/Artificial_intelligence",
"wordcount": 27623,
"timestamp": "2026-09-06T15:12:23Z"
}
FieldType
titlestring
snippetstring
pageidnumber
urlstring
wordcountnumber
timestampstring

Results are exportable from Apify Console or the API as JSON, CSV, Excel, or XML.

How to run it

In Apify Console — open the Actor, fill in the input form, click Start, then download the results from the Dataset tab.

With the JavaScript client

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('cloud9_ai/wikipedia-scraper').call({
"mode": "search",
"query": "artificial intelligence",
"language": "en",
"maxResults": 10
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

With the Python client

from apify_client import ApifyClient
client = ApifyClient('YOUR_APIFY_TOKEN')
run = client.actor('cloud9_ai/wikipedia-scraper').call(run_input={
"mode": "search",
"query": "artificial intelligence",
"language": "en",
"maxResults": 10
})
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item)

With the API — POST https://api.apify.com/v2/acts/cloud9_ai~wikipedia-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN with the input JSON as the body.

Notes and limits

  • No API key, account, or login is needed — just the input above.
  • maxResults caps how much a single run collects, which is also what caps the run's cost.
  • Requests are paced and failed requests are retried automatically, so runs stay inside the source's rate limits.
  • Only publicly available data is collected. How you use the output is your responsibility, including the source's terms of use and any applicable data-protection law.

Support

Found a bug, or need a field that isn't in the output? Open an issue on the Issues tab of this Actor in Apify Console. Issues there are read and answered.

License

Apache-2.0