Wikipedia Article Scraper
Pricing
from $3.00 / 1,000 results
Wikipedia Article Scraper
Scrape Wikipedia articles by search keyword or exact title. Returns summaries, full article text, categories, and links. Supports 300+ languages.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
cloud9
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
0
Monthly active users
19 days ago
Last modified
Categories
Share
Scrape Wikipedia articles by search keyword or exact title. Returns summaries, full article text, categories, and links. Supports 300+ languages.
Use cases
- Search Wikipedia and get structured results for a RAG pipeline
- Build a glossary or entity list for a domain
- Enrich records with canonical article URLs
- Measure article length and freshness as a coverage signal
- Bootstrap reference content for an app
Input
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
mode | string | No | "search" | search: search articles by keyword | article: get full article by title | summary: get short summary by title Allowed: search, article, summary. |
query | string | No | "artificial intelligence" | Keyword to search, or exact Wikipedia article title |
language | string | No | "en" | Wikipedia language code (en, ja, de, fr, es, zh, etc.) |
maxResults | integer | No | 10 | Maximum number of search results (only for search mode) |
Example input
{"mode": "search","query": "artificial intelligence","language": "en","maxResults": 10}
Output
The exact fields depend on the mode you run. This is real output from an actual run of this Actor:
{"title": "Artificial intelligence","snippet": "Artificial intelligence (AI) is the capability of computational systems to perform tasks typically associated with human intelligence, such as learnin…","pageid": 1164,"url": "https://en.wikipedia.org/wiki/Artificial_intelligence","wordcount": 27623,"timestamp": "2026-09-06T15:12:23Z"}
| Field | Type |
|---|---|
title | string |
snippet | string |
pageid | number |
url | string |
wordcount | number |
timestamp | string |
Results are exportable from Apify Console or the API as JSON, CSV, Excel, or XML.
How to run it
In Apify Console — open the Actor, fill in the input form, click Start, then download the results from the Dataset tab.
With the JavaScript client
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });const run = await client.actor('cloud9_ai/wikipedia-scraper').call({"mode": "search","query": "artificial intelligence","language": "en","maxResults": 10});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
With the Python client
from apify_client import ApifyClientclient = ApifyClient('YOUR_APIFY_TOKEN')run = client.actor('cloud9_ai/wikipedia-scraper').call(run_input={"mode": "search","query": "artificial intelligence","language": "en","maxResults": 10})for item in client.dataset(run['defaultDatasetId']).iterate_items():print(item)
With the API — POST https://api.apify.com/v2/acts/cloud9_ai~wikipedia-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN with the input JSON as the body.
Notes and limits
- No API key, account, or login is needed — just the input above.
maxResultscaps how much a single run collects, which is also what caps the run's cost.- Requests are paced and failed requests are retried automatically, so runs stay inside the source's rate limits.
- Only publicly available data is collected. How you use the output is your responsibility, including the source's terms of use and any applicable data-protection law.
Support
Found a bug, or need a field that isn't in the output? Open an issue on the Issues tab of this Actor in Apify Console. Issues there are read and answered.
License
Apache-2.0