Wikipedia Scraper: Articles, Summaries & Pageviews
Pricing
Pay per event
Wikipedia Scraper: Articles, Summaries & Pageviews
Scrape Wikipedia in any of 300+ languages: full-text article search, clean summaries with description, extract, image and coordinates, plus monthly pageview counts per article. Export CSV, Excel, JSON or XML. No login or API key. Pay only for saved rows.
Pricing
Pay per event
Rating
0.0
(0)
Developer
RecordsData
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
17 hours ago
Last modified
Categories
Share
๐ Wikipedia Scraper: Articles, Summaries & Pageviews
Wikipedia Scraper by PunkRecordsData extracts Wikipedia articles from any of 300+ language editions: full-text search results, clean summaries (description, extract, image, coordinates) and monthly pageview counts per article. Enter a keyword or article titles, get structured rows, export CSV, Excel, JSON or XML. No login, no API key. Verified 2026-10-03 against live Wikipedia. Priced per result.
Wikipedia Scraper turns a keyword or a list of article titles into a dataset built from Wikimedia's official APIs. In a test run on the English edition, the search "quantum computing" returned 10 article rows, each with a clean extract, thumbnail and last-modified date. Pageviews are measured too: the Spanish article "Alan Turing" had 33,124 views in September 2026. It is built for content and SEO teams, NLP and RAG builders, PR analysts and researchers.
๐ What does Wikipedia Scraper do?
- Search Wikipedia articles by keyword and get each hit as a full summary row.
- Fetch summaries for exact article titles or pasted Wikipedia URLs.
- Download monthly pageview counts for any article, up to 120 months back.
- Switch language with one field (en, es, de, fr, pt, ja and every other edition).
- Get lighter search rows (snippet, size, word count) when you only need a ranked list.
- Export Wikipedia data to CSV, Excel, JSON or XML, or pull it by API.
๐ What data can you extract from Wikipedia?
| Field | Description |
|---|---|
recordType | article, search-result or pageviews |
language | Wikipedia edition code, for example en |
title | Article title |
description | One-line Wikidata description |
extract | Clean plain-text lead summary, up to 1,500 characters |
pageId | Wikipedia page ID |
articleType | Page type, for example standard or disambiguation |
thumbnailUrl | Lead image (only when the article has one) |
latitude, longitude | Coordinates (only for places) |
wordCount | Article word count (search hits) |
sizeBytes, snippet | Size and highlighted snippet (light search rows) |
month, views | Month (YYYY-MM) and view count (pageview rows) |
lastModified | Last edit timestamp |
url | Canonical article URL |
scrapedAt | UTC timestamp of the scrape |
Fields that do not exist for an article are left out of the row instead of being filled with placeholders.
๐งพ Sample output
Real row from a run with the search "quantum computing" on the English edition:
{"recordType": "article","language": "en","title": "Quantum computing","description": "Computer hardware technology that uses quantum mechanics","extract": "A quantum computer is a computer that represents and processes information using quantum states. Quantum computations exploit phenomena such as superposition, interference, and entanglement. Quantum computers have the potential to complete some calculations exponentially faster than classical computers. For example, a large-scale quantum computer could break widely used encryption schemes and aid physicists in performing physical simulations. However, current hardware implementations of quantum computation are largely experimental and suitable for only certain specialized tasks.","pageId": 25220,"articleType": "standard","thumbnailUrl": "https://thumb.wikimedia.org/wikipedia/commons/thumb/b/b9/IBM_Quantum_Computer_Demo_at_ITUWTSA_2024%2C_Delhi_2.jpg/330px-IBM_Quantum_Computer_Demo_at_ITUWTSA_2024%2C_Delhi_2.jpg","lastModified": "2026-10-02T15:55:36Z","url": "https://en.wikipedia.org/wiki/Quantum_computing","scrapedAt": "2026-10-04T05:04:51.552Z","wordCount": 13459}
A pageview row looks like this:
{"recordType": "pageviews","language": "es","title": "Alan Turing","month": "2026-09","views": 33124,"url": "https://es.wikipedia.org/wiki/Alan_Turing","scrapedAt": "2026-10-04T05:05:32.511Z"}
๐ฐ How much does it cost to scrape Wikipedia?
Pay per event, charged only for rows that are saved. Prices at the Free tier (Apify subscription tiers get lower rates):
| Event | What you pay for | Price |
|---|---|---|
article-record | One article summary row | $0.012 |
search-record | One light search row (summaries off) | $0.010 |
pageview-record | One month of pageviews for an article | $0.010 |
So 1,000 article summaries cost about $12. Error rows, "article not found" rows and empty searches are never charged. If your max charge per run is reached, the run stops cleanly with the status "Stopped at your max charge limit". Free users get a 10-row preview.
๐ How to scrape Wikipedia in 3 steps
- Open the actor and set the Language (default
en) and a Search query, or paste Article titles. - Optionally switch on Monthly pageviews and choose how many months back.
- Click Start, then download the dataset as CSV, Excel, JSON or XML.
โ๏ธ Input
| Field | Default | Meaning |
|---|---|---|
language | en | Wikipedia edition code |
searchQuery | artificial intelligence | Full-text search keywords |
fetchSummariesForResults | true | Enrich each hit into a full summary row |
articleTitles | [] | Exact titles or Wikipedia URLs |
includePageviews | false | Add monthly pageview rows for the titles list |
pageviewMonths | 12 | Months of history, 1 to 120 |
maxItems | 10 | Total rows cap (free plan limited to 10) |
{"language": "es","searchQuery": "","articleTitles": ["Alan Turing"],"includePageviews": true,"pageviewMonths": 12,"maxItems": 50}
Pageviews are returned per article from the titles list. The current, still-running month is included, so pageviewMonths: 3 can return 4 monthly rows.
๐ฆ Output
Each row is one article, search hit or pageview month. The dataset has an Articles table view with thumbnail, title, description, extract, views, URL and timestamps, plus the full JSON. If a search has no matches, the run finishes as succeeded with the message "No results for this search. Nothing was charged." If Wikipedia is unreachable and nothing was saved, the run fails with a clear message instead of finishing silently.
โ๏ธ Wikipedia scraper vs alternatives
Compared on the Apify Store (prices read from the public store listings, 2026-10-03):
| This actor | automation-lab | fatihtahta | crawlerbros | |
|---|---|---|---|---|
| Per-article price | $0.012 | $0.00115 + $0.001 start | $0.00499 per record | $0.001 per item + $0.005 start |
| Monthly pageviews | Yes, up to 120 months | Not listed | Not listed | Not listed |
| Search hits enriched with image and coordinates | Yes | Not verified | Not verified | Not verified |
We are the most expensive per article. What the extra buys: pageview time series and search rows that already carry description, extract, thumbnail and coordinates, so you skip a second step. If you only need bulk plain article text at the lowest price, the cheaper actors fit better.
๐ผ Use cases
- SEO and content research: find the articles around a topic and see their real monthly traffic.
- NLP and RAG corpora: build multilingual datasets of clean extracts by topic.
- Brand and PR monitoring: track pageview curves for people, companies and products.
- Education and reference tools: summaries with images and coordinates in any language.
๐ Run via API, schedule and integrations
curl -X POST "https://api.apify.com/v2/acts/recordsdata~wikipedia-articles-scraper/runs?token=<YOUR_TOKEN>" \-H "Content-Type: application/json" \-d '{"language":"en","searchQuery":"quantum computing","maxItems":10}'
Use the Apify client for JavaScript or Python, schedule monthly pageview refreshes in the Console, or connect Zapier, Make, n8n and Google Sheets. The actor is also callable from AI agents through the Apify MCP server.
๐ก๏ธ Is it legal to scrape Wikipedia?
The actor reads Wikimedia's official public APIs with an identified user agent and a polite request delay, and it only returns public article data, no personal accounts. Wikipedia text is licensed CC BY-SA, so credit Wikipedia when you republish it. Check your own use case for compliance.
โ Frequently asked questions
Does Wikipedia Scraper need an API key or login?
No. It uses public Wikimedia endpoints. You only need an Apify account.
Which Wikipedia languages are supported?
Every edition. Set language to the subdomain code, for example es, de, fr, pt or ja.
Where do the pageview numbers come from?
From Wikimedia's official pageview API, monthly totals for all access types and agents.
How long are the extracts?
The lead-section summary as plain text, cut at 1,500 characters.
Why did I get 0 results?
The search had no matching articles in that language edition, or the titles do not exist there. Try another spelling or language code. Nothing is charged in that case.
Why is there an error row for my title?
When an exact title is not found, one row with an error message is saved so you can see which title failed. Error rows are free.
Are lat/long and thumbnail always present?
No. Coordinates exist only for places and thumbnails only for articles with a lead image, so those fields are omitted otherwise.
Can I limit my spend?
Yes. Set maxItems and the Apify max charge per run. The run stops cleanly when either is reached.
๐ Want more research data? Other PunkRecordsData scrapers
๐ฌ Support
Found a bug or need a missing field? Open the Issues tab on this actor's page or write to contact.punkrecordsdata@gmail.com.
Last updated: 2026-10-03