Wikipedia Scraper: Articles, Summaries & Pageviews avatar

Wikipedia Scraper: Articles, Summaries & Pageviews

Pricing

Pay per event

Go to Apify Store
Wikipedia Scraper: Articles, Summaries & Pageviews

Wikipedia Scraper: Articles, Summaries & Pageviews

Scrape Wikipedia in any of 300+ languages: full-text article search, clean summaries with description, extract, image and coordinates, plus monthly pageview counts per article. Export CSV, Excel, JSON or XML. No login or API key. Pay only for saved rows.

Pricing

Pay per event

Rating

0.0

(0)

Developer

RecordsData

RecordsData

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

16 hours ago

Last modified

Share

PunkRecordsData

๐ŸŒ Wikipedia Scraper: Articles, Summaries & Pageviews

Wikipedia Scraper by PunkRecordsData extracts Wikipedia articles from any of 300+ language editions: full-text search results, clean summaries (description, extract, image, coordinates) and monthly pageview counts per article. Enter a keyword or article titles, get structured rows, export CSV, Excel, JSON or XML. No login, no API key. Verified 2026-10-03 against live Wikipedia. Priced per result.

Wikipedia Scraper turns a keyword or a list of article titles into a dataset built from Wikimedia's official APIs. In a test run on the English edition, the search "quantum computing" returned 10 article rows, each with a clean extract, thumbnail and last-modified date. Pageviews are measured too: the Spanish article "Alan Turing" had 33,124 views in September 2026. It is built for content and SEO teams, NLP and RAG builders, PR analysts and researchers.

๐Ÿ“‹ What does Wikipedia Scraper do?

  • Search Wikipedia articles by keyword and get each hit as a full summary row.
  • Fetch summaries for exact article titles or pasted Wikipedia URLs.
  • Download monthly pageview counts for any article, up to 120 months back.
  • Switch language with one field (en, es, de, fr, pt, ja and every other edition).
  • Get lighter search rows (snippet, size, word count) when you only need a ranked list.
  • Export Wikipedia data to CSV, Excel, JSON or XML, or pull it by API.

๐Ÿ“Š What data can you extract from Wikipedia?

FieldDescription
recordTypearticle, search-result or pageviews
languageWikipedia edition code, for example en
titleArticle title
descriptionOne-line Wikidata description
extractClean plain-text lead summary, up to 1,500 characters
pageIdWikipedia page ID
articleTypePage type, for example standard or disambiguation
thumbnailUrlLead image (only when the article has one)
latitude, longitudeCoordinates (only for places)
wordCountArticle word count (search hits)
sizeBytes, snippetSize and highlighted snippet (light search rows)
month, viewsMonth (YYYY-MM) and view count (pageview rows)
lastModifiedLast edit timestamp
urlCanonical article URL
scrapedAtUTC timestamp of the scrape

Fields that do not exist for an article are left out of the row instead of being filled with placeholders.

๐Ÿงพ Sample output

Real row from a run with the search "quantum computing" on the English edition:

{
"recordType": "article",
"language": "en",
"title": "Quantum computing",
"description": "Computer hardware technology that uses quantum mechanics",
"extract": "A quantum computer is a computer that represents and processes information using quantum states. Quantum computations exploit phenomena such as superposition, interference, and entanglement. Quantum computers have the potential to complete some calculations exponentially faster than classical computers. For example, a large-scale quantum computer could break widely used encryption schemes and aid physicists in performing physical simulations. However, current hardware implementations of quantum computation are largely experimental and suitable for only certain specialized tasks.",
"pageId": 25220,
"articleType": "standard",
"thumbnailUrl": "https://thumb.wikimedia.org/wikipedia/commons/thumb/b/b9/IBM_Quantum_Computer_Demo_at_ITUWTSA_2024%2C_Delhi_2.jpg/330px-IBM_Quantum_Computer_Demo_at_ITUWTSA_2024%2C_Delhi_2.jpg",
"lastModified": "2026-10-02T15:55:36Z",
"url": "https://en.wikipedia.org/wiki/Quantum_computing",
"scrapedAt": "2026-10-04T05:04:51.552Z",
"wordCount": 13459
}

A pageview row looks like this:

{
"recordType": "pageviews",
"language": "es",
"title": "Alan Turing",
"month": "2026-09",
"views": 33124,
"url": "https://es.wikipedia.org/wiki/Alan_Turing",
"scrapedAt": "2026-10-04T05:05:32.511Z"
}

๐Ÿ’ฐ How much does it cost to scrape Wikipedia?

Pay per event, charged only for rows that are saved. Prices at the Free tier (Apify subscription tiers get lower rates):

EventWhat you pay forPrice
article-recordOne article summary row$0.012
search-recordOne light search row (summaries off)$0.010
pageview-recordOne month of pageviews for an article$0.010

So 1,000 article summaries cost about $12. Error rows, "article not found" rows and empty searches are never charged. If your max charge per run is reached, the run stops cleanly with the status "Stopped at your max charge limit". Free users get a 10-row preview.

๐Ÿš€ How to scrape Wikipedia in 3 steps

  1. Open the actor and set the Language (default en) and a Search query, or paste Article titles.
  2. Optionally switch on Monthly pageviews and choose how many months back.
  3. Click Start, then download the dataset as CSV, Excel, JSON or XML.

โš™๏ธ Input

FieldDefaultMeaning
languageenWikipedia edition code
searchQueryartificial intelligenceFull-text search keywords
fetchSummariesForResultstrueEnrich each hit into a full summary row
articleTitles[]Exact titles or Wikipedia URLs
includePageviewsfalseAdd monthly pageview rows for the titles list
pageviewMonths12Months of history, 1 to 120
maxItems10Total rows cap (free plan limited to 10)
{
"language": "es",
"searchQuery": "",
"articleTitles": ["Alan Turing"],
"includePageviews": true,
"pageviewMonths": 12,
"maxItems": 50
}

Pageviews are returned per article from the titles list. The current, still-running month is included, so pageviewMonths: 3 can return 4 monthly rows.

๐Ÿ“ฆ Output

Each row is one article, search hit or pageview month. The dataset has an Articles table view with thumbnail, title, description, extract, views, URL and timestamps, plus the full JSON. If a search has no matches, the run finishes as succeeded with the message "No results for this search. Nothing was charged." If Wikipedia is unreachable and nothing was saved, the run fails with a clear message instead of finishing silently.

โš–๏ธ Wikipedia scraper vs alternatives

Compared on the Apify Store (prices read from the public store listings, 2026-10-03):

This actorautomation-labfatihtahtacrawlerbros
Per-article price$0.012$0.00115 + $0.001 start$0.00499 per record$0.001 per item + $0.005 start
Monthly pageviewsYes, up to 120 monthsNot listedNot listedNot listed
Search hits enriched with image and coordinatesYesNot verifiedNot verifiedNot verified

We are the most expensive per article. What the extra buys: pageview time series and search rows that already carry description, extract, thumbnail and coordinates, so you skip a second step. If you only need bulk plain article text at the lowest price, the cheaper actors fit better.

๐Ÿ’ผ Use cases

  • SEO and content research: find the articles around a topic and see their real monthly traffic.
  • NLP and RAG corpora: build multilingual datasets of clean extracts by topic.
  • Brand and PR monitoring: track pageview curves for people, companies and products.
  • Education and reference tools: summaries with images and coordinates in any language.

๐Ÿ”Œ Run via API, schedule and integrations

curl -X POST "https://api.apify.com/v2/acts/recordsdata~wikipedia-articles-scraper/runs?token=<YOUR_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"language":"en","searchQuery":"quantum computing","maxItems":10}'

Use the Apify client for JavaScript or Python, schedule monthly pageview refreshes in the Console, or connect Zapier, Make, n8n and Google Sheets. The actor is also callable from AI agents through the Apify MCP server.

The actor reads Wikimedia's official public APIs with an identified user agent and a polite request delay, and it only returns public article data, no personal accounts. Wikipedia text is licensed CC BY-SA, so credit Wikipedia when you republish it. Check your own use case for compliance.

โ“ Frequently asked questions

Does Wikipedia Scraper need an API key or login?

No. It uses public Wikimedia endpoints. You only need an Apify account.

Which Wikipedia languages are supported?

Every edition. Set language to the subdomain code, for example es, de, fr, pt or ja.

Where do the pageview numbers come from?

From Wikimedia's official pageview API, monthly totals for all access types and agents.

How long are the extracts?

The lead-section summary as plain text, cut at 1,500 characters.

Why did I get 0 results?

The search had no matching articles in that language edition, or the titles do not exist there. Try another spelling or language code. Nothing is charged in that case.

Why is there an error row for my title?

When an exact title is not found, one row with an error message is saved so you can see which title failed. Error rows are free.

Are lat/long and thumbnail always present?

No. Coordinates exist only for places and thumbnails only for articles with a lead image, so those fields are omitted otherwise.

Can I limit my spend?

Yes. Set maxItems and the Apify max charge per run. The run stops cleanly when either is reached.

๐Ÿ”— Want more research data? Other PunkRecordsData scrapers

๐Ÿ’ฌ Support

Found a bug or need a missing field? Open the Issues tab on this actor's page or write to contact.punkrecordsdata@gmail.com.

Last updated: 2026-10-03