# Wikipedia Scraper: Articles, Summaries & Pageviews (`recordsdata/wikipedia-articles-scraper`) Actor

Scrape Wikipedia in any of 300+ languages: full-text article search, clean summaries with description, extract, image and coordinates, plus monthly pageview counts per article. Export CSV, Excel, JSON or XML. No login or API key. Pay only for saved rows.

- **URL**: https://apify.com/recordsdata/wikipedia-articles-scraper.md
- **Developed by:** [RecordsData](https://apify.com/recordsdata) (community)
- **Categories:** Education, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<p align="center">
  <img src="https://api.apify.com/v2/key-value-stores/AAm3a1h3Z9nYfrvh9/records/banner" alt="PunkRecordsData" width="100%" />
</p>

## 🌐 Wikipedia Scraper: Articles, Summaries & Pageviews

> **Wikipedia Scraper by PunkRecordsData extracts Wikipedia articles from any of 300+ language editions: full-text search results, clean summaries (description, extract, image, coordinates) and monthly pageview counts per article. Enter a keyword or article titles, get structured rows, export CSV, Excel, JSON or XML. No login, no API key. Verified 2026-10-03 against live Wikipedia. Priced per result.**

Wikipedia Scraper turns a keyword or a list of article titles into a dataset built from Wikimedia's official APIs. In a test run on the English edition, the search "quantum computing" returned 10 article rows, each with a clean extract, thumbnail and last-modified date. Pageviews are measured too: the Spanish article "Alan Turing" had 33,124 views in September 2026. It is built for content and SEO teams, NLP and RAG builders, PR analysts and researchers.

### 📋 What does Wikipedia Scraper do?

- Search Wikipedia articles by keyword and get each hit as a full summary row.
- Fetch summaries for exact article titles or pasted Wikipedia URLs.
- Download monthly pageview counts for any article, up to 120 months back.
- Switch language with one field (en, es, de, fr, pt, ja and every other edition).
- Get lighter search rows (snippet, size, word count) when you only need a ranked list.
- Export Wikipedia data to CSV, Excel, JSON or XML, or pull it by API.

### 📊 What data can you extract from Wikipedia?

| Field | Description |
|---|---|
| `recordType` | `article`, `search-result` or `pageviews` |
| `language` | Wikipedia edition code, for example `en` |
| `title` | Article title |
| `description` | One-line Wikidata description |
| `extract` | Clean plain-text lead summary, up to 1,500 characters |
| `pageId` | Wikipedia page ID |
| `articleType` | Page type, for example `standard` or `disambiguation` |
| `thumbnailUrl` | Lead image (only when the article has one) |
| `latitude`, `longitude` | Coordinates (only for places) |
| `wordCount` | Article word count (search hits) |
| `sizeBytes`, `snippet` | Size and highlighted snippet (light search rows) |
| `month`, `views` | Month (`YYYY-MM`) and view count (pageview rows) |
| `lastModified` | Last edit timestamp |
| `url` | Canonical article URL |
| `scrapedAt` | UTC timestamp of the scrape |

Fields that do not exist for an article are left out of the row instead of being filled with placeholders.

### 🧾 Sample output

Real row from a run with the search "quantum computing" on the English edition:

```json
{
  "recordType": "article",
  "language": "en",
  "title": "Quantum computing",
  "description": "Computer hardware technology that uses quantum mechanics",
  "extract": "A quantum computer is a computer that represents and processes information using quantum states. Quantum computations exploit phenomena such as superposition, interference, and entanglement. Quantum computers have the potential to complete some calculations exponentially faster than classical computers. For example, a large-scale quantum computer could break widely used encryption schemes and aid physicists in performing physical simulations. However, current hardware implementations of quantum computation are largely experimental and suitable for only certain specialized tasks.",
  "pageId": 25220,
  "articleType": "standard",
  "thumbnailUrl": "https://thumb.wikimedia.org/wikipedia/commons/thumb/b/b9/IBM_Quantum_Computer_Demo_at_ITUWTSA_2024%2C_Delhi_2.jpg/330px-IBM_Quantum_Computer_Demo_at_ITUWTSA_2024%2C_Delhi_2.jpg",
  "lastModified": "2026-10-02T15:55:36Z",
  "url": "https://en.wikipedia.org/wiki/Quantum_computing",
  "scrapedAt": "2026-10-04T05:04:51.552Z",
  "wordCount": 13459
}
```

A pageview row looks like this:

```json
{
  "recordType": "pageviews",
  "language": "es",
  "title": "Alan Turing",
  "month": "2026-09",
  "views": 33124,
  "url": "https://es.wikipedia.org/wiki/Alan_Turing",
  "scrapedAt": "2026-10-04T05:05:32.511Z"
}
```

### 💰 How much does it cost to scrape Wikipedia?

Pay per event, charged only for rows that are saved. Prices at the Free tier (Apify subscription tiers get lower rates):

| Event | What you pay for | Price |
|---|---|---|
| `article-record` | One article summary row | $0.012 |
| `search-record` | One light search row (summaries off) | $0.010 |
| `pageview-record` | One month of pageviews for an article | $0.010 |

So 1,000 article summaries cost about $12. Error rows, "article not found" rows and empty searches are never charged. If your max charge per run is reached, the run stops cleanly with the status "Stopped at your max charge limit". Free users get a 10-row preview.

### 🚀 How to scrape Wikipedia in 3 steps

1. Open the actor and set the **Language** (default `en`) and a **Search query**, or paste **Article titles**.
2. Optionally switch on **Monthly pageviews** and choose how many months back.
3. Click **Start**, then download the dataset as CSV, Excel, JSON or XML.

### ⚙️ Input

| Field | Default | Meaning |
|---|---|---|
| `language` | `en` | Wikipedia edition code |
| `searchQuery` | `artificial intelligence` | Full-text search keywords |
| `fetchSummariesForResults` | `true` | Enrich each hit into a full summary row |
| `articleTitles` | `[]` | Exact titles or Wikipedia URLs |
| `includePageviews` | `false` | Add monthly pageview rows for the titles list |
| `pageviewMonths` | `12` | Months of history, 1 to 120 |
| `maxItems` | `10` | Total rows cap (free plan limited to 10) |

```json
{
  "language": "es",
  "searchQuery": "",
  "articleTitles": ["Alan Turing"],
  "includePageviews": true,
  "pageviewMonths": 12,
  "maxItems": 50
}
```

Pageviews are returned per article from the titles list. The current, still-running month is included, so `pageviewMonths: 3` can return 4 monthly rows.

### 📦 Output

Each row is one article, search hit or pageview month. The dataset has an **Articles** table view with thumbnail, title, description, extract, views, URL and timestamps, plus the full JSON. If a search has no matches, the run finishes as succeeded with the message "No results for this search. Nothing was charged." If Wikipedia is unreachable and nothing was saved, the run fails with a clear message instead of finishing silently.

### ⚖️ Wikipedia scraper vs alternatives

Compared on the Apify Store (prices read from the public store listings, 2026-10-03):

| | This actor | automation-lab | fatihtahta | crawlerbros |
|---|---|---|---|---|
| Per-article price | $0.012 | $0.00115 + $0.001 start | $0.00499 per record | $0.001 per item + $0.005 start |
| Monthly pageviews | Yes, up to 120 months | Not listed | Not listed | Not listed |
| Search hits enriched with image and coordinates | Yes | Not verified | Not verified | Not verified |

We are the most expensive per article. What the extra buys: pageview time series and search rows that already carry description, extract, thumbnail and coordinates, so you skip a second step. If you only need bulk plain article text at the lowest price, the cheaper actors fit better.

### 💼 Use cases

- **SEO and content research:** find the articles around a topic and see their real monthly traffic.
- **NLP and RAG corpora:** build multilingual datasets of clean extracts by topic.
- **Brand and PR monitoring:** track pageview curves for people, companies and products.
- **Education and reference tools:** summaries with images and coordinates in any language.

### 🔌 Run via API, schedule and integrations

```bash
curl -X POST "https://api.apify.com/v2/acts/recordsdata~wikipedia-articles-scraper/runs?token=<YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"language":"en","searchQuery":"quantum computing","maxItems":10}'
```

Use the Apify client for JavaScript or Python, schedule monthly pageview refreshes in the Console, or connect Zapier, Make, n8n and Google Sheets. The actor is also callable from AI agents through the Apify MCP server.

### 🛡️ Is it legal to scrape Wikipedia?

The actor reads Wikimedia's official public APIs with an identified user agent and a polite request delay, and it only returns public article data, no personal accounts. Wikipedia text is licensed CC BY-SA, so credit Wikipedia when you republish it. Check your own use case for compliance.

### ❓ Frequently asked questions

#### Does Wikipedia Scraper need an API key or login?

No. It uses public Wikimedia endpoints. You only need an Apify account.

#### Which Wikipedia languages are supported?

Every edition. Set `language` to the subdomain code, for example `es`, `de`, `fr`, `pt` or `ja`.

#### Where do the pageview numbers come from?

From Wikimedia's official pageview API, monthly totals for all access types and agents.

#### How long are the extracts?

The lead-section summary as plain text, cut at 1,500 characters.

#### Why did I get 0 results?

The search had no matching articles in that language edition, or the titles do not exist there. Try another spelling or language code. Nothing is charged in that case.

#### Why is there an error row for my title?

When an exact title is not found, one row with an `error` message is saved so you can see which title failed. Error rows are free.

#### Are lat/long and thumbnail always present?

No. Coordinates exist only for places and thumbnails only for articles with a lead image, so those fields are omitted otherwise.

#### Can I limit my spend?

Yes. Set `maxItems` and the Apify max charge per run. The run stops cleanly when either is reached.

### 🔗 Want more research data? Other PunkRecordsData scrapers

- [arXiv Research Papers Scraper](https://apify.com/recordsdata/arxiv-research-papers-scraper)
- [Hacker News Search Scraper](https://apify.com/recordsdata/hackernews-search-scraper)
- [PubMed Articles Scraper](https://apify.com/recordsdata/pubmed-articles-scraper)

### 💬 Support

Found a bug or need a missing field? Open the Issues tab on this actor's page or write to contact.punkrecordsdata@gmail.com.

Last updated: 2026-10-03

# Actor input Schema

## `language` (type: `string`):

Wikipedia language edition code, for example en, es, de, fr, pt or ja. Any of the 300+ editions works; the code is the subdomain of wikipedia.org.

## `searchQuery` (type: `string`):

Full-text article search keywords, for example artificial intelligence. Each matching article becomes one row. Leave empty to use only the article titles list.

## `fetchSummariesForResults` (type: `boolean`):

When true, every search hit is enriched with the clean summary (description, extract, image, coordinates) and billed as an article summary. When false, lighter search rows (snippet, size, word count) are returned instead.

## `articleTitles` (type: `array`):

Exact article titles or full Wikipedia URLs to fetch summaries for, for example Alan Turing. Combine with pageviews to get traffic history per article.

## `includePageviews` (type: `boolean`):

When true, adds one row per month with view counts for each article in the titles list (Wikimedia pageview API, all platforms and agents).

## `pageviewMonths` (type: `integer`):

How many months of pageview history to return per article, from 1 up to 120 (10 years).

## `maxItems` (type: `integer`):

Maximum number of rows to return (articles, search results and pageview months together). Free users are limited to 10 rows; paid users can set up to 1,000,000.

## Actor input object example

```json
{
  "language": "en",
  "searchQuery": "artificial intelligence",
  "fetchSummariesForResults": true,
  "articleTitles": [],
  "includePageviews": false,
  "pageviewMonths": 12,
  "maxItems": 10
}
```

# Actor output Schema

## `overview` (type: `string`):

Key fields per row

## `fullData` (type: `string`):

Complete dataset with all fields

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "artificial intelligence",
    "articleTitles": [],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("recordsdata/wikipedia-articles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "artificial intelligence",
    "articleTitles": [],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("recordsdata/wikipedia-articles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "artificial intelligence",
  "articleTitles": [],
  "maxItems": 10
}' |
apify call recordsdata/wikipedia-articles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,recordsdata/wikipedia-articles-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lTUH481mf98HjfGTA/builds/j8hMOlHlQ7RiQ6egA/openapi.json
