# Wikipedia Article Watch & Scraper API (`mouadapi/wikipedia-articles`) Actor

Wikipedia scraper and API: current article summaries, full text, sections and pageviews, or only the articles that changed since your last run. Never charged for failed or unchanged rows. Attribution on every row; articles about people are left out.

- **URL**: https://apify.com/mouadapi/wikipedia-articles.md
- **Developed by:** [COMPASS DEV](https://apify.com/mouadapi) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.60 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Wikipedia article watch and scraper API** that returns **Wikipedia articles** from the official **Wikipedia API** as
flat rows — each article's current summary, full text, sections, categories and pageviews, or **only what changed since
your last run** (new revisions, summary changes, sections added or removed, moves and deletions) — and never charges for
failed results or unchanged articles.

**Watch or export:** give a watch list name (`stateName`) to watch your articles for changes; without one, every call
returns the articles' current content (export). `mode` overrides this, and `mode: "watch"` without a name uses the watch
list `default`.

Built for AI agents and knowledge bases that must stay current: run it on a schedule and feed only the changes to your
index. **Licence and attribution on every row** (CC BY-SA 4.0 + the article's history page). **Articles about people are
left out.** **You are never charged for unchanged, missing or failed articles** — never charged for failed results.

**Why this one:**

- **Flat change fields:** each changed article is one row that says what changed (`changeType`, `sectionsAdded`,
  `sectionsRemoved`, `summaryBefore` / `summaryAfter`, `sizeDelta`, `previousTitle`), with no diff to parse.
- **Attribution on every row:** the CC BY-SA 4.0 licence and the article's history page, ready to cite.
- **People left out:** articles about individuals are never returned.
- **Unchanged articles are free:** a watch run where nothing changed costs only Apify's small Actor-start charge.

### Quick start

- **One call, current content:** `{"articles": ["Bitcoin", "Photosynthesis"]}` returns both articles (export mode), every
  time you call it. `{"searchQuery": ["photosynthesis"], "maxResults": 2}` returns the first 2 search results.
- **Watch for changes:** click **Start**. The form is filled in with three articles, watch mode, a watch list name and
  "also output unchanged articles":

  ```json
  {
    "articles": ["Artificial intelligence", "https://en.wikipedia.org/wiki/Climate_change", "Bitcoin"],
    "mode": "watch",
    "stateName": "example-watchlist",
    "includeUnchanged": true
  }
  ```

  The first run is the **baseline**: one row per article (`changeType: "baseline"`), with the summary and the section
  list. Run it again later (or add an Apify schedule): only articles edited since then come back with what changed; with
  `includeUnchanged`, the others come back as free `unchanged` rows.

### Use cases

- **Keep a knowledge base or RAG index current.** Watch the articles your index is built on; re-embed only the ones that
  come back changed (`summaryAfter`, `sectionsAdded`, `sectionsRemoved`).
- **Monitor articles about your topics, products or places.** A daily or weekly schedule tells you when an article was
  edited, moved or deleted, with the size change and the new lead text.
- **Export articles for analysis.** Export mode returns summaries, full text, section titles, categories and 30-day
  pageviews for a list of titles or a search.

### What it does

| Mode | Use it for | Input |
|---|---|---|
| **watch** (a watch list name, or `mode: "watch"`) | Only what changed since the last run with the same `stateName` | `articles` + `stateName` |
| **export** (no watch list name, or `mode: "export"`) | Each article's current content, every run | `articles` and/or `searchQuery` |

For every article: title, page ID, URL, summary (lead section as plain text), Wikidata ID, latest revision ID and time,
size, and — on request — full text, section titles, categories and daily pageviews for the last 30 days. Watch mode always
keeps the section list, because "sections added or removed" is one of the changes it reports.

**Watch mode, change types** (`changeType`):

| Value | Meaning | Charged? |
|---|---|---|
| `baseline` | First time this article is seen under this `stateName` | Yes |
| `new_revision` | Edited since the last run (same lead, same sections) | Yes |
| `summary_changed` | The lead section's text changed (`summaryBefore` / `summaryAfter`) | Yes |
| `sections_changed` | Sections were added or removed (`sectionsAdded` / `sectionsRemoved`) | Yes |
| `moved` | The page was renamed (`previousTitle`) | Yes |
| `deleted` | The title no longer leads to an article (a `no_data` row, once) | **No** |
| `unchanged` | Not edited since the last run; output only with `includeUnchanged` | **No** |

If nothing is new or changed, a watch run returns **one free `no_data` row** that says so ("No new or changed articles since
the last run (N unchanged, not returned)"), so a working run never comes back empty.

### Input

Watch three articles (the first run is the baseline):

```json
{ "mode": "watch", "stateName": "my-kb-articles", "articles": ["Large language model", "https://fr.wikipedia.org/wiki/Paris"] }
```

Export the first 20 search results with their full text:

```json
{ "mode": "export", "searchQuery": ["renewable energy"], "maxResults": 20, "includeFullText": true }
```

| Field | Default | Description |
|---|---|---|
| `articles` | — | Required (or `searchQuery`). Titles or Wikipedia URLs from English, French or German Wikipedia (`https://de.wikipedia.org/wiki/Berlin`) |
| `mode` | watch if `stateName` is given, else export | `watch` (only changes) or `export` (current content); a search without articles is an export |
| `stateName` | — | The name of your watch list, kept between runs in your own Apify storage; giving one turns on watch mode (`mode: "watch"` without a name uses `default`) |
| `includeUnchanged` | `false` | Watch mode: also output unchanged articles (free) |
| `searchQuery` | — | Export mode: search text, one query per line |
| `language` | `en` | Language for titles and searches: `en`, `fr` or `de` |
| `maxResults` | `10` | Export mode: most articles per search |
| `includeFullText` | `false` | Whole article as plain text |
| `includeSections` | `false` | Export mode: section titles (watch mode always has them) |
| `includeCategories` | `false` | Visible categories |
| `includePageviews` | `false` | Daily user pageviews for the last 30 full days, summed and per day |
| `maxItems` | `100` | Most articles checked in one run (listed articles plus search results; at most 1,000) |

**Field names from other tools:** `articleTitles`, `articleUrls`, `titles`, `urls`, `startUrls` (→ `articles`);
`searchQueries` (→ `searchQuery`); `maxResultsPerSearch`, `maxArticlesPerQuery`, `maxSearchResults` (→ `maxResults`);
`includeFullContent` (→ `includeFullText`).

### Output

A baseline row (watch mode; a real row from run `rz3dIbCdCxwUtIduD`, summary and the 41 section titles shortened):

```json
{
  "status": "ok",
  "error": null,
  "attempts": 1,
  "input": "Earth",
  "reason": null,
  "title": "Earth",
  "language": "en",
  "pageId": 9228,
  "url": "https://en.wikipedia.org/wiki/Earth",
  "summary": "Earth is the third planet from the Sun and the only astronomical object known to harbor life. …",
  "fullText": null,
  "sections": ["Etymology", "Natural history", "Formation", "After formation", "…"],
  "categories": null,
  "wikidataId": "Q2",
  "lastRevisionId": 1377391917,
  "lastEditedAt": "2026-09-29T04:44:36Z",
  "sizeBytes": 226005,
  "pageviews30d": null,
  "pageviewsDaily": null,
  "changeType": "baseline",
  "previousRevisionId": null,
  "previousTitle": null,
  "sizeDelta": null,
  "sectionsAdded": null,
  "sectionsRemoved": null,
  "summaryBefore": null,
  "summaryAfter": null,
  "license": "CC BY-SA 4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by-sa/4.0/",
  "attributionUrl": "https://en.wikipedia.org/w/index.php?title=Earth&action=history",
  "scrapedAt": "2026-09-29T21:07:06.863Z"
}
```

On a later run, a changed article has `changeType`, `previousRevisionId`, `sizeDelta`, `sectionsAdded`,
`sectionsRemoved`, `summaryBefore` and `summaryAfter` filled in.

| `status` | Meaning | Charged? |
|---|---|---|
| `ok` | Article returned: an export row, a baseline row or a changed article | Yes |
| `ok` (`changeType: "unchanged"`) | Watch mode with `includeUnchanged`: not edited since the last run | **No** |
| `no_data` | `reason`: `missing`, `redirect_to_missing`, `disambiguation`, `not_an_article`, `deleted`, or `biography` (left out: the article is about a person; only title, Wikidata ID and reason are returned) | **No** |
| `failed` | Invalid input, or no answer after retries (`error` says why) | **No** |

The key-value store holds `RUN_REPORT` (counts, charged and free rows, requests, pauses). The watch list lives in a named
store in your own account: `wikipedia-articles-state-<stateName>`.

### Pricing

Pay per article returned (event `article`): export rows, baseline rows and changed articles.

| Apify plan | Per article | Per 1,000 articles |
|---|---|---|
| Free | $0.0010 | $1.00 |
| Bronze | $0.0008 | $0.80 |
| Silver | $0.0007 | $0.70 |
| Gold (and Platinum, Diamond) | $0.0006 | $0.60 |

- **Never charged for failed results**, missing articles, disambiguation pages, biographies left out, deletions or
  unchanged articles. A watch run where nothing changed costs only Apify's small Actor-start charge.
- Your maximum charge per run is respected: the run stops before it, and outputs only what it could charge.
- No usage fees on top: the price per article covers the platform's compute.

### Use it from AI agents

- **MCP:** add the Actor through the Apify MCP server (`https://mcp.apify.com?actors=mouadapi/wikipedia-articles`), then ask
  e.g. *"Which of these Wikipedia articles changed since yesterday, and what changed?"* Each row says `ok`, `no_data` or
  `failed`, and `changeType` names the change.
- **No state needed:** `{"articles": ["Bitcoin"]}` or `{"searchQuery": ["photosynthesis"]}` returns current content
  every time (export). Add a `stateName` only when you want the changes since the last call with that name.
- **API:** one call returns the rows:

```bash
curl -X POST "https://api.apify.com/v2/acts/mouadapi~wikipedia-articles/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" -d '{"stateName": "agent-kb", "articles": ["Large language model"]}'
```

- **x402 payments:** the Actor is pay-per-event only, with no usage fees, limited permissions and no Standby mode, so
  agents can pay per article with x402.
- Flat rows with the licence and attribution URL on each, ready to cite.

The same call from JavaScript (the Apify client):

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('mouadapi/wikipedia-articles').call({ articles: ['Earth'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

### Limits

- Up to 1,000 articles per run (`maxItems`, default 100). **A 1,000-article export takes about 9 minutes**, because we
  follow Wikimedia's bot policy: one request at a time, at most 4 a second, and a 5-second pause after any slow answer.
  Small lists take seconds.
- Official Wikimedia APIs only: the MediaWiki Action API and the pageviews API. No HTML
  pages, no dumps, no following links.
- Current content only: watch mode compares with what it saved at the last run; it never downloads old revisions.
- Section titles and full text come from Wikipedia's plain-text extract: no infoboxes, tables, images or references.
- No editor names, IP addresses or edit comments, ever.
- English, French and German Wikipedia only: the languages where articles about people can be recognised by their
  categories. Other languages are refused (free failed row).
- Articles about people are left out (recognised by categories such as "Living people", "1879 births", "Naissance en …",
  "Geboren …", "Frau"), as are search results about people. An article with no categories at all is left out too,
  because we can't rule out that it is about a person (free).

### Known issues

- The people check relies on Wikipedia's categories. On a test of 180 known articles (30 people and 30 other articles
  per language) it was right every time, but a person article that is missing its birth, death or gender categories
  would not be recognised.
- Unchanged rows (`includeUnchanged`) repeat the summary and sections saved at the last change; they don't carry full
  text, categories or pageviews.
- If Wikipedia answers "too many requests" or reports server lag, the run pauses 1, 2 and then 4 minutes, and stops if
  it still can't continue (unchecked articles get free `failed` rows).

### FAQ

**How much does it cost?** $1.00 per 1,000 articles returned on the Free plan, down to $0.60 on Gold, plus Apify's small
Actor-start charge per run. Watching 100 articles daily where 5 change a day costs about $0.005 a day on the Free plan.

**Am I charged for articles that didn't change?** No. Unchanged, missing, deleted and left-out articles and failed checks
are free.

**Why does a large run take minutes?** Wikimedia asks bots to send one request at a time and to wait 5 seconds after
any slow answer, so a 1,000-article export takes about 9 minutes. We follow those rules, so your runs don't strain
Wikipedia and aren't blocked.

**Is this allowed?** Wikipedia's text is licensed CC BY-SA 4.0, which allows commercial reuse with attribution; every row
carries the licence and the history-page URL for attribution. The Actor uses Wikimedia's documented APIs with an
identifying User-Agent, one request at a time.

Also by the same author: [DNS Lookup & SSL Certificate Checker](https://apify.com/mouadapi/dns-ssl-checker) and
[Woolworths Price Scraper & Monitor](https://apify.com/mouadapi/woolworths-price-monitor).

# Actor input Schema

## `articles` (type: `array`):

Required (or give searchQuery instead). Article titles or Wikipedia URLs, one per line (English, French or German Wikipedia: https://fr.wikipedia.org/wiki/Paris). Titles use the language below.

## `searchQuery` (type: `array`):

Search text, one query per line; each finds up to maxResults articles. A search without articles runs as an export. Articles about people are left out.

## `mode` (type: `string`):

watch: remember your articles and output only what changed since the last run with the same watch list name (stateName). export: output each article's current content every run. If you leave it empty: a search without articles is an export, a watch list name means watch, and nothing means export (the same call returns the articles every time).

## `stateName` (type: `string`):

Name of your watch list. Giving a name turns on watch mode (unless mode says otherwise). The run remembers each article under this name in your own Apify storage and compares with it next time. Watch mode without a name uses the list "default".

## `includeUnchanged` (type: `boolean`):

Watch mode: also output articles that did not change (changeType "unchanged", never charged). Off: only changed articles are output.

## `language` (type: `string`):

Wikipedia language for titles and searches: en, fr or de (the languages where articles about people can be recognised and left out). URLs keep their own language.

## `maxResults` (type: `integer`):

Export mode: most articles per search query (default 10).

## `includeFullText` (type: `boolean`):

Add the whole article as plain text (default off). Costs one extra request per article.

## `includeSections` (type: `boolean`):

Export mode: add the list of section titles (watch mode always includes them). Costs one extra request per article.

## `includeCategories` (type: `boolean`):

Add the article's visible categories.

## `includePageviews` (type: `boolean`):

Add daily user pageviews for the last 30 full days, summed and per day. Costs one extra request per article.

## `maxItems` (type: `integer`):

Most articles checked in one run (listed articles plus search results; default 100, at most 1,000). Articles beyond it are listed in the log and not checked.

## Actor input object example

```json
{
  "articles": [
    "Artificial intelligence",
    "https://en.wikipedia.org/wiki/Climate_change",
    "Bitcoin"
  ],
  "mode": "watch",
  "stateName": "example-watchlist",
  "includeUnchanged": true,
  "language": "en",
  "maxResults": 10,
  "includeSections": false,
  "includeCategories": false,
  "includePageviews": false,
  "maxItems": 100
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset with the article rows (status ok / no\_data / failed)

## `runReport` (type: `string`):

Summary of the run (counts, charged and free rows, stop reason)

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "articles": [
        "Artificial intelligence",
        "https://en.wikipedia.org/wiki/Climate_change",
        "Bitcoin"
    ],
    "mode": "watch",
    "stateName": "example-watchlist",
    "includeUnchanged": true,
    "maxResults": 10,
    "includeFullText": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("mouadapi/wikipedia-articles").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "articles": [
        "Artificial intelligence",
        "https://en.wikipedia.org/wiki/Climate_change",
        "Bitcoin",
    ],
    "mode": "watch",
    "stateName": "example-watchlist",
    "includeUnchanged": True,
    "maxResults": 10,
    "includeFullText": False,
}

# Run the Actor and wait for it to finish
run = client.actor("mouadapi/wikipedia-articles").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "articles": [
    "Artificial intelligence",
    "https://en.wikipedia.org/wiki/Climate_change",
    "Bitcoin"
  ],
  "mode": "watch",
  "stateName": "example-watchlist",
  "includeUnchanged": true,
  "maxResults": 10,
  "includeFullText": false
}' |
apify call mouadapi/wikipedia-articles --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mouadapi/wikipedia-articles"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dZ1GgHravBQRNPukw/builds/8IUTEI0SwM4fXEfsn/openapi.json
