Wikipedia Article Watch & Scraper API
Pricing
from $0.60 / 1,000 articles
Wikipedia Article Watch & Scraper API
Wikipedia scraper and API: current article summaries, full text, sections and pageviews, or only the articles that changed since your last run. Never charged for failed or unchanged rows. Attribution on every row; articles about people are left out.
Pricing
from $0.60 / 1,000 articles
Rating
0.0
(0)
Developer
COMPASS DEV
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Wikipedia article watch and scraper API that returns Wikipedia articles from the official Wikipedia API as flat rows — each article's current summary, full text, sections, categories and pageviews, or only what changed since your last run (new revisions, summary changes, sections added or removed, moves and deletions) — and never charges for failed results or unchanged articles.
Watch or export: give a watch list name (stateName) to watch your articles for changes; without one, every call
returns the articles' current content (export). mode overrides this, and mode: "watch" without a name uses the watch
list default.
Built for AI agents and knowledge bases that must stay current: run it on a schedule and feed only the changes to your index. Licence and attribution on every row (CC BY-SA 4.0 + the article's history page). Articles about people are left out. You are never charged for unchanged, missing or failed articles — never charged for failed results.
Why this one:
- Flat change fields: each changed article is one row that says what changed (
changeType,sectionsAdded,sectionsRemoved,summaryBefore/summaryAfter,sizeDelta,previousTitle), with no diff to parse. - Attribution on every row: the CC BY-SA 4.0 licence and the article's history page, ready to cite.
- People left out: articles about individuals are never returned.
- Unchanged articles are free: a watch run where nothing changed costs only Apify's small Actor-start charge.
Quick start
-
One call, current content:
{"articles": ["Bitcoin", "Photosynthesis"]}returns both articles (export mode), every time you call it.{"searchQuery": ["photosynthesis"], "maxResults": 2}returns the first 2 search results. -
Watch for changes: click Start. The form is filled in with three articles, watch mode, a watch list name and "also output unchanged articles":
{"articles": ["Artificial intelligence", "https://en.wikipedia.org/wiki/Climate_change", "Bitcoin"],"mode": "watch","stateName": "example-watchlist","includeUnchanged": true}The first run is the baseline: one row per article (
changeType: "baseline"), with the summary and the section list. Run it again later (or add an Apify schedule): only articles edited since then come back with what changed; withincludeUnchanged, the others come back as freeunchangedrows.
Use cases
- Keep a knowledge base or RAG index current. Watch the articles your index is built on; re-embed only the ones that
come back changed (
summaryAfter,sectionsAdded,sectionsRemoved). - Monitor articles about your topics, products or places. A daily or weekly schedule tells you when an article was edited, moved or deleted, with the size change and the new lead text.
- Export articles for analysis. Export mode returns summaries, full text, section titles, categories and 30-day pageviews for a list of titles or a search.
What it does
| Mode | Use it for | Input |
|---|---|---|
watch (a watch list name, or mode: "watch") | Only what changed since the last run with the same stateName | articles + stateName |
export (no watch list name, or mode: "export") | Each article's current content, every run | articles and/or searchQuery |
For every article: title, page ID, URL, summary (lead section as plain text), Wikidata ID, latest revision ID and time, size, and — on request — full text, section titles, categories and daily pageviews for the last 30 days. Watch mode always keeps the section list, because "sections added or removed" is one of the changes it reports.
Watch mode, change types (changeType):
| Value | Meaning | Charged? |
|---|---|---|
baseline | First time this article is seen under this stateName | Yes |
new_revision | Edited since the last run (same lead, same sections) | Yes |
summary_changed | The lead section's text changed (summaryBefore / summaryAfter) | Yes |
sections_changed | Sections were added or removed (sectionsAdded / sectionsRemoved) | Yes |
moved | The page was renamed (previousTitle) | Yes |
deleted | The title no longer leads to an article (a no_data row, once) | No |
unchanged | Not edited since the last run; output only with includeUnchanged | No |
If nothing is new or changed, a watch run returns one free no_data row that says so ("No new or changed articles since
the last run (N unchanged, not returned)"), so a working run never comes back empty.
Input
Watch three articles (the first run is the baseline):
{ "mode": "watch", "stateName": "my-kb-articles", "articles": ["Large language model", "https://fr.wikipedia.org/wiki/Paris"] }
Export the first 20 search results with their full text:
{ "mode": "export", "searchQuery": ["renewable energy"], "maxResults": 20, "includeFullText": true }
| Field | Default | Description |
|---|---|---|
articles | — | Required (or searchQuery). Titles or Wikipedia URLs from English, French or German Wikipedia (https://de.wikipedia.org/wiki/Berlin) |
mode | watch if stateName is given, else export | watch (only changes) or export (current content); a search without articles is an export |
stateName | — | The name of your watch list, kept between runs in your own Apify storage; giving one turns on watch mode (mode: "watch" without a name uses default) |
includeUnchanged | false | Watch mode: also output unchanged articles (free) |
searchQuery | — | Export mode: search text, one query per line |
language | en | Language for titles and searches: en, fr or de |
maxResults | 10 | Export mode: most articles per search |
includeFullText | false | Whole article as plain text |
includeSections | false | Export mode: section titles (watch mode always has them) |
includeCategories | false | Visible categories |
includePageviews | false | Daily user pageviews for the last 30 full days, summed and per day |
maxItems | 100 | Most articles checked in one run (listed articles plus search results; at most 1,000) |
Field names from other tools: articleTitles, articleUrls, titles, urls, startUrls (→ articles);
searchQueries (→ searchQuery); maxResultsPerSearch, maxArticlesPerQuery, maxSearchResults (→ maxResults);
includeFullContent (→ includeFullText).
Output
A baseline row (watch mode; a real row from run rz3dIbCdCxwUtIduD, summary and the 41 section titles shortened):
{"status": "ok","error": null,"attempts": 1,"input": "Earth","reason": null,"title": "Earth","language": "en","pageId": 9228,"url": "https://en.wikipedia.org/wiki/Earth","summary": "Earth is the third planet from the Sun and the only astronomical object known to harbor life. …","fullText": null,"sections": ["Etymology", "Natural history", "Formation", "After formation", "…"],"categories": null,"wikidataId": "Q2","lastRevisionId": 1377391917,"lastEditedAt": "2026-09-29T04:44:36Z","sizeBytes": 226005,"pageviews30d": null,"pageviewsDaily": null,"changeType": "baseline","previousRevisionId": null,"previousTitle": null,"sizeDelta": null,"sectionsAdded": null,"sectionsRemoved": null,"summaryBefore": null,"summaryAfter": null,"license": "CC BY-SA 4.0","licenseUrl": "https://creativecommons.org/licenses/by-sa/4.0/","attributionUrl": "https://en.wikipedia.org/w/index.php?title=Earth&action=history","scrapedAt": "2026-09-29T21:07:06.863Z"}
On a later run, a changed article has changeType, previousRevisionId, sizeDelta, sectionsAdded,
sectionsRemoved, summaryBefore and summaryAfter filled in.
status | Meaning | Charged? |
|---|---|---|
ok | Article returned: an export row, a baseline row or a changed article | Yes |
ok (changeType: "unchanged") | Watch mode with includeUnchanged: not edited since the last run | No |
no_data | reason: missing, redirect_to_missing, disambiguation, not_an_article, deleted, or biography (left out: the article is about a person; only title, Wikidata ID and reason are returned) | No |
failed | Invalid input, or no answer after retries (error says why) | No |
The key-value store holds RUN_REPORT (counts, charged and free rows, requests, pauses). The watch list lives in a named
store in your own account: wikipedia-articles-state-<stateName>.
Pricing
Pay per article returned (event article): export rows, baseline rows and changed articles.
| Apify plan | Per article | Per 1,000 articles |
|---|---|---|
| Free | $0.0010 | $1.00 |
| Bronze | $0.0008 | $0.80 |
| Silver | $0.0007 | $0.70 |
| Gold (and Platinum, Diamond) | $0.0006 | $0.60 |
- Never charged for failed results, missing articles, disambiguation pages, biographies left out, deletions or unchanged articles. A watch run where nothing changed costs only Apify's small Actor-start charge.
- Your maximum charge per run is respected: the run stops before it, and outputs only what it could charge.
- No usage fees on top: the price per article covers the platform's compute.
Use it from AI agents
- MCP: add the Actor through the Apify MCP server (
https://mcp.apify.com?actors=mouadapi/wikipedia-articles), then ask e.g. "Which of these Wikipedia articles changed since yesterday, and what changed?" Each row saysok,no_dataorfailed, andchangeTypenames the change. - No state needed:
{"articles": ["Bitcoin"]}or{"searchQuery": ["photosynthesis"]}returns current content every time (export). Add astateNameonly when you want the changes since the last call with that name. - API: one call returns the rows:
curl -X POST "https://api.apify.com/v2/acts/mouadapi~wikipedia-articles/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" -d '{"stateName": "agent-kb", "articles": ["Large language model"]}'
- x402 payments: the Actor is pay-per-event only, with no usage fees, limited permissions and no Standby mode, so agents can pay per article with x402.
- Flat rows with the licence and attribution URL on each, ready to cite.
The same call from JavaScript (the Apify client):
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('mouadapi/wikipedia-articles').call({ articles: ['Earth'] });const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
Limits
- Up to 1,000 articles per run (
maxItems, default 100). A 1,000-article export takes about 9 minutes, because we follow Wikimedia's bot policy: one request at a time, at most 4 a second, and a 5-second pause after any slow answer. Small lists take seconds. - Official Wikimedia APIs only: the MediaWiki Action API and the pageviews API. No HTML pages, no dumps, no following links.
- Current content only: watch mode compares with what it saved at the last run; it never downloads old revisions.
- Section titles and full text come from Wikipedia's plain-text extract: no infoboxes, tables, images or references.
- No editor names, IP addresses or edit comments, ever.
- English, French and German Wikipedia only: the languages where articles about people can be recognised by their categories. Other languages are refused (free failed row).
- Articles about people are left out (recognised by categories such as "Living people", "1879 births", "Naissance en …", "Geboren …", "Frau"), as are search results about people. An article with no categories at all is left out too, because we can't rule out that it is about a person (free).
Known issues
- The people check relies on Wikipedia's categories. On a test of 180 known articles (30 people and 30 other articles per language) it was right every time, but a person article that is missing its birth, death or gender categories would not be recognised.
- Unchanged rows (
includeUnchanged) repeat the summary and sections saved at the last change; they don't carry full text, categories or pageviews. - If Wikipedia answers "too many requests" or reports server lag, the run pauses 1, 2 and then 4 minutes, and stops if
it still can't continue (unchecked articles get free
failedrows).
FAQ
How much does it cost? $1.00 per 1,000 articles returned on the Free plan, down to $0.60 on Gold, plus Apify's small Actor-start charge per run. Watching 100 articles daily where 5 change a day costs about $0.005 a day on the Free plan.
Am I charged for articles that didn't change? No. Unchanged, missing, deleted and left-out articles and failed checks are free.
Why does a large run take minutes? Wikimedia asks bots to send one request at a time and to wait 5 seconds after any slow answer, so a 1,000-article export takes about 9 minutes. We follow those rules, so your runs don't strain Wikipedia and aren't blocked.
Is this allowed? Wikipedia's text is licensed CC BY-SA 4.0, which allows commercial reuse with attribution; every row carries the licence and the history-page URL for attribution. The Actor uses Wikimedia's documented APIs with an identifying User-Agent, one request at a time.
Also by the same author: DNS Lookup & SSL Certificate Checker and Woolworths Price Scraper & Monitor.