Wikipedia Article Watch & Scraper API avatar

Wikipedia Article Watch & Scraper API

Pricing

from $0.60 / 1,000 articles

Go to Apify Store
Wikipedia Article Watch & Scraper API

Wikipedia Article Watch & Scraper API

Wikipedia scraper and API: current article summaries, full text, sections and pageviews, or only the articles that changed since your last run. Never charged for failed or unchanged rows. Attribution on every row; articles about people are left out.

Pricing

from $0.60 / 1,000 articles

Rating

0.0

(0)

Developer

COMPASS DEV

COMPASS DEV

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Wikipedia article watch and scraper API that returns Wikipedia articles from the official Wikipedia API as flat rows — each article's current summary, full text, sections, categories and pageviews, or only what changed since your last run (new revisions, summary changes, sections added or removed, moves and deletions) — and never charges for failed results or unchanged articles.

Watch or export: give a watch list name (stateName) to watch your articles for changes; without one, every call returns the articles' current content (export). mode overrides this, and mode: "watch" without a name uses the watch list default.

Built for AI agents and knowledge bases that must stay current: run it on a schedule and feed only the changes to your index. Licence and attribution on every row (CC BY-SA 4.0 + the article's history page). Articles about people are left out. You are never charged for unchanged, missing or failed articles — never charged for failed results.

Why this one:

  • Flat change fields: each changed article is one row that says what changed (changeType, sectionsAdded, sectionsRemoved, summaryBefore / summaryAfter, sizeDelta, previousTitle), with no diff to parse.
  • Attribution on every row: the CC BY-SA 4.0 licence and the article's history page, ready to cite.
  • People left out: articles about individuals are never returned.
  • Unchanged articles are free: a watch run where nothing changed costs only Apify's small Actor-start charge.

Quick start

  • One call, current content: {"articles": ["Bitcoin", "Photosynthesis"]} returns both articles (export mode), every time you call it. {"searchQuery": ["photosynthesis"], "maxResults": 2} returns the first 2 search results.

  • Watch for changes: click Start. The form is filled in with three articles, watch mode, a watch list name and "also output unchanged articles":

    {
    "articles": ["Artificial intelligence", "https://en.wikipedia.org/wiki/Climate_change", "Bitcoin"],
    "mode": "watch",
    "stateName": "example-watchlist",
    "includeUnchanged": true
    }

    The first run is the baseline: one row per article (changeType: "baseline"), with the summary and the section list. Run it again later (or add an Apify schedule): only articles edited since then come back with what changed; with includeUnchanged, the others come back as free unchanged rows.

Use cases

  • Keep a knowledge base or RAG index current. Watch the articles your index is built on; re-embed only the ones that come back changed (summaryAfter, sectionsAdded, sectionsRemoved).
  • Monitor articles about your topics, products or places. A daily or weekly schedule tells you when an article was edited, moved or deleted, with the size change and the new lead text.
  • Export articles for analysis. Export mode returns summaries, full text, section titles, categories and 30-day pageviews for a list of titles or a search.

What it does

ModeUse it forInput
watch (a watch list name, or mode: "watch")Only what changed since the last run with the same stateNamearticles + stateName
export (no watch list name, or mode: "export")Each article's current content, every runarticles and/or searchQuery

For every article: title, page ID, URL, summary (lead section as plain text), Wikidata ID, latest revision ID and time, size, and — on request — full text, section titles, categories and daily pageviews for the last 30 days. Watch mode always keeps the section list, because "sections added or removed" is one of the changes it reports.

Watch mode, change types (changeType):

ValueMeaningCharged?
baselineFirst time this article is seen under this stateNameYes
new_revisionEdited since the last run (same lead, same sections)Yes
summary_changedThe lead section's text changed (summaryBefore / summaryAfter)Yes
sections_changedSections were added or removed (sectionsAdded / sectionsRemoved)Yes
movedThe page was renamed (previousTitle)Yes
deletedThe title no longer leads to an article (a no_data row, once)No
unchangedNot edited since the last run; output only with includeUnchangedNo

If nothing is new or changed, a watch run returns one free no_data row that says so ("No new or changed articles since the last run (N unchanged, not returned)"), so a working run never comes back empty.

Input

Watch three articles (the first run is the baseline):

{ "mode": "watch", "stateName": "my-kb-articles", "articles": ["Large language model", "https://fr.wikipedia.org/wiki/Paris"] }

Export the first 20 search results with their full text:

{ "mode": "export", "searchQuery": ["renewable energy"], "maxResults": 20, "includeFullText": true }
FieldDefaultDescription
articles—Required (or searchQuery). Titles or Wikipedia URLs from English, French or German Wikipedia (https://de.wikipedia.org/wiki/Berlin)
modewatch if stateName is given, else exportwatch (only changes) or export (current content); a search without articles is an export
stateName—The name of your watch list, kept between runs in your own Apify storage; giving one turns on watch mode (mode: "watch" without a name uses default)
includeUnchangedfalseWatch mode: also output unchanged articles (free)
searchQuery—Export mode: search text, one query per line
languageenLanguage for titles and searches: en, fr or de
maxResults10Export mode: most articles per search
includeFullTextfalseWhole article as plain text
includeSectionsfalseExport mode: section titles (watch mode always has them)
includeCategoriesfalseVisible categories
includePageviewsfalseDaily user pageviews for the last 30 full days, summed and per day
maxItems100Most articles checked in one run (listed articles plus search results; at most 1,000)

Field names from other tools: articleTitles, articleUrls, titles, urls, startUrls (→ articles); searchQueries (→ searchQuery); maxResultsPerSearch, maxArticlesPerQuery, maxSearchResults (→ maxResults); includeFullContent (→ includeFullText).

Output

A baseline row (watch mode; a real row from run rz3dIbCdCxwUtIduD, summary and the 41 section titles shortened):

{
"status": "ok",
"error": null,
"attempts": 1,
"input": "Earth",
"reason": null,
"title": "Earth",
"language": "en",
"pageId": 9228,
"url": "https://en.wikipedia.org/wiki/Earth",
"summary": "Earth is the third planet from the Sun and the only astronomical object known to harbor life. …",
"fullText": null,
"sections": ["Etymology", "Natural history", "Formation", "After formation", "…"],
"categories": null,
"wikidataId": "Q2",
"lastRevisionId": 1377391917,
"lastEditedAt": "2026-09-29T04:44:36Z",
"sizeBytes": 226005,
"pageviews30d": null,
"pageviewsDaily": null,
"changeType": "baseline",
"previousRevisionId": null,
"previousTitle": null,
"sizeDelta": null,
"sectionsAdded": null,
"sectionsRemoved": null,
"summaryBefore": null,
"summaryAfter": null,
"license": "CC BY-SA 4.0",
"licenseUrl": "https://creativecommons.org/licenses/by-sa/4.0/",
"attributionUrl": "https://en.wikipedia.org/w/index.php?title=Earth&action=history",
"scrapedAt": "2026-09-29T21:07:06.863Z"
}

On a later run, a changed article has changeType, previousRevisionId, sizeDelta, sectionsAdded, sectionsRemoved, summaryBefore and summaryAfter filled in.

statusMeaningCharged?
okArticle returned: an export row, a baseline row or a changed articleYes
ok (changeType: "unchanged")Watch mode with includeUnchanged: not edited since the last runNo
no_datareason: missing, redirect_to_missing, disambiguation, not_an_article, deleted, or biography (left out: the article is about a person; only title, Wikidata ID and reason are returned)No
failedInvalid input, or no answer after retries (error says why)No

The key-value store holds RUN_REPORT (counts, charged and free rows, requests, pauses). The watch list lives in a named store in your own account: wikipedia-articles-state-<stateName>.

Pricing

Pay per article returned (event article): export rows, baseline rows and changed articles.

Apify planPer articlePer 1,000 articles
Free$0.0010$1.00
Bronze$0.0008$0.80
Silver$0.0007$0.70
Gold (and Platinum, Diamond)$0.0006$0.60
  • Never charged for failed results, missing articles, disambiguation pages, biographies left out, deletions or unchanged articles. A watch run where nothing changed costs only Apify's small Actor-start charge.
  • Your maximum charge per run is respected: the run stops before it, and outputs only what it could charge.
  • No usage fees on top: the price per article covers the platform's compute.

Use it from AI agents

  • MCP: add the Actor through the Apify MCP server (https://mcp.apify.com?actors=mouadapi/wikipedia-articles), then ask e.g. "Which of these Wikipedia articles changed since yesterday, and what changed?" Each row says ok, no_data or failed, and changeType names the change.
  • No state needed: {"articles": ["Bitcoin"]} or {"searchQuery": ["photosynthesis"]} returns current content every time (export). Add a stateName only when you want the changes since the last call with that name.
  • API: one call returns the rows:
curl -X POST "https://api.apify.com/v2/acts/mouadapi~wikipedia-articles/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" -d '{"stateName": "agent-kb", "articles": ["Large language model"]}'
  • x402 payments: the Actor is pay-per-event only, with no usage fees, limited permissions and no Standby mode, so agents can pay per article with x402.
  • Flat rows with the licence and attribution URL on each, ready to cite.

The same call from JavaScript (the Apify client):

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('mouadapi/wikipedia-articles').call({ articles: ['Earth'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Limits

  • Up to 1,000 articles per run (maxItems, default 100). A 1,000-article export takes about 9 minutes, because we follow Wikimedia's bot policy: one request at a time, at most 4 a second, and a 5-second pause after any slow answer. Small lists take seconds.
  • Official Wikimedia APIs only: the MediaWiki Action API and the pageviews API. No HTML pages, no dumps, no following links.
  • Current content only: watch mode compares with what it saved at the last run; it never downloads old revisions.
  • Section titles and full text come from Wikipedia's plain-text extract: no infoboxes, tables, images or references.
  • No editor names, IP addresses or edit comments, ever.
  • English, French and German Wikipedia only: the languages where articles about people can be recognised by their categories. Other languages are refused (free failed row).
  • Articles about people are left out (recognised by categories such as "Living people", "1879 births", "Naissance en …", "Geboren …", "Frau"), as are search results about people. An article with no categories at all is left out too, because we can't rule out that it is about a person (free).

Known issues

  • The people check relies on Wikipedia's categories. On a test of 180 known articles (30 people and 30 other articles per language) it was right every time, but a person article that is missing its birth, death or gender categories would not be recognised.
  • Unchanged rows (includeUnchanged) repeat the summary and sections saved at the last change; they don't carry full text, categories or pageviews.
  • If Wikipedia answers "too many requests" or reports server lag, the run pauses 1, 2 and then 4 minutes, and stops if it still can't continue (unchecked articles get free failed rows).

FAQ

How much does it cost? $1.00 per 1,000 articles returned on the Free plan, down to $0.60 on Gold, plus Apify's small Actor-start charge per run. Watching 100 articles daily where 5 change a day costs about $0.005 a day on the Free plan.

Am I charged for articles that didn't change? No. Unchanged, missing, deleted and left-out articles and failed checks are free.

Why does a large run take minutes? Wikimedia asks bots to send one request at a time and to wait 5 seconds after any slow answer, so a 1,000-article export takes about 9 minutes. We follow those rules, so your runs don't strain Wikipedia and aren't blocked.

Is this allowed? Wikipedia's text is licensed CC BY-SA 4.0, which allows commercial reuse with attribution; every row carries the licence and the history-page URL for attribution. The Actor uses Wikimedia's documented APIs with an identifying User-Agent, one request at a time.

Also by the same author: DNS Lookup & SSL Certificate Checker and Woolworths Price Scraper & Monitor.