Wikipedia & Wikidata API — Articles, Search & Pageviews avatar

Wikipedia & Wikidata API — Articles, Search & Pageviews

Pricing

from $0.36 / 1,000 articles

Go to Apify Store
Wikipedia & Wikidata API — Articles, Search & Pageviews

Wikipedia & Wikidata API — Articles, Search & Pageviews

Wikipedia search, article text (intro or full), descriptions, categories and Wikidata ids; daily human pageviews and a day's top 1,000; Wikidata entities with flattened properties. Any language edition, no API key, no browser. Diagnostics and empty runs are free.

Pricing

from $0.36 / 1,000 articles

Rating

0.0

(0)

Developer

Insight Solutions

Insight Solutions

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 hours ago

Last modified

Share

Wikipedia and Wikidata as clean rows, keyless, in any language edition. Search articles; fetch an article's plain text (the intro or the whole thing) with its short description, categories, Wikidata id, image, sections, links and recent revisions; get daily human pageviews for any title, or a day's top 1,000; look up Wikidata entities by id or by name with labels, descriptions, aliases, sitelinks and the properties you choose, flattened into readable values.

Read straight from Wikimedia's public APIs — the MediaWiki Action API, the pageviews API and Wikidata — with a descriptive User-Agent and a pace far below what Wikimedia allows. No key, no login, no browser, and no page HTML: article text comes out of Wikipedia's own plain-text extracts.

At a glance

Input — this is the Store prefill; paste it and run:

{ "language": "en", "queries": ["web scraping"], "titles": ["Web scraping", "Large language model"],
"categories": [], "maxResults": 20, "articleContent": "intro", "fetchArticlesForSearch": false,
"includeCategories": true, "includeSections": false, "includeLinks": false, "pageviews": "last30",
"topViews": [], "wikidataIds": ["Q665452"], "wikidataSearch": [], "includeRawClaims": false,
"maxConcurrency": 4, "maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }

Twenty search hits for "web scraping", the intros of two articles, their last thirty days of pageviews and the Wikidata entity for web scraping — about six requests and a few seconds.

Output — one row per result, four kinds of paid row in one schema. The fields you will use most are text, description and categories on an article; snippet on a search-result; totalViews, averageDaily and series on a pageviews row; label, description and properties on an entity (full list under Output reference). Every row carries license and attribution. Anything that could not be read comes back as a free diagnostic row (ok: false, errorType, error) instead of a charge.

Price — $0.60 per 1,000 articles, $0.20 per 1,000 search results, $0.30 per 1,000 pageviews rows, $0.50 per 1,000 Wikidata entities (+ $0.001 per run); diagnostics and empty runs free; no API key, no browser, limited permissions, works over the Apify MCP server and with x402.

From code — client.actor("insight.solutions/wikipedia-api").call(run_input={…}) with apify-client, or

POST https://api.apify.com/v2/acts/insight.solutions~wikipedia-api/run-sync-get-dataset-items
.


What you get

One row per result. Same columns on every row, null where a column does not apply. Abridged real rows:

An article (articleContent: "intro"):

{
"rowType": "article",
"input": "Web scraping",
"source": "en.wikipedia.org",
"license": "CC BY-SA 4.0",
"attribution": "Wikipedia contributors",
"title": "Web scraping",
"pageId": 2696619,
"url": "https://en.wikipedia.org/wiki/Web_scraping",
"description": "Method of extracting data from websites",
"text": "Web scraping, web harvesting, or web data extraction is data scraping used for extracting data from websites. Web scraping software may directly access the World Wide Web …",
"textTruncated": false,
"textChars": 2809,
"wordCount": 436,
"wikidataId": "Q665452",
"categories": ["Web scraping"],
"lastModified": "2026-09-30T16:28:17Z",
"lastRevisionId": 1375681792,
"pageLengthBytes": 34836,
"isDisambiguation": false
}

A search result:

{
"rowType": "search-result",
"query": "large language model",
"rank": 1,
"title": "Large language model",
"pageId": 73248112,
"url": "https://en.wikipedia.org/wiki/Large_language_model",
"snippet": "A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially",
"size": 136413,
"wordCount": 13579,
"lastModified": "2026-09-29T11:25:26Z",
"totalHits": 91421
}

A pageviews row (human traffic only):

{
"rowType": "pageviews",
"source": "wikimedia.org",
"title": "Web scraping",
"from": "2026-08-31",
"to": "2026-09-29",
"days": 29,
"totalViews": 161717,
"averageDaily": 5576.4,
"maxDay": { "date": "2026-09-02", "views": 48385 },
"series": [{ "date": "2026-09-01", "views": 4677 }, { "date": "2026-09-02", "views": 48385 }],
"agent": "user"
}

An entity:

{
"rowType": "entity",
"source": "wikidata.org",
"license": "CC0",
"attribution": "Wikidata contributors",
"entityId": "Q42",
"label": "Douglas Adams",
"description": "British science fiction writer and humorist (1952–2001)",
"aliases": ["Douglas Noël Adams", "Douglas Noel Adams", "Douglas N. Adams"],
"url": "https://www.wikidata.org/wiki/Q42",
"wikipediaUrl": "https://en.wikipedia.org/wiki/Douglas_Adams",
"sitelinkCount": 132,
"claimCount": 354,
"properties": {
"instanceOf": [{ "id": "Q5", "label": "human" }],
"officialWebsite": ["https://douglasadams.com"],
"image": ["https://commons.wikimedia.org/wiki/Special:FilePath/Douglas_adams_portrait.jpg"],
"country": null
},
"modified": "2026-09-29T09:06:14Z"
}

Quick start

You wantInput
Search results plus each hit's intro{ "queries": ["retrieval augmented generation"], "maxResults": 20, "fetchArticlesForSearch": true }
Full articles as plain text for RAG{ "titles": ["Web scraping", "Large language model"], "articleContent": "full", "maxTextChars": 200000 }
A whole category's articles{ "categories": ["Category:Machine learning"], "maxResults": 200 }
A 90-day pageviews trend for a list of titles{ "titles": ["ChatGPT", "Claude (language model)"], "articleContent": "none", "pageviews": "last90" }
The most-read articles of a day{ "topViews": ["2026-09-28"], "maxResults": 100 }
Company facts from Wikidata{ "wikidataSearch": ["Apify"], "wikidataIds": ["Q95", "Q312"] }
German Wikipedia{ "language": "de", "titles": ["Screen Scraping"] }

Input

FieldTypeDefaultWhat it does
languagestringenThe edition, as its subdomain: en, de, fr, es, ja, simple, zh-yue …
queriesstring[][]Full-text searches; one search-result row per hit
titlesstring[][]Titles (Web scraping, Web_scraping) or URLs; a URL's own language wins
categoriesstring[][]Category:Web scraping (prefix optional) → its articles, as article rows. Subcategories are not descended into
maxResultsint50Per query, category, top-views day and Wikidata search; max 10,000
articleContentenumintrointro (lead section), full (whole article), none (no article rows)
fetchArticlesForSearchboolfalseAlso an article row for every search hit
maxTextCharsint50,000Cap on text; textTruncated says when it bit
includeCategoriesbooltrueVisible categories, up to 50, hidden maintenance ones excluded
includeSectionsboolfalseSection outline; +1 request per article
includeLinksboolfalseArticle links, external links, backlink count; +1 request per article
includeRevisionsboolfalseLast 20 revisions — id, time, size, summary; never the editor; +1 request per article
pageviewsenumnonelast30, last90 (both end yesterday) or range
pageviewsFrom / pageviewsTostring—For range: YYYYMMDD or YYYY-MM-DD, from 2015-07-01
topViewsstring[][]Dates whose top list to return (YYYY-MM-DD)
keepSpecialPagesboolfalseKeep Main_Page, Special: and other non-article pages in a top list
wikidataIdsstring[][]Q42, P31 or wikidata.org/wiki/Q42 links
wikidataSearchstring[][]Names to look up; up to maxResults entities each
entityPropertiesstring[]13 defaultsProperty ids or bundled names; see Entity properties
includeRawClaimsboolfalseEvery statement as Wikidata's own JSON, in claimsRaw
maxConcurrencyint4Requests in flight across the whole run
maxRunSecsint240Time budget, 30–3600
proxyConfigurationobject{ "useApifyProxy": true }Datacenter; every endpoint answers it

At least one of queries, titles, categories, topViews, wikidataIds or wikidataSearch is needed.


Output reference

rowTypeEventWhat it is
articlearticleOne article: title, pageId, url, canonicalUrl, description, text, textTruncated, textChars, wordCount, wikidataId, categories, imageUrl, thumbnailUrl, lastModified, lastRevisionId, pageLengthBytes, isDisambiguation, redirectedFrom, coordinates, and with the options sections, links, linksCount, linksCountIsPartial, externalLinks, linksHereCount, lastRevisions. query and rank when it came from a search
search-resultsearch-resultOne hit: title, pageId, url, snippet (plain text), size, wordCount, lastModified, totalHits, query, rank
pageviewspageviewsOne title over a window: from, to, days, totalViews, averageDaily, maxDay, series, agent: "user" — or one top-list entry: date, rank, views, title, url
entityentityOne Wikidata entity: entityId, label, description, aliases, url, wikipediaUrl, sitelinkCount, claimCount, properties, claimsRaw, modified, and query/rank when found by name
diagnosticfreeAnything that could not be read, one row each

Every row carries ok, rowType, input, error, errorType, scrapedAt, source, sourceUrl (the request that produced it), license and attribution.

A few columns, precisely. textChars and wordCount describe the whole extract Wikipedia returned, even when maxTextChars cut text. lastModified on an article is MediaWiki's touched — the last time the page changed or was re-rendered, so it can be newer than the last edit. description comes from the article's short description; for the rare article that has none, the Actor asks Wikipedia's page summary instead, which is also the only source of coordinates. Search paging can return the same article twice when the index moves between pages; a repeat is skipped, so one query never bills one article twice. A redirect and its target, or two spellings of one title, are one article and one charge.

Entity properties

The 13 defaults are instanceOf (P31), country (P17), locatedIn (P131), headquarters (P159), inception (P571), officialWebsite (P856), employees (P1128), industry (P452), founder (P112), ceo (P169), legalForm (P1454), coordinates (P625) and image (P18). Sixty properties have names bundled (src/properties.json); any other id works too and is keyed by its id.

Values are flattened by type: an item becomes { id, label } with the label resolved in your language; a date keeps its real precision ("1952-03-11", "2015-06", or just "1974" when Wikidata only asserts a year); a quantity is { amount, unit, unitLabel }; a coordinate { latitude, longitude }; an image or logo a Commons URL that always serves the current file. Deprecated statements are never shown, and preferred ones win over normal ones. Every selected key is present on every row, null when the entity does not have it.

Labels are looked up in your language, then Wikidata's multilingual mul label, then English — many names (Q42's among them) now live only in mul.

Diagnostic rows

errorType is one of invalid-input, not-found, no-results, rate-limited, blocked, http, timeout, deadline, budget or upstream-format. None is ever charged for. A run whose entries were usable but produced only these — a title that does not exist, a search with no hits, Wikimedia refusing a request — finishes SUCCEEDED with zero results, a "0 results … Nothing was charged." status message and a free row per entry saying why; nothing is billed, start fee included. A run finishes FAILED only when there was nothing usable to attempt (no input, an invalid language, every entry invalid) or the Actor itself hit an error.


Languages

Any Wikipedia edition works: set language to its subdomain (de for de.wikipedia.org, zh-yue, simple …). Search, categories, the top list and pageviews use that edition; a title given as a URL uses the edition in its URL, so one run can mix editions. Wikidata labels, descriptions and aliases follow language with the mul → English fallback described above, and wikipediaUrl links to that edition's article when there is one, otherwise the English one.


Licensing

Wikipedia text is licensed CC BY-SA 4.0. You may reuse it, commercially too, if you attribute it (credit "Wikipedia contributors" and link the article — url on every row) and share alike (release what you build from the text under the same licence). Every Wikipedia-derived row carries license: "CC BY-SA 4.0" and attribution: "Wikipedia contributors" so the terms travel with the data. Pageview counts are published by the Wikimedia Foundation alongside that content and carry the same fields.

Wikidata is CC0 — no conditions — and entity rows say so (license: "CC0", attribution: "Wikidata contributors").

This Actor sells retrieval and shaping, not the content. Editor usernames are never collected; an article about a person is encyclopedic content, and no contact details of any kind — email, phone, address or personal social handles from Wikidata — are ever extracted, even into claimsRaw.


What you are never charged for

  • Every diagnostic row.
  • Titles, categories, ids and names that do not exist or have no data.
  • Anything Wikimedia refused (rate-limited, blocked) or that timed out.
  • A search hit the index repeated across pages, and a second spelling or redirect of an article already returned.
  • The extra requests behind an article (page summary, sections, links, revisions): they are included in the article's price.
  • A run that returns nothing at all: it finishes SUCCEEDED with zero results and bills nothing, start fee included.

Pricing

EventWhat it isFREEBRONZESILVERGOLD
actor-startOnce per run, only after a paid row$0.001$0.001$0.001$0.001
articleOne article$0.0006$0.0006$0.00048$0.00036
search-resultOne search hit$0.0002$0.0002$0.00016$0.00012
pageviewsOne title's pageviews, or one top-list entry$0.0003$0.0003$0.00024$0.00018
entityOne Wikidata entity$0.0005$0.0005$0.0004$0.0003
RunCost
The prefill — 20 hits, 2 intros, 2 pageviews rows, 1 entity$0.0073
100 search results + 100 full articles + their 30-day pageviews$0.111
1,000 article intros$0.601
30-day pageviews for 500 titles$0.151
100 Wikidata entities$0.051

Charging is charge-after-push: rows are in your dataset before the event is recorded. ACTOR_MAX_TOTAL_CHARGE_USD is respected — the run stops fetching when the budget cannot cover the next row, adds a free budget row, and finishes SUCCEEDED saying so.


Proxy and politeness

The default is { "useApifyProxy": true } — Apify's datacenter pool. Every Wikimedia endpoint answered from it in the capture probe, and the page summary also answered with no proxy at all; residential buys nothing here.

Every request identifies itself as insight-solutions-wikipedia-api/0.1.0 (+https://apify.com/insight.solutions), as Wikimedia's API etiquette asks. Requests start at least 100 ms apart across the whole run — at most 10 a second, against the 50 (Action API) and 200 (REST) per second Wikimedia tolerates from one client — with at most maxConcurrency in flight, and successive pages of one search or category pause 250–600 ms. A 429 or 403 rotates the proxy session once; a second refusal becomes a free rate-limited or blocked row rather than a retry loop.


Use it from an AI agent, or from code

One JSON object in, one flat array out. The Actor runs with limited permissions, uses pay-per-event pricing and never enters Standby, so it works over the Apify MCP server and with x402 agentic payments.

curl -X POST "https://api.apify.com/v2/acts/insight.solutions~wikipedia-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"queries":["vector database"],"maxResults":10,"fetchArticlesForSearch":true,"articleContent":"full"}'
# pip install apify-client
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("insight.solutions/wikipedia-api").call(run_input={
"titles": ["Retrieval-augmented generation", "Vector database"],
"articleContent": "full",
"pageviews": "last30",
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row["rowType"] == "article":
print(row["title"], row["wordCount"], row["url"], row["license"])
elif row["rowType"] == "pageviews":
print(row["title"], row["totalViews"], row["averageDaily"])

FAQ

Do I need an API key? No. Every Wikimedia API this Actor reads is public and unauthenticated, and it sends no cookies and no credentials of anybody's.

Why is the text missing tables and infoboxes? Because it is Wikipedia's own plain-text extract, which is prose: tables, infoboxes, references and templates are dropped. Structured facts are better taken from the article's Wikidata entity — wikidataId is on every article row.

Why do the pageviews differ from the page's own statistics tool? This Actor counts human traffic only (agent: user). Tools that default to "all agents" include crawlers and self-identified bots, which on some articles are a large share.

Why does a 30-day window have 29 days? days counts days with data; the API leaves out days it has nothing for. The window itself ends yesterday, because a day's counts are published the next day.

Can I get more than 1,000 search results? Yes, up to maxResults (10,000), 50 per request. Wikipedia's search reports totalHits on every row so you know how many exist.

Where do the entity labels come from? One extra wbgetentities pass per run resolves every item a selected property points at — up to 500 ids, 50 per request — in your language.

What happens if the format changes? If a response stops parsing, the run returns a free upstream-format row naming the request rather than a wrong value. A run that returned no real row finishes SUCCEEDED with zero results and bills nothing.

Is this affiliated with Wikipedia or the Wikimedia Foundation? No. It reads their public APIs, following their API etiquette, and links every row back to its source.


Limitations

  • Extracts drop tables, infoboxes and references, and are plain text only.
  • Full text is one request per article; intros batch twenty. A thousand full articles take a few minutes.
  • Pageviews are user agent only — crawlers and bots are excluded — and the newest available day is yesterday.
  • Categories are not recursive: a category's own articles, not its subcategories'.
  • coordinates is filled only for articles without a short description, where the page summary is consulted.
  • The top-list filter knows English namespace names plus the special and project namespaces of the largest editions; set keepSpecialPages and filter yourself for other editions.
  • Label resolution stops at 500 ids per run; beyond that, label is null and the id is still there.
  • The upstream format may change. Wikimedia versions its APIs carefully, but during this Actor's capture probe one REST endpoint (page/related, not used here) already answered "This API endpoint is being decommissioned". A response that stops parsing becomes a free upstream-format row, never a wrong value.

Our other Actors

Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.

Video, audio & social

News, documents & the web

Business, finance & jobs

Apps & games