Wikipedia & Wikidata API — Articles, Search & Pageviews
Pricing
from $0.36 / 1,000 articles
Wikipedia & Wikidata API — Articles, Search & Pageviews
Wikipedia search, article text (intro or full), descriptions, categories and Wikidata ids; daily human pageviews and a day's top 1,000; Wikidata entities with flattened properties. Any language edition, no API key, no browser. Diagnostics and empty runs are free.
Pricing
from $0.36 / 1,000 articles
Rating
0.0
(0)
Developer
Insight Solutions
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
8 hours ago
Last modified
Categories
Share
Wikipedia and Wikidata as clean rows, keyless, in any language edition. Search articles; fetch an article's plain text (the intro or the whole thing) with its short description, categories, Wikidata id, image, sections, links and recent revisions; get daily human pageviews for any title, or a day's top 1,000; look up Wikidata entities by id or by name with labels, descriptions, aliases, sitelinks and the properties you choose, flattened into readable values.
Read straight from Wikimedia's public APIs — the MediaWiki Action API, the pageviews API and Wikidata — with a descriptive User-Agent and a pace far below what Wikimedia allows. No key, no login, no browser, and no page HTML: article text comes out of Wikipedia's own plain-text extracts.
At a glance
Input — this is the Store prefill; paste it and run:
{ "language": "en", "queries": ["web scraping"], "titles": ["Web scraping", "Large language model"],"categories": [], "maxResults": 20, "articleContent": "intro", "fetchArticlesForSearch": false,"includeCategories": true, "includeSections": false, "includeLinks": false, "pageviews": "last30","topViews": [], "wikidataIds": ["Q665452"], "wikidataSearch": [], "includeRawClaims": false,"maxConcurrency": 4, "maxRunSecs": 240, "proxyConfiguration": { "useApifyProxy": true } }
Twenty search hits for "web scraping", the intros of two articles, their last thirty days of pageviews and the Wikidata entity for web scraping — about six requests and a few seconds.
Output — one row per result, four kinds of paid row in one schema. The
fields you will use most are text, description and categories on an
article; snippet on a search-result; totalViews, averageDaily and
series on a pageviews row; label, description and properties on an
entity (full list under Output reference). Every row carries license and
attribution. Anything that could not be read comes back as a free diagnostic
row (ok: false, errorType, error) instead of a charge.
Price — $0.60 per 1,000 articles, $0.20 per 1,000 search results, $0.30 per 1,000 pageviews rows, $0.50 per 1,000 Wikidata entities (+ $0.001 per run); diagnostics and empty runs free; no API key, no browser, limited permissions, works over the Apify MCP server and with x402.
From code —
client.actor("insight.solutions/wikipedia-api").call(run_input={…}) with
apify-client, or
POST https://api.apify.com/v2/acts/insight.solutions~wikipedia-api/run-sync-get-dataset-itemsWhat you get
One row per result. Same columns on every row, null where a column does not
apply. Abridged real rows:
An article (articleContent: "intro"):
{"rowType": "article","input": "Web scraping","source": "en.wikipedia.org","license": "CC BY-SA 4.0","attribution": "Wikipedia contributors","title": "Web scraping","pageId": 2696619,"url": "https://en.wikipedia.org/wiki/Web_scraping","description": "Method of extracting data from websites","text": "Web scraping, web harvesting, or web data extraction is data scraping used for extracting data from websites. Web scraping software may directly access the World Wide Web …","textTruncated": false,"textChars": 2809,"wordCount": 436,"wikidataId": "Q665452","categories": ["Web scraping"],"lastModified": "2026-09-30T16:28:17Z","lastRevisionId": 1375681792,"pageLengthBytes": 34836,"isDisambiguation": false}
A search result:
{"rowType": "search-result","query": "large language model","rank": 1,"title": "Large language model","pageId": 73248112,"url": "https://en.wikipedia.org/wiki/Large_language_model","snippet": "A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially","size": 136413,"wordCount": 13579,"lastModified": "2026-09-29T11:25:26Z","totalHits": 91421}
A pageviews row (human traffic only):
{"rowType": "pageviews","source": "wikimedia.org","title": "Web scraping","from": "2026-08-31","to": "2026-09-29","days": 29,"totalViews": 161717,"averageDaily": 5576.4,"maxDay": { "date": "2026-09-02", "views": 48385 },"series": [{ "date": "2026-09-01", "views": 4677 }, { "date": "2026-09-02", "views": 48385 }],"agent": "user"}
An entity:
{"rowType": "entity","source": "wikidata.org","license": "CC0","attribution": "Wikidata contributors","entityId": "Q42","label": "Douglas Adams","description": "British science fiction writer and humorist (1952–2001)","aliases": ["Douglas Noël Adams", "Douglas Noel Adams", "Douglas N. Adams"],"url": "https://www.wikidata.org/wiki/Q42","wikipediaUrl": "https://en.wikipedia.org/wiki/Douglas_Adams","sitelinkCount": 132,"claimCount": 354,"properties": {"instanceOf": [{ "id": "Q5", "label": "human" }],"officialWebsite": ["https://douglasadams.com"],"image": ["https://commons.wikimedia.org/wiki/Special:FilePath/Douglas_adams_portrait.jpg"],"country": null},"modified": "2026-09-29T09:06:14Z"}
Quick start
| You want | Input |
|---|---|
| Search results plus each hit's intro | { "queries": ["retrieval augmented generation"], "maxResults": 20, "fetchArticlesForSearch": true } |
| Full articles as plain text for RAG | { "titles": ["Web scraping", "Large language model"], "articleContent": "full", "maxTextChars": 200000 } |
| A whole category's articles | { "categories": ["Category:Machine learning"], "maxResults": 200 } |
| A 90-day pageviews trend for a list of titles | { "titles": ["ChatGPT", "Claude (language model)"], "articleContent": "none", "pageviews": "last90" } |
| The most-read articles of a day | { "topViews": ["2026-09-28"], "maxResults": 100 } |
| Company facts from Wikidata | { "wikidataSearch": ["Apify"], "wikidataIds": ["Q95", "Q312"] } |
| German Wikipedia | { "language": "de", "titles": ["Screen Scraping"] } |
Input
| Field | Type | Default | What it does |
|---|---|---|---|
language | string | en | The edition, as its subdomain: en, de, fr, es, ja, simple, zh-yue … |
queries | string[] | [] | Full-text searches; one search-result row per hit |
titles | string[] | [] | Titles (Web scraping, Web_scraping) or URLs; a URL's own language wins |
categories | string[] | [] | Category:Web scraping (prefix optional) → its articles, as article rows. Subcategories are not descended into |
maxResults | int | 50 | Per query, category, top-views day and Wikidata search; max 10,000 |
articleContent | enum | intro | intro (lead section), full (whole article), none (no article rows) |
fetchArticlesForSearch | bool | false | Also an article row for every search hit |
maxTextChars | int | 50,000 | Cap on text; textTruncated says when it bit |
includeCategories | bool | true | Visible categories, up to 50, hidden maintenance ones excluded |
includeSections | bool | false | Section outline; +1 request per article |
includeLinks | bool | false | Article links, external links, backlink count; +1 request per article |
includeRevisions | bool | false | Last 20 revisions — id, time, size, summary; never the editor; +1 request per article |
pageviews | enum | none | last30, last90 (both end yesterday) or range |
pageviewsFrom / pageviewsTo | string | — | For range: YYYYMMDD or YYYY-MM-DD, from 2015-07-01 |
topViews | string[] | [] | Dates whose top list to return (YYYY-MM-DD) |
keepSpecialPages | bool | false | Keep Main_Page, Special: and other non-article pages in a top list |
wikidataIds | string[] | [] | Q42, P31 or wikidata.org/wiki/Q42 links |
wikidataSearch | string[] | [] | Names to look up; up to maxResults entities each |
entityProperties | string[] | 13 defaults | Property ids or bundled names; see Entity properties |
includeRawClaims | bool | false | Every statement as Wikidata's own JSON, in claimsRaw |
maxConcurrency | int | 4 | Requests in flight across the whole run |
maxRunSecs | int | 240 | Time budget, 30–3600 |
proxyConfiguration | object | { "useApifyProxy": true } | Datacenter; every endpoint answers it |
At least one of queries, titles, categories, topViews, wikidataIds or
wikidataSearch is needed.
Output reference
rowType | Event | What it is |
|---|---|---|
article | article | One article: title, pageId, url, canonicalUrl, description, text, textTruncated, textChars, wordCount, wikidataId, categories, imageUrl, thumbnailUrl, lastModified, lastRevisionId, pageLengthBytes, isDisambiguation, redirectedFrom, coordinates, and with the options sections, links, linksCount, linksCountIsPartial, externalLinks, linksHereCount, lastRevisions. query and rank when it came from a search |
search-result | search-result | One hit: title, pageId, url, snippet (plain text), size, wordCount, lastModified, totalHits, query, rank |
pageviews | pageviews | One title over a window: from, to, days, totalViews, averageDaily, maxDay, series, agent: "user" — or one top-list entry: date, rank, views, title, url |
entity | entity | One Wikidata entity: entityId, label, description, aliases, url, wikipediaUrl, sitelinkCount, claimCount, properties, claimsRaw, modified, and query/rank when found by name |
diagnostic | free | Anything that could not be read, one row each |
Every row carries ok, rowType, input, error, errorType, scrapedAt,
source, sourceUrl (the request that produced it), license and
attribution.
A few columns, precisely. textChars and wordCount describe the whole
extract Wikipedia returned, even when maxTextChars cut text. lastModified
on an article is MediaWiki's touched — the last time the page changed or was
re-rendered, so it can be newer than the last edit. description comes from the
article's short description; for the rare article that has none, the Actor asks
Wikipedia's page summary instead, which is also the only source of
coordinates. Search paging can return the same article twice when the index
moves between pages; a repeat is skipped, so one query never bills one article
twice. A redirect and its target, or two spellings of one title, are one article
and one charge.
Entity properties
The 13 defaults are instanceOf (P31), country (P17), locatedIn (P131),
headquarters (P159), inception (P571), officialWebsite (P856),
employees (P1128), industry (P452), founder (P112), ceo (P169),
legalForm (P1454), coordinates (P625) and image (P18). Sixty properties
have names bundled (src/properties.json); any other id works too and is keyed
by its id.
Values are flattened by type: an item becomes { id, label } with the label
resolved in your language; a date keeps its real precision ("1952-03-11",
"2015-06", or just "1974" when Wikidata only asserts a year); a quantity is
{ amount, unit, unitLabel }; a coordinate { latitude, longitude }; an image
or logo a Commons URL that always serves the current file. Deprecated statements
are never shown, and preferred ones win over normal ones. Every selected key is
present on every row, null when the entity does not have it.
Labels are looked up in your language, then Wikidata's multilingual mul label,
then English — many names (Q42's among them) now live only in mul.
Diagnostic rows
errorType is one of invalid-input, not-found, no-results,
rate-limited, blocked, http, timeout, deadline, budget or
upstream-format. None is ever charged for. A run whose entries were usable but
produced only these — a title that does not exist, a search with no hits,
Wikimedia refusing a request — finishes SUCCEEDED with zero results, a
"0 results … Nothing was charged." status message and a free row per entry
saying why; nothing is billed, start fee included. A run finishes FAILED
only when there was nothing usable to attempt (no input, an invalid language,
every entry invalid) or the Actor itself hit an error.
Languages
Any Wikipedia edition works: set language to its subdomain (de for
de.wikipedia.org, zh-yue, simple …). Search, categories, the top list and
pageviews use that edition; a title given as a URL uses the edition in its URL,
so one run can mix editions. Wikidata labels, descriptions and aliases follow
language with the mul → English fallback described above, and
wikipediaUrl links to that edition's article when there is one, otherwise the
English one.
Licensing
Wikipedia text is licensed CC BY-SA 4.0. You may reuse it, commercially
too, if you attribute it (credit "Wikipedia contributors" and link the
article — url on every row) and share alike (release what you build from
the text under the same licence). Every Wikipedia-derived row carries
license: "CC BY-SA 4.0" and attribution: "Wikipedia contributors" so the
terms travel with the data. Pageview counts are published by the Wikimedia
Foundation alongside that content and carry the same fields.
Wikidata is CC0 — no conditions — and entity rows say so
(license: "CC0", attribution: "Wikidata contributors").
This Actor sells retrieval and shaping, not the content. Editor usernames are
never collected; an article about a person is encyclopedic content, and no
contact details of any kind — email, phone, address or personal social handles
from Wikidata — are ever extracted, even into claimsRaw.
What you are never charged for
- Every
diagnosticrow. - Titles, categories, ids and names that do not exist or have no data.
- Anything Wikimedia refused (
rate-limited,blocked) or that timed out. - A search hit the index repeated across pages, and a second spelling or redirect of an article already returned.
- The extra requests behind an article (page summary, sections, links, revisions): they are included in the article's price.
- A run that returns nothing at all: it finishes SUCCEEDED with zero results and bills nothing, start fee included.
Pricing
| Event | What it is | FREE | BRONZE | SILVER | GOLD |
|---|---|---|---|---|---|
actor-start | Once per run, only after a paid row | $0.001 | $0.001 | $0.001 | $0.001 |
article | One article | $0.0006 | $0.0006 | $0.00048 | $0.00036 |
search-result | One search hit | $0.0002 | $0.0002 | $0.00016 | $0.00012 |
pageviews | One title's pageviews, or one top-list entry | $0.0003 | $0.0003 | $0.00024 | $0.00018 |
entity | One Wikidata entity | $0.0005 | $0.0005 | $0.0004 | $0.0003 |
| Run | Cost |
|---|---|
| The prefill — 20 hits, 2 intros, 2 pageviews rows, 1 entity | $0.0073 |
| 100 search results + 100 full articles + their 30-day pageviews | $0.111 |
| 1,000 article intros | $0.601 |
| 30-day pageviews for 500 titles | $0.151 |
| 100 Wikidata entities | $0.051 |
Charging is charge-after-push: rows are in your dataset before the event is
recorded. ACTOR_MAX_TOTAL_CHARGE_USD is respected — the run stops fetching
when the budget cannot cover the next row, adds a free budget row, and
finishes SUCCEEDED saying so.
Proxy and politeness
The default is { "useApifyProxy": true } — Apify's datacenter pool. Every
Wikimedia endpoint answered from it in the capture probe, and the page summary
also answered with no proxy at all; residential buys nothing here.
Every request identifies itself as
insight-solutions-wikipedia-api/0.1.0 (+https://apify.com/insight.solutions),
as Wikimedia's API etiquette asks. Requests start at least 100 ms apart across
the whole run — at most 10 a second, against the 50 (Action API) and 200 (REST)
per second Wikimedia tolerates from one client — with at most maxConcurrency
in flight, and successive pages of one search or category pause 250–600 ms. A
429 or 403 rotates the proxy session once; a second refusal becomes a free
rate-limited or blocked row rather than a retry loop.
Use it from an AI agent, or from code
One JSON object in, one flat array out. The Actor runs with limited permissions, uses pay-per-event pricing and never enters Standby, so it works over the Apify MCP server and with x402 agentic payments.
curl -X POST "https://api.apify.com/v2/acts/insight.solutions~wikipedia-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"queries":["vector database"],"maxResults":10,"fetchArticlesForSearch":true,"articleContent":"full"}'
# pip install apify-clientfrom apify_client import ApifyClientclient = ApifyClient("<APIFY_TOKEN>")run = client.actor("insight.solutions/wikipedia-api").call(run_input={"titles": ["Retrieval-augmented generation", "Vector database"],"articleContent": "full","pageviews": "last30",})for row in client.dataset(run["defaultDatasetId"]).iterate_items():if row["rowType"] == "article":print(row["title"], row["wordCount"], row["url"], row["license"])elif row["rowType"] == "pageviews":print(row["title"], row["totalViews"], row["averageDaily"])
FAQ
Do I need an API key? No. Every Wikimedia API this Actor reads is public and unauthenticated, and it sends no cookies and no credentials of anybody's.
Why is the text missing tables and infoboxes?
Because it is Wikipedia's own plain-text extract, which is prose: tables,
infoboxes, references and templates are dropped. Structured facts are better
taken from the article's Wikidata entity — wikidataId is on every article row.
Why do the pageviews differ from the page's own statistics tool?
This Actor counts human traffic only (agent: user). Tools that default to
"all agents" include crawlers and self-identified bots, which on some articles
are a large share.
Why does a 30-day window have 29 days?
days counts days with data; the API leaves out days it has nothing for. The
window itself ends yesterday, because a day's counts are published the next day.
Can I get more than 1,000 search results?
Yes, up to maxResults (10,000), 50 per request. Wikipedia's search reports
totalHits on every row so you know how many exist.
Where do the entity labels come from?
One extra wbgetentities pass per run resolves every item a selected property
points at — up to 500 ids, 50 per request — in your language.
What happens if the format changes?
If a response stops parsing, the run returns a free upstream-format row naming
the request rather than a wrong value. A run that returned no real row finishes
SUCCEEDED with zero results and bills nothing.
Is this affiliated with Wikipedia or the Wikimedia Foundation? No. It reads their public APIs, following their API etiquette, and links every row back to its source.
Limitations
- Extracts drop tables, infoboxes and references, and are plain text only.
- Full text is one request per article; intros batch twenty. A thousand full articles take a few minutes.
- Pageviews are
useragent only — crawlers and bots are excluded — and the newest available day is yesterday. - Categories are not recursive: a category's own articles, not its subcategories'.
coordinatesis filled only for articles without a short description, where the page summary is consulted.- The top-list filter knows English namespace names plus the special and
project namespaces of the largest editions; set
keepSpecialPagesand filter yourself for other editions. - Label resolution stops at 500 ids per run; beyond that,
labelis null and the id is still there. - The upstream format may change. Wikimedia versions its APIs carefully, but
during this Actor's capture probe one REST endpoint (
page/related, not used here) already answered "This API endpoint is being decommissioned". A response that stops parsing becomes a freeupstream-formatrow, never a wrong value.
Our other Actors
Every Insight Solutions Actor is pay-per-result with no browser, no login and no API key, and every one of them returns free diagnostic rows instead of billing for failures. Prices are per 1,000 results.
Video, audio & social
- YouTube Transcript API — captions as timed segments, text, SRT or VTT, with language fallback and translation.
- YouTube Comments API — comments and replies with likes, pinned and hearted flags, newest or top sort.
- YouTube Channel API — a channel's videos, Shorts and live streams, plus YouTube search.
- Podcast Search, Episodes & Charts API — Apple Podcasts search, charts and full episode feeds.
- Bluesky Scraper — profiles, posts, followers and follows from the public AT Protocol API.
- Telegram Channel Scraper — posts, views and channel stats from public Telegram channels.
- Substack Scraper — posts with full free text, comments and publication profiles.
- Hacker News API — stories, comments, users, front page and a structured "Who is hiring?" parser from the official HN APIs.
- Discourse Forum API — topics, posts and categories from any Discourse community via its own JSON endpoints, usernames only.
News, documents & the web
- Google News Search, Topics & Real Article URLs — news search and topic feeds with the publisher's real URL decoded.
- Website to Markdown — Content Extractor for LLMs & RAG — any site as clean Markdown, text and heading-aware chunks.
- Internet Archive API — archive.org search, item metadata, files and reviews.
- Wayback Machine Toolkit — archived URL inventories, snapshots and text diffs between dates.
- Website Technology Detector — the tech stack behind any site, with the evidence for each detection.
- Domain Intelligence API — DNS, RDAP registration, TLS certificate and HTTP facts in one row per domain.
- SEO Page Audit — sitemap crawl with on-page checks, structured data and broken-link reports.
- Keyword Suggestions API — Google, YouTube, Bing, Amazon and eBay autocomplete with alphabet and question expansions.
- Website Contact Extractor — emails, phone numbers and social profiles from any list of websites.
- Web Search Results API — Bing and DuckDuckGo organic results with snippets, no key, no browser.
- Company Enrichment API — a domain in, a company profile out: firmographics, contacts, tech stack, DNS and hiring signal.
- Company Dossier API — one company in, twelve sections out: profile, tech, contacts, DNS, open roles, news, SEC filings, federal awards, recalls, YC batch and apps.
- Press Releases API — GlobeNewswire and PR Newswire releases plus any newsroom feed, by keyword, company, ticker or subject.
- Federal Register API — rules, proposed rules, notices and the Public Inspection desk with dockets, comment deadlines and CFR references.
- Academic Papers Search API — OpenAlex, Crossref, arXiv and PubMed in one row per paper: abstract, citations, open-access PDF, authors and venue.
- RSS & Atom Feed Monitor — any RSS, Atom or JSON feed (or an OPML file) in, only the new items out, with keyword filters and a webhook.
- Website Change Monitor — watch any pages, diff the text between runs, get change rows with added/removed lines, keyword alerts and a webhook.
Business, finance & jobs
- Congress & Insider Trades API — STOCK Act periodic transaction reports and SEC Form 4 insider trades in one schema.
- Federal Contracts, Grants & Lobbying API — SAM.gov opportunities, USAspending awards, Grants.gov notices and Senate lobbying filings in one schema.
- SEC EDGAR API — filings, XBRL financials and full-text search by ticker or CIK.
- Clinical Trials & FDA API — ClinicalTrials.gov studies plus openFDA recalls, labels, approvals, 510(k)s and adverse-event reports.
- Product & Vehicle Recalls API — CPSC, NHTSA, FDA and USDA recalls, vehicle complaints and ratings, plus a VIN decoder.
- Y Combinator Companies, Batches & Founders — the YC directory with founders and social links, filterable by batch, industry and hiring status.
- Career Site Jobs API — jobs straight from Greenhouse, Lever, Ashby, Workable and 10+ other ATS career sites.
- New Job Postings Monitor — new, closed and changed postings on the career sites you watch.
- Hiring Signals API — Open Roles & Hiring Surge by Company — one row per company per run: open roles, what opened and closed, department and seniority breakdowns, and a hiring-surge flag.
- Remote Jobs API — RemoteOK, Remotive, We Work Remotely, Himalayas, Jobicy and more in one schema, deduplicated.
- Shopify Products API — any Shopify store's catalogue, variants, prices and stock signals.
- Shopify Store Monitor — price drops, sales, restocks, sell-outs and new products on any Shopify store, one row per change.
- Public Tenders API — EU TED, UK Find a Tender and Contracts Finder notices by keyword, CPV code, country, stage and deadline.
- Nonprofit & IRS 990 Lookup API — search US nonprofits and get EIN, NTEE code and multi-year Form 990 financials.
- OpenStreetMap Places API — businesses and points of interest by category and area from OpenStreetMap: name, address, coordinates, website, phone, opening hours.
Apps & games
- App Store & Google Play Reviews API — reviews from both stores with ratings, versions and developer replies.
- App Store Top Charts & App Search API — Apple top charts by country and genre, plus app search and details.
- App Store Keyword Rank Tracker — where any app ranks for any keyword on the App Store and Google Play, with rank changes and ASO suggestions.
- Steam Reviews API — Steam reviews with playtime, helpfulness and game details.
- Steam Game Data API — prices, tags, review scores, live player counts and top charts.