Wikipedia Scraper — Search, Summaries & Full Text avatar

Wikipedia Scraper — Search, Summaries & Full Text

Pricing

from $0.70 / 1,000 result items

Go to Apify Store
Wikipedia Scraper — Search, Summaries & Full Text

Wikipedia Scraper — Search, Summaries & Full Text

Search Wikipedia and get article summaries or full plain-text content in any language: title, extract, description, thumbnail, coordinates, categories, links count, last edit. Built for research and RAG agents. No API key.

Pricing

from $0.70 / 1,000 result items

Rating

0.0

(0)

Developer

Samat Makatov

Samat Makatov

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 hours ago

Last modified

Share

Wikipedia Scraper — summaries, full articles, Wikidata facts & pageviews in any language

Turn a list of names, topics or page titles into structured Wikipedia data: relevance-ranked search or exact title lookup in any of the ~340 language editions, lead summary or the complete plain-text article split into sections, categories, links, inter-language links, images, the linked Wikidata record (industry, headquarters, CEO, revenue, employees, ticker, ISIN, LEI, website, social handles…) and pageview statistics for attention tracking. Built for due-diligence and KYB teams, market researchers, content/SEO teams, and RAG pipelines.

No API key, no proxy, no browser — official Wikimedia REST and Action APIs.

Use cases

  • Company / person due diligenceincludeWikidata: true gives industry, HQ, founders, CEO, parent, subsidiaries, employees, revenue (with year), ticker, ISIN, LEI, OpenCorporates & Crunchbase ids, official site and social links, plus the encyclopedic summary.
  • Entity enrichment for CRM / lead lists — feed 500 company names via queries, keep fields: ["query","title","description","extract","wikidataId","url"].
  • Attention / trend monitoringincludePageviews: true, pageviewsDays: 90, includeDailyPageviews: true on brands, competitors or topics; alert on spikes.
  • RAG / knowledge base ingestionmode: "full", includeSections: true gives clean plain text with a table of contents; maxTextChars caps size.
  • Multilingual localisation researchincludeLangLinks: true lists which languages cover a topic; run the same queries with lang: "de", "kk", "zh" for local perspectives.
  • Topic mapping / SEO content clustersincludeLinks: true, includeCategories: true on seed articles to build a related-topics graph.

Input

FieldTypeDefaultAllowed values / notes
queriesstring[]Required. One entry per lookup: free text, an entity name or an exact page title. Up to a few thousand per run.
langstringenAny Wikipedia language code (en, de, fr, es, ru, zh, ja, ar, kk, uz, simple, zh-yue…). See Reference.
searchModestringtexttext (relevance search like the site search box), title (prefix match on titles), nearmatch (best single title match, tolerant to case/diacritics), exact (treat the query as the page title; follows redirects).
resultsPerQueryinteger11–50 top hits saved per query (ignored for exact).
skipDisambiguationbooleanfalseSkip disambiguation list pages (always flagged via isDisambiguation).
modestringsummarysummary (lead paragraph via REST summary), full (complete plain-text article → fullText), metadata (no text).
maxTextCharsinteger0Truncate fullText on a word boundary (0 = whole article); fullTextChars keeps the real length.
includeSectionsbooleanfalseTable of contents [{level, title, chars}] (chars only in full mode; anchor in other modes).
includeCategoriesbooleanfalseVisible categories (≤100).
includeLinksbooleanfalseTitles of linked articles.
maxLinksinteger501–500 cap for includeLinks.
includeLangLinksbooleanfalselanguages (codes), languagesCount, langLinks [{lang, title, url}].
includeImagesbooleanfalseUp to 20 images with URL and caption (thumbnail / originalImage are always included).
includeWikidatabooleanfalseCurated structured facts from the linked Wikidata item (see Reference → Wikidata properties). +1–2 requests per page.
includePageviewsbooleanfalseHuman pageviews, all platforms, for the last pageviewsDays: total, avgPerDay, peakDay, peakViews.
pageviewsDaysinteger301–365, window ends yesterday (UTC).
includeDailyPageviewsbooleanfalseAdds pageviews.daily [{date, views}].
dedupebooleantrueSave each page once even if several queries resolve to it.
fieldsstring[][]Keep only these output fields.

Reference

Language editions (lang)

Any live edition works; these are the built-in known codes (unknown codes only log a warning):

CodeLanguageCodeLanguageCodeLanguage
enEnglishdeGermanfrFrench
esSpanishitItalianptPortuguese
ruRussianjaJapanesezhChinese
arArabicnlDutchplPolish
svSwedishukUkrainianviVietnamese
trTurkishfaPersiankoKorean
idIndonesiancsCzechheHebrew
huHungarianfiFinnishno / nnNorwegian
daDanishroRomanianelGreek
bgBulgariansrSerbianhrCroatian
skSlovakslSlovenianltLithuanian
lvLatvianetEstonianthThai
hiHindibnBengalitaTamil
teTelugumlMalayalammrMarathi
urUrdumsMalaycaCatalan
euBasqueglGaliciankkKazakh
uzUzbekkyKyrgyztgTajik
tkTurkmenazAzerbaijanikaGeorgian
hyArmenianmnMongolianbeBelarusian
ttTatarbaBashkirsqAlbanian
mkMacedonianbsBosnianisIcelandic
gaIrishcyWelshafAfrikaans
swSwahiliamAmharichaHausa
yoYorubaneNepalisiSinhala
myBurmesekmKhmerloLao
tlTagaloglaLatineoEsperanto
simpleSimple Englishzh-yueCantonesecebCebuano
warWarayarzEgyptian ArabicckbCentral Kurdish
kuKurdishpsPashtosdSindhi
paPunjabiguGujaratiknKannada
orOdiaasAssamese

Full list: https://meta.wikimedia.org/wiki/List_of_Wikipedias

Wikidata properties extracted (includeWikidata)

Output keyPropertyOutput keyProperty
instanceOfP31countryP17
headquartersP159inception / dissolvedP571 / P576
foundedByP112ceo / chairpersonP169 / P488
industryP452legalFormP1454
ownedBy / parentOrganizationP127 / P749subsidiariesP355
employees ({value, year})P1128revenue / netProfit / totalAssets ({amount, currencyItem, year})P2139 / P2295 / P2403
websiteP856coordinatesP625
tickerSymbol / stockExchangeP249 / P414isin / leiP946 / P1278
openCorporatesId / crunchbaseIdP1320 / P2088rorId / gridIdP6782 / P2427
twitter / linkedinCompany / facebook / youtubeChannelP2002 / P4264 / P2013 / P2397image / logoP18 / P154
population / capitalP1082 / P36dateOfBirth / dateOfDeathP569 / P570
citizenship / occupation / employerP27 / P106 / P108educatedAt / positionHeld / memberOfP69 / P39 / P463
ownerOfP1830phone / emailP1329 / P968
streetAddress / postalCodeP6375 / P281officialName / shortNameP1448 / P1813

Item values (people, places, companies, currencies) are returned as labels in the edition language, falling back to English, then Wikidata's language-neutral mul label (used by most person items since 2024), then any language. The original Q-ids are kept in wikidata.entityIds (e.g. { "ceo": "Q106028933", "foundedBy": ["Q483382", "Q332591", "Q19837"] }) for joins.

Preferred-rank claims win, deprecated ones are dropped, time-series claims (revenue, employees) return the latest value with its year. Item values are resolved to labels in lang, falling back to English.

Examples

KYB / due diligence on a list of companies

{ "queries": ["Kaspi.kz", "Air Astana", "Halyk Bank", "KazMunayGas"], "searchMode": "nearmatch", "includeWikidata": true, "includeLangLinks": true, "fields": ["query", "title", "description", "extract", "wikidata", "languagesCount", "url"] }

Brand attention tracker (run weekly)

{ "queries": ["Tesla, Inc.", "BYD Auto", "Rivian"], "searchMode": "exact", "mode": "metadata", "includePageviews": true, "pageviewsDays": 90, "includeDailyPageviews": true }

Full articles for a RAG index, with table of contents

{ "queries": ["Stablecoin", "Central bank digital currency", "Payment system"], "searchMode": "exact", "mode": "full", "includeSections": true, "includeCategories": true, "maxTextChars": 60000 }

Topic research in Russian — top 5 hits per query

{ "queries": ["агропромышленный комплекс Казахстана", "цифровизация сельского хозяйства"], "lang": "ru", "resultsPerQuery": 5, "skipDisambiguation": true }

Related-topics graph for content planning

{ "queries": ["Agentic AI"], "searchMode": "exact", "includeLinks": true, "maxLinks": 200, "includeCategories": true, "mode": "metadata" }

Output

One item per page found. Queries with no result and failed queries are not items and are not charged — they are listed in the SUMMARY record and the run's status message. Example (trimmed):

{
"query": "Astana International Financial Centre",
"rank": 1,
"found": true,
"lang": "en",
"pageId": 57219556,
"title": "Astana International Financial Centre",
"requestedTitle": null,
"description": "Financial hub in Astana, Kazakhstan",
"isDisambiguation": false,
"extract": "The Astana International Financial Centre (AIFC) is a financial hub in Astana, Kazakhstan …",
"snippet": "The Astana International Financial Centre (AIFC) is a financial hub …",
"thumbnail": "https://upload.wikimedia.org/wikipedia/commons/thumb/…/320px-AIFC.jpg",
"originalImage": "https://upload.wikimedia.org/wikipedia/commons/…/AIFC.jpg",
"coordinates": { "lat": 51.09, "lon": 71.42 },
"url": "https://en.wikipedia.org/wiki/Astana_International_Financial_Centre",
"lastEdited": "2026-08-30T11:02:41Z",
"revisionId": 1312345678,
"wordCount": 1874,
"wikidataId": "Q28155597",
"wikidata": { "id": "Q28155597", "instanceOf": "financial centre", "country": "Kazakhstan", "inception": "2018-01-01", "headquarters": "Astana", "website": "https://aifc.kz", "twitter": "AIFC_KZ", "url": "https://www.wikidata.org/wiki/Q28155597", "entityIds": { "instanceOf": "Q338313", "country": "Q232", "headquarters": "Q1520" } },
"pageviews": { "days": 30, "from": "2026-08-14", "to": "2026-09-12", "total": 2140, "avgPerDay": 71.3, "peakDay": "2026-09-02", "peakViews": 188 },
"fetchedAt": "2026-09-13T08:05:12.345Z"
}
FieldTypeMeaning
query, rank, foundstring, int, boolWhich input produced the row and its search rank.
lang, pageId, title, requestedTitle, redirectedEdition, stable page id, canonical title; requestedTitle when a redirect/search changed it.
description, extract, snippetstringWikidata short description, lead text, search snippet.
isDisambiguationboolPage is a disambiguation list.
thumbnail, originalImage, images[]Lead image; images with includeImages.
coordinatesobject{lat, lon} for places.
url, lastEdited, revisionId, wordCount, pageLengthBytesPage metadata.
fullText, fullTextChars, sections[]mode: "full" / includeSections.
categories[], links[], languages[], languagesCount, langLinks[]Optional graph data.
wikidataId, wikidata{}Q-id and curated facts (includeWikidata); item values as labels, their Q-ids in wikidata.entityIds.
pageviews{}Attention metrics (includePageviews).
fetchedAtISOProvenance.

A SUMMARY record in the key-value store: { lang, queries, saved, notFound (count), notFoundCount, notFoundInputs[], errors[{input, error}], errorCount, mode, searchMode }.

Use it from code / agents

curl -X POST "https://api.apify.com/v2/acts/yadroo~wikipedia-search/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"queries":["Kaspi.kz","Halyk Bank"],"searchMode":"nearmatch","includeWikidata":true}'
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('yadroo/wikipedia-search').call({ queries: ['Stablecoin'], mode: 'full', includeSections: true });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("yadroo/wikipedia-search").call(run_input={"queries": ["Tesla, Inc."], "searchMode": "exact", "includePageviews": True, "pageviewsDays": 90})
items = client.dataset(run["defaultDatasetId"]).list_items().items

MCP: add https://mcp.apify.com to Claude / Cursor / any MCP client and call the yadroo/wikipedia-search tool with the same JSON input.

Pricing

Pay per event: $0.001 per run start + $0.001 per page. Typical runs: 10 company lookups with Wikidata ≈ $0.011; 500-name enrichment ≈ $0.5; full-article ingestion of 100 pages ≈ $0.1.

Limits & FAQ

  • Rate limits — Wikimedia asks for ≤200 req/s overall and a descriptive User-Agent; the actor uses 1–6 requests per page (depending on toggles) with a 150 ms pause, and retries 429/5xx with backoff.
  • Freshness — live from Wikipedia at run time; pageviews are published with ~1 day delay (window ends yesterday).
  • Not found — a query with no page is not an item (you are not charged for it); it is listed in SUMMARY.notFoundInputs. A query-level API error goes to SUMMARY.errors; the run fails only when no query could be answered at all.
  • Disambiguation — flagged via isDisambiguation; use skipDisambiguation or searchMode: "text" with resultsPerQuery: 3 for ambiguous names.
  • Text size — full articles can exceed 200 KB; use maxTextChars for LLM budgets.
  • Wikidata coverage — only properties in the curated table are extracted; other claims are not returned (roadmap: wikidataProperties input for arbitrary P-ids).
  • Roadmap — revision history / last editors, geosearch by coordinates, Wikipedia "on this day" and trending feeds.

Made by Yadroo. Sibling actors: openalex-works, arxiv-papers, openlibrary-books, domain-intel, sanctions-screen.