Wikipedia Articles: search and full text by keyword or title avatar

Wikipedia Articles: search and full text by keyword or title

Pricing

from $0.98 / 1,000 article delivereds

Go to Apify Store
Wikipedia Articles: search and full text by keyword or title

Wikipedia Articles: search and full text by keyword or title

Wikipedia articles by keyword, exact title or URL, up to 50 per query: the full article text cleaned of footnote markers and reference lists, section titles, short description, thumbnail, page id, last edit date and the CC BY-SA license. Any language edition. Pay per delivered article.

Pricing

from $0.98 / 1,000 article delivereds

Rating

0.0

(0)

Developer

Steadydata Team

Steadydata Team

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

21 hours ago

Last modified

Share

Wikipedia articles by keyword, exact title or URL, up to 50 per query: the full article text cleaned of footnote markers and reference lists, section titles, short description, thumbnail, page id, last edit date and the CC BY-SA license. Any language edition. Pay per delivered article.

Why this scraper

  • Only delivered results are charged. A query that matches nothing and a title that does not exist come back as clear error records at no cost.
  • The text is clean, and that is the work. Wikipedia's own HTML carries the footnote markers inside the sentences and the reference list at the end. Measured on the Stroopwafel article: left in, you get "...often caramel.[3][4]" and 1.339 characters of reference lines; taken out, 3.526 characters of readable prose. Both are removed by markup and not by English words, so it works the same on every language edition.
  • Three kinds of input, one actor. A keyword searches Wikipedia and returns the articles it ranks; an exact title or a full article URL returns that one article. A URL also decides the language edition by itself.
  • Everything you need to cite it. Every row carries the article URL, the page id, the revision it came from, when it was last edited, and the CC BY-SA licence with its link.

Who this is for

Anyone who needs encyclopaedia text as rows instead of pages: feeding a retrieval index or an AI agent, building a glossary, enriching a dataset with a short description per subject, or tracking when articles on a topic were last changed. Paste keywords, titles or URLs in queries (up to 100 per run), pick the language edition, and every delivered article comes back as one row with its text, its section titles and its metadata.

Who this is not for

This is the article text, not the wiki source: templates, infobox tables and the reference list are not in text, and neither are images, categories or the links between articles. It reads one language edition at a time, so the same subject in three languages means three queries (or three URLs). Very long articles come back in full, so a run of fifty of them is a large dataset. And it reads what Wikipedia publishes: pages that do not exist, or that exist only on another edition, come back as a free ARTICLE_NOT_FOUND error row.

Input fields

FieldTypeRequired or defaultWhat it does
querieslist of textrequiredOne per row, up to 100: a keyword to search for, an exact article title, or a Wikipedia article URL. A URL or exact title returns that one article; a keyword returns the articles Wikipedia ranks for it.
languagetextenTwo-letter code of the Wikipedia edition to read, for example en, nl, de or es. A full article URL always wins over this setting.
maxArticlesPerQuerynumber5Cost ceiling per keyword, in Wikipedia's own ranking order. One delivered article is one charged event. A title or URL always returns one article.
includeTexttrue/falsetrueOn, every article carries its full cleaned text and its section titles, which costs one extra request per article. Off returns the search fields only and is much faster for a wide scan.

Input example

{
"queries": [
"stroopwafel",
"https://en.wikipedia.org/wiki/Web_scraping",
"artificial intelligence"
],
"language": "en",
"maxArticlesPerQuery": 5,
"includeText": true
}

Output example

FieldTypeWhat it holdsExample
querytextThe search term, title or URL this row was built from, so a row can always be traced back.https://en.wikipedia.org/wiki/Web_scraping
positionnumberThe rank Wikipedia's own search gave this article for the query; 1 for a title or URL.1
titletextThe article title as Wikipedia shows it.Web scraping
keytextThe article key used in its URL, with underscores instead of spaces.Web_scraping
pageIdnumberWikipedia's numeric page id, stable across renames, the key to join other data on.2696619
urltextThe article page on the chosen language edition.https://en.wikipedia.org/wiki/Web_scraping
languagetextThe language edition this article came from, as its two-letter code.en
descriptiontextThe one-line description Wikidata holds for this article; empty when there is none.Dutch cookie with caramel filling
excerpttextThe matching snippet from Wikipedia's search, without its highlight markup. Empty for a title or URL.A stroopwafel (Dutch pronunciation: [ˈstroːpˌʋaːfəl] ; li...
texttextThe readable article text, with the footnote markers and the reference list removed. Empty when the text was not requested or the article could not be read.Web scraping, web harvesting, or web data extraction is d...
textCharsnumberHow many characters the delivered text holds, so a row can be filtered on length.19632
sectionslistThe section titles of the article, in the order they appear.["History", "Techniques"]
thumbnailUrltextThe image Wikipedia shows next to this article in search results; empty when it has none.https://thumb.wikimedia.org/wikipedia/commons/thumb/9/9a/...
lastEditedtextWhen the article was last changed, in UTC.2026-09-19T09:13:04Z
revisionIdnumberThe revision this text came from, so a later run can be compared with this one.1375681745
licensetextThe licence the article text is published under. Reuse is allowed WITH attribution.CC BY-SA 4.0
licenseUrltextThe full licence text, to link to when you republish any of this.https://creativecommons.org/licenses/by-sa/4.0/

Error codes: EMPTY_QUERY, NO_ARTICLES, ARTICLE_NOT_FOUND, BLOCKED, FETCH_FAILED.

A real row, from the run of 01-10-2026 on the example input above, with the text shortened here so the page stays readable:

{
"query": "stroopwafel",
"position": 1,
"title": "Stroopwafel",
"key": "Stroopwafel",
"pageId": 1621210,
"url": "https://en.wikipedia.org/wiki/Stroopwafel",
"language": "en",
"description": "Dutch cookie with caramel filling",
"excerpt": "A stroopwafel (Dutch pronunciation: [ˈstroːpˌʋaːfəl] ; lit. 'syrup waffle') is a thin, round biscuit made from two layers of sweet baked dough held together",
"text": "A stroopwafel (Dutch pronunciation: [ˈstroːpˌʋaːfəl] ⓘ; lit. 'syrup waffle') is a thin, round biscuit made from two layers of sweet baked dough held together by a treacle/syrup filling, often caramel. First made in the city of Gouda in South Holland, stroopwafels are a well-known Dutch treat popular throughout the Netherlands. ...",
"textChars": 3526,
"sections": ["Description", "Etymology", "History", "Variants", "Gallery", "See also"],
"thumbnailUrl": "https://thumb.wikimedia.org/wikipedia/commons/thumb/9/9a/Stroopwafel.jpg/60px-Stroopwafel.jpg",
"lastEdited": "2026-07-19T13:38:40Z",
"revisionId": 1364952958,
"license": "CC BY-SA 4.0",
"licenseUrl": "https://creativecommons.org/licenses/by-sa/4.0/",
"status": "ok"
}

Pricing

Pay per event: one article-delivered event per delivered result. No charge for inputs that fail, no separate platform-usage surcharge.

FAQ

Which languages can I read? Every Wikipedia edition. Set language to its two-letter code (en, nl, de, es, fr, pl, ja, and so on), or paste a full article URL, which decides the edition by itself. A keyword is searched inside that edition, so search in the language you set.

Do I need an API key or an account? No. This reads Wikimedia's official public REST API, which needs neither.

May I republish the text I get? Yes, under the same terms Wikipedia itself uses: the text is CC BY-SA 4.0, so reuse including commercial reuse is allowed WITH attribution and share-alike. That is why every row carries the article URL, the licence name and the licence link. Check the licence text for what attribution has to look like in your case.

What exactly is removed from the text? The footnote markers inside the sentences and the reference list at the end, plus tables. Headings, paragraphs, lists and the "See also" entries stay. The section titles come along separately in sections, so you can split the text yourself.

How fast is it, and what does a run cost? Measured on 01-10-2026: a search costs one request, every article one more. Twenty-four articles with full text took 23 seconds and $0.0027 of platform usage, so about $0.11 per 1,000 articles on top of the per-article price.

Can I get the search results without the full text? Yes, switch includeText off. Then no article is fetched at all, only the search, which is a lot faster and lighter for a wide scan. You still get the title, description, excerpt, URL, page id and thumbnail per article.

Is personal data collected? No. This delivers encyclopaedia articles. It does not read user pages, talk pages, edit histories or contributor names, and it holds no accounts or personal profiles.

What happens when the source changes? Sources change from time to time; that is the nature of this work. The actor is monitored daily and fixed fast, and while it is broken you are not charged, because only delivered results cost anything.