# Wikipedia Articles: search and full text by keyword or title (`steadydata/wikipedia-articles`) Actor

Wikipedia articles by keyword, exact title or URL, up to 50 per query: the full article text cleaned of footnote markers and reference lists, section titles, short description, thumbnail, page id, last edit date and the CC BY-SA license. Any language edition. Pay per delivered article.

- **URL**: https://apify.com/steadydata/wikipedia-articles.md
- **Developed by:** [Steadydata Team](https://apify.com/steadydata) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.98 / 1,000 article delivereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Wikipedia Articles: search and full text by keyword or title

Wikipedia articles by keyword, exact title or URL, up to 50 per query: the full article text cleaned of footnote markers and reference lists, section titles, short description, thumbnail, page id, last edit date and the CC BY-SA license. Any language edition. Pay per delivered article.

### Why this scraper

- **Only delivered results are charged.** A query that matches nothing and a title that does
  not exist come back as clear error records at no cost.
- **The text is clean, and that is the work.** Wikipedia's own HTML carries the footnote
  markers inside the sentences and the reference list at the end. Measured on the Stroopwafel
  article: left in, you get "...often caramel.\[3]\[4]" and 1.339 characters of reference lines;
  taken out, 3.526 characters of readable prose. Both are removed by markup and not by English
  words, so it works the same on every language edition.
- **Three kinds of input, one actor.** A keyword searches Wikipedia and returns the articles it
  ranks; an exact title or a full article URL returns that one article. A URL also decides the
  language edition by itself.
- **Everything you need to cite it.** Every row carries the article URL, the page id, the
  revision it came from, when it was last edited, and the CC BY-SA licence with its link.

### Who this is for

Anyone who needs encyclopaedia text as rows instead of pages: feeding a retrieval index or an
AI agent, building a glossary, enriching a dataset with a short description per subject, or
tracking when articles on a topic were last changed. Paste keywords, titles or URLs in
`queries` (up to 100 per run), pick the language edition, and every delivered article comes
back as one row with its text, its section titles and its metadata.

### Who this is not for

This is the article text, not the wiki source: templates, infobox tables and the reference
list are not in `text`, and neither are images, categories or the links between articles. It
reads one language edition at a time, so the same subject in three languages means three
queries (or three URLs). Very long articles come back in full, so a run of fifty of them is a
large dataset. And it reads what Wikipedia publishes: pages that do not exist, or that exist
only on another edition, come back as a free `ARTICLE_NOT_FOUND` error row.

### Input fields

| Field | Type | Required or default | What it does |
|---|---|---|---|
| `queries` | list of text | required | One per row, up to 100: a keyword to search for, an exact article title, or a Wikipedia article URL. A URL or exact title returns that one article; a keyword returns the articles Wikipedia ranks for it. |
| `language` | text | en | Two-letter code of the Wikipedia edition to read, for example en, nl, de or es. A full article URL always wins over this setting. |
| `maxArticlesPerQuery` | number | 5 | Cost ceiling per keyword, in Wikipedia's own ranking order. One delivered article is one charged event. A title or URL always returns one article. |
| `includeText` | true/false | true | On, every article carries its full cleaned text and its section titles, which costs one extra request per article. Off returns the search fields only and is much faster for a wide scan. |

### Input example

```json
{
    "queries": [
        "stroopwafel",
        "https://en.wikipedia.org/wiki/Web_scraping",
        "artificial intelligence"
    ],
    "language": "en",
    "maxArticlesPerQuery": 5,
    "includeText": true
}
```

### Output example

| Field | Type | What it holds | Example |
|---|---|---|---|
| `query` | text | The search term, title or URL this row was built from, so a row can always be traced back. | `https://en.wikipedia.org/wiki/Web_scraping` |
| `position` | number | The rank Wikipedia's own search gave this article for the query; 1 for a title or URL. | `1` |
| `title` | text | The article title as Wikipedia shows it. | `Web scraping` |
| `key` | text | The article key used in its URL, with underscores instead of spaces. | `Web_scraping` |
| `pageId` | number | Wikipedia's numeric page id, stable across renames, the key to join other data on. | `2696619` |
| `url` | text | The article page on the chosen language edition. | `https://en.wikipedia.org/wiki/Web_scraping` |
| `language` | text | The language edition this article came from, as its two-letter code. | `en` |
| `description` | text | The one-line description Wikidata holds for this article; empty when there is none. | `Dutch cookie with caramel filling` |
| `excerpt` | text | The matching snippet from Wikipedia's search, without its highlight markup. Empty for a title or URL. | `A stroopwafel (Dutch pronunciation: [ˈstroːpˌʋaːfəl] ; li...` |
| `text` | text | The readable article text, with the footnote markers and the reference list removed. Empty when the text was not requested or the article could not be read. | `Web scraping, web harvesting, or web data extraction is d...` |
| `textChars` | number | How many characters the delivered text holds, so a row can be filtered on length. | `19632` |
| `sections` | list | The section titles of the article, in the order they appear. | `["History", "Techniques"]` |
| `thumbnailUrl` | text | The image Wikipedia shows next to this article in search results; empty when it has none. | `https://thumb.wikimedia.org/wikipedia/commons/thumb/9/9a/...` |
| `lastEdited` | text | When the article was last changed, in UTC. | `2026-09-19T09:13:04Z` |
| `revisionId` | number | The revision this text came from, so a later run can be compared with this one. | `1375681745` |
| `license` | text | The licence the article text is published under. Reuse is allowed WITH attribution. | `CC BY-SA 4.0` |
| `licenseUrl` | text | The full licence text, to link to when you republish any of this. | `https://creativecommons.org/licenses/by-sa/4.0/` |

Error codes: `EMPTY_QUERY`, `NO_ARTICLES`, `ARTICLE_NOT_FOUND`, `BLOCKED`, `FETCH_FAILED`.

A real row, from the run of 01-10-2026 on the example input above, with the text shortened here
so the page stays readable:

```json
{
    "query": "stroopwafel",
    "position": 1,
    "title": "Stroopwafel",
    "key": "Stroopwafel",
    "pageId": 1621210,
    "url": "https://en.wikipedia.org/wiki/Stroopwafel",
    "language": "en",
    "description": "Dutch cookie with caramel filling",
    "excerpt": "A stroopwafel (Dutch pronunciation: [ˈstroːpˌʋaːfəl] ; lit. 'syrup waffle') is a thin, round biscuit made from two layers of sweet baked dough held together",
    "text": "A stroopwafel (Dutch pronunciation: [ˈstroːpˌʋaːfəl] ⓘ; lit. 'syrup waffle') is a thin, round biscuit made from two layers of sweet baked dough held together by a treacle/syrup filling, often caramel. First made in the city of Gouda in South Holland, stroopwafels are a well-known Dutch treat popular throughout the Netherlands. ...",
    "textChars": 3526,
    "sections": ["Description", "Etymology", "History", "Variants", "Gallery", "See also"],
    "thumbnailUrl": "https://thumb.wikimedia.org/wikipedia/commons/thumb/9/9a/Stroopwafel.jpg/60px-Stroopwafel.jpg",
    "lastEdited": "2026-07-19T13:38:40Z",
    "revisionId": 1364952958,
    "license": "CC BY-SA 4.0",
    "licenseUrl": "https://creativecommons.org/licenses/by-sa/4.0/",
    "status": "ok"
}
```

### Related actors from steadydata

- [wikipedia-pageviews](https://apify.com/steadydata/wikipedia-pageviews): how much attention the same article gets, per day
- [webpage-to-markdown](https://apify.com/steadydata/webpage-to-markdown): the readable text of any other page
- [arxiv-papers](https://apify.com/steadydata/arxiv-papers): research papers on the same subject, with abstracts

### Pricing

Pay per event: one `article-delivered` event per delivered result. No charge for inputs
that fail, no separate platform-usage surcharge.

### FAQ

**Which languages can I read?**
Every Wikipedia edition. Set `language` to its two-letter code (`en`, `nl`, `de`, `es`, `fr`,
`pl`, `ja`, and so on), or paste a full article URL, which decides the edition by itself. A
keyword is searched inside that edition, so search in the language you set.

**Do I need an API key or an account?**
No. This reads Wikimedia's official public REST API, which needs neither.

**May I republish the text I get?**
Yes, under the same terms Wikipedia itself uses: the text is CC BY-SA 4.0, so reuse including
commercial reuse is allowed WITH attribution and share-alike. That is why every row carries the
article URL, the licence name and the licence link. Check the licence text for what attribution
has to look like in your case.

**What exactly is removed from the text?**
The footnote markers inside the sentences and the reference list at the end, plus tables.
Headings, paragraphs, lists and the "See also" entries stay. The section titles come along
separately in `sections`, so you can split the text yourself.

**How fast is it, and what does a run cost?**
Measured on 01-10-2026: a search costs one request, every article one more. Twenty-four
articles with full text took 23 seconds and $0.0027 of platform usage, so about $0.11 per 1,000
articles on top of the per-article price.

**Can I get the search results without the full text?**
Yes, switch `includeText` off. Then no article is fetched at all, only the search, which is a
lot faster and lighter for a wide scan. You still get the title, description, excerpt, URL,
page id and thumbnail per article.

**Is personal data collected?**
No. This delivers encyclopaedia articles. It does not read user pages, talk pages, edit
histories or contributor names, and it holds no accounts or personal profiles.

**What happens when the source changes?**
Sources change from time to time; that is the nature of this work. The actor is
monitored daily and fixed fast, and while it is broken you are not charged, because
only delivered results cost anything.

# Changelog

This Actor's version history is a separate document: https://apify.com/steadydata/wikipedia-articles/changelog.md

# Actor input Schema

## `queries` (type: `array`):

One per row, up to 100: a keyword to search for, an exact article title, or a Wikipedia article URL. A URL or exact title returns that one article; a keyword returns the articles Wikipedia ranks for it.

## `language` (type: `string`):

Two-letter code of the Wikipedia edition to read, for example en, nl, de or es. A full article URL always wins over this setting.

## `maxArticlesPerQuery` (type: `integer`):

Cost ceiling per keyword, in Wikipedia's own ranking order. One delivered article is one charged event. A title or URL always returns one article.

## `includeText` (type: `boolean`):

On, every article carries its full cleaned text and its section titles, which costs one extra request per article. Off returns the search fields only and is much faster for a wide scan.

## Actor input object example

```json
{
  "queries": [
    "stroopwafel",
    "https://en.wikipedia.org/wiki/Web_scraping",
    "artificial intelligence"
  ],
  "language": "en",
  "maxArticlesPerQuery": 5,
  "includeText": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "stroopwafel",
        "https://en.wikipedia.org/wiki/Web_scraping",
        "artificial intelligence"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("steadydata/wikipedia-articles").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "stroopwafel",
        "https://en.wikipedia.org/wiki/Web_scraping",
        "artificial intelligence",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("steadydata/wikipedia-articles").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "stroopwafel",
    "https://en.wikipedia.org/wiki/Web_scraping",
    "artificial intelligence"
  ]
}' |
apify call steadydata/wikipedia-articles --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,steadydata/wikipedia-articles"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bGFxeEuv3BO3IUumg/builds/qLT9QokrvfGQsgYSg/openapi.json
