# Wikisource Texts Scraper (`parseforge/wikisource-texts-scraper`) Actor

Scrapes full texts from Wikisource by exact title or full-text search. Returns each work as a flat row with title, language edition, and complete body in plain text or raw wikitext.

- **URL**: https://apify.com/parseforge/wikisource-texts-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Automation, Other
- **Stats:** 2 total users, 1 monthly users, 93.3% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.81 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner.jpg)](https://apify.com/parseforge?fpr=vmoqkp)

### Wikisource Texts Scraper

**Scrape full texts from Wikisource by title or search, in plain text or raw wikitext.** Every record includes the page title, language edition, and complete body. No login or API key. Export to CSV, JSON, Excel, or XML.

Wikisource hosts millions of free public-domain and source texts, but downloading them one by one is slow and the MediaWiki API returns markup you have to clean yourself. This Actor reads the public pages directly, fetches exact titles or full-text search results, and returns each work as one clean record. Choose plain text for analysis or raw wikitext for archival.

| Who uses it | What they scrape Wikisource for |
|---|---|
| Digital humanities researchers | Build a corpus of historical texts for computational analysis |
| Librarians and archivists | Harvest public-domain works for a digital collection |
| NLP engineers | Create training datasets from clean, out-of-copyright text |
| Literary scholars | Compare editions of a work across language editions |
| Content curators | Gather source texts for a reading app or website |

### What it does

This Actor collects Wikisource texts by exact page title or full-text search, and returns each one as a flat row with the title, language edition, and complete body.

- 📚 **Title or search:** fetch exact pages like 'The Raven (Poe)' or run a full-text search across the whole edition.
- 🌍 **15 language editions:** English, French, German, Spanish, Italian, Russian, Polish, Portuguese, Chinese, Latin, Greek, Dutch, Swedish, Arabic, and Hebrew.
- 🧹 **Plain text mode:** markup stripped, ready for analysis or reading.
- 📝 **Raw wikitext mode:** original MediaWiki markup, including templates and formatting.
- 🔢 **Up to a million texts per run:** set a hard cap so a broad search never runs away.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with Wikisource data

**📖 Build a research corpus.**

A digital humanities researcher enters a search term like 'sonnet' and collects every matching poem in plain text for stylometric analysis.

**🌐 Compare translations.**

A literary scholar fetches the same title from the English, French, and German editions to study how a classic was translated.

**🤖 Create NLP training data.**

An NLP engineer scrapes thousands of public-domain novels in plain text to train a language model without licensing issues.

**🗄️ Archive source texts.**

A librarian harvests a list of exact titles in raw wikitext to preserve the original formatting and templates.

**📱 Populate a reading app.**

A content curator pulls a set of classic short stories and exports them as JSON to load into a mobile reading interface.

### Why choose this scraper

|  | What you get |
|---|---|
| **No API key or login** | Reads the public Wikisource pages directly, so you can start a run immediately. |
| **Clean text or raw markup** | Pick plain text for analysis or wikitext for archival, per run. |
| **One row per work** | Every record is flat and consistent, ready for CSV, JSON, Excel, or XML export. |
| **Multi-language** | Query any of 15 language editions from the same Actor. |

### How it compares

No other Store actor targets Wikisource the same way, so the honest comparison is with the alternatives teams actually weigh.

| | Wikisource Texts Scraper | Build it in-house | By hand |
|---|---|---|---|
| Setup | Run it now, zero config | Days of engineering | None, but hours per pull |
| When Wikisource changes | Maintained for you | You fix it | You re-learn the page |
| Proxies, retries, anti-bot | Built in | Your problem | Browser only |
| Output | Fixed JSON schema, CSV/Excel export | Whatever you build | Copy-paste |
| Cost | Pay per result | Engineering time | Analyst hours |

### Configure the run

Drive the Actor from exact page titles, a full-text search term, or both together, and set the language edition and text mode before the run starts. The Input tab lists every parameter.

A first run with the defaults:

```json
{
  "titles": [
    "The Raven (Poe)",
    "The Bells (Poe)"
  ],
  "maxItems": 10
}
```

A larger pull:

```json
{
  "titles": [
    "The Raven (Poe)",
    "The Bells (Poe)"
  ],
  "maxItems": 200
}
```

### Pricing

Pay-per-result: **$0.004 per result** collected. You pay only for the results written to your dataset.

| Results collected | Approximate cost |
|---|---|
| 100 results | $0.40 |
| 1,000 results | $4.00 |
| 10,000 results | $40.00 |

New Apify accounts start with $5 in free credit.

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Open the [Wikisource Texts Scraper](https://apify.com/parseforge/wikisource-texts-scraper?fpr=vmoqkp).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to Wikisource through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/wikisource-texts-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check that the exact title matches the Wikisource page name, including parentheses and capitalization. For search, try a broader term or a different language edition.

**Why is the text full of markup?**

You are in Raw wikitext mode. Switch the Text mode input to Plain text and run again.

**Why did the run stop before my search was complete?**

The Maximum texts limit was reached. Increase the limit and run again.

**Why do I get a different language than I expected?**

Each language edition is a separate subdomain. Make sure the Language edition input is set to the one you want.

**Why is a title missing even though it exists on Wikisource?**

The page name may include a disambiguation suffix or different punctuation. Copy the exact title from the Wikisource URL.

### FAQ

| Question | Answer |
|---|---|
| Do I need a Wikimedia API key? | No. The Actor reads the public Wikisource pages directly, so there is no registration or authentication step. |
| What is the difference between plain text and raw wikitext? | Plain text strips all MediaWiki markup and returns readable text. Raw wikitext returns the original source, including templates, links, and formatting codes. |
| Can I fetch a specific page by its exact title? | Yes. Add one or more exact titles to the Page titles field, for example 'The Raven (Poe)' or 'Pride and Prejudice'. Each title becomes one record. |
| How does the search field work? | It runs a full-text search across the selected language edition and returns matching pages up to the Maximum texts limit. |
| Can I combine exact titles and a search term in one run? | Yes. The Actor fetches all exact titles first, then adds search results, and the Maximum texts cap applies to the total. |
| Which language editions are supported? | English, French, German, Spanish, Italian, Russian, Polish, Portuguese, Chinese, Latin, Greek, Dutch, Swedish, Arabic, and Hebrew. |
| What does each record contain? | Each record includes the page title, the language edition, and the full text body in the mode you selected. The exact field list is shown in the sample output. |
| Is there a limit on how many texts I can collect? | You can set Maximum texts from 1 up to 1,000,000 per run. The default is 10. |
| Can I export the results? | Yes. The dataset can be exported to CSV, JSON, Excel, or XML from the Apify platform. |
| Is the text in the public domain? | Wikisource hosts texts that are free to use, but individual works may have different licenses. Check the page's own license information before redistribution. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Wikimedia Foundation, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `titles` (type: `array`):

Exact Wikisource page titles to fetch, for example 'The Raven (Poe)' or 'Pride and Prejudice'. Each title becomes one record. Leave empty if you are using the Search field instead.

## `search` (type: `string`):

Full-text search across Wikisource. Matching pages are fetched as records, up to Max items. Used on its own or combined with explicit titles.

## `language` (type: `string`):

Which Wikisource language edition to query. Each language lives on its own subdomain.

## `mode` (type: `string`):

'Plain text' returns readable text with markup stripped. 'Raw wikitext' returns the original MediaWiki markup, including templates and formatting.

## `maxItems` (type: `integer`):

Maximum number of texts to collect per run.

## Actor input object example

```json
{
  "titles": [
    "The Raven (Poe)",
    "The Bells (Poe)"
  ],
  "language": "en",
  "mode": "text",
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "titles": [
        "The Raven (Poe)",
        "The Bells (Poe)"
    ],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/wikisource-texts-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "titles": [
        "The Raven (Poe)",
        "The Bells (Poe)",
    ],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/wikisource-texts-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "titles": [
    "The Raven (Poe)",
    "The Bells (Poe)"
  ],
  "maxItems": 10
}' |
apify call parseforge/wikisource-texts-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/wikisource-texts-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lxlVYwr6rwamXD1Nh/builds/XJfp6QmAjbPi62Lh7/openapi.json
