# Article Text Extractor: Clean Text and Markdown from URLs (`madrasco/article-text-extractor`) Actor

Main text of the article and blog URLs you list, as plain text and Markdown, with title, author, date and language. Article-body F1 0.959 on a held-out half of a public benchmark. One request per URL, robots.txt obeyed; blocked pages reported, not bypassed.

- **URL**: https://apify.com/madrasco/article-text-extractor.md
- **Developed by:** [Jack Valmadre](https://apify.com/madrasco) (community)
- **Categories:** AI, Developer tools, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 article extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Article Text Extractor

Give it the URLs of articles or blog posts; get back the main text as plain text and Markdown, with the title, author, publication date and language the page states. The extractor aims to leave out menus, sidebars, comments and footers (see the accuracy figures below for how well that works on a public test set).

### What it does

- One HTTP GET per URL you name. Links on the page are **not** followed; nothing else is crawled.
- robots.txt is always obeyed (read once per site, also for redirect targets; Crawl-delay honoured up to 10 seconds); at most one request at a time per site.
- Pages behind a bot check, a login or an access-denied answer are reported as `blocked` and not retried. No proxies, no browser, no paywall bypass.
- Main-text extraction uses the open-source [trafilatura](https://github.com/adbar/trafilatura) library (Apache-2.0) with settings we chose on a public benchmark (below).

### What it does not do

- It does not run JavaScript, so pages that only load their text with JavaScript come back as `no-article`.
- It does not log in, solve captchas or get around paywalls.
- Author, date and language are taken as the page states them and are not checked; the date can be the page's first publication date rather than its last update.
- You are responsible for having the right to use the text of the pages you choose.

### Accuracy (measured by us, September 2026)

On the public [scrapinghub article-extraction benchmark](https://github.com/scrapinghub/article-extraction-benchmark) (181 news and blog pages saved in 2019, each with a reference copy of its article text), we chose the extraction settings on one half of the pages and then measured the other half (93 pages): **article-body F1 0.959** (± 0.012; precision 0.950, recall 0.968). F1 here is the word overlap between our text and the reference text, averaged over pages, computed with the benchmark's own scoring script (trafilatura 2.2.0, run 2026-09-26).

What this number does and does not tell you:

- It is the benchmark's own measure on its own pages. Your pages will differ: the pages are from 2019 and mostly English-language news, so modern layouts, other languages and JavaScript-heavy sites are not covered.
- Only the article body is scored; title, author, date and language are not.
- On the same 93 pages, trafilatura's default settings also score 0.959; our settings trade a little recall for precision (less stray text), they do not raise the average. The number describes the open-source extractor this actor runs, not something unique to it.

### Input

| Field | Default | Meaning |
|---|---|---|
| `urls` | - | Article URLs, one per line (up to 1,000 per run). |
| `startUrls`, `articleUrls`, `links`, `url` | - | Other names for the same thing, for API callers and AI agents. All are merged and duplicates removed. |
| `includeMarkdown` | `true` | Also return Markdown. |
| `includeText` | `true` | Return plain text. |
| `timeoutSecs` | `20` | Give up on a page after this many seconds. |
| `maxConcurrency` | `5` | Different sites fetched in parallel. |

### Use with AI agents

Any of the URL field names works (`{"urls": ["https://example.com/post"]}` or `{"url": "..."}`). Each row carries `status`, so an agent can tell an extracted article (`ok`) from `no-article`, `blocked`, `disallowed-by-robots`, `error` and `skipped`.

### Output (one dataset row per URL)

`url`, `finalUrl` (after redirects), `httpStatus`, `status`, `error`, `title`, `author`, `date` (YYYY-MM-DD), `language` (as declared by the page, e.g. `en-GB`), `siteName`, `text`, `markdown`, `wordCount`, `fetchedAt`. The run's `OUTPUT` record counts URLs by status.

### Pricing

Pay per event: US$0.001 per article extracted (`article-extracted`, rows with status `ok`), plus a start fee of US$0.00005 per GB of run memory, charged once per run (at most US$0.00005 at the default 512 MB). Rows with any other status (blocked, disallowed by robots.txt, no article found, failed) are not charged. There is no extra charge for Apify platform usage. Apify shows the price before you run, and you can set a maximum cost per run: once no further article fits, the remaining URLs get a free `skipped` row.

### Privacy

The actor stores nothing beyond the run's own dataset. It reads only the pages you name and returns what those pages show; it does not look up, enrich or combine information about people.

### Support

Questions and bug reports: open an issue in the **Issues** tab on this actor's page. We aim to respond within 14 days. Replies are written with AI assistance; a human owner can be reached on request.

### About

Published by Madrasco and built and maintained with AI assistance.

# Actor input Schema

## `urls` (type: `array`):

Article or blog-post pages to extract. Each URL is fetched once; links on the page are not followed. Up to 1,000 URLs per run; duplicates are removed. From the API or an AI agent you can also pass startUrls, articleUrls, links or a single url.

## `startUrls` (type: `array`):

Same as Article URLs, in the Apify request-list format (\[{"url": "https://..."}] or plain strings). Merged with the other URL fields.

## `url` (type: `string`):

One article URL. Merged with the other URL fields.

## `includeMarkdown` (type: `boolean`):

Also return the article as Markdown (headings, lists, emphasis, tables).

## `includeText` (type: `boolean`):

Return the article as plain text.

## `timeoutSecs` (type: `integer`):

Give up on a page after this long (reported as an error, not charged).

## `maxConcurrency` (type: `integer`):

How many different sites to fetch at once. Never more than one request at a time to the same host.

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Web_scraping"
  ],
  "includeMarkdown": true,
  "includeText": true,
  "timeoutSecs": 20,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

One dataset row per URL: title, author, date, language, text, Markdown, word count and status.

## `summary` (type: `string`):

Counts of URLs given, processed and skipped, by status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Web_scraping"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("madrasco/article-text-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://en.wikipedia.org/wiki/Web_scraping"] }

# Run the Actor and wait for it to finish
run = client.actor("madrasco/article-text-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/Web_scraping"
  ]
}' |
apify call madrasco/article-text-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,madrasco/article-text-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4Pgl5eJ7s5axlAnkn/builds/h1c614lcgypbaVsoe/openapi.json
