# Article Extractor (`deriverge/article-extractor`) Actor

Clean text and Markdown of news and blog articles, with the headline, publish date, section and tags. Give it article links, an RSS feed or a whole site; for feeds and sites it extracts the newest articles.

- **URL**: https://apify.com/deriverge/article-extractor.md
- **Developed by:** [deriverge s.r.o.](https://apify.com/deriverge) (community)
- **Categories:** News, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Article Extractor

Article Extractor turns an article page into one row: the article text without menus and ads, optional Markdown with headings, links and images, and the metadata publishers add for search engines, such as the headline, publish and update dates, section, tags, main image and paywall flag. Paste an article link, an RSS or Atom feed, or a site address and click **Start**; for a feed or a site it extracts the newest articles. Pages that build their text with JavaScript return no article. It costs $1.00 per 1,000 articles on the Free plan ($0.50 on Business), and the $5 of monthly credit in Apify's Free plan covers about 5,000 articles.

### What data does it return?

| Field | Description |
|---|---|
| `title`, `description` | Headline, and the article's short description, usually from its structured data or meta tags. |
| `text` | The article as plain text, with a blank line between paragraphs where the page marks them. |
| `markdown` | The article as Markdown with headings, lists, links and images, when **Output format** includes Markdown. `null` otherwise. |
| `publishedAt`, `modifiedAt` | Publish and last update time, ISO 8601 UTC, from the page's structured data or meta tags. |
| `siteName`, `section`, `tags`, `language` | Publication name, section such as `Politics`, tags such as `Programming`, and the language code, such as `en-GB`. |
| `image` | Link to the main image of the article. |
| `isAccessibleForFree` | `false` when the publisher marks the article as paywalled, `true` when it marks it as free to read, `null` when the page does not say. |
| `wordCount`, `readingTimeMinutes` | Words in `text`, and the reading time at 230 words per minute. |
| `url`, `canonicalUrl`, `sourcePage` | The page that was read, the canonical address it declares, and the feed or site it was found through. |
| `status`, `error` | `ok` (extracted and charged), `not_article`, `blocked` or `error`, with the reason in `error`. |
| `isNew` | With a watch name, `true` for an article that no earlier run with that name returned. |

### How to extract articles

1. Add addresses to **Articles, feeds or websites**, one per line: an article such as `https://blog.apify.com/what-is-web-scraping/`, a feed such as `https://feeds.bbci.co.uk/news/rss.xml`, or a site or section such as `https://blog.apify.com`.
2. For feeds and sites, set **Maximum articles per feed or website**, 20 by default, taken in the order the feed lists them.
3. Choose the **Output format**: plain text, or plain text and Markdown.
4. Click **Start**. Articles appear in the **Output** tab and can be exported as CSV, Excel or JSON or read through the API.

An API input that follows a news feed and a blog with Markdown and returns only articles it has not returned before:

```json
{
  "urls": ["https://feeds.bbci.co.uk/news/rss.xml", "https://blog.apify.com"],
  "maxArticlesPerSite": 10,
  "format": "markdown",
  "newOnly": true,
  "watchName": "tech-news"
}
```

### Example output

One article from a run on 30 September 2026 with Markdown on. `text` and `markdown` are shortened here; the full row holds the whole article of 1,363 words. The run came before that day's paragraph fix, which is why `text` still runs two paragraphs together in `web scraping.You could`.

```json
{
  "url": "https://blog.apify.com/what-is-web-scraping/",
  "canonicalUrl": "https://blog.apify.com/what-is-web-scraping/",
  "sourcePage": null,
  "status": "ok",
  "error": null,
  "title": "What is web scraping?",
  "description": "The basics of web scraping: what it is, how it works, real-world use cases, and how to get started.",
  "siteName": "Apify Blog",
  "language": "en",
  "publishedAt": "2024-10-15T14:32:00.000Z",
  "modifiedAt": "2026-07-07T12:57:19.000Z",
  "section": null,
  "tags": [
    "Programming"
  ],
  "image": "https://storage.ghost.io/c/f2/6e/f26ec999-9a90-4aee-a0d4-9b3ca2bb668f/content/images/2024/04/what-is-web-scraping-complete-guide.png",
  "isAccessibleForFree": null,
  "wordCount": 1363,
  "readingTimeMinutes": 6,
  "text": "Web scraping is the process of automatically extracting data from a website. You use a program called a web scraper to access a web page, interpret the data, and extract what you need. The data is saved in a structured format such as an Excel file, JSON, or XML so that you can use it in spreadsheets or apps. It's also easy to confuse with crawling: see web crawling vs. web scraping.You could do this manually by copying and pasting, but scraping is typically performed using an automated tool ...",
  "markdown": "Web scraping is the process of automatically extracting data from a website. You use a program called a web scraper to access a web page, interpret the data, and extract what you need. The data is saved in a structured format such as an Excel file, JSON, or XML so that you can use it in spreadsheets or apps. It's also easy to confuse with crawling: see [web crawling vs. web scraping](https://blog.apify.com/web-crawling-vs-web-scraping).\n\n![What is web scraping: diagram showing data going from website through scraping platform to structured data](https://storage.ghost.io/c/f2/6e/f26ec999-9a90-4aee-a0d4-9b3ca2bb668f/content/images/2023/09/what-is-web-scraping-websites-web-scraper-structured-data-1.png)\n\n...",
  "isNew": null,
  "checkedAt": "2026-09-30T00:34:16.088Z"
}
```

### How much does it cost to extract articles?

| | Free plan | Starter | Scale | Business |
|---|---|---|---|---|
| 1,000 articles | $1.00 | $0.80 | $0.65 | $0.50 |

You pay only for the events in the table. There is no start fee, and compute time and proxies are included.

1,000 articles cost $1.00 on the Free plan and $0.50 on Business, with or without Markdown. Pages without an article, blocked pages, failed downloads and articles skipped by the new-only mode are not charged.

### Limits

- When `text` comes from the article body in the page's structured data, it has only the line breaks the publisher put there, and `markdown` then holds the same plain text.
- Pages that load their text with JavaScript give a `not_article` row.
- Paywalls are not bypassed: a paywalled article gives only the part the page shows without a subscription.
- For a homepage or section page, the actor uses the feed the page links to, or tries common feed paths such as `/feed`, `/rss` and `/rss.xml`. Without any feed, the page gives a `not_article` row.
- A site that answers with a bot protection page gives a `blocked` row.
- The new-only memory keeps each article for 180 days.

### Following a site or feed

Give the run a **Watch name** such as `tech-news`, or save it as a task, turn on **Return only articles new since the last run** and schedule it. Each run then extracts only articles it has not returned before, so a daily run becomes a feed of full texts. Without new-only mode, a watch name still marks every article with `isNew`.

### FAQ

#### Is it legal to extract articles?

The actor reads pages that anyone can open without logging in, much like a browser's reader mode, and leaves out author names. The articles remain their publishers' work, so check the publisher's terms before you republish any text. You are responsible for how you use the results.

#### How is the article text found?

With Mozilla's Readability library, the code behind the reader view in Firefox, run on the page's HTML. When the page's structured data carries the full article body and it is about as long as that result, the structured version is used instead.

#### What if a feed lists only headlines?

It works the same way. The actor opens every article the feed links to, up to **Maximum articles per feed or website**, and extracts the text from the article page itself.

### Related scrapers

- [Google News Scraper](https://apify.com/deriverge/google-news-scraper)
- [Substack Scraper](https://apify.com/deriverge/substack-scraper)

### Support

This actor is built and maintained by deriverge s.r.o., a software company based in the Czech Republic. If a run fails or a field you need is missing, please open an issue in the **Issues** tab or write to us at info@deriverge.com. We respond in English and Czech. Runs can be scheduled in Apify Console or started from the **API** tab, which has examples for Python, JavaScript and cURL and works with Make, Zapier, n8n and the Apify MCP server. If the actor saves you time, a short review helps other people find it.

# Changelog

This Actor's version history is a separate document: https://apify.com/deriverge/article-extractor/changelog.md

# Actor input Schema

## `urls` (type: `array`):

One per line: an article link, an RSS or Atom feed, or a website or section (its feed is found and its newest articles are extracted).

## `maxArticlesPerSite` (type: `integer`):

For feeds and websites: how many of the newest articles to extract.

## `format` (type: `string`):

Plain text always; add Markdown with headings, lists, links and images when you need it.

## `newOnly` (type: `boolean`):

With a watch name or a saved task, return only articles you have not received before. Schedule it for a news feed of full texts.

## `watchName` (type: `string`):

Name of the snapshot used by the new-only mode, for example "tech-news". Runs from a saved task get one automatically.

## `maxItems` (type: `integer`):

Hard cap on returned rows in the run.

## Actor input object example

```json
{
  "urls": [
    "https://blog.apify.com/what-is-web-scraping/"
  ],
  "maxArticlesPerSite": 20,
  "format": "text",
  "newOnly": false,
  "maxItems": 100000
}
```

# Actor output Schema

## `articles` (type: `string`):

One row per article with clean text (and Markdown), dates, section, tags and image.

## `summary` (type: `string`):

Articles extracted, pages without an article, blocked and failed pages.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://blog.apify.com/what-is-web-scraping/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("deriverge/article-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://blog.apify.com/what-is-web-scraping/"] }

# Run the Actor and wait for it to finish
run = client.actor("deriverge/article-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://blog.apify.com/what-is-web-scraping/"
  ]
}' |
apify call deriverge/article-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,deriverge/article-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zYmFpARLyhTkdFe5N/builds/YpZffDFPkab0AatLM/openapi.json
