# Article Extractor: Clean Text, Author and Date from News URLs (`pistachio_implementation/article-extractor`) Actor

Turn news and blog article URLs into clean text or Markdown with title, author, publish date, site, language, tags, lead image and word count. Fast plain HTTP, no browser, flat price per article, failed URLs free. Built for media monitoring, research and LLM pipelines.

- **URL**: https://apify.com/pistachio\_implementation/article-extractor.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** News, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Article Extractor: Clean Text, Author and Date from News URLs

Give this actor a list of news, blog or magazine article URLs and get back one clean row per article: the main text without menus, ads and sidebars, plus title, author, publish and update dates, site name, language, section, tags, lead image, word count and reading time. Markdown and cleaned HTML are available too, ready for LLM, RAG and summarization pipelines.

It uses Mozilla Readability (the engine behind Firefox Reader View) for the text and reads the publisher's own metadata (Open Graph, article tags, JSON LD and time tags) for the facts, so dates and authors come from the source rather than guesses. It fetches each page once over plain HTTP, with no browser, which keeps it fast and cheap. You pay per article that was extracted, plus a tiny start fee per run; failed URLs cost nothing.

### What you can use it for

- Media monitoring: turn a list of article links into a clean table of headlines, authors and dates.
- Feed articles into an LLM for summaries, classification or sentiment.
- Build a RAG knowledge base from blog posts and documentation pages as Markdown.
- Research: collect word counts, publish dates and tags across many publications.
- Content audits: check which of your own posts lack an author, a date or a lead image.

### Input

| Field | What it does | Default |
|---|---|---|
| Article URLs | The pages to extract. The actor reads only these; it does not follow links | required |
| Include plain text | Main text with paragraph breaks | on |
| Include Markdown | Main text as Markdown with headings, lists, links and images | off |
| Include cleaned HTML | The cleaned article HTML | off |
| Maximum text length | Cut text, Markdown and HTML to this many characters (0 means no limit) | 0 |
| Respect robots.txt | Skip pages a site closes to automated tools (skipped pages are free) | on |
| Parallel sites | How many URLs to work on at once | 5 |

Example input:

```json
{
  "urls": [
    "https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
    "https://blog.apify.com/what-is-web-scraping/"
  ],
  "includeMarkdown": true,
  "maxTextLength": 20000
}
```

### Output

One row per URL. Successful rows have `success: true`; failed rows have `success: false` and an `error` that says why (HTTP error, bot check, not an article, robots.txt). Download as JSON, CSV or Excel, or read it through the API.

```json
{
  "url": "https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
  "finalUrl": "https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
  "canonicalUrl": "https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
  "success": true,
  "statusCode": 200,
  "title": "PNOE’s new face mask wants to make lab-grade breath testing a self-serve affair",
  "author": "Connie Loizos",
  "publishedAt": "2026-09-27T01:40:30.000Z",
  "modifiedAt": "2026-09-27T01:40:40.000Z",
  "siteName": "TechCrunch",
  "description": "PNOĒ, the Malden, Mass.-based startup whose breath-analyzing mask...",
  "language": "en-US",
  "section": "Hardware",
  "tags": [],
  "leadImageUrl": "https://techcrunch.com/wp-content/uploads/2026/09/PNOE-2.0.png?resize=1200,675",
  "wordCount": 1053,
  "readingTimeMinutes": 5,
  "isAccessibleForFree": null,
  "likelyArticle": true,
  "excerpt": "At first glance, the newest device from PNOĒ looks like something a comic-book villain might wear...",
  "textTruncated": true,
  "text": "At first glance, the newest device from PNOĒ looks like something a comic-book villain might wear. The mask, which covers the nose and mouth...",
  "markdown": "At first glance, the newest device from PNOĒ looks like something a comic-book villain might wear...",
  "extractedAt": "2026-09-27T06:40:01.577Z"
}
```

### Pricing

Pay per event, no subscription, no charge for platform usage on top.

| Event | Price |
|---|---|
| Article extracted | $0.0015 ($1.50 per 1,000 articles) |
| Actor start | $0.00005 per run (Apify's standard start event) |

Failed URLs, pages that are not articles, bot checks and robots.txt skips are free. You can set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

### Limits

- Plain HTTP only. Pages that build their text with JavaScript after loading, or that sit behind a bot check or a login, come back as a free error row instead of text.
- Some publishers refuse automated readers or tools run on Apify, either in robots.txt or at their server (in testing on 27 Sep 2026: the robots.txt files of BBC and Le Monde close them to Apify crawlers, The Guardian answered 403 and NPR stopped answering). The actor respects that and returns a free error row; it does not disguise itself.
- Paywalled articles return only what the publisher shows to anonymous readers. `isAccessibleForFree` tells you when the publisher marks an article as paid.
- Home pages, section pages and video pages are rejected as "not an article" when they have under 100 words of main text.
- Requests to the same site are spaced one second apart, so 1,000 URLs from one site take about 17 minutes; URLs spread across many sites run in parallel.
- Up to 5,000 URLs per run.

### FAQ

**Which sites work?** Most news sites, blogs, magazines, documentation pages and Wikipedia. It works best on server rendered pages, which is most publishing sites.

**Where does the author come from?** From the page's JSON LD, then its author meta tags, then the byline Readability finds. If none exists the field is empty rather than guessed.

**Can I use the text in my product?** The actor extracts what you point it at. You are responsible for having the right to use the content, as with any reader tool. The robots.txt option is on by default.

**Does it crawl a whole site?** No. It reads exactly the URLs you give it. Use a sitemap or RSS tool to collect article links first, then pass them in.

**Why did a URL fail?** The `error` field says: HTTP status, timeout, bot check, not HTML, too little text, or robots.txt. Failed rows are never charged.

# Actor input Schema

## `urls` (type: `array`):

News, blog or magazine article URLs. One row comes back per URL. The actor reads only these pages; it does not follow links.

## `includeText` (type: `boolean`):

Return the main article text as plain text with paragraph breaks.

## `includeMarkdown` (type: `boolean`):

Also return the main article as Markdown (headings, lists, links and images kept). Handy for LLM and RAG pipelines.

## `includeHtml` (type: `boolean`):

Also return the cleaned article HTML without menus, ads and sidebars.

## `maxTextLength` (type: `integer`):

Cut text, Markdown and HTML to this many characters. 0 means no limit.

## `respectRobotsTxt` (type: `boolean`):

Skip pages the site's robots.txt closes to automated tools. Skipped pages are free.

## `maxConcurrency` (type: `integer`):

How many URLs to work on at once. Requests to the same site are always spaced one second apart.

## Actor input object example

```json
{
  "urls": [
    "https://blog.apify.com/what-is-web-scraping/",
    "https://en.wikipedia.org/wiki/Web_scraping"
  ],
  "includeText": true,
  "includeMarkdown": false,
  "includeHtml": false,
  "maxTextLength": 0,
  "respectRobotsTxt": true,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://blog.apify.com/what-is-web-scraping/",
        "https://en.wikipedia.org/wiki/Web_scraping"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/article-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://blog.apify.com/what-is-web-scraping/",
        "https://en.wikipedia.org/wiki/Web_scraping",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/article-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://blog.apify.com/what-is-web-scraping/",
    "https://en.wikipedia.org/wiki/Web_scraping"
  ]
}' |
apify call pistachio_implementation/article-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/article-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ugl23yRVSPHONQRJh/builds/GUDmiOn34o3QA1Pz8/openapi.json
