# Article Scraper & Text Extractor – Clean Markdown for LLM/RAG (`forevertools/article-extractor`) Actor

Article scraper & web page text extraction: turn any list of URLs into clean article text and Markdown (HTML to Markdown) with title, author, date, site, language and word count. Boilerplate removed; content extraction for LLM/RAG. Works as a news article & blog scraper. No browser, fast.

- **URL**: https://apify.com/forevertools/article-extractor.md
- **Developed by:** [Forever Tools](https://apify.com/forevertools) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 article extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Article Scraper & Clean Text Extractor to Markdown for LLM/RAG

Paste a list of URLs and get the **main content of every page as clean plain text and Markdown**, with
**title, author, published date, site name, language, description and word count**. Menus, sidebars, footers,
cookie notices, scripts and styles are removed (using Mozilla Readability, the engine behind Firefox Reader View),
so what you get is the article itself, ready to chunk and embed.

Use it to: feed web pages into an LLM or RAG pipeline, build a knowledge base from docs or blog posts, collect
articles for summarisation or classification, archive readable copies of pages, or get word counts and metadata
for a list of URLs.

It does not use a browser, so it's fast and cheap. The trade-off: it reads the HTML the server sends and does
not run JavaScript (see Limitations).

### Features

- **Bulk**: any number of URLs, 5 in parallel by default (up to 20). If one URL fails, the error goes in its row
  and the run keeps going.
- **Markdown for LLMs**: headings, lists, links (made absolute), bold/italic, code blocks and simple tables become
  GitHub-flavoured Markdown. Complex tables (merged cells, no header row, e.g. Wikipedia infoboxes) become
  `Label: value` lines instead of raw HTML.
- **Plain text** with paragraph breaks kept, plus a `wordCount`.
- **Metadata** from JSON-LD (schema.org Article/NewsArticle/BlogPosting), OpenGraph and standard meta tags, with
  Readability's guesses as a fallback. Dates are normalised to ISO 8601 when they parse.
- **Retries** with backoff for timeouts, network errors and HTTP 408/425/429/5xx.
- **Encoding aware**: the charset from the HTTP header or `<meta charset>` is used to decode the page.
- Optional **cleaned HTML** of the main content (`includeHtml`).

### Input

```json
{
  "urls": ["https://en.wikipedia.org/wiki/Sourdough", "paulgraham.com/greatwork.html"],
  "includeHtml": false,
  "maxConcurrency": 5,
  "maxRetries": 2,
  "timeoutSecs": 30
}
```

### Output (one dataset row per URL)

```json
{
  "url": "https://paulgraham.com/greatwork.html",
  "finalUrl": "https://paulgraham.com/greatwork.html",
  "statusCode": 200,
  "title": "How to Do Great Work",
  "author": null,
  "publishedDate": null,
  "siteName": null,
  "language": null,
  "description": "If you collected lists of techniques for doing great work in a lot of different fields, what would the intersection look like? ...",
  "text": "July 2023\nIf you collected lists of techniques for doing great work ...",
  "markdown": "![How to Do Great Work](https://...gif)July 2023\n\nIf you collected lists of techniques ...",
  "wordCount": 11801,
  "error": null
}
```

- Fields the page doesn't provide are `null` (the example page has no author or language tags).
- `finalUrl` is the URL after redirects. `statusCode` is the final HTTP status.
- `error` is `null` on success. On failure (HTTP 4xx/5xx, DNS error, timeout, non-HTML content such as a PDF, invalid
  URL) the row still appears with `error` set, e.g. `"HTTP 404"` or `"fetch failed: ENOTFOUND"`, and empty content.
- With `includeHtml: true` an extra `html` field holds the cleaned main-content HTML.

The Console has two views: **Articles** (metadata table) and **Content** (Markdown). Export as JSON, CSV or Excel,
or fetch via the API.

### Pricing

Pay per event: **$1 per 1,000 articles** ($0.001 per successfully extracted URL). Failed URLs (error rows) are
**not charged**. Apify platform usage is billed separately as usual. If you set a maximum charge for the run, the
actor stops when it is reached.

### FAQ

**Does it work on any website?** It works on pages whose content is in the HTML the server returns: most news sites,
blogs, documentation, Wikipedia and similar. It won't see content that only appears after JavaScript runs.

**What about home pages or listing pages?** They are extracted, but a home page has no single "article", so the
result may be a list of headlines, a section of the page, or some navigation text. The actor is built for article
and documentation pages.

**Can it get past paywalls, logins or CAPTCHAs?** No. It sees what a logged-out visitor sees. Bot-protected sites
may return an error or a challenge page.

**Is the Markdown good for chunking?** Yes, that's the goal: headings are `#`-style, and there are no scripts, styles or
layout tables. Split on headings or paragraphs.

**Why is `title` sometimes "Page - Site"?** The OpenGraph/`<title>` value is returned as the site publishes it.

### Limitations

- No JavaScript rendering. Single-page apps and pages that load content client-side may come back empty or partial.
- Only HTML (and plain text) pages. PDFs, images and other files get an error row.
- Pages over 15 MB are skipped with an error.
- Which part of a page counts as "main content" is decided by Readability's heuristics. They are good on articles and
  can be wrong on unusual layouts.
- Respect the terms of the sites you extract from.

### Support

Open an issue on the actor's Issues tab and include the URL and your input. Issues are answered asynchronously.

Not affiliated with Mozilla or any website you extract from. Built and maintained with AI assistance.

### Related tools

Other actors by the same developer (same flat pay-per-result pricing, no subscription):

- [Apple App Store Reviews Scraper (Multi-Country)](https://apify.com/forevertools/apple-app-store-reviews)
- [Company Jobs Scraper: Workday, Greenhouse, Lever, Ashby](https://apify.com/forevertools/ats-company-jobs)
- [Bulk Domain Checker — WHOIS/RDAP, DNS, SPF/DMARC, SSL Expiry](https://apify.com/forevertools/domain-whois-dns-ssl)
- [Bulk PageSpeed Insights & Core Web Vitals Checker](https://apify.com/forevertools/pagespeed-core-web-vitals)
- [PDF to Text Extractor (Bulk, with Metadata)](https://apify.com/forevertools/pdf-to-text-extractor)
- [Website SEO Audit Crawler](https://apify.com/forevertools/website-seo-audit)
- [Sitemap Extractor & Bulk URL Status Checker](https://apify.com/forevertools/sitemap-url-status-checker)
- [Website Tech Stack Detector (CMS, Framework, Analytics)](https://apify.com/forevertools/website-tech-stack-detector)
- [Website Screenshot – Bulk Full Page PNG, JPEG & PDF](https://apify.com/forevertools/website-screenshot)

### Integrations

Run it from the Apify API, a schedule, or no-code tools: the Apify apps for **Zapier**, **Make** and **n8n** can start any public actor ("Run Actor") and read its dataset. AI agents can call it through the **Apify MCP server**.

# Actor input Schema

## `urls` (type: `array`):

Article or page URLs to extract. Domains without a scheme get https://. Duplicates are processed once.

## `includeHtml` (type: `boolean`):

Also return the cleaned main-content HTML in an `html` field.

## `maxConcurrency` (type: `integer`):

URLs fetched in parallel.

## `maxRetries` (type: `integer`):

Retries per URL for network errors, timeouts and HTTP 408/425/429/5xx. Other 4xx errors are not retried.

## `timeoutSecs` (type: `integer`):

Per-attempt request timeout.

## Actor input object example

```json
{
  "urls": [
    "https://en.wikipedia.org/wiki/Sourdough",
    "https://paulgraham.com/greatwork.html"
  ],
  "includeHtml": false,
  "maxConcurrency": 5,
  "maxRetries": 2,
  "timeoutSecs": 30
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://en.wikipedia.org/wiki/Sourdough",
        "https://paulgraham.com/greatwork.html"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("forevertools/article-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://en.wikipedia.org/wiki/Sourdough",
        "https://paulgraham.com/greatwork.html",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("forevertools/article-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://en.wikipedia.org/wiki/Sourdough",
    "https://paulgraham.com/greatwork.html"
  ]
}' |
apify call forevertools/article-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,forevertools/article-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/tUSBoIEWX3hvdaIP4/builds/Iyh7yhspdlSpwDSTO/openapi.json
