Web Article Extractor — Clean Reader Mode Text & Metadata avatar

Web Article Extractor — Clean Reader Mode Text & Metadata

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Web Article Extractor — Clean Reader Mode Text & Metadata

Web Article Extractor — Clean Reader Mode Text & Metadata

Turn any web page into clean readable text — headings, paragraphs and lists in order. Plain text, Markdown or HTML output, ideal for AI/LLM input. Bulk URLs.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Maged

Maged

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

14 hours ago

Last modified

Share

What does Web Article Extractor do?

Web Article Extractor turns any web page into clean, readable text, like your browser's reader mode. It strips menus, ads and clutter and keeps the headings, paragraphs and lists in their original order. Choose plain text, Markdown or HTML output, which makes it ideal for AI/LLM pipelines, summarization, translation and content archiving.

It runs on the Apify platform, so you can process many URLs in one run, call it through the API, schedule it, and send the results to Google Sheets, Zapier, Make or your own app.

Why use Web Article Extractor?

  • LLM and RAG input: get clean Markdown to feed ChatGPT, Claude or a vector database, without HTML noise.
  • Summarization and translation: extract the text first, then process only the content that matters.
  • Content research: collect competitor articles, blog posts and documentation as text.
  • Archiving and offline reading: save articles as clean text or simple HTML.
  • Works on modern sites: pages are loaded like a real visitor, so JavaScript-rendered content is captured.

How to extract article text from a URL

  1. Click Try for free (or Start).
  2. Paste a Page URL, or several under Page URLs.
  3. Choose the Output format: Plain text, Markdown or HTML.
  4. Click Start and download the results as JSON, CSV or Excel.

Input

FieldDescription
Page URLA single page to extract
Page URLsSeveral pages, one per line (each becomes one result)
Output formatplain, markdown or html
{
"target_urls": [
"https://jamesclear.com/five-step-creative-process",
"https://en.wikipedia.org/wiki/Web_scraping"
],
"output_format": "markdown"
}

Output

Each page is one row. You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

{
"url": "https://jamesclear.com/five-step-creative-process",
"title": "The 5-Step Creative Process",
"content": "# The 5-Step Creative Process\n\nIn 1940, an advertising executive named James Webb Young...",
"format": "markdown",
"wordCount": 1534,
"headings": ["The 5-Step Creative Process", "Step 1: Gather new material"],
"error": null,
"scrapedAt": "2026-10-02T09:00:00+00:00"
}

Data fields

FieldDescription
urlThe page that was read
titlePage title
contentClean article content in the chosen format
formatOutput format used
wordCountNumber of words extracted
headingsThe page's headings, in order (handy as an outline)
errorWhy a page could not be read, if it failed

How many results will I get?

You're charged per result, so the cost follows the number of rows below. Your plan's rates are on the Pricing tab.

InputResults
1 URL1 row
100 URLs100 rows

Tips

  • Use Markdown for AI. It keeps the structure (headings, lists) that helps language models understand the text.
  • Check wordCount. A very low count usually means the page is a login wall, a paywall or mostly images.
  • Whole-site conversion: to convert entire websites to Markdown, see Web Page to Markdown Converter Pro.

FAQ

Does it work on paywalled pages? It reads what a normal visitor sees. Content hidden behind a login or paywall isn't available.

Are images included? No. The output is text-focused. Use Website Image Scraper to collect images.

Found a bug or need a custom solution? Open an issue in the Issues tab. Custom versions are available on request.


⭐ Found this Actor useful? A quick review on the Store helps other users find it and keeps it maintained.