Article Extractor avatar

Article Extractor

Pricing

from $0.50 / 1,000 articles

Go to Apify Store
Article Extractor

Article Extractor

Clean text and Markdown of news and blog articles, with the headline, publish date, section and tags. Give it article links, an RSS feed or a whole site; for feeds and sites it extracts the newest articles.

Pricing

from $0.50 / 1,000 articles

Rating

0.0

(0)

Developer

deriverge s.r.o.

deriverge s.r.o.

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

18 hours ago

Last modified

Share

Article Extractor turns an article page into one row: the article text without menus and ads, optional Markdown with headings, links and images, and the metadata publishers add for search engines, such as the headline, publish and update dates, section, tags, main image and paywall flag. Paste an article link, an RSS or Atom feed, or a site address and click Start; for a feed or a site it extracts the newest articles. Pages that build their text with JavaScript return no article. It costs $1.00 per 1,000 articles on the Free plan ($0.50 on Business), and the $5 of monthly credit in Apify's Free plan covers about 5,000 articles.

What data does it return?

FieldDescription
title, descriptionHeadline, and the article's short description, usually from its structured data or meta tags.
textThe article as plain text, with a blank line between paragraphs where the page marks them.
markdownThe article as Markdown with headings, lists, links and images, when Output format includes Markdown. null otherwise.
publishedAt, modifiedAtPublish and last update time, ISO 8601 UTC, from the page's structured data or meta tags.
siteName, section, tags, languagePublication name, section such as Politics, tags such as Programming, and the language code, such as en-GB.
imageLink to the main image of the article.
isAccessibleForFreefalse when the publisher marks the article as paywalled, true when it marks it as free to read, null when the page does not say.
wordCount, readingTimeMinutesWords in text, and the reading time at 230 words per minute.
url, canonicalUrl, sourcePageThe page that was read, the canonical address it declares, and the feed or site it was found through.
status, errorok (extracted and charged), not_article, blocked or error, with the reason in error.
isNewWith a watch name, true for an article that no earlier run with that name returned.

How to extract articles

  1. Add addresses to Articles, feeds or websites, one per line: an article such as https://blog.apify.com/what-is-web-scraping/, a feed such as https://feeds.bbci.co.uk/news/rss.xml, or a site or section such as https://blog.apify.com.
  2. For feeds and sites, set Maximum articles per feed or website, 20 by default, taken in the order the feed lists them.
  3. Choose the Output format: plain text, or plain text and Markdown.
  4. Click Start. Articles appear in the Output tab and can be exported as CSV, Excel or JSON or read through the API.

An API input that follows a news feed and a blog with Markdown and returns only articles it has not returned before:

{
"urls": ["https://feeds.bbci.co.uk/news/rss.xml", "https://blog.apify.com"],
"maxArticlesPerSite": 10,
"format": "markdown",
"newOnly": true,
"watchName": "tech-news"
}

Example output

One article from a run on 30 September 2026 with Markdown on. text and markdown are shortened here; the full row holds the whole article of 1,363 words. The run came before that day's paragraph fix, which is why text still runs two paragraphs together in web scraping.You could.

{
"url": "https://blog.apify.com/what-is-web-scraping/",
"canonicalUrl": "https://blog.apify.com/what-is-web-scraping/",
"sourcePage": null,
"status": "ok",
"error": null,
"title": "What is web scraping?",
"description": "The basics of web scraping: what it is, how it works, real-world use cases, and how to get started.",
"siteName": "Apify Blog",
"language": "en",
"publishedAt": "2024-10-15T14:32:00.000Z",
"modifiedAt": "2026-07-07T12:57:19.000Z",
"section": null,
"tags": [
"Programming"
],
"image": "https://storage.ghost.io/c/f2/6e/f26ec999-9a90-4aee-a0d4-9b3ca2bb668f/content/images/2024/04/what-is-web-scraping-complete-guide.png",
"isAccessibleForFree": null,
"wordCount": 1363,
"readingTimeMinutes": 6,
"text": "Web scraping is the process of automatically extracting data from a website. You use a program called a web scraper to access a web page, interpret the data, and extract what you need. The data is saved in a structured format such as an Excel file, JSON, or XML so that you can use it in spreadsheets or apps. It's also easy to confuse with crawling: see web crawling vs. web scraping.You could do this manually by copying and pasting, but scraping is typically performed using an automated tool ...",
"markdown": "Web scraping is the process of automatically extracting data from a website. You use a program called a web scraper to access a web page, interpret the data, and extract what you need. The data is saved in a structured format such as an Excel file, JSON, or XML so that you can use it in spreadsheets or apps. It's also easy to confuse with crawling: see [web crawling vs. web scraping](https://blog.apify.com/web-crawling-vs-web-scraping).\n\n![What is web scraping: diagram showing data going from website through scraping platform to structured data](https://storage.ghost.io/c/f2/6e/f26ec999-9a90-4aee-a0d4-9b3ca2bb668f/content/images/2023/09/what-is-web-scraping-websites-web-scraper-structured-data-1.png)\n\n...",
"isNew": null,
"checkedAt": "2026-09-30T00:34:16.088Z"
}

How much does it cost to extract articles?

Free planStarterScaleBusiness
1,000 articles$1.00$0.80$0.65$0.50

You pay only for the events in the table. There is no start fee, and compute time and proxies are included.

1,000 articles cost $1.00 on the Free plan and $0.50 on Business, with or without Markdown. Pages without an article, blocked pages, failed downloads and articles skipped by the new-only mode are not charged.

Limits

  • When text comes from the article body in the page's structured data, it has only the line breaks the publisher put there, and markdown then holds the same plain text.
  • Pages that load their text with JavaScript give a not_article row.
  • Paywalls are not bypassed: a paywalled article gives only the part the page shows without a subscription.
  • For a homepage or section page, the actor uses the feed the page links to, or tries common feed paths such as /feed, /rss and /rss.xml. Without any feed, the page gives a not_article row.
  • A site that answers with a bot protection page gives a blocked row.
  • The new-only memory keeps each article for 180 days.

Following a site or feed

Give the run a Watch name such as tech-news, or save it as a task, turn on Return only articles new since the last run and schedule it. Each run then extracts only articles it has not returned before, so a daily run becomes a feed of full texts. Without new-only mode, a watch name still marks every article with isNew.

FAQ

The actor reads pages that anyone can open without logging in, much like a browser's reader mode, and leaves out author names. The articles remain their publishers' work, so check the publisher's terms before you republish any text. You are responsible for how you use the results.

How is the article text found?

With Mozilla's Readability library, the code behind the reader view in Firefox, run on the page's HTML. When the page's structured data carries the full article body and it is about as long as that result, the structured version is used instead.

What if a feed lists only headlines?

It works the same way. The actor opens every article the feed links to, up to Maximum articles per feed or website, and extracts the text from the article page itself.

Support

This actor is built and maintained by deriverge s.r.o., a software company based in the Czech Republic. If a run fails or a field you need is missing, please open an issue in the Issues tab or write to us at info@deriverge.com. We respond in English and Czech. Runs can be scheduled in Apify Console or started from the API tab, which has examples for Python, JavaScript and cURL and works with Make, Zapier, n8n and the Apify MCP server. If the actor saves you time, a short review helps other people find it.