Article Extractor
Pricing
from $0.50 / 1,000 articles
Article Extractor
Clean text and Markdown of news and blog articles, with the headline, publish date, section and tags. Give it article links, an RSS feed or a whole site; for feeds and sites it extracts the newest articles.
Pricing
from $0.50 / 1,000 articles
Rating
0.0
(0)
Developer
deriverge s.r.o.
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
18 hours ago
Last modified
Categories
Share
Article Extractor turns an article page into one row: the article text without menus and ads, optional Markdown with headings, links and images, and the metadata publishers add for search engines, such as the headline, publish and update dates, section, tags, main image and paywall flag. Paste an article link, an RSS or Atom feed, or a site address and click Start; for a feed or a site it extracts the newest articles. Pages that build their text with JavaScript return no article. It costs $1.00 per 1,000 articles on the Free plan ($0.50 on Business), and the $5 of monthly credit in Apify's Free plan covers about 5,000 articles.
What data does it return?
| Field | Description |
|---|---|
title, description | Headline, and the article's short description, usually from its structured data or meta tags. |
text | The article as plain text, with a blank line between paragraphs where the page marks them. |
markdown | The article as Markdown with headings, lists, links and images, when Output format includes Markdown. null otherwise. |
publishedAt, modifiedAt | Publish and last update time, ISO 8601 UTC, from the page's structured data or meta tags. |
siteName, section, tags, language | Publication name, section such as Politics, tags such as Programming, and the language code, such as en-GB. |
image | Link to the main image of the article. |
isAccessibleForFree | false when the publisher marks the article as paywalled, true when it marks it as free to read, null when the page does not say. |
wordCount, readingTimeMinutes | Words in text, and the reading time at 230 words per minute. |
url, canonicalUrl, sourcePage | The page that was read, the canonical address it declares, and the feed or site it was found through. |
status, error | ok (extracted and charged), not_article, blocked or error, with the reason in error. |
isNew | With a watch name, true for an article that no earlier run with that name returned. |
How to extract articles
- Add addresses to Articles, feeds or websites, one per line: an article such as
https://blog.apify.com/what-is-web-scraping/, a feed such ashttps://feeds.bbci.co.uk/news/rss.xml, or a site or section such ashttps://blog.apify.com. - For feeds and sites, set Maximum articles per feed or website, 20 by default, taken in the order the feed lists them.
- Choose the Output format: plain text, or plain text and Markdown.
- Click Start. Articles appear in the Output tab and can be exported as CSV, Excel or JSON or read through the API.
An API input that follows a news feed and a blog with Markdown and returns only articles it has not returned before:
{"urls": ["https://feeds.bbci.co.uk/news/rss.xml", "https://blog.apify.com"],"maxArticlesPerSite": 10,"format": "markdown","newOnly": true,"watchName": "tech-news"}
Example output
One article from a run on 30 September 2026 with Markdown on. text and markdown are shortened here; the full row holds the whole article of 1,363 words. The run came before that day's paragraph fix, which is why text still runs two paragraphs together in web scraping.You could.
{"url": "https://blog.apify.com/what-is-web-scraping/","canonicalUrl": "https://blog.apify.com/what-is-web-scraping/","sourcePage": null,"status": "ok","error": null,"title": "What is web scraping?","description": "The basics of web scraping: what it is, how it works, real-world use cases, and how to get started.","siteName": "Apify Blog","language": "en","publishedAt": "2024-10-15T14:32:00.000Z","modifiedAt": "2026-07-07T12:57:19.000Z","section": null,"tags": ["Programming"],"image": "https://storage.ghost.io/c/f2/6e/f26ec999-9a90-4aee-a0d4-9b3ca2bb668f/content/images/2024/04/what-is-web-scraping-complete-guide.png","isAccessibleForFree": null,"wordCount": 1363,"readingTimeMinutes": 6,"text": "Web scraping is the process of automatically extracting data from a website. You use a program called a web scraper to access a web page, interpret the data, and extract what you need. The data is saved in a structured format such as an Excel file, JSON, or XML so that you can use it in spreadsheets or apps. It's also easy to confuse with crawling: see web crawling vs. web scraping.You could do this manually by copying and pasting, but scraping is typically performed using an automated tool ...","markdown": "Web scraping is the process of automatically extracting data from a website. You use a program called a web scraper to access a web page, interpret the data, and extract what you need. The data is saved in a structured format such as an Excel file, JSON, or XML so that you can use it in spreadsheets or apps. It's also easy to confuse with crawling: see [web crawling vs. web scraping](https://blog.apify.com/web-crawling-vs-web-scraping).\n\n\n\n...","isNew": null,"checkedAt": "2026-09-30T00:34:16.088Z"}
How much does it cost to extract articles?
| Free plan | Starter | Scale | Business | |
|---|---|---|---|---|
| 1,000 articles | $1.00 | $0.80 | $0.65 | $0.50 |
You pay only for the events in the table. There is no start fee, and compute time and proxies are included.
1,000 articles cost $1.00 on the Free plan and $0.50 on Business, with or without Markdown. Pages without an article, blocked pages, failed downloads and articles skipped by the new-only mode are not charged.
Limits
- When
textcomes from the article body in the page's structured data, it has only the line breaks the publisher put there, andmarkdownthen holds the same plain text. - Pages that load their text with JavaScript give a
not_articlerow. - Paywalls are not bypassed: a paywalled article gives only the part the page shows without a subscription.
- For a homepage or section page, the actor uses the feed the page links to, or tries common feed paths such as
/feed,/rssand/rss.xml. Without any feed, the page gives anot_articlerow. - A site that answers with a bot protection page gives a
blockedrow. - The new-only memory keeps each article for 180 days.
Following a site or feed
Give the run a Watch name such as tech-news, or save it as a task, turn on Return only articles new since the last run and schedule it. Each run then extracts only articles it has not returned before, so a daily run becomes a feed of full texts. Without new-only mode, a watch name still marks every article with isNew.
FAQ
Is it legal to extract articles?
The actor reads pages that anyone can open without logging in, much like a browser's reader mode, and leaves out author names. The articles remain their publishers' work, so check the publisher's terms before you republish any text. You are responsible for how you use the results.
How is the article text found?
With Mozilla's Readability library, the code behind the reader view in Firefox, run on the page's HTML. When the page's structured data carries the full article body and it is about as long as that result, the structured version is used instead.
What if a feed lists only headlines?
It works the same way. The actor opens every article the feed links to, up to Maximum articles per feed or website, and extracts the text from the article page itself.
Related scrapers
Support
This actor is built and maintained by deriverge s.r.o., a software company based in the Czech Republic. If a run fails or a field you need is missing, please open an issue in the Issues tab or write to us at info@deriverge.com. We respond in English and Czech. Runs can be scheduled in Apify Console or started from the API tab, which has examples for Python, JavaScript and cURL and works with Make, Zapier, n8n and the Apify MCP server. If the actor saves you time, a short review helps other people find it.