Article & News Extractor (clean text, author, date, markdown) avatar

Article & News Extractor (clean text, author, date, markdown)

Pricing

from $1.20 / 1,000 results

Go to Apify Store
Article & News Extractor (clean text, author, date, markdown)

Article & News Extractor (clean text, author, date, markdown)

Article & News Extractor returns clean article text, title, author(s), publish/modified date, tags and images from any news or blog URL — one row per URL, as Markdown, plain text or HTML.

Pricing

from $1.20 / 1,000 results

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What is Article & News Extractor?

Article & News Extractor is an Apify Actor that turns any news article or blog post URL into clean, structured data: title, byline, publish/modified date, tags, images and the full body as Markdown, plain text or HTML. Point it at blog.apify.com, a BBC News story, a TechCrunch post or any article-shaped page and it strips the navigation, ads, comments and related-articles clutter, leaving the words a human — or an LLM — actually came for. It reads Schema.org Article/NewsArticle JSON-LD first, falls back to Open Graph/Twitter meta tags and a Readability-style content extraction, and needs no browser or proxy, so it is fast and cheap even for thousands of URLs. Runs from the Apify Console, the API, schedules, or the Apify MCP server for AI agents.

What data does Article & News Extractor extract?

One row per URL, with every field present (null when the source does not publish it):

FieldTypeDescription
url, finalUrl, statusCodestring, numberURL requested, URL after redirects, HTTP status
title, subtitlestringHeadline and dek/subheading
author, authorsstring, arrayByline joined with commas, and each author name separately
publishedAt, modifiedAtstringFirst-publish and last-modified timestamps
publisher, siteName, sectionstringPublishing organisation, site/brand name, editorial section
tagsarrayKeywords/tags the page declares
langstringPage language code, e.g. en
descriptionstringMeta/OG summary
mainImage, imagesstring, arrayLead image, and up to 20 other image URLs from the body
wordCount, readingTimeMinutesnumberBody length and estimated reading time at 200 wpm
markdown, text, htmlstringBody in the format(s) outputFormat asks for; the others are null
canonicalstringThe page's canonical URL
isPaywalledbooleanTrue when the page signals a paywall
extractionMethodstringHow the body was found: json-ld, readability or heuristic
error, scrapedAtstringWhy a row failed, and when it was fetched

How to use Article & News Extractor

  1. Paste the article URLs you want into Article URLs — one row comes back per URL.
  2. Pick an Output format: markdown (default, best for LLM/RAG ingestion), text, html, or all for every format at once.
  3. Click Start. Export the dataset as JSON, CSV, Excel or HTML, or pull it via the API.

Example input

{
"urls": ["https://blog.apify.com/best-web-scraping-tools/", "https://techcrunch.com/2026/09/09/apple-unveils-its-first-foldable-the-iphone-duo/"],
"outputFormat": "markdown",
"includeImages": true,
"maxConcurrency": 5
}

Example output

{
"url": "https://blog.apify.com/best-web-scraping-tools/",
"finalUrl": "https://blog.apify.com/best-web-scraping-tools/",
"statusCode": 200,
"title": "The best web scraping tools for 2026",
"author": "Theo Vasilis",
"authors": ["Theo Vasilis"],
"publishedAt": "2026-06-19T08:00:00.000Z",
"publisher": "Apify Blog",
"siteName": "Apify Blog",
"wordCount": 5312,
"readingTimeMinutes": 27,
"markdown": "Finding accurate, up-to-date data at scale has become essential...",
"text": null,
"html": null,
"isPaywalled": false,
"extractionMethod": "readability",
"error": null,
"scrapedAt": "2026-09-12T21:00:00.000Z"
}

You can download the dataset in various formats such as JSON, HTML, CSV or Excel.

Input parameters

ParameterTypeDefaultDescription
urlsarray["https://blog.apify.com/best-web-scraping-tools/"]Article URLs to extract, one row each
outputFormatstringmarkdownmarkdown, text, html, or all
includeImagesbooleantrueReturn mainImage/images and keep <img> in the body
maxConcurrencyinteger5URLs fetched in parallel (1-50)

Pricing

Article & News Extractor uses pay-per-event pricing: $0.002 per extracted article, i.e. $2 per 1,000 articles, plus a negligible actor-start fee. Every article is a single HTTP GET with no browser and no proxy, so runs of thousands of URLs stay cheap. Set Maximum cost per run and the Actor trims the URL list to what the budget covers instead of overspending.

Article & News Extractor vs. Diffbot and browser-based scrapers

Diffbot's Article API charges per API call at enterprise pricing tiers and requires a subscription; browser-based article scrapers spin up a full Chromium instance per page, which costs more compute and is slower for pages that do not need JavaScript rendering. Article & News Extractor reads the same signals server-rendered pages already publish — JSON-LD, Open Graph, a readable DOM — with a single lightweight HTTP request, at a fraction of the per-page cost, and returns Markdown ready for an LLM prompt or a RAG index rather than a raw API payload to reshape.

Using Article & News Extractor with AI agents and MCP

Article & News Extractor is pay-per-event with limited permissions — the two requirements for an Actor to be callable through the Apify MCP server at mcp.apify.com. An agent passes a list of urls and gets back clean Markdown per article — ready to drop straight into a prompt, a RAG index, a news-monitoring dashboard or a newsletter pipeline — without writing or maintaining a site-specific scraper. The same run works from n8n, Make, Zapier and LangChain through Apify's integrations.

FAQ

What if a page has no JSON-LD or usable body text? The row still returns whatever Open Graph/Twitter metadata exists, extractionMethod: null, markdown/text/html all null, and error: "Could not find article content on this page" — the run still succeeds for every other URL.

Does this work on JavaScript-rendered sites? Only for what is present in the initial HTML response. Most news and blog CMSes server-render the article body and JSON-LD for SEO, so this covers the overwhelming majority of article pages without needing a browser.

How is isPaywalled detected? First from an explicit isAccessibleForFree flag in the page's JSON-LD; failing that, from a short body combined with wording like "subscribe to continue reading". It is a heuristic, not a guarantee — always check wordCount on suspect rows.

Is this legal to run? Yes. It reads exactly the HTML, meta tags and structured data a page already serves to browsers and search engines. You are responsible for complying with the target site's terms of use for your own use case.

Can I export to CSV or Excel? Yes, from the Output tab or the API, with a ready-made Overview view.

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data

Support and feedback

Found an article that extracts poorly, or a site whose byline/date format is not recognised? Open an issue on the Issues tab.