Article Extractor: Clean Text, Author and Date from News URLs avatar

Article Extractor: Clean Text, Author and Date from News URLs

Pricing

Pay per event

Go to Apify Store
Article Extractor: Clean Text, Author and Date from News URLs

Article Extractor: Clean Text, Author and Date from News URLs

Turn news and blog article URLs into clean text or Markdown with title, author, publish date, site, language, tags, lead image and word count. Fast plain HTTP, no browser, flat price per article, failed URLs free. Built for media monitoring, research and LLM pipelines.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Hay Equipos

Hay Equipos

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

an hour ago

Last modified

Share

Give this actor a list of news, blog or magazine article URLs and get back one clean row per article: the main text without menus, ads and sidebars, plus title, author, publish and update dates, site name, language, section, tags, lead image, word count and reading time. Markdown and cleaned HTML are available too, ready for LLM, RAG and summarization pipelines.

It uses Mozilla Readability (the engine behind Firefox Reader View) for the text and reads the publisher's own metadata (Open Graph, article tags, JSON LD and time tags) for the facts, so dates and authors come from the source rather than guesses. It fetches each page once over plain HTTP, with no browser, which keeps it fast and cheap. You pay per article that was extracted, plus a tiny start fee per run; failed URLs cost nothing.

What you can use it for

  • Media monitoring: turn a list of article links into a clean table of headlines, authors and dates.
  • Feed articles into an LLM for summaries, classification or sentiment.
  • Build a RAG knowledge base from blog posts and documentation pages as Markdown.
  • Research: collect word counts, publish dates and tags across many publications.
  • Content audits: check which of your own posts lack an author, a date or a lead image.

Input

FieldWhat it doesDefault
Article URLsThe pages to extract. The actor reads only these; it does not follow linksrequired
Include plain textMain text with paragraph breakson
Include MarkdownMain text as Markdown with headings, lists, links and imagesoff
Include cleaned HTMLThe cleaned article HTMLoff
Maximum text lengthCut text, Markdown and HTML to this many characters (0 means no limit)0
Respect robots.txtSkip pages a site closes to automated tools (skipped pages are free)on
Parallel sitesHow many URLs to work on at once5

Example input:

{
"urls": [
"https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
"https://blog.apify.com/what-is-web-scraping/"
],
"includeMarkdown": true,
"maxTextLength": 20000
}

Output

One row per URL. Successful rows have success: true; failed rows have success: false and an error that says why (HTTP error, bot check, not an article, robots.txt). Download as JSON, CSV or Excel, or read it through the API.

{
"url": "https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
"finalUrl": "https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
"canonicalUrl": "https://techcrunch.com/2026/09/26/pnoes-new-face-mask-wants-to-make-lab-grade-breath-testing-a-self-serve-affair/",
"success": true,
"statusCode": 200,
"title": "PNOE’s new face mask wants to make lab-grade breath testing a self-serve affair",
"author": "Connie Loizos",
"publishedAt": "2026-09-27T01:40:30.000Z",
"modifiedAt": "2026-09-27T01:40:40.000Z",
"siteName": "TechCrunch",
"description": "PNOĒ, the Malden, Mass.-based startup whose breath-analyzing mask...",
"language": "en-US",
"section": "Hardware",
"tags": [],
"leadImageUrl": "https://techcrunch.com/wp-content/uploads/2026/09/PNOE-2.0.png?resize=1200,675",
"wordCount": 1053,
"readingTimeMinutes": 5,
"isAccessibleForFree": null,
"likelyArticle": true,
"excerpt": "At first glance, the newest device from PNOĒ looks like something a comic-book villain might wear...",
"textTruncated": true,
"text": "At first glance, the newest device from PNOĒ looks like something a comic-book villain might wear. The mask, which covers the nose and mouth...",
"markdown": "At first glance, the newest device from PNOĒ looks like something a comic-book villain might wear...",
"extractedAt": "2026-09-27T06:40:01.577Z"
}

Pricing

Pay per event, no subscription, no charge for platform usage on top.

EventPrice
Article extracted$0.0015 ($1.50 per 1,000 articles)
Actor start$0.00005 per run (Apify's standard start event)

Failed URLs, pages that are not articles, bot checks and robots.txt skips are free. You can set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

Limits

  • Plain HTTP only. Pages that build their text with JavaScript after loading, or that sit behind a bot check or a login, come back as a free error row instead of text.
  • Some publishers refuse automated readers or tools run on Apify, either in robots.txt or at their server (in testing on 27 Sep 2026: the robots.txt files of BBC and Le Monde close them to Apify crawlers, The Guardian answered 403 and NPR stopped answering). The actor respects that and returns a free error row; it does not disguise itself.
  • Paywalled articles return only what the publisher shows to anonymous readers. isAccessibleForFree tells you when the publisher marks an article as paid.
  • Home pages, section pages and video pages are rejected as "not an article" when they have under 100 words of main text.
  • Requests to the same site are spaced one second apart, so 1,000 URLs from one site take about 17 minutes; URLs spread across many sites run in parallel.
  • Up to 5,000 URLs per run.

FAQ

Which sites work? Most news sites, blogs, magazines, documentation pages and Wikipedia. It works best on server rendered pages, which is most publishing sites.

Where does the author come from? From the page's JSON LD, then its author meta tags, then the byline Readability finds. If none exists the field is empty rather than guessed.

Can I use the text in my product? The actor extracts what you point it at. You are responsible for having the right to use the content, as with any reader tool. The robots.txt option is on by default.

Does it crawl a whole site? No. It reads exactly the URLs you give it. Use a sitemap or RSS tool to collect article links first, then pass them in.

Why did a URL fail? The error field says: HTTP status, timeout, bot check, not HTML, too little text, or robots.txt. Failed rows are never charged.