News Article Extractor
Pricing
from $2.00 / 1,000 results
News Article Extractor
Extracts clean article text, headline, authors, publish date, and images from any news site — or discovers the articles itself from a homepage
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Mina
Maintained by CommunityActor stats
0
Bookmarked
4
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
Turn any news article URL into clean structured data — headline, full text, authors, publish date, and images, with the ads, navigation, and cookie banners stripped out. Point it at a news site and it finds the articles for you.
What does News Article Extractor do?
Give it article links, or give it a news site and let it discover the articles itself. For each article you get the headline, the full body text with boilerplate removed, the authors, the publication date, the lead image, and the page's own metadata — description, keywords, tags, language, and canonical URL.
It works on any news site rather than a fixed list, so you are not limited to a set of supported publishers.
Why use News Article Extractor?
- Media monitoring — track coverage of a brand, person, or topic across many outlets.
- AI and RAG pipelines — clean article text is far better training and retrieval input than raw HTML.
- Content research — study how competitors cover a subject, and how often.
- Archiving — keep a durable, readable copy of reporting that may change or disappear.
- Newsletters and aggregation — pull the day's stories from a set of sources on a schedule.
Running it on Apify adds scheduling, an API, integrations (Sheets, Slack, Zapier, S3), automatic proxy fallback, and run monitoring.
How to use News Article Extractor
- Paste article links into Article URLs, or news site homepages into News sites — or both.
- Set Max articles.
- Click Start.
Results appear in the Output tab as they are extracted, and download as JSON, CSV, Excel, or XML.
Input
| Field | Description |
|---|---|
article_urls | Direct links to articles, one per line. |
site_urls | Homepage or section URLs. The Actor finds the article links on each page (skipping section, video and tag pages) and extracts them, sharing Max articles fairly across the sites. A discovered page that isn't an article is skipped and not billed. |
max_articles | Total cap across every site and URL given. |
language | Two-letter hint such as en or ar. Leave empty to detect it — set it only when detection gets a site wrong. |
include_article_html | Also return the cleaned HTML of the article body, keeping links and formatting. |
max_concurrency | How many pages to fetch at once. Lower it if a publisher rate-limits you. |
proxy_country | Comma-separated country codes for the residential proxy. Pages are fetched directly first, then through Apify's datacenter proxy; residential exits are used only for publishers that refuse both. |
Output
{"url": "https://techcrunch.com/2026/08/20/example-story/","canonical_url": "https://techcrunch.com/2026/08/20/example-story/","title": "Apollo Atomics wants to make nuclear power cheap","text": "A nuclear start-up has raised forty million dollars...","word_count": 498,"authors": ["Jane Roe"],"published_at": "2026-08-20T10:30:00+00:00","language": "en","site_name": "TechCrunch","description": "The company plans to complete regulatory review this year.","keywords": ["nuclear", "energy"],"tags": ["energy"],"top_image": "https://techcrunch.com/wp-content/uploads/lead.jpg","images": ["https://techcrunch.com/wp-content/uploads/inline.jpg"],"videos": []}
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
Data fields
| Field | Description |
|---|---|
url, canonical_url | The link scraped, and the publisher's own canonical version of it. |
title, text, word_count | Headline and clean body text. |
authors, published_at | Byline and publication date, as an ISO timestamp where the page declares one. |
language, site_name | Detected language and the publication's name. |
description, keywords, tags | The page's own metadata. |
top_image, images, videos | Lead image, other images in the article, and embedded video URLs. |
article_html | Cleaned HTML of the body, when you ask for it. |
How much does it cost to extract articles?
Cost tracks the number of articles you extract, so Max articles is the main lever on what a run costs.
Tips
- Section pages beat homepages. Pointing at
example.com/technology/gives more focused results than the homepage. - Lower the concurrency to 1 or 2 if a publisher starts refusing you. Five works well across mixed sources.
- Schedule it against the same sites daily and the dataset becomes a rolling archive.
- Check
FAILED_INPUTS. Links that could not be read are not written to the dataset or charged; each is listed with the reason in theFAILED_INPUTSrecord of the run's key-value store, so nothing disappears silently.
FAQ and support
Which news sites does it support? Any of them. There is no supported-site list to check against.
Why is an article missing from the results? Some articles are behind a paywall, and a few sites load their text in a way that cannot be read. Those links are not charged; they are listed with the reason in the FAILED_INPUTS record of the run's key-value store, so you can see exactly which links need another approach.
Why is the publish date sometimes empty? Not every publisher declares one in a machine-readable way. When it is missing from the page, it is missing from the output — the Actor does not guess.
Is scraping news articles legal? Article text is usually copyrighted. Extracting it for analysis, monitoring, or internal research is common practice, but republishing it is a different matter. You are responsible for how you use the output, including copyright and each publisher's terms. Consult a lawyer if you are unsure.
Found a bug or want a field that is missing? Open an issue on the Actor's Issues tab.
SEO Keywords
news article extractor, article text scraper, news scraping api, extract article content, news aggregator data, media monitoring tool, article headline author and date, boilerplate removal scraper, news homepage crawler, press coverage tracking