Google News & RSS Scraper — Articles, Publishers, Dates avatar

Google News & RSS Scraper — Articles, Publishers, Dates

Pricing

Pay per usage

Go to Apify Store
Google News & RSS Scraper — Articles, Publishers, Dates

Google News & RSS Scraper — Articles, Publishers, Dates

Aggregate and parse RSS/Atom feeds from any source. Extract articles with titles, descriptions, authors, dates, images. Optionally fetch full article content. Perfect for news monitoring and AI pipelines. $0.0005/article.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Ken Digital

Ken Digital

Maintained by Community

Actor stats

0

Bookmarked

30

Total users

4

Monthly active users

8 days ago

Last modified

Categories

Share

Search Google News by keyword, language and country, and parse any RSS or Atom feed, into clean article rows: title, publisher, link, publication date, summary and image. Built for news monitoring, media tracking and feeding LLM pipelines.

Google News in one line

{
"googleNewsQueries": ["openai", "\"electric vehicles\" site:reuters.com"],
"googleNewsLanguage": "en",
"googleNewsCountry": "US",
"googleNewsTimeRange": "1d"
}

Each query returns up to about 100 recent articles, with the publisher in source and its site in sourceUrl; the publisher suffix is removed from title. Google News operators work inside the query: quotes, site:, intitle:, -exclude. Mix Google News queries and regular feed URLs in the same run.

link for Google News articles is Google's redirect URL, which opens the original article in a browser. Full-text extraction (fetchFullContent) applies to regular feeds only.

Features

  • Google News search — keyword, language, country and time-range filters, publisher on every row
  • Multi-format support — RSS 2.0, RSS 1.0 (RDF), and Atom feeds
  • Namespace handling — Parses media:, dc:, content:encoded, and standard Atom namespaces
  • Full content extraction — Optionally follows article links and strips HTML to plain text
  • Robust parsing — Handles encoding issues, CDATA blocks, malformed dates, and missing fields
  • Structured output — Consistent schema across all feed types

Output Schema

Each article in the dataset contains:

FieldTypeDescription
feedUrlstringSource feed URL
feedTitlestringFeed/channel title
titlestringArticle title
linkstringArticle permalink (Google redirect URL for Google News)
sourcestringPublisher name, when the feed provides it (always for Google News)
sourceUrlstringPublisher website
descriptionstringArticle summary (HTML stripped, max 5000 chars)
authorstringAuthor name
publishedDatestringISO 8601 publication date
categoriesstring[]Tags/categories from the feed
imageUrlstringThumbnail or featured image URL
guidstringUnique identifier (GUID or permalink)
fullContentstringFull article text (only when fetchFullContent is enabled)

Input Parameters

{
"feedUrls": [
"https://feeds.bbci.co.uk/news/rss.xml",
"https://rss.nytimes.com/services/xml/rss/index.xml",
"https://hnrss.org/frontpage"
],
"maxArticles": 100,
"fetchFullContent": false
}
  • googleNewsQueries — Google News searches, one feed per query
  • googleNewsLanguage (default: en), googleNewsCountry (default: US) — edition to search
  • googleNewsTimeRange (default: any) — 1h, 1d, 7d or 30d
  • feedUrls — Array of RSS/Atom feed URLs to aggregate. At least one of feedUrls or googleNewsQueries is required.
  • maxArticles (default: 100) — Maximum articles for the whole run, across all feeds and queries. Set to 0 for unlimited.
  • fetchFullContent (default: false) — Follow article links and extract full text content

Use Cases

📡 News Monitoring & Media Intelligence

Track coverage across dozens of news outlets. Monitor specific topics by aggregating topic-specific RSS feeds from major publishers. Feed results into sentiment analysis or trend detection pipelines.

📋 Content Curation & Newsletters

Aggregate content from niche blogs, industry publications, and thought leaders into a single dataset. Use as the data source for automated newsletter generation or content recommendation systems.

🔍 Competitive Intelligence

Subscribe to competitor blogs, press release feeds, and industry news. Get structured alerts when new content is published. Combine with keyword filtering for targeted monitoring.

📊 Research & Dataset Building

Build timestamped article datasets for NLP research, media studies, or training data collection. The consistent schema makes downstream processing straightforward.

🤖 AI Pipeline Input

Use as a data source for LLM-powered summarization, classification, or knowledge base updates. The structured output integrates directly with vector databases and RAG pipelines.

⏰ Scheduled Monitoring

Run on a schedule (hourly, daily) with Apify's scheduling feature. Combine with deduplication logic downstream to maintain a continuously updated article database.

Technical Notes

  • Uses Python stdlib xml.etree.ElementTree for XML parsing (no lxml dependency)
  • HTTP requests via httpx with async support and configurable timeouts
  • Handles BOM-prefixed feeds and common encoding edge cases
  • Date parsing supports RFC 822 (RSS) and ISO 8601 (Atom) formats
  • Full content extraction removes <script>, <style>, and <noscript> blocks before stripping HTML

Pricing

$0.001 per article parsed and pushed to the dataset.

Example Feeds to Get Started

FeedURL
BBC Newshttps://feeds.bbci.co.uk/news/rss.xml
Hacker Newshttps://hnrss.org/frontpage
TechCrunchhttps://techcrunch.com/feed/
ArXiv CS.AIhttp://arxiv.org/rss/cs.AI
Reddit r/technologyhttps://www.reddit.com/r/technology/.rss

🔗 More Scrapers by Ken Digital

ScraperWhat it doesPrice
YouTube Channel ScraperVideos, stats, metadataFree (pay for compute only)
France Job ScraperWTTJ + France Travail + Hellowork$0.003/job
France Real Estate Scraper5 sources + DVF price analysis$0.005/listing
Website Content CrawlerHTML → Markdown for AI/RAG$0.002/page
Google Trends ScraperKeywords, regions, related queries$0.005/keyword
GitHub Repo ScraperStars, forks, languages, topicsFree (pay for compute only)
RSS News AggregatorMulti-source feed parsing$0.001/article
Instagram Profile ScraperFollowers, bio, posts$0.005/profile
Google Maps ScraperBusinesses, reviews, contactsFree (pay for compute only)
TikTok ScraperVideos, likes, sharesFree (pay for compute only)
Google SERP ScraperSearch results, PAA, snippets$0.003/search
Trustpilot ScraperReviews, ratings, sentiment$0.002/review

👉 View all scrapers

🔗 Quick Integration

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("joyouscam35875/rss-news-aggregator").call(run_input={...})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('joyouscam35875/rss-news-aggregator').call({...});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

No-code: Make / Zapier / n8n

Search for this actor in the Apify connector. No code needed.

Feedback

Missing a field, hit a bug, or need an input this actor doesn't take yet? Open an issue on the actor's Issues tab. If the Google News & RSS Scraper saved you time, a short review on the Store page helps other people find it.