News & RSS Scraper: Google News to JSON avatar

News & RSS Scraper: Google News to JSON

Pricing

from $1.40 / 1,000 articles

Go to Apify Store
News & RSS Scraper: Google News to JSON

News & RSS Scraper: Google News to JSON

Scrape news and blog articles from Google News searches, Google News topics, and any RSS 2.0, RSS 1.0, Atom or JSON Feed URL. One identical row per article: title, link, publisher, publish date, author, summary, categories, images and optional full article text for RAG pipelines.

Pricing

from $1.40 / 1,000 articles

Rating

0.0

(0)

Developer

Axiora Solutions

Axiora Solutions

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Categories

Share

News & RSS Feed Scraper — Google News, RSS, Atom, JSON Feed

Google News and RSS feed scraper that turns any source into one clean, analysis-ready article list. Give it Google News searches and topics, RSS 2.0, RSS 1.0/RDF, Atom 1.0 or JSON Feed URLs — or just a bare website domain — and it unifies every one of them into a single identical schema. No API key or login is needed, and the fastest way to try it is to leave the prefilled sources and click Start.

What you get

  • title, titleWithoutPublisher and url for every article, with Google News publisher names recovered.
  • publishedAt in ISO 8601, plus updatedAt and ageHours for monitoring runs.
  • contentText and contentChars — the article body when the feed already ships it, at zero extra cost.
  • Optional articleText from the page, with articleTextStatus and articleTextSelector for auditing quality.
  • author, summary, categories, enclosures, imageUrl, publisherName and publisherDomain.
  • Stable articleUid and contentHash keys for incremental indexing and deduplication.

Quick start

  1. Open the Actor and leave the prefilled gnews:web scraping and https://techcrunch.com/feed/ in Sources.
  2. Add any mix of gnews:, gnews-topic: or feed/website URLs to the same Sources list.
  3. Set Max articles per source and any filters, then click Start.
  4. Read the Articles, Text for AI and By source dataset tabs.
{
"sources": ["gnews:\"artificial intelligence\" when:2d", "https://techcrunch.com/feed/"],
"maxItemsPerSource": 50,
"publishedAfter": "48 hours",
"deduplicateBy": "title"
}

Example output

One representative dataset row, with the body already present in the feed:

{
"ok": true,
"errorCode": null,
"articleUid": "a41d7c9b2f0e5538",
"sourceLabel": "https://techcrunch.com/feed/",
"sourceType": "rss-2.0",
"feedUrl": "https://techcrunch.com/feed/",
"feedTitle": "TechCrunch",
"feedLanguage": "en-US",
"title": "Startup raises $40M to automate data pipelines",
"titleWithoutPublisher": "Startup raises $40M to automate data pipelines",
"url": "https://techcrunch.com/2026/10/01/startup-raises-40m/",
"urlResolution": "direct",
"publisherName": null,
"publisherDomain": "techcrunch.com",
"guid": "https://techcrunch.com/?p=2884412",
"publishedAt": "2026-10-01T14:22:00.000Z",
"ageHours": 21.6,
"author": "Jane Doe",
"summary": "The round was led by...",
"contentText": "The round was led by... full article body from the feed ...",
"contentChars": 4820,
"articleText": "The round was led by... extracted from the page ...",
"articleTextChars": 5102,
"articleTextStatus": "ok",
"articleTextSelector": "article",
"canonicalUrl": "https://techcrunch.com/2026/10/01/startup-raises-40m/",
"categories": ["Startups", "Funding"],
"imageUrl": "https://techcrunch.com/wp-content/uploads/round.jpg",
"contentHash": "3fb7a1c08d2e4956",
"scrapedAt": "2026-10-02T12:00:00.000Z"
}

One input list, four kinds of source

Everything goes in the Sources array:

EntryWhat it does
gnews:web scrapingGoogle News search, up to 100 articles
gnews:"series A" when:7dGoogle News search with its own operators (when:, site:, quotes)
gnews:topGoogle News top stories for your language and country
gnews-topic:TECHNOLOGYA Google News topic: WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE, HEALTH
https://techcrunch.com/feed/Any RSS, Atom or JSON feed, read directly
blog.cloudflare.comA plain website — the Actor finds the feed for you

Feed auto-discovery reads <link rel="alternate"> declarations first, then tries ten conventional paths (/feed, /rss.xml, /atom.xml, /index.xml, /feed.json, …). You paste a domain; you get articles.

What makes this feed scraper different

  • 🧩 Four formats, one schema — publishedAt is always ISO 8601 whether the feed used RFC 822, RFC 3339 or a JSON Feed timestamp. categories, enclosures and author are normalised the same way.
  • 🪪 Honest Google News URL handling — Google News items link to a consent-gated redirect page that cannot be resolved without a browser. Most scrapers hand you that URL and say nothing. This Actor sets urlResolution: "google-news-redirect" and gives you publisherName and a clean titleWithoutPublisher so attribution still works.
  • 🧠 Feed content is often the whole article — many publishers ship the full body inside the feed. contentText captures it with zero extra requests, so check contentChars before paying the time cost of full-text extraction.
  • 📄 Real article extraction when you need it — optional full text uses a readability-style pass over article, [itemprop=articleBody], .entry-content and similar containers, and reports articleTextSelector so you can audit quality. articleTextStatus always explains a missing body instead of leaving a silent null.
  • 🔁 Deduplication you control — by canonical URL (tracking parameters stripped), by normalised title (collapses the same story syndicated across twenty outlets), by feed GUID, or off.
  • 🎯 Filters cut your bill — publishedAfter, keywords and excludeKeywords run before anything is written, so filtered articles are never charged.
  • 🛟 One dead feed never fails the run — it becomes an ok: false row with a stable errorCode.

Running on Apify adds scheduling, webhooks, monitoring, API and SDK access, and one-click export to JSON, CSV, Excel, Google Sheets and 20+ integrations.

How to use it

  1. Put your searches, topics, feeds or domains in Sources.
  2. Set Google News language and Google News country if you want a non-US edition (de-DE + DE, en-GB + GB, pt-BR + BR).
  3. Set Max articles per source and Max articles for the whole run.
  4. Add Published after (24 hours, 7 days) for monitoring runs.
  5. Turn on Fetch full article text only if contentChars from a test run shows the feeds are truncated.
  6. Click Start, then use the Articles, Text for AI and By source dataset tabs.

How do I monitor brand mentions?

Use gnews:"your brand" plus gnews:"your brand" when:1d, set Published after to 24 hours, deduplicate by Normalised title, and schedule the Actor hourly or daily. Diff on articleUid downstream to get only new stories.

How much does it cost to scrape news feeds?

Pricing is pay per event with exactly one event:

EventWhat triggers itBilled
ArticleOne article written to the datasetper article
Actor startOnce per run, platform feeper run

Full-text extraction is included in the per-article price. Turning it on costs you run time, not money. Articles removed by filters or de-duplication are not billed, because they are never written. A source that fails is not billed.

500 articles is 500 billed events — that is the whole calculation. Compute, bandwidth and storage are included; there is no separate platform-usage charge on top.

Set Max cost per run in the run options for a hard ceiling. Higher Apify plans get progressively lower per-article pricing through Apify Store tier discounts.

Evaluating? Set Max articles for the whole run to 20 with one source.

Example input

{
"sources": [
"gnews:\"artificial intelligence\" when:2d",
"gnews-topic:TECHNOLOGY",
"https://techcrunch.com/feed/",
"blog.cloudflare.com"
],
"googleNewsLanguage": "en-US",
"googleNewsCountry": "US",
"maxItemsPerSource": 50,
"maxItemsTotal": 300,
"publishedAfter": "48 hours",
"keywords": ["funding", "launch"],
"excludeKeywords": ["sponsored"],
"deduplicateBy": "title",
"includeArticleText": true,
"articleTextMaxChars": 12000
}

A Google News output row

Note urlResolution and the recovered publisher:

{
"ok": true,
"articleUid": "b72e1f904ac35d10",
"sourceLabel": "gnews:web scraping",
"sourceType": "google-news",
"title": "She Code Africa and Apify announce hackathon winners - TechCabal",
"titleWithoutPublisher": "She Code Africa and Apify announce hackathon winners",
"url": "https://news.google.com/rss/articles/CBMipgFBVV95cUxNOFlQWk1F",
"urlResolution": "google-news-redirect",
"publisherName": "TechCabal",
"publisherDomain": null,
"publishedAt": "2026-10-01T09:11:00.000Z",
"articleTextStatus": "skippedGoogleNewsRedirect"
}

Use cases

  • Brand and competitor monitoring — Google News searches plus de-duplication by title gives one row per story instead of twenty.
  • Deal and funding signals — gnews:"series A" when:7d on a schedule, filtered by your sector keywords.
  • RAG and LLM ingestion — contentText and articleText are clean plain text with stable articleUid and contentHash keys for incremental indexing.
  • Newsroom and research dashboards — mix twenty publisher feeds and five Google News topics in one dataset.
  • Podcast catalogues — podcast RSS feeds expose audio in enclosures with type and byte length.
  • Content gap analysis — pull competitor blog feeds by domain and compare categories and publishing cadence.
ActorUse it for
Substack Newsletter Archive ScraperNewsletter archives with engagement metrics, which RSS does not expose
ATS Job ScraperHiring signals for the companies appearing in your news feed
SEC EDGAR APIThe filings behind the financial headlines

Frequently asked questions

Why is the Google News URL not the publisher's URL?

Because Google does not let it be. A news.google.com/rss/articles/... link responds with a consent redirect and then a JavaScript application; there is no Location header and no readable link in the HTML. Resolving it requires a real browser, which would multiply the cost of every run. This Actor is explicit about it: urlResolution is set to google-news-redirect, publisherName is recovered from the title, and full-text extraction reports skippedGoogleNewsRedirect instead of silently returning nothing.

If you need canonical publisher URLs and article bodies, add the publishers' own feeds as sources. That path returns urlResolution: "direct" and works with full-text extraction.

Which feed formats are supported?

RSS 2.0, RSS 1.0 / RDF, Atom 1.0 and JSON Feed 1.x. Namespaces are stripped, so dc:creator, content:encoded and media:* extensions are picked up as well.

Do I need full-text extraction?

Often not. Run once without it and look at contentChars — many publishers put the whole article in the feed. Only enable it when contentChars is small compared with the real article.

How many articles does Google News return?

Up to 100 per search, around 70 per topic and around 40 for top stories. That is Google's limit. Use several narrower searches, or add when:1d and schedule the Actor more often, to cover more ground.

Can I use Google News search operators?

Yes. when:7d, site:example.com, quoted phrases and -exclusions all pass straight through. gnews:"series A" when:7d -crypto is a valid source.

Feeds exist to be consumed by software; that is their entire purpose. This Actor sends ordinary HTTP requests and identifies itself. You remain responsible for how you use and republish the content, including the publisher's copyright and licensing terms.

Can I run this on a schedule?

Yes, and it is the intended pattern. Set Published after to a short window, pick a de-duplication strategy, and diff on articleUid.

Something looks wrong — how do I report it?

Open the Issues tab on this Actor page with the source entry and the field you expected. Feed-format edge cases are treated as bugs and get fixed in the normaliser.


Runnable examples and how-to guides for these Actors: github.com/batow133/axiora-apify-actors