Google News Scraper | Keywords, Topics, Real URLs & Text avatar

Google News Scraper | Keywords, Topics, Real URLs & Text

Pricing

from $2.24 / 1,000 articles

Go to Apify Store
Google News Scraper | Keywords, Topics, Real URLs & Text

Google News Scraper | Keywords, Topics, Real URLs & Text

Scrape Google News by keyword, topic, location or RSS URL in any language/country. Clean title, publisher, REAL article URL (redirect decoded), publish time, age, related coverage, optional full text. Time filters, dedupe, monitor mode for new articles only. $3.20 per 1,000.

Pricing

from $2.24 / 1,000 articles

Rating

0.0

(0)

Developer

Mr Zack

Mr Zack

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

3

Monthly active users

7 days ago

Last modified

Share

Google News Scraper — Keywords, Topics, Real URLs, Full Text & Monitoring

Pull news from Google News by keyword, topic (Business, Technology, Sports…), location or any Google News RSS URL — in any language and country edition — and get clean JSON: title, publisher, real publisher URL (Google's redirect links are decoded, not left as news.google.com/...), publish time, age in hours, snippet or related coverage, and optionally the full article text. Time filters (last hour → 30 days), cross-feed deduplication, and a monitor mode that only returns articles you haven't seen in earlier runs. No browser, no API key — $3.20 per 1,000 articles.

Who uses this

  • Newsletters, digests & AI summarizers — queries + timeRange: 24h + includeArticleText on a daily schedule → ready-to-summarize corpus with source URLs.
  • Brand / competitor / crisis monitoring — onlyNew on an hourly schedule: you only pay for fresh mentions.
  • Traders & crypto desks — bitcoin, "federal reserve", site:reuters.com tesla every 15 minutes; ageHours tells you what just broke.
  • Researchers & local media — topic and location feeds in 40+ editions (language: id, country: ID → Indonesian news; de/DE, pt/BR, ja/JP…).

Input

FieldDescription
queriesKeywords; Google operators work ("exact", -word, site:, intitle:).
topicstop, world, nation, business, technology, entertainment, science, sports, health.
locationsCity / region / country names → local news feeds.
feedUrlsAny Google News RSS URL (or a generic RSS feed).
timeRangeany / 1h / 6h / 24h / 7d / 30d.
publishedAfter, publishedBeforeExact date window (YYYY-MM-DD, UTC). Google after:/before: on searches + post-fetch filter on every other feed; overrides timeRange.
language, countryEdition, e.g. en+US, id+ID, de+DE.
maxArticlesPerFeedUp to 100 (Google's feed limit).
sourcesKeep only these publisher domains (max 10). Searches become one feed per source (site:) → up to 100 items per source.
excludeSourcesDrop these publisher domains or source names (max 50) from every feed.
excludeWordsDrop articles containing these words/phrases (max 20). On searches = Google -word (matches the whole article); everywhere = title + snippet filter.
proxyConfigurationOptional. Empty = direct, with automatic switch to Apify datacenter proxy if Google refuses the IP.
resolveUrlsDecode Google redirect links to publisher URLs (default on, included in price).
includeArticleText, maxTextArticlesFetch publisher pages and extract the readable body.
dedupeDrop the same story appearing under several queries/topics.
onlyNewMonitor mode — remember seen articles per feed set; later runs push only new ones.
includeFeedSummaryFree per-feed summary row.
{ "queries": ["bitcoin", "\"federal reserve\" rates"], "topics": ["business"], "timeRange": "24h", "language": "en", "country": "US", "includeArticleText": true, "maxTextArticles": 50, "onlyNew": true }

Output

article

{ "type": "article", "title": "Bitcoin network used by exchanges hit by $320 million exploit. Hackers claim they're the 'good guys'",
"source": "CoinDesk", "sourceUrl": "https://www.coindesk.com", "url": "https://www.coindesk.com/markets/2026/09/07/bitcoin-network-used-by-exchanges-hit-by-usd320-million-exploit", "domain": "coindesk.com", "urlResolved": true,
"googleUrl": "https://news.google.com/rss/articles/CBMi2wFBVV95cUxN…", "googleId": "CBMi2wFBVV95cUxN…",
"publishedAt": "2026-09-07T10:01:06.000Z", "ageHours": 3.9, "snippet": null, "snippetSource": null, "relatedCoverage": [],
"feedKind": "query", "feed": "bitcoin", "position": 2, "language": "en", "country": "US",
"articleText": "…full body…", "wordCount": 1549, "scrapedAt": "2026-09-07T13:57:43Z" }

Top-stories items include relatedCoverage[] { title, source, googleUrl, googleId } (the other outlets covering the same story).

feed-summary (free): articles, sources, topSources[], oldest, newest, urlsResolved, withText, fetched, pushed.

Feeds that fail (blocked, malformed) are pushed as type: "failed" and never charged.

Pricing (pay per event)

EventPrice
Article (with resolved URL)$0.0032
Article with full text$0.006
Actor start$0.001

1,000 articles ≈ $3.20; with full text ≈ $6. Text is billed only when a readable body was actually extracted — paywalled or JavaScript-only pages stay at the Article price and carry articleTextError.

Notes

  • Google News feeds return at most 100 items per query/topic; split broad topics into several queries for more.
  • publishedAt comes from Google's feed; ageHours is computed at scrape time.
  • Publisher-URL decoding uses Google's own redirect mechanism; if Google changes it, rows fall back to googleUrl with urlResolved: false (still charged as Article, still clickable).
  • Related on this profile: YouTube Channel & Shorts Scraper, Hyperliquid Smart Money Tracker, Amazon / TikTok Shop / AliExpress product scrapers.

Changelog

  • 0.1.9 (25 Sep 2026) — Clean article text. Some publishers double-encode HTML entities in their body copy, which surfaced as S&P inside articleText; the text is now decoded fully. No new fields, no price change.

  • 0.1.8 (24 Sep 2026) — Reliability: A mistyped feed URL or a search with no articles now ends the run successfully with a plain explanation instead of a false "Google News blocked us" failure, and a short Google block gets a patient retry. Long runs stop just before the timeout and keep every article collected so far. No price or field change.

  • 0.1.7 (23 Sep 2026) — Second pass for publisher URLs. Google occasionally returns no publisher link for individual articles on the first attempt (20 of 93 on one run, with no block in sight). Those rows are now retried once before the run ends, so more rows carry a real url. No new fields, no price change.

  • 0.1.6 (22 Sep 2026) — No more unexplained empty articleText. With includeArticleText: true the text is fetched for the first maxTextArticles rows (default 50); rows beyond that cap used to come back with an empty articleText and no reason, which looked like a broken column on runs with 90+ articles. Those rows now carry articleTextError: "not fetched: maxTextArticles cap (50) reached — raise maxTextArticles …", unresolved Google URLs say so too, and the run's status message reports N/M with article text (K skipped by cap). Rows without text are still billed at the cheaper article price, never article-text.

  • 0.1.5 (21 Sep 2026) — snippet is no longer empty on the full-text path. Google's RSS has not carried a snippet for keyword/topic feeds for years (the description is just title + publisher), so snippet was null on every row unless the feed was Top Stories. With includeArticleText: true the publisher page is fetched anyway, so the row now takes the page's own og:description / meta description as snippet — also on paywalled pages where the body cannot be read. New field snippetSource (rss, page-meta or null). Nothing else changed.

  • 0.1.4 (19 Sep 2026) — Typo guard on input. Apify accepts input fields an Actor does not know without complaining, so a misspelled option (e.g. maxItem instead of maxItems) used to produce a successful run with the setting silently inactive. The run log now warns for every unknown field and suggests the closest real one.

  • 0.1.3 (18 Sep 2026) — Fix: sources without a query returned 0 articles. A run with only sources (e.g. ["reuters.com"]) fell back to the top-stories demo feed and then filtered it down to nothing. It now runs one site: feed per source (up to 100 items each), so "everything Reuters published today" is a one-field input. Behaviour with queries/topics unchanged.

  • 0.1.2 (16 Sep 2026) — includeArticleText: publisher pages that answer 401/403/429 to the datacenter IP are retried once through the proxy (measured: 21 of 40 text failures in a real run were IP blocks, not paywalls). Failed extractions are still never charged.

  • 0.1.1 (16 Sep 2026) — Reliability: V8 heap capped at 70 % of run memory so garbage is collected before the container limit (prevents out-of-memory kills on big pages). No output change.

  • 0.1.0 (14 Sep 2026) — Date window (publishedAfter/publishedBefore), source allow/deny lists (sources, excludeSources), excludeWords, proxyConfiguration + automatic proxy fallback when Google News blocks the run's IP (previously such a run failed). Filtered rows are dropped before billing. New summary fields skippedFiltered, proxyUsed, proxySwitches. Existing fields and defaults unchanged.

  • 0.0.4 (10 Sep 2026) — Maintenance: billing safety (chargeSafely — a failed charge can no longer kill a run that already has data), test suite wiring, README.

  • 0.0.1 (7 Sep 2026) — Launch.

Found this useful? A review helps more than you'd think

If this Actor saved you time, a short review on the Reviews tab of this Store page takes 30 seconds and is the only signal other buyers have before they spend anything. If something is broken instead, open a ticket on the Issues tab — parser bugs and field requests get fixed.

Related Actors from the same account