Google News Scraper | Keywords, Topics, Real URLs & Text
Pricing
from $2.24 / 1,000 articles
Google News Scraper | Keywords, Topics, Real URLs & Text
Scrape Google News by keyword, topic, location or RSS URL in any language/country. Clean title, publisher, REAL article URL (redirect decoded), publish time, age, related coverage, optional full text. Time filters, dedupe, monitor mode for new articles only. $3.20 per 1,000.
Pricing
from $2.24 / 1,000 articles
Rating
0.0
(0)
Developer
Mr Zack
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
3
Monthly active users
7 days ago
Last modified
Categories
Share
Google News Scraper — Keywords, Topics, Real URLs, Full Text & Monitoring
Pull news from Google News by keyword, topic (Business, Technology, Sports…), location or any Google News RSS URL — in any language and country edition — and get clean JSON: title, publisher, real publisher URL (Google's redirect links are decoded, not left as news.google.com/...), publish time, age in hours, snippet or related coverage, and optionally the full article text. Time filters (last hour → 30 days), cross-feed deduplication, and a monitor mode that only returns articles you haven't seen in earlier runs. No browser, no API key — $3.20 per 1,000 articles.
Who uses this
- Newsletters, digests & AI summarizers —
queries+timeRange: 24h+includeArticleTexton a daily schedule → ready-to-summarize corpus with source URLs. - Brand / competitor / crisis monitoring —
onlyNewon an hourly schedule: you only pay for fresh mentions. - Traders & crypto desks —
bitcoin,"federal reserve",site:reuters.com teslaevery 15 minutes;ageHourstells you what just broke. - Researchers & local media — topic and location feeds in 40+ editions (
language: id, country: ID→ Indonesian news;de/DE,pt/BR,ja/JP…).
Input
| Field | Description |
|---|---|
queries | Keywords; Google operators work ("exact", -word, site:, intitle:). |
topics | top, world, nation, business, technology, entertainment, science, sports, health. |
locations | City / region / country names → local news feeds. |
feedUrls | Any Google News RSS URL (or a generic RSS feed). |
timeRange | any / 1h / 6h / 24h / 7d / 30d. |
publishedAfter, publishedBefore | Exact date window (YYYY-MM-DD, UTC). Google after:/before: on searches + post-fetch filter on every other feed; overrides timeRange. |
language, country | Edition, e.g. en+US, id+ID, de+DE. |
maxArticlesPerFeed | Up to 100 (Google's feed limit). |
sources | Keep only these publisher domains (max 10). Searches become one feed per source (site:) → up to 100 items per source. |
excludeSources | Drop these publisher domains or source names (max 50) from every feed. |
excludeWords | Drop articles containing these words/phrases (max 20). On searches = Google -word (matches the whole article); everywhere = title + snippet filter. |
proxyConfiguration | Optional. Empty = direct, with automatic switch to Apify datacenter proxy if Google refuses the IP. |
resolveUrls | Decode Google redirect links to publisher URLs (default on, included in price). |
includeArticleText, maxTextArticles | Fetch publisher pages and extract the readable body. |
dedupe | Drop the same story appearing under several queries/topics. |
onlyNew | Monitor mode — remember seen articles per feed set; later runs push only new ones. |
includeFeedSummary | Free per-feed summary row. |
{ "queries": ["bitcoin", "\"federal reserve\" rates"], "topics": ["business"], "timeRange": "24h", "language": "en", "country": "US", "includeArticleText": true, "maxTextArticles": 50, "onlyNew": true }
Output
article
{ "type": "article", "title": "Bitcoin network used by exchanges hit by $320 million exploit. Hackers claim they're the 'good guys'","source": "CoinDesk", "sourceUrl": "https://www.coindesk.com", "url": "https://www.coindesk.com/markets/2026/09/07/bitcoin-network-used-by-exchanges-hit-by-usd320-million-exploit", "domain": "coindesk.com", "urlResolved": true,"googleUrl": "https://news.google.com/rss/articles/CBMi2wFBVV95cUxN…", "googleId": "CBMi2wFBVV95cUxN…","publishedAt": "2026-09-07T10:01:06.000Z", "ageHours": 3.9, "snippet": null, "snippetSource": null, "relatedCoverage": [],"feedKind": "query", "feed": "bitcoin", "position": 2, "language": "en", "country": "US","articleText": "…full body…", "wordCount": 1549, "scrapedAt": "2026-09-07T13:57:43Z" }
Top-stories items include relatedCoverage[] { title, source, googleUrl, googleId } (the other outlets covering the same story).
feed-summary (free): articles, sources, topSources[], oldest, newest, urlsResolved, withText, fetched, pushed.
Feeds that fail (blocked, malformed) are pushed as type: "failed" and never charged.
Pricing (pay per event)
| Event | Price |
|---|---|
| Article (with resolved URL) | $0.0032 |
| Article with full text | $0.006 |
| Actor start | $0.001 |
1,000 articles ≈ $3.20; with full text ≈ $6. Text is billed only when a readable body was actually extracted — paywalled or JavaScript-only pages stay at the Article price and carry articleTextError.
Notes
- Google News feeds return at most 100 items per query/topic; split broad topics into several queries for more.
publishedAtcomes from Google's feed;ageHoursis computed at scrape time.- Publisher-URL decoding uses Google's own redirect mechanism; if Google changes it, rows fall back to
googleUrlwithurlResolved: false(still charged as Article, still clickable). - Related on this profile: YouTube Channel & Shorts Scraper, Hyperliquid Smart Money Tracker, Amazon / TikTok Shop / AliExpress product scrapers.
Changelog
-
0.1.9 (25 Sep 2026) — Clean article text. Some publishers double-encode HTML entities in their body copy, which surfaced as
S&PinsidearticleText; the text is now decoded fully. No new fields, no price change. -
0.1.8 (24 Sep 2026) — Reliability: A mistyped feed URL or a search with no articles now ends the run successfully with a plain explanation instead of a false "Google News blocked us" failure, and a short Google block gets a patient retry. Long runs stop just before the timeout and keep every article collected so far. No price or field change.
-
0.1.7 (23 Sep 2026) — Second pass for publisher URLs. Google occasionally returns no publisher link for individual articles on the first attempt (20 of 93 on one run, with no block in sight). Those rows are now retried once before the run ends, so more rows carry a real
url. No new fields, no price change. -
0.1.6 (22 Sep 2026) — No more unexplained empty
articleText. WithincludeArticleText: truethe text is fetched for the firstmaxTextArticlesrows (default 50); rows beyond that cap used to come back with an emptyarticleTextand no reason, which looked like a broken column on runs with 90+ articles. Those rows now carryarticleTextError: "not fetched: maxTextArticles cap (50) reached — raise maxTextArticles …", unresolved Google URLs say so too, and the run's status message reportsN/M with article text (K skipped by cap). Rows without text are still billed at the cheaperarticleprice, neverarticle-text. -
0.1.5 (21 Sep 2026) —
snippetis no longer empty on the full-text path. Google's RSS has not carried a snippet for keyword/topic feeds for years (the description is just title + publisher), sosnippetwas null on every row unless the feed was Top Stories. WithincludeArticleText: truethe publisher page is fetched anyway, so the row now takes the page's ownog:description/ meta description assnippet— also on paywalled pages where the body cannot be read. New fieldsnippetSource(rss,page-metaor null). Nothing else changed. -
0.1.4 (19 Sep 2026) — Typo guard on input. Apify accepts input fields an Actor does not know without complaining, so a misspelled option (e.g.
maxIteminstead ofmaxItems) used to produce a successful run with the setting silently inactive. The run log now warns for every unknown field and suggests the closest real one. -
0.1.3 (18 Sep 2026) — Fix:
sourceswithout a query returned 0 articles. A run with onlysources(e.g.["reuters.com"]) fell back to the top-stories demo feed and then filtered it down to nothing. It now runs onesite:feed per source (up to 100 items each), so "everything Reuters published today" is a one-field input. Behaviour with queries/topics unchanged. -
0.1.2 (16 Sep 2026) —
includeArticleText: publisher pages that answer 401/403/429 to the datacenter IP are retried once through the proxy (measured: 21 of 40 text failures in a real run were IP blocks, not paywalls). Failed extractions are still never charged. -
0.1.1 (16 Sep 2026) — Reliability: V8 heap capped at 70 % of run memory so garbage is collected before the container limit (prevents out-of-memory kills on big pages). No output change.
-
0.1.0 (14 Sep 2026) — Date window (
publishedAfter/publishedBefore), source allow/deny lists (sources,excludeSources),excludeWords,proxyConfiguration+ automatic proxy fallback when Google News blocks the run's IP (previously such a run failed). Filtered rows are dropped before billing. New summary fieldsskippedFiltered,proxyUsed,proxySwitches. Existing fields and defaults unchanged. -
0.0.4 (10 Sep 2026) — Maintenance: billing safety (
chargeSafely— a failed charge can no longer kill a run that already has data), test suite wiring, README. -
0.0.1 (7 Sep 2026) — Launch.
Found this useful? A review helps more than you'd think
If this Actor saved you time, a short review on the Reviews tab of this Store page takes 30 seconds and is the only signal other buyers have before they spend anything. If something is broken instead, open a ticket on the Issues tab — parser bugs and field requests get fixed.
Related Actors from the same account