Google News Scraper — Headlines by Keyword, Topic, Country avatar

Google News Scraper — Headlines by Keyword, Topic, Country

Pricing

from $1.05 / 1,000 result items

Go to Apify Store
Google News Scraper — Headlines by Keyword, Topic, Country

Google News Scraper — Headlines by Keyword, Topic, Country

Google News for AI agents and monitoring: keyword search with site:/when: operators, topic feeds (World, Business, Technology, Science, Health...), local news by city, top stories in 85 verified country/language editions. Title, source, time, Google link, optional real publisher URL.

Pricing

from $1.05 / 1,000 result items

Rating

0.0

(0)

Developer

Samat Makatov

Samat Makatov

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Pull Google News headlines without an API key, browser or proxy: keyword search (with Google's site:, when:, quotes and OR operators), section feeds (World, Business, Technology, Science, Health…), local news for any city, or the top stories of one of 86 verified country/language editions. Every row carries title, publisher, publisher site, publish time and the Google link. Built for monitoring, research and lead-generation agents that need cheap, fresh news at scale.

Use cases

  • Brand / competitor monitoring — run queries: ["Kaspi", "Halyk Bank"] every hour with sinceHours: 2, push new rows to Slack.
  • Sales trigger alerts — "<company> acquisition OR funding OR layoffs" restricted to sites: ["reuters.com","bloomberg.com"].
  • Local market watch — locations: ["Almaty","Astana"] for city-level news in the local language.
  • Sector digest — topic: "TECHNOLOGY" per edition (DE, JP, BR…) to compare what each market talks about.
  • Crisis / reputation tracking — excludeKeywords to drop noise, dedupe to collapse syndicated copies, sourceUrl to group coverage by publisher.
  • Dataset building — 100 headlines per query × N queries, stable ids, ISO timestamps.

Input

FieldTypeDefaultNotes
modeselectautoauto, search, topic, geo, topStories. Auto = topic > locations > queries/sites > top stories.
queriesstring[]—One feed per query. Google operators allowed: "exact", OR, -word, site:x.com, intitle:, when:3d.
topicselect""WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SCIENCE, SPORTS, HEALTH.
locationsstring[]—City/region names for Google's local feed, e.g. Almaty, Texas.
editionselect""Verified COUNTRY:lang edition (see Reference). Overrides country/lang.
countrystringUSISO-3166 alpha-2 (gl).
langstringenLanguage code (hl): en, de, ru, es-419, zh-Hans…
sinceHoursint168Client-side time filter; in search mode also sent as when: so the 100-item window is fresh. 0 = off.
useWhenOperatorbooltrueSet false to keep Google's default relevance window and filter only client-side.
sitesstring[]—Search mode: (site:a OR site:b). Works with empty queries = latest from those publishers.
excludeSitesstring[]—Search mode: -site:.
includeKeywordsstring[]—Keep rows whose title+snippet contain any (case-insensitive).
excludeKeywordsstring[]—Drop rows containing any.
limitint50Per feed, max 100 (Google's RSS cap).
maxItemsint500Hard cap per run.
sortselectnewestnewest, oldest, feed (Google order).
dedupebooltrueSkip repeated stories across feeds (normalized title / Google id).
decodeUrlsboolfalseResolve news.google.com/rss/articles/… to the publisher URL (articleUrl), 2 requests per article — only where news.google.com/robots.txt allows it. As of 2026-10-02 it disallows both paths the resolver needs (/rss/articles/ and /_/), so the actor does not call them and articleUrl stays null except for the rare links that embed the publisher URL. See Why is articleUrl null? below.
maxDecodeint20Most articles to resolve over the network per run (max 200). Every try counts, resolved or not.
includeSnippetbooltrueSnippet is whatever Google adds beyond the headline (often nothing).
fieldsstring[]—Keep only these output fields, in your order. id is always kept. Case, snake_case and the RSS names link, pubDate, guid, description, publisher are understood; an unknown name is left out and named in the status.

Reference

Verified editions (edition = COUNTRY:lang)

86 verified editions, checked on 2026-09-13: requesting each of these returns the feed of that edition without a redirect. edition is a select list: Apify refuses any other value before the run starts. For other countries use country + lang (free text).

RegionEditions
EnglishUS:en GB:en IE:en CA:en AU:en NZ:en IN:en PK:en SG:en MY:en PH:en IL:en ZA:en NG:en KE:en GH:en TZ:en UG:en ZW:en BW:en NA:en ET:en
EuropeDE:de AT:de CH:de CH:fr FR:fr BE:fr BE:nl NL:nl IT:it ES:es PT:pt-150 PL:pl CZ:cs SK:sk HU:hu RO:ro BG:bg RS:sr SI:sl LT:lt LV:lv EE:et FI:fi SE:sv NO:no GR:el TR:tr UA:uk UA:ru RU:ru
AmericasUS:es-419 CA:fr MX:es-419 AR:es-419 CL:es-419 CO:es-419 PE:es-419 VE:es-419 CU:es-419 BR:pt-419
AsiaIN:hi IN:bn IN:ta IN:te IN:ml IN:mr IN:gu IN:pa BD:bn ID:id TH:th VN:vi JP:ja KR:ko CN:zh-Hans TW:zh-Hant HK:zh-Hant
MENA / Africa (other)IL:he SA:ar AE:ar EG:ar LB:ar MA:fr SN:fr

Not an edition of its own (Google redirects — the actor warns and stores effectiveEdition): KZ:*, BY:*, UZ:*, AZ:*, AM:* → RU:ru; DK:da → NO:no; HR:hr, LK, NP, GE, IR, most of Central America → US:en; Gulf/Maghreb Arabic → EG:ar; French Africa → FR:fr. For Kazakhstan use edition: "RU:ru" + locations: ["Алматы"] or queries with Kazakh terms.

Search operators (inside queries)

OperatorExampleMeaning
quotes"central bank"exact phrase
ORKaspi OR Halykeither
-tesla -stockexclude
site:site:reuters.compublisher (or use sites)
when:when:12h, when:3d, when:2mtime window (auto-added from sinceHours)
intitle:intitle:layoffsword must be in headline

Examples

Hourly brand monitor (fresh only)

{ "queries": ["Kaspi", "Halyk Bank", "Freedom Finance"], "edition": "RU:ru", "sinceHours": 2, "limit": 100, "dedupe": true }

Deal-flow alerts from tier-1 press

{ "queries": ["fintech acquisition", "fintech funding round"], "sites": ["reuters.com", "bloomberg.com", "ft.com", "techcrunch.com"], "sinceHours": 24 }

Local news for two cities

{ "mode": "geo", "locations": ["Almaty", "Astana"], "edition": "RU:ru", "sinceHours": 48, "limit": 40 }

Compare tech agendas across markets (run once per edition)

{ "topic": "TECHNOLOGY", "edition": "JP:ja", "limit": 30, "sinceHours": 24 }

Latest from specific publishers, no keyword

{ "sites": ["forbes.kz", "kursiv.media"], "country": "KZ", "lang": "ru", "sinceHours": 72, "limit": 50 }

Output

One row per article:

{
"id": "CBMinwFBVV95cUxQcmhHODdl…",
"feed": "Kaspi",
"feedType": "search",
"title": "Фондовый рынок Казахстана оказался под давлением внешних факторов",
"source": "Forbes.kz",
"sourceUrl": "https://forbes.kz",
"url": "https://news.google.com/rss/articles/CBMinwFBVV95cUxQcmhHODdl…?oc=5",
"articleUrl": null,
"publishedAt": "2026-09-11T15:30:24.000Z",
"snippet": null,
"edition": "KZ:ru",
"effectiveEdition": "RU:ru",
"country": "KZ",
"lang": "ru",
"feedUrl": "https://news.google.com/rss/search?q=Kaspi%20(site%3Areuters.com…)&hl=ru-KZ&gl=KZ&ceid=KZ%3Aru",
"fetchedAt": "2026-09-12T23:36:10.112Z"
}
FieldDescription
idGoogle's article id (stable across runs; use for dedupe in your DB).
feed, feedTypeWhich query / topic / location produced the row; search, topic, geo, top.
title, source, sourceUrlHeadline without the " - Publisher" suffix; publisher name and site.
urlGoogle redirect link (always present).
articleUrlPublisher URL when decodeUrls is on, robots.txt allows Google's resolver and the budget lasts; else null (as of 2026-10-02: null unless the Google link embeds the URL).
publishedAtISO-8601 UTC.
snippetExtra text Google adds beyond the headline, or null.
edition, effectiveEditionRequested vs. actually served edition.
feedUrl, fetchedAtSource feed and fetch time.

The key-value store record SUMMARY holds { items, feeds, feedsRead, feedsNotRead, edition, effectiveEdition, decodedUrls, urlResolving, errors[], notes[], status, stoppedBy }. decodedUrls counts resolved publisher URLs only; urlResolving (with decodeUrls) has attempted, resolved, failed, notTried, stoppedBy and what robots.txt said about the resolver paths (robotsTxt) and this run's feed paths (robotsTxtFeeds).

With fields, each row holds id plus the fields you listed, in your order, e.g. "fields": ["publishedAt", "source", "title"] → { "id": …, "publishedAt": …, "source": …, "title": … }.

Use it from code / agents

curl -X POST "https://api.apify.com/v2/acts/yadroo~google-news-search/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"queries":["Kaspi"],"edition":"RU:ru","sinceHours":24}'
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('yadroo/google-news-search').call({ queries: ['x402 protocol'], sinceHours: 72 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("yadroo/google-news-search").call(run_input={"topic": "BUSINESS", "edition": "DE:de", "limit": 30})
items = client.dataset(run["defaultDatasetId"]).list_items().items

MCP: add https://mcp.apify.com to Claude / Cursor / any MCP client and call the yadroo/google-news-search tool with the same JSON input.

Pricing

Pay per event: $0.001 per run start + $0.0015 per article. Store discounts: Bronze −10 %, Silver −20 %, Gold and above −30 % on the article price; the start event is the same on every plan; platform usage is included. Typical runs: 3 queries × 50 articles = 150 rows ≈ $0.23; an hourly monitor keeping only the last 2 h usually returns 0–20 rows ≈ $0.001–0.03. Resolving URLs adds time, not price.

Your Maximum cost per run is respected: the run saves only the articles it pays for and ends with "Stopped at your spending limit: N rows delivered". A run close to its timeout stops starting new feeds, saves what it has read and ends with "Stopped before the run timeout".

Limits & FAQ

  • How fresh? Google's RSS lags the web UI by a few minutes. sinceHours + when: keep results within your window.
  • How many? Max 100 items per feed (Google's cap); topic/geo/top feeds return 30–70. Use several narrower queries for more.
  • Rate limits. The actor pauses ~0.7 s between feeds and ~0.4 s between URL resolutions. A feed answered with HTTP 429/5xx is tried 3 times with backoff (honouring Retry-After), then recorded in SUMMARY.errors and named in the status; after 3 refused or throttled feeds in a row the run stops asking Google and says how many feeds it did not read. If every feed it tried failed, the run fails with the reason.
  • Why is articleUrl null? As of 2026-10-02 news.google.com/robots.txt disallows /rss/articles/ and /_/, the two paths Google's URL resolver needs, so the actor does not call them (it reads robots.txt on every run that would, and resolves again if Google allows it). Other reasons: decodeUrls is off; the maxDecode budget ran out; news.google.com/robots.txt disallows Google's resolver paths (the actor checks it first); Google refused a request (HTTP 403/429 — resolving then stops for the rest of the run); 3 resolutions in a row failed (Google may have changed its resolver); or the run was about to time out. Articles are saved either way, and the status names the reason. The Google link in url still works.
  • Why do I get Russian news for KZ? Google has no Kazakhstan edition; it serves RU:ru. Use locations or Kazakh/Russian keywords.
  • Invalid topic / empty geo? Google answers with an HTML page; the actor reports a clear error for that feed instead of pushing garbage.
  • Terms. Google's feed header states it is provided for personal, non-commercial feed readers. You are responsible for how you use the data; keep request volumes modest.
  • Roadmap. Publisher-level aggregation (count per source), optional full-text extraction via a sibling actor.

Made by Yadroo · Sibling actors: rss-to-json · crypto-news · hackernews-search · youtube-channel-feed · wikipedia-search