# Google News Scraper — Headlines by Keyword, Topic, Country (`yadroo/google-news-search`) Actor

Google News for AI agents and monitoring: keyword search with site:/when: operators, topic feeds (World, Business, Technology, Science, Health...), local news by city, top stories in 85 verified country/language editions. Title, source, time, Google link, optional real publisher URL.

- **URL**: https://apify.com/yadroo/google-news-search.md
- **Developed by:** [Samat Makatov](https://apify.com/yadroo) (community)
- **Categories:** News, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.05 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Pull Google News headlines without an API key, browser or proxy: keyword search (with Google's `site:`, `when:`, quotes and `OR` operators), section feeds (World, Business, Technology, Science, Health…), local news for any city, or the top stories of one of **86 verified country/language editions**. Every row carries title, publisher, publisher site, publish time and the Google link. Built for monitoring, research and lead-generation agents that need cheap, fresh news at scale.

### Use cases

- **Brand / competitor monitoring** — run `queries: ["Kaspi", "Halyk Bank"]` every hour with `sinceHours: 2`, push new rows to Slack.
- **Sales trigger alerts** — `"<company> acquisition OR funding OR layoffs"` restricted to `sites: ["reuters.com","bloomberg.com"]`.
- **Local market watch** — `locations: ["Almaty","Astana"]` for city-level news in the local language.
- **Sector digest** — `topic: "TECHNOLOGY"` per edition (DE, JP, BR…) to compare what each market talks about.
- **Crisis / reputation tracking** — `excludeKeywords` to drop noise, `dedupe` to collapse syndicated copies, `sourceUrl` to group coverage by publisher.
- **Dataset building** — 100 headlines per query × N queries, stable ids, ISO timestamps.

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `mode` | select | `auto` | `auto`, `search`, `topic`, `geo`, `topStories`. Auto = topic > locations > queries/sites > top stories. |
| `queries` | string\[] | — | One feed per query. Google operators allowed: `"exact"`, `OR`, `-word`, `site:x.com`, `intitle:`, `when:3d`. |
| `topic` | select | `""` | `WORLD`, `NATION`, `BUSINESS`, `TECHNOLOGY`, `ENTERTAINMENT`, `SCIENCE`, `SPORTS`, `HEALTH`. |
| `locations` | string\[] | — | City/region names for Google's local feed, e.g. `Almaty`, `Texas`. |
| `edition` | select | `""` | Verified `COUNTRY:lang` edition (see Reference). Overrides `country`/`lang`. |
| `country` | string | `US` | ISO-3166 alpha-2 (gl). |
| `lang` | string | `en` | Language code (hl): `en`, `de`, `ru`, `es-419`, `zh-Hans`… |
| `sinceHours` | int | `168` | Client-side time filter; in search mode also sent as `when:` so the 100-item window is fresh. `0` = off. |
| `useWhenOperator` | bool | `true` | Set `false` to keep Google's default relevance window and filter only client-side. |
| `sites` | string\[] | — | Search mode: `(site:a OR site:b)`. Works with empty queries = latest from those publishers. |
| `excludeSites` | string\[] | — | Search mode: `-site:`. |
| `includeKeywords` | string\[] | — | Keep rows whose title+snippet contain any (case-insensitive). |
| `excludeKeywords` | string\[] | — | Drop rows containing any. |
| `limit` | int | `50` | Per feed, max 100 (Google's RSS cap). |
| `maxItems` | int | `500` | Hard cap per run. |
| `sort` | select | `newest` | `newest`, `oldest`, `feed` (Google order). |
| `dedupe` | bool | `true` | Skip repeated stories across feeds (normalized title / Google id). |
| `decodeUrls` | bool | `false` | Resolve `news.google.com/rss/articles/…` to the publisher URL (`articleUrl`), 2 requests per article — only where news.google.com/robots.txt allows it. **As of 2026-10-02 it disallows both paths the resolver needs (`/rss/articles/` and `/_/`), so the actor does not call them and `articleUrl` stays `null`** except for the rare links that embed the publisher URL. See *Why is `articleUrl` null?* below. |
| `maxDecode` | int | `20` | Most articles to resolve over the network per run (max 200). Every try counts, resolved or not. |
| `includeSnippet` | bool | `true` | Snippet is whatever Google adds beyond the headline (often nothing). |
| `fields` | string\[] | — | Keep only these output fields, in your order. `id` is always kept. Case, snake\_case and the RSS names `link`, `pubDate`, `guid`, `description`, `publisher` are understood; an unknown name is left out and named in the status. |

### Reference

#### Verified editions (`edition` = `COUNTRY:lang`)

86 verified editions, checked on 2026-09-13: requesting each of these returns the feed of that edition without a redirect. `edition` is a select list: Apify refuses any other value before the run starts. For other countries use `country` + `lang` (free text).

| Region | Editions |
|---|---|
| English | `US:en` `GB:en` `IE:en` `CA:en` `AU:en` `NZ:en` `IN:en` `PK:en` `SG:en` `MY:en` `PH:en` `IL:en` `ZA:en` `NG:en` `KE:en` `GH:en` `TZ:en` `UG:en` `ZW:en` `BW:en` `NA:en` `ET:en` |
| Europe | `DE:de` `AT:de` `CH:de` `CH:fr` `FR:fr` `BE:fr` `BE:nl` `NL:nl` `IT:it` `ES:es` `PT:pt-150` `PL:pl` `CZ:cs` `SK:sk` `HU:hu` `RO:ro` `BG:bg` `RS:sr` `SI:sl` `LT:lt` `LV:lv` `EE:et` `FI:fi` `SE:sv` `NO:no` `GR:el` `TR:tr` `UA:uk` `UA:ru` `RU:ru` |
| Americas | `US:es-419` `CA:fr` `MX:es-419` `AR:es-419` `CL:es-419` `CO:es-419` `PE:es-419` `VE:es-419` `CU:es-419` `BR:pt-419` |
| Asia | `IN:hi` `IN:bn` `IN:ta` `IN:te` `IN:ml` `IN:mr` `IN:gu` `IN:pa` `BD:bn` `ID:id` `TH:th` `VN:vi` `JP:ja` `KR:ko` `CN:zh-Hans` `TW:zh-Hant` `HK:zh-Hant` |
| MENA / Africa (other) | `IL:he` `SA:ar` `AE:ar` `EG:ar` `LB:ar` `MA:fr` `SN:fr` |

Not an edition of its own (Google redirects — the actor warns and stores `effectiveEdition`): `KZ:*`, `BY:*`, `UZ:*`, `AZ:*`, `AM:*` → `RU:ru`; `DK:da` → `NO:no`; `HR:hr`, `LK`, `NP`, `GE`, `IR`, most of Central America → `US:en`; Gulf/Maghreb Arabic → `EG:ar`; French Africa → `FR:fr`. For Kazakhstan use `edition: "RU:ru"` + `locations: ["Алматы"]` or queries with Kazakh terms.

#### Search operators (inside `queries`)

| Operator | Example | Meaning |
|---|---|---|
| quotes | `"central bank"` | exact phrase |
| `OR` | `Kaspi OR Halyk` | either |
| `-` | `tesla -stock` | exclude |
| `site:` | `site:reuters.com` | publisher (or use `sites`) |
| `when:` | `when:12h`, `when:3d`, `when:2m` | time window (auto-added from `sinceHours`) |
| `intitle:` | `intitle:layoffs` | word must be in headline |

### Examples

**Hourly brand monitor (fresh only)**

```json
{ "queries": ["Kaspi", "Halyk Bank", "Freedom Finance"], "edition": "RU:ru", "sinceHours": 2, "limit": 100, "dedupe": true }
```

**Deal-flow alerts from tier-1 press**

```json
{ "queries": ["fintech acquisition", "fintech funding round"], "sites": ["reuters.com", "bloomberg.com", "ft.com", "techcrunch.com"], "sinceHours": 24 }
```

**Local news for two cities**

```json
{ "mode": "geo", "locations": ["Almaty", "Astana"], "edition": "RU:ru", "sinceHours": 48, "limit": 40 }
```

**Compare tech agendas across markets** (run once per edition)

```json
{ "topic": "TECHNOLOGY", "edition": "JP:ja", "limit": 30, "sinceHours": 24 }
```

**Latest from specific publishers, no keyword**

```json
{ "sites": ["forbes.kz", "kursiv.media"], "country": "KZ", "lang": "ru", "sinceHours": 72, "limit": 50 }
```

### Output

One row per article:

```json
{
  "id": "CBMinwFBVV95cUxQcmhHODdl…",
  "feed": "Kaspi",
  "feedType": "search",
  "title": "Фондовый рынок Казахстана оказался под давлением внешних факторов",
  "source": "Forbes.kz",
  "sourceUrl": "https://forbes.kz",
  "url": "https://news.google.com/rss/articles/CBMinwFBVV95cUxQcmhHODdl…?oc=5",
  "articleUrl": null,
  "publishedAt": "2026-09-11T15:30:24.000Z",
  "snippet": null,
  "edition": "KZ:ru",
  "effectiveEdition": "RU:ru",
  "country": "KZ",
  "lang": "ru",
  "feedUrl": "https://news.google.com/rss/search?q=Kaspi%20(site%3Areuters.com…)&hl=ru-KZ&gl=KZ&ceid=KZ%3Aru",
  "fetchedAt": "2026-09-12T23:36:10.112Z"
}
```

| Field | Description |
|---|---|
| `id` | Google's article id (stable across runs; use for dedupe in your DB). |
| `feed`, `feedType` | Which query / topic / location produced the row; `search`, `topic`, `geo`, `top`. |
| `title`, `source`, `sourceUrl` | Headline without the " - Publisher" suffix; publisher name and site. |
| `url` | Google redirect link (always present). |
| `articleUrl` | Publisher URL when `decodeUrls` is on, robots.txt allows Google's resolver and the budget lasts; else `null` (as of 2026-10-02: `null` unless the Google link embeds the URL). |
| `publishedAt` | ISO-8601 UTC. |
| `snippet` | Extra text Google adds beyond the headline, or `null`. |
| `edition`, `effectiveEdition` | Requested vs. actually served edition. |
| `feedUrl`, `fetchedAt` | Source feed and fetch time. |

The key-value store record `SUMMARY` holds `{ items, feeds, feedsRead, feedsNotRead, edition, effectiveEdition, decodedUrls, urlResolving, errors[], notes[], status, stoppedBy }`. `decodedUrls` counts resolved publisher URLs only; `urlResolving` (with `decodeUrls`) has `attempted`, `resolved`, `failed`, `notTried`, `stoppedBy` and what robots.txt said about the resolver paths (`robotsTxt`) and this run's feed paths (`robotsTxtFeeds`).

With `fields`, each row holds `id` plus the fields you listed, in your order, e.g. `"fields": ["publishedAt", "source", "title"]` → `{ "id": …, "publishedAt": …, "source": …, "title": … }`.

### Use it from code / agents

```bash
curl -X POST "https://api.apify.com/v2/acts/yadroo~google-news-search/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"queries":["Kaspi"],"edition":"RU:ru","sinceHours":24}'
```

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('yadroo/google-news-search').call({ queries: ['x402 protocol'], sinceHours: 72 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

```python
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("yadroo/google-news-search").call(run_input={"topic": "BUSINESS", "edition": "DE:de", "limit": 30})
items = client.dataset(run["defaultDatasetId"]).list_items().items
```

MCP: add `https://mcp.apify.com` to Claude / Cursor / any MCP client and call the `yadroo/google-news-search` tool with the same JSON input.

### Pricing

Pay per event: **$0.001 per run start + $0.0015 per article**. Store discounts: Bronze −10 %, Silver −20 %, Gold and above −30 % on the article price; the start event is the same on every plan; platform usage is included. Typical runs: 3 queries × 50 articles = 150 rows ≈ $0.23; an hourly monitor keeping only the last 2 h usually returns 0–20 rows ≈ $0.001–0.03. Resolving URLs adds time, not price.

Your **Maximum cost per run** is respected: the run saves only the articles it pays for and ends with "Stopped at your spending limit: N rows delivered". A run close to its timeout stops starting new feeds, saves what it has read and ends with "Stopped before the run timeout".

### Limits & FAQ

- **How fresh?** Google's RSS lags the web UI by a few minutes. `sinceHours` + `when:` keep results within your window.
- **How many?** Max 100 items per feed (Google's cap); topic/geo/top feeds return 30–70. Use several narrower queries for more.
- **Rate limits.** The actor pauses ~0.7 s between feeds and ~0.4 s between URL resolutions. A feed answered with HTTP 429/5xx is tried 3 times with backoff (honouring Retry-After), then recorded in `SUMMARY.errors` and named in the status; after 3 refused or throttled feeds in a row the run stops asking Google and says how many feeds it did not read. If every feed it tried failed, the run fails with the reason.
- **Why is `articleUrl` null?** As of 2026-10-02 news.google.com/robots.txt disallows `/rss/articles/` and `/_/`, the two paths Google's URL resolver needs, so the actor does not call them (it reads robots.txt on every run that would, and resolves again if Google allows it). Other reasons: `decodeUrls` is off; the `maxDecode` budget ran out; news.google.com/robots.txt disallows Google's resolver paths (the actor checks it first); Google refused a request (HTTP 403/429 — resolving then stops for the rest of the run); 3 resolutions in a row failed (Google may have changed its resolver); or the run was about to time out. Articles are saved either way, and the status names the reason. The Google link in `url` still works.
- **Why do I get Russian news for KZ?** Google has no Kazakhstan edition; it serves `RU:ru`. Use `locations` or Kazakh/Russian keywords.
- **Invalid topic / empty geo?** Google answers with an HTML page; the actor reports a clear error for that feed instead of pushing garbage.
- **Terms.** Google's feed header states it is provided for personal, non-commercial feed readers. You are responsible for how you use the data; keep request volumes modest.
- **Roadmap.** Publisher-level aggregation (count per source), optional full-text extraction via a sibling actor.

***

Made by **Yadroo** · Sibling actors: [rss-to-json](https://apify.com/yadroo/rss-to-json) · [crypto-news](https://apify.com/yadroo/crypto-news) · [hackernews-search](https://apify.com/yadroo/hackernews-search) · [youtube-channel-feed](https://apify.com/yadroo/youtube-channel-feed) · [wikipedia-search](https://apify.com/yadroo/wikipedia-search)

# Actor input Schema

## `mode` (type: `string`):

Which Google News feed to read. `auto` picks `topic` when a topic is set, else `geo` when locations are set, else `search` when queries are set, else `topStories`.

## `queries` (type: `array`):

One feed per query (mode `search`). Google operators work inside a query: quotes for exact phrase, `OR`, `-word`, `site:domain.com`, `when:7d`, `intitle:`. Each query returns up to 100 articles.

## `topic` (type: `string`):

Google News section feed for the chosen edition (mode `topic`). Ignores queries.

## `locations` (type: `array`):

City / region names for Google's local-news feed (mode `geo`), e.g. `Almaty`, `Berlin`, `Texas`. One feed per location; use the edition language for best results.

## `edition` (type: `string`):

Verified Google News edition. Overrides `country` and `lang` when set.

## `country` (type: `string`):

ISO-3166 alpha-2 country, e.g. `US`, `DE`, `KZ`. Used when `edition` is empty.

## `lang` (type: `string`):

Language code, e.g. `en`, `de`, `ru`, `es-419`, `zh-Hans`. Used when `edition` is empty.

## `sinceHours` (type: `integer`):

Client-side filter on publish time; in `search` mode it is also sent to Google as a `when:` operator so the 100-item window is spent on fresh news. 0 = no time filter.

## `useWhenOperator` (type: `boolean`):

Search mode only. Turn off to filter purely on the client (Google then returns its default relevance window).

## `sites` (type: `array`):

Search mode: restrict to publisher domains, e.g. `reuters.com`, `bloomberg.com` (joined with OR). Works without queries too — then you get the site's latest indexed articles.

## `excludeSites` (type: `array`):

Search mode: publisher domains to drop via `-site:`.

## `includeKeywords` (type: `array`):

Case-insensitive client-side filter applied to title + snippet (useful with topic/geo feeds that have no query).

## `excludeKeywords` (type: `array`):

Case-insensitive client-side exclusion on title + snippet.

## `limit` (type: `integer`):

Google RSS returns at most 100 items per feed.

## `maxItems` (type: `integer`):

Hard cap on dataset rows for the run — protects cost when many queries are given.

## `sort` (type: `string`):

Order used before the per-feed limit is applied.

## `dedupe` (type: `boolean`):

Skip repeated stories across all feeds of the run (normalized title or identical Google id).

## `decodeUrls` (type: `boolean`):

Google links point to news.google.com redirects. Turn on to resolve the publisher URL into `articleUrl` (2 extra requests per article, about 1 s each). The actor first checks news.google.com/robots.txt and does not call Google's resolver where it is disallowed — as of 2026-10-02 it disallows both paths the resolver needs (/rss/articles/ and /\_/), so `articleUrl` stays null except for the rare links that embed the publisher URL. Resolving also stops for the rest of the run at the first refusal from Google (HTTP 403/429) or after 3 failures in a row; articles are always saved, and the status says why `articleUrl` is null.

## `maxDecode` (type: `integer`):

Most articles to resolve over the network per run (each try counts, resolved or not). Articles beyond it keep `articleUrl: null`. Links that already contain the publisher URL are resolved without a request and do not count.

## `includeSnippet` (type: `boolean`):

Google feeds carry little beyond the headline; the snippet is what remains after stripping title and source (often empty).

## `fields` (type: `array`):

Keep only these fields, in this order, e.g. \["publishedAt", "source", "title", "url"]. `id` is always kept (first unless you place it). Field names: id, feed, feedType, title, source, sourceUrl, url, articleUrl, publishedAt, snippet, edition, effectiveEdition, country, lang, feedUrl, fetchedAt. Letter case, spaces, snake\_case and the RSS names link / pubDate / guid / description / publisher are understood; an unknown name is left out and named in the status. Empty = every field.

## Actor input object example

```json
{
  "mode": "auto",
  "queries": [
    "Kazakhstan fintech",
    "x402 protocol"
  ],
  "topic": "",
  "edition": "",
  "country": "US",
  "lang": "en",
  "sinceHours": 168,
  "useWhenOperator": true,
  "limit": 50,
  "maxItems": 500,
  "sort": "newest",
  "dedupe": true,
  "decodeUrls": false,
  "maxDecode": 20,
  "includeSnippet": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "Kazakhstan fintech",
        "x402 protocol"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("yadroo/google-news-search").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "Kazakhstan fintech",
        "x402 protocol",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("yadroo/google-news-search").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "Kazakhstan fintech",
    "x402 protocol"
  ]
}' |
apify call yadroo/google-news-search --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,yadroo/google-news-search"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/m6S333LkovjQcWR7r/builds/6KnMSPoGxP4ilzVJY/openapi.json
