# Google News Monitor - Full Text & Real URLs (`datagrit/google-news-rss-monitor`) Actor

Monitor Google News for keywords, site: queries and topics in any country; decoded publisher URLs, optional full text, only-new mode.

- **URL**: https://apify.com/datagrit/google-news-rss-monitor.md
- **Developed by:** [datagrit](https://apify.com/datagrit) (community)
- **Categories:** News, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Google News Monitor do?

Google News Monitor watches Google News for your keywords, site: searches and topic sections in any country and language, and returns every article as a clean row: headline, source, publication time, the **real publisher URL** decoded from the Google News link and, if you want it, the **full article text**, author and image. Switch on **Only new articles** and schedule it, and each run returns just the articles you have not received yet. It is built for brand and competitor monitoring, PR and media tracking, market news for trading or research, and news feeds for AI agents and n8n, Make or Zapier workflows.

### Use cases

- **Brand and competitor monitoring**: track mentions of your company, products and competitors in several countries at once, with duplicates across queries merged into one row.
- **PR and media clipping**: collect every article from the last day or hour with the publisher's own URL, ready for a clipping report or a spreadsheet.
- **Company research before a call or deal**: pull recent coverage of a company, including the full text, and feed it to an LLM for a summary.
- **Market and industry news**: follow sectors, tickers or regulators with Google News operators such as "exact phrase", OR, -exclude and site:reuters.com.
- **News datasets and alerts for AI agents**: a typed output schema with a description of every field, list inputs and short runs make it easy to call from agents and automation tools.

### Sample output

```json
{
  "query": "openai",
  "matchedQueries": ["openai"],
  "title": "FTC launches broad investigation into Anthropic, OpenAI",
  "sourceName": "washingtonpost.com",
  "sourceDomain": "washingtonpost.com",
  "sourceHomepage": "https://www.washingtonpost.com",
  "publishedAt": "2026-09-30T17:44:14.000Z",
  "hoursSincePublished": 0.6,
  "originalUrl": "https://www.washingtonpost.com/technology/2026/09/30/ftc-launches-broad-investigation-into-anthropic-openai/",
  "urlStatus": "resolved",
  "googleNewsUrl": "https://news.google.com/rss/articles/CBMirAFBVV95cUxPNERSOFU3...?oc=5",
  "articleId": "CBMirAFBVV95cUxPNERSOFU3...",
  "snippet": "The probe reflects a focus by Trump administration officials on using existing laws to police AI.",
  "language": "en",
  "country": "US",
  "matchedEditions": ["US:en", "GB:en"],
  "fullText": "The Federal Trade Commission has opened a broad investigation into the safety of artificial intelligence systems...",
  "fullTextStatus": "partial",
  "author": "Ian Duncan",
  "imageUrl": "https://www.washingtonpost.com/wp-apps/imrs.php?src=...",
  "found": true,
  "scrapedAt": "2026-09-30T18:20:00.000Z"
}
```

Every field is described in the dataset schema of the Actor. The **Articles** view shows the main columns and the **Full text** view shows author, snippet, text and image.

### How much does it cost?

You pay per article returned. Pricing depends on your Apify plan: a small fee when a run starts, then a price per article that is lower on paid plans. Decoded URLs and full text are included in the article price, with no extra events. You can set a maximum spend on a run and the Actor stops when the limit is reached. The status row of a run that finds nothing, articles dropped by the time range or site: checks, and articles already delivered in only-new mode are never charged. The Apify free plan includes monthly credit you can use to try it.

### Input

- **Search queries**: one Google News search per line. All Google News operators work, for example `"electric vehicles" -tesla`, `openai site:reuters.com` or `nvidia when:12h`.
- **Topic sections**: TOP, WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE or HEALTH.
- **Editions**: countries and languages as COUNTRY:language codes, for example US:en, GB:en, DE:de, FR:fr or BR:pt-419 (UK:en is read as GB:en). Each query is read in each edition. Google silently serves another edition for a pair it does not offer (PL:en returns the US:en edition), so the Actor compares the edition Google actually served with the one you asked for, skips a mismatched edition without returning or charging its articles, and names it in the run status. If none of the requested editions exists, the run fails.
- **Time range**: past hour, 6 or 12 hours, day, 3 days, week, month or year.
- **Maximum articles**: the most articles one run returns, newest first.
- **Only articles new since the last run**, **Decode original article URLs** (on by default) and **Full text, author and image** (off by default).

With an empty input the Actor runs a free example for "artificial intelligence" so you can see the output.

### How does URL decoding work?

Google News links point to news.google.com and hide the publisher URL in an encoded article ID. The Actor decodes each returned article through the same public endpoint the Google News website uses: one request for the article page and one decoding request per 10 articles, about 1.1 requests per article. The result goes into originalUrl. The run status reports how many links were decoded and how many decoded URLs are on the source's own domain ("Original URL resolved for 20 of 20 articles, on the source's own domain for 20 of 20"). If Google stops decoding links, or if most decoded URLs point away from the article's source, the run fails with a clear message before any article is returned or charged, instead of returning rows with missing or wrong URLs.

### FAQ

#### How many articles can one query return?

A Google News RSS feed returns about 100 articles at most (between 70 and 110 in our tests on 30 September and 1 October 2026). When a feed reaches that cap, older matches are cut off, and the run status names the capped query and edition. Use a shorter time range, a when: operator, more specific keywords, or split the query to get complete coverage. Per-feed counts are saved in the FEED\_STATS record of the run.

#### Why is the time range checked again?

Google News does not always respect its own date operator. On 30 September 2026 the query site:bbc.co.uk when:1d returned 8 articles published between 2011 and 2025 among its 100 results. The Actor checks the publication time of every article against the time range and against when:, after: and before: in the query (after: and before: with one day of tolerance, because Google does not state the time zone of these dates), drops the articles outside and counts them in the status. site: queries are checked the same way against the source domain and the decoded publisher URL.

#### Why is fullText empty or short for some articles?

Full text is read with plain HTTP, without a browser. Paywalled sites return a teaser or refuse the request, and some sites render text with scripts. fullTextStatus tells you what happened for each article (ok, partial, empty, blocked, failed), and the run status gives the totals.

#### How often should I run it?

For monitoring, schedule it every hour or every few hours with **Only articles new since the last run** on. The memory is kept per combination of queries, topics, editions and time range, so different monitors never hide each other's articles.

#### How long does a run take?

Reading a feed takes about a second. Decoding URLs adds about 1.1 requests per article, spaced to stay polite to Google, so 20 articles take about half a minute and 100 articles one to two minutes. Full text adds about a second per article, four publishers at a time.

#### Is it legal to use this data?

The Actor reads public Google News feeds and public article pages without logging in and without bypassing any access control. Article texts are protected by the publishers' copyright, and how you use them is your responsibility. This is not legal advice.

### Related Actors

Other public-data Actors from the same publisher, for example company job postings with salaries from Greenhouse and Ashby, are listed on the Store profile.

# Changelog

This Actor's version history is a separate document: https://apify.com/datagrit/google-news-rss-monitor/changelog.md

# Actor input Schema

## `queries` (type: `array`):

Google News searches, one per line. Google News operators work: "exact phrase", OR, -exclude, site:reuters.com, intitle:word, when:1d, after:2026-09-01 and before:2026-09-30. Each query is read in every edition below. An article found by several queries or editions is returned once, with all of them in matchedQueries and matchedEditions. Leave empty when you only monitor topics.

## `topics` (type: `array`):

Optional Google News sections to monitor in every edition: TOP (top stories), WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE or HEALTH. Rows from a section have the query topic:NAME, for example topic:TECHNOLOGY. Unknown names are skipped and listed in the run status.

## `editions` (type: `array`):

Google News editions to read, as COUNTRY:language codes: US:en, GB:en, IN:en, AU:en, CA:en, CA:fr, DE:de, FR:fr, ES:es, IT:it, NL:nl, PL:pl, JP:ja, BR:pt-419, MX:es-419 and other editions Google News offers. en-US style codes are accepted and UK is read as GB. Each edition is one feed per query, so 3 queries x 2 editions read 6 feeds. Google serves another edition for a pair it does not offer (for example PL:en returns US:en); such an edition is skipped, returns no articles, is not charged and is named in the run status. If none of the editions exists, the run fails.

## `timeRange` (type: `string`):

Keep only articles published within this period before the run. The Actor adds the matching when: operator to every query that has no when:, after: or before: of its own, and then checks the publication date of every article again against the time range and the query's own when:, after: and before: (the last two with one day of tolerance), because Google News sometimes returns older articles (on 30 September 2026, 8 of 100 results of site:bbc.co.uk when:1d were from 2011 to 2025). Topic sections are filtered by the same date check.

## `maxItems` (type: `integer`):

Most articles to return in one run, newest first, across all queries, topics and editions. One Google News feed holds about 100 articles at most, so this is the practical maximum per query and edition.

## `onlyNew` (type: `boolean`):

Return only articles that an earlier run with the same queries, topics, editions and time range has not returned yet. The memory is kept per combination of those settings in your account, so two monitors with different queries never hide each other's articles, and only articles that were actually returned are remembered. Use it with a schedule to get a stream of fresh news.

## `resolveUrls` (type: `boolean`):

Google News links point to news.google.com, not to the publisher. When on, the Actor decodes every returned article to its original publisher URL (originalUrl) and checks site: queries against it too. This takes one request per article plus one per 10 articles to Google News (about 1.1 per article), so a run of 100 articles takes one to two minutes longer. Always on when Full text is on.

## `fullText` (type: `boolean`):

Open each article on the publisher's site and extract the article text, author, main image and summary (snippet). Plain HTTP, no browser: paywalled and script-rendered pages return partial or no text, which fullTextStatus reports per article. Adds roughly one second per article.

## `proxyConfiguration` (type: `object`):

Optional proxy for requests to Google News and publishers. Not needed in normal use.

## Actor input object example

```json
{
  "queries": [
    "openai",
    "\"electric vehicles\""
  ],
  "topics": [],
  "editions": [
    "US:en"
  ],
  "timeRange": "1d",
  "maxItems": 20,
  "onlyNew": false,
  "resolveUrls": true,
  "fullText": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All returned articles as a dataset.

## `feedStats` (type: `string`):

Items read per query and edition, feed cap flags and feed errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "openai",
        "\"electric vehicles\""
    ],
    "editions": [
        "US:en"
    ],
    "timeRange": "1d",
    "maxItems": 20,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagrit/google-news-rss-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": [
        "openai",
        "\"electric vehicles\"",
    ],
    "editions": ["US:en"],
    "timeRange": "1d",
    "maxItems": 20,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("datagrit/google-news-rss-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "openai",
    "\\"electric vehicles\\""
  ],
  "editions": [
    "US:en"
  ],
  "timeRange": "1d",
  "maxItems": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call datagrit/google-news-rss-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagrit/google-news-rss-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kaaoxUfTUuQOoy0J8/builds/5knTvFDha9R1yjOzL/openapi.json
