# Google News Scraper — Articles & Monitor (`zenomastro/google-news-intelligence`) Actor

Google News scraper for brand, competitor and market monitoring. Search by country, language and date; filter publishers, resolve URLs and optionally extract article text. Free analytics report publisher/domain share, daily/query volume, enrichment coverage and tracked keyword mentions.

- **URL**: https://apify.com/zenomastro/google-news-intelligence.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** Marketing, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 news articles

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper API — Articles, Coverage & Monitor

### Why use this Actor?

Google News scraper for brand, competitor and market monitoring. Search by country, language and date; filter publishers, resolve URLs and optionally extract article text. Free analytics report publisher/domain share, daily/query volume, enrichment coverage and tracked keyword mentions.

### Features

- **Search queries** — Search queries. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Language** — Language. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Country code** — Country code. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Maximum articles per query** — Maximum articles per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **From date (YYYY-MM-DD)** — From date (YYYY-MM-DD). Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **To date (YYYY-MM-DD, exclusive)** — To date (YYYY-MM-DD, exclusive). Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Split date ranges by day to reduce RSS caps** — Split date ranges by day to reduce RSS caps. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Only these publisher domains** — Only these publisher domains. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Exclude publisher domains** — Exclude publisher domains. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Try to resolve Google News links to publisher URLs** — Try to resolve Google News links to publisher URLs. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.
- **Optional persistent monitor key** — Use the same key on scheduled runs to mark already-seen articles.
- **Emit only newly seen articles** — Emit only newly seen articles. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

### Use cases

- Brand monitoring.
- Competitor news tracking.
- Pr research.
- News datasets and alerts.

### Example input

```json
{
  "queries": [
    "OpenAI"
  ],
  "language": "en",
  "country": "US",
  "maxItemsPerQuery": 100,
  "splitDateRangeByDay": true,
  "resolveArticleUrls": false
}
```

### Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

### FAQ

**What is this Actor for?**\
It is designed for brand monitoring, competitor news tracking, PR research.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

### Search keywords

google news scraper, google news scraper python, google news scraper api, google news scraper free, google news scraper apify, google news scraper npm, google news scraping, google news scraping python, google news rss scraper, google news web scraper, news, google, articles, search

Search Google News without a browser and turn RSS results into clean, scheduled monitoring datasets.

The Actor accepts multiple queries, country/language editions, date windows and publisher-domain filters. It automatically deduplicates overlapping RSS windows, can split a historical date range into daily windows to reduce Google's roughly 100-item feed ceiling, and can keep a persistent seen-set for scheduled brand/news monitoring.

### Main advantages

- Multiple queries in one run.
- Country and language localization.
- Optional historical `after:` / `before:` windows.
- Daily window splitting for better coverage of high-volume topics.
- Include/exclude publisher-domain filters before billing.
- Deduplication across overlapping windows and queries.
- Optional publisher article URL resolution.
- Persistent `monitorKey` with `isNew` and `onlyNew`.
- Retries, timeouts, bounded input sizes and clean error rows.
- Pay only for article records that are actually delivered when PPE is enabled.

### Example

```json
{
  "queries":["OpenAI","Claude AI"],
  "language":"en",
  "country":"US",
  "maxItemsPerQuery":100,
  "includeDomains":["reuters.com","techcrunch.com"],
  "monitorKey":"ai-competitors",
  "onlyNew":true
}
```

Each article contains the query, title, publication time, publisher, publisher website/domain, Google News URL, cleaned description and optional resolved article URL.

This Actor intentionally uses Google's lightweight public RSS surface instead of a headless browser. That keeps runtime and platform cost low enough for hourly or daily monitoring schedules. Google News RSS is an undocumented public endpoint and can change, so network/shape failures are isolated into diagnostic rows rather than silently dropping the whole run.

Use public news data responsibly and respect publisher terms when following article URLs.

### Extended capabilities

- Search Google News across bounded date windows with deduplication and persistent monitoring.
- Optionally fetch destination articles and extract canonical URL, dates, author, image, text, and word count.
- Preserve source metadata while isolating per-article enrichment failures.

# Actor input Schema

## `queries` (type: `array`):

Search queries. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `language` (type: `string`):

Language. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `country` (type: `string`):

Country code. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `maxItemsPerQuery` (type: `integer`):

Maximum articles per query. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `dateFrom` (type: `string`):

From date (YYYY-MM-DD). Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `dateTo` (type: `string`):

To date (YYYY-MM-DD, exclusive). Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `splitDateRangeByDay` (type: `boolean`):

Split date ranges by day to reduce RSS caps. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `includeDomains` (type: `array`):

Only these publisher domains. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `excludeDomains` (type: `array`):

Exclude publisher domains. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `resolveArticleUrls` (type: `boolean`):

Try to resolve Google News links to publisher URLs. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `monitorKey` (type: `string`):

Use the same key on scheduled runs to mark already-seen articles.

## `onlyNew` (type: `boolean`):

Emit only newly seen articles. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `requestTimeoutSecs` (type: `integer`):

Request timeout. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `retries` (type: `integer`):

Retries. Configure this value to control the Actor run; bounded defaults are chosen for reliable production use.

## `enrichArticles` (type: `boolean`):

Resolve publisher URLs and extract cleaned article text plus canonical URL, author, title, description, image and publication metadata when public HTML is available.

## `maxArticleEnrichments` (type: `integer`):

Hard cap on publisher article fetches per run.

## `maxArticleChars` (type: `integer`):

Maximum cleaned article text kept per enriched article.

## `analyticsKeywords` (type: `array`):

Optional topics, brands, products or phrases counted in the free summary across article title, description and enriched article text.

## Actor input object example

```json
{
  "queries": [
    "OpenAI"
  ],
  "language": "en",
  "country": "US",
  "maxItemsPerQuery": 100,
  "splitDateRangeByDay": true,
  "includeDomains": [],
  "excludeDomains": [],
  "resolveArticleUrls": false,
  "onlyNew": false,
  "requestTimeoutSecs": 25,
  "retries": 2,
  "enrichArticles": false,
  "maxArticleEnrichments": 100,
  "maxArticleChars": 100000,
  "analyticsKeywords": []
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/google-news-intelligence").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/google-news-intelligence").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call zenomastro/google-news-intelligence --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/google-news-intelligence"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Q8IcXVkBV8u71ldg1/builds/M0UlmasNScNFznlwm/openapi.json
