# Google News Scraper – Headlines, Real Article URLs & Alerts (`gazidev/google-news-scraper`) Actor

Scrape Google News headlines by keyword, topic, location, language and country. Resolves original publisher URLs, optional og:image/description metadata, monitor mode for new articles only, dedupe. Fast HTTP-only RSS scraper, $1 per 1,000 articles. Headlines, links and metadata only.

- **URL**: https://apify.com/gazidev/google-news-scraper.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper – Headlines, Real Article URLs & Alerts

Scrape **Google News** headlines for any keyword, topic, location, language and country, and get the **real publisher URL** for each one instead of a `news.google.com` redirect link. It reads Google News RSS feeds over plain HTTP (no browser, no login), so runs are fast and cheap.

- **Any query, topic or edition.** Google News operators work (`"exact phrase"`, `OR`, `-exclude`, `site:`, `when:`), as do topics, local news, top stories and any language/country edition, with time windows from 1 hour to 30 days.
- **Real article URLs and alerts.** Google News links are decoded to the publisher URL, duplicates across queries are removed, and monitor mode returns only articles you haven't seen yet.
- **$1 per 1,000 articles.** Optional og:image/description metadata costs +$0.002 per article. The Actor returns headlines, links and metadata only, never full article text.

### Quick start

This is the prefilled input: 5 headlines in a few seconds, for about $0.005.

```json
{ "queries": ["artificial intelligence"], "maxItemsPerQuery": 5 }
```

#### Sample output

| title | source\_name | published\_at | url\_resolved |
|---|---|---|---|
| ‘That’s so AI!’ What gen Alpha’s biggest insult tells us | The Guardian | 2026-09-24 | true |
| US, China agree to cut tariffs on $30 billion worth of goods, set up channel for AI incidents | The Hill | 2026-09-26 | true |
| Agentes no autorizados de OpenAI atacaron tres sitios web distintos del Gobierno de EE.UU. | CNN en Español | 2026-09-26 | true |

#### Price comparison (Apify Store, September 2026)

| Actor | Price per 1,000 articles | Monthly users |
|---|---|---|
| **Google News Scraper (this Actor)** | **$1** (+ $0.0001 per run) | new |
| data\_xplorer/google-news-scraper-fast | $4 | 557 |

**Related Actors:** [Bluesky Scraper](https://apify.com/gazidev/bluesky-scraper), [Substack Scraper](https://apify.com/gazidev/substack-scraper) and [Polymarket Scraper & API](https://apify.com/gazidev/polymarket-data).

### What can you use it for?

- **Media monitoring & PR**: track mentions of your brand, competitors or executives in every country
- **News alerts**: schedule monitor mode every hour and send only new headlines to Slack, email or a webhook
- **Market & finance research**: headlines about tickers, sectors or commodities in real time
- **AI agents & RAG**: give an LLM fresh, dated headlines with publisher links to cite
- **Local news**: collect regional headlines by city or state
- **Datasets**: build headline datasets across languages for NLP and trend analysis

### Input example

```json
{
  "queries": ["artificial intelligence", "\"OpenAI\" OR \"Anthropic\" -stock"],
  "topics": ["TECHNOLOGY"],
  "language": "en",
  "country": "US",
  "timeWindow": "1d",
  "maxItemsPerQuery": 50,
  "resolveOriginalUrl": true,
  "fetchArticleMetadata": false,
  "monitorMode": false
}
```

| Field | Description |
|---|---|
| `queries` | Keywords (one per line). Google News search operators work. |
| `topics` | `WORLD`, `NATION`, `BUSINESS`, `TECHNOLOGY`, `ENTERTAINMENT`, `SPORTS`, `SCIENCE`, `HEALTH`, or a topic ID / `news.google.com/topics/...` URL. |
| `locations` | City/region names for local headlines, e.g. `Berlin`, `Texas`. |
| `includeTopStories` | Also scrape the edition's top stories. |
| `language` / `country` | Edition, e.g. `en`/`US`, `en`/`GB`, `de`/`DE`, `tr`/`TR`, `es-419`/`MX`, `pt-419`/`BR`, `ja`/`JP`. |
| `timeWindow` | `any`, `1h`, `1d`, `7d`, `30d`. |
| `maxItemsPerQuery` | 1–100 (Google News RSS returns up to ~100 per feed). |
| `dedupe` | Remove duplicates across queries (article ID, headline+source, publisher URL). Default on. |
| `resolveOriginalUrl` | Decode to the publisher URL. Default on. |
| `fetchArticleMetadata` | Read og:\* metadata from the publisher page (+$0.002 per enriched article). |
| `monitorMode` / `monitorStateKey` | Only new articles since previous runs; state lives in the named key-value store `google-news-monitor`. |
| `proxyConfiguration` | Optional Apify datacenter proxy (normally not needed). |

### Output example

```json
{
  "title": "‘That’s so AI!’ What gen Alpha’s biggest insult tells us",
  "source_name": "The Guardian",
  "source_url": "https://www.theguardian.com",
  "published_at": "2026-09-24T04:00:00Z",
  "article_url": "https://www.theguardian.com/society/2026/sep/24/thats-so-ai-what-gen-alphas-biggest-insult-tells-us",
  "url_resolved": true,
  "google_news_url": "https://news.google.com/rss/articles/CBMioAFBVV95cUxQ...?oc=5",
  "snippet": null,
  "image_url": null,
  "query": "artificial intelligence",
  "feed_type": "search",
  "rank": 1,
  "language": "en",
  "country": "US",
  "article_id": "CBMioAFBVV95cUxQ...",
  "scraped_at": "2026-09-27T01:12:42Z",
  "related_coverage": null
}
```

With `fetchArticleMetadata: true`, items also carry `meta_title`, `meta_description`, `meta_image`,
`meta_published_time`, `meta_site_name`, `meta_author` and `metadata_fetched`. `snippet` and `image_url`
are then filled from og:description / og:image. Top stories include `related_coverage`, a list of other
outlets covering the same story. A run summary (resolution stats, per-feed counts, errors) is saved to
the `OUTPUT` record of the default key-value store.

### Pricing

Pay per event, so you only pay for articles you get:

| Event | Price |
|---|---|
| Article | **$0.001** ($1 per 1,000) |
| Metadata enrichment (optional, only when found) | +$0.002 |
| Actor start | $0.0001 |

Duplicates, already-seen articles in monitor mode, failed feeds and blocked publisher pages cost nothing.
The Actor stops cleanly when your run's maximum charge is reached.

| | This Actor | Market leader (Apify Store, Sept 2026) |
|---|---|---|
| Price per 1,000 articles | **$1** | $4 |
| Original publisher URL | ✅ decoded | varies |
| Monitor mode (only new) | ✅ built-in | – |
| Dedupe across queries | ✅ | – |

### Use with AI agents / Apify MCP

The Actor works as a tool for Claude, ChatGPT, Cursor and other MCP clients through the
[Apify MCP server](https://mcp.apify.com). Add it by name and let the agent call it with
`{"queries": ["<topic>"], "timeWindow": "1d", "maxItemsPerQuery": 20}`. The output is flat JSON with ISO
dates and publisher links the agent can cite. The input is small and every field has a default, so
agents rarely send a bad request. For recurring briefings, schedule it with `monitorMode: true` so each
run returns only new headlines.

### FAQ

**How many articles can I get per query?** Google News RSS returns up to about 100 articles per feed.
For more coverage, split a topic into narrower queries (e.g. add `site:` or `when:` / `after:` operators)
or run several time windows. Results are deduplicated across queries.

**How does original URL resolution work?** Google News links are encoded. The Actor decodes them with
the same public endpoints news.google.com uses in the browser. Resolution is best-effort. If Google changes
or throttles it, you still get the article with the Google News URL in `article_url` and
`url_resolved: false`. The run doesn't fail.

**Do I need a proxy?** Usually not. Tests resolved 600 articles in one run without a proxy. If you
run very large jobs and see `url_resolved: false`, enable the Apify datacenter proxy or lower `maxConcurrency`.

**Does it return the full article text?** No. It returns headlines, links, dates and page metadata
(og:title/description/image) only. It never returns full article text.

**Why is some metadata missing?** Some publishers block automated requests (HTTP 403) or have no og tags.
The reason is shown in `metadata_error`, and you're only charged for metadata when it's found.

**How does monitor mode remember what I've seen?** It stores article IDs, headline+source keys and
publisher URLs in the named key-value store `google-news-monitor` (entries expire after 45 days).
Use `monitorStateKey` to run several independent monitors.

### Legal / responsible use

This Actor returns **headlines, links and page metadata only**, never full article content. You are
responsible for complying with Google's terms (including the terms shown in Google News RSS feeds),
publishers' terms and copyright, and the laws that apply to how you store and use the data.

# Actor input Schema

## `queries` (type: `array`):

Keywords to search on Google News, one per line. Google News operators work: "exact phrase", OR, -exclude, site:reuters.com, intitle:, before:/after:, when:1d.

## `topics` (type: `array`):

Google News topic sections: WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE, HEALTH (one per line). You can also paste a topic ID or a news.google.com/topics/... URL.

## `locations` (type: `array`):

City, region or country names for local headlines, e.g. "Berlin", "Texas".

## `includeTopStories` (type: `boolean`):

Also scrape the edition's front-page top stories.

## `language` (type: `string`):

Google News language code, e.g. en, de, fr, es, tr, ja, pt-419, zh-Hans.

## `country` (type: `string`):

Two-letter country code of the edition, e.g. US, GB, DE, FR, TR, IN, JP, BR.

## `timeWindow` (type: `string`):

Only articles published within this period (adds Google's when: operator to searches and filters all feeds by date).

## `maxItemsPerQuery` (type: `integer`):

Google News RSS returns up to ~100 articles per feed. Use narrower queries or time windows for more coverage.

## `dedupe` (type: `boolean`):

Skip articles already returned by another query/feed in this run (same article ID, same headline+source, or same publisher URL).

## `resolveOriginalUrl` (type: `boolean`):

Decode Google News redirect links into the real article URL (best-effort; if it fails the Google News URL is kept and url\_resolved=false).

## `fetchArticleMetadata` (type: `boolean`):

Visit each publisher page and read its og:title, og:description, og:image and published time. Never the article text. Extra $0.002 per article where metadata is found.

## `includeRelatedCoverage` (type: `boolean`):

For clustered top stories, include the related articles (title, source, link) Google groups with the headline.

## `monitorMode` (type: `boolean`):

Only return articles not seen in previous runs. State is kept in the named key-value store 'google-news-monitor'. Ideal for scheduled runs / alerts.

## `monitorStateKey` (type: `string`):

Name of this monitor. Leave empty to derive it from your queries, topics, language and country. Use different keys for independent monitors.

## `maxConcurrency` (type: `integer`):

Parallel HTTP requests. Lower it if URL resolution gets rate-limited.

## `proxyConfiguration` (type: `object`):

Not needed normally. Enable Apify datacenter proxy if Google starts rate-limiting (e.g. very large runs).

## Actor input object example

```json
{
  "queries": [
    "artificial intelligence"
  ],
  "includeTopStories": false,
  "language": "en",
  "country": "US",
  "timeWindow": "any",
  "maxItemsPerQuery": 5,
  "dedupe": true,
  "resolveOriginalUrl": true,
  "fetchArticleMetadata": false,
  "includeRelatedCoverage": true,
  "monitorMode": false,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `articles` (type: `string`):

No description

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "artificial intelligence"
    ],
    "maxItemsPerQuery": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["artificial intelligence"],
    "maxItemsPerQuery": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "artificial intelligence"
  ],
  "maxItemsPerQuery": 5
}' |
apify call gazidev/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/google-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/emcpuIPG1R1rSgoFw/builds/xaG32jpvxcNgGMHpb/openapi.json
