# Google News Scraper (`labrat011/google-news-scraper`) Actor

Scrape Google News by keyword or topic in any language and country. Real publisher URLs (not news.google.com redirects), source, date and related coverage. Go past Google's 100-result cap with a date range: one feed per day.

- **URL**: https://apify.com/labrat011/google-news-scraper.md
- **Developed by:** [mick\_](https://apify.com/labrat011) (community)
- **Categories:** News, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 articles

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<img src="https://apify-image-uploads-prod.s3.us-east-1.amazonaws.com/wCP1WauwRX2Gr3Gir-actor-TGHoMTjtg0yEek5bO-JUZrWZ4unL-google-news-scraper.png" alt="Google News Scraper logo" width="120">

## Google News Scraper

Scrape Google News by keyword or topic, in any language and country. You get the **real publisher URL** for every article (not a news.google.com redirect), plus title, source, publish time, and the related coverage Google groups with each top story.

| At a glance | |
|---|---|
| **You give it** | Keywords or topics, plus optional language, country and date range |
| **You get** | One row per article: title, real publisher URL, source, publish time, related coverage |
| **Price** | $0.005 per run + $0.002 per article (Free plan, lower on paid plans) |
| **Speed** | 30 articles with resolved URLs in about 6 seconds |
| **Needs** | Nothing: no login, no API key, no proxy |

### What you get

- **Real article URLs.** Google News RSS links are encoded redirects. This actor resolves each one to the publisher's own link, so you can open, fetch or deduplicate articles directly. In a side-by-side test on the same query, a popular alternative returned 0 of 30 real URLs; this actor returned 30 of 30.
- **Past the 100-result cap.** Google returns about 100 articles per search. Set a date range and the actor searches one day at a time, about 100 articles per day: a 3-day range for "tesla" returned 299 unique articles, all with resolved URLs, in about 35 seconds (under 5 seconds with `resolveUrls` off).
- **Keyword search with Google operators:** `"exact phrase"`, `site:reuters.com`, `-exclude`, `OR`.
- **Topic headlines:** Top stories, World, Nation, Business, Technology, Entertainment, Sports, Science, Health.
- **Any edition:** language and country, for example `de` + `DE`, `en` + `GB`, `pt-BR` + `BR`.
- **Related coverage:** for clustered top stories, the other outlets' headlines on the same story.
- **Deduplication** across queries and days in the same run.
- **Light and fast:** 30 articles with resolved URLs took 6 seconds and $0.0003 of platform usage.

### Use cases

- **Brand and competitor monitoring.** Every new article that mentions your company, delivered to Slack.
- **PR and media lists.** Which outlets and domains cover your space, and how often.
- **Market and investment research.** Daily news flow for a ticker or sector, with dates, for backtests and dashboards.
- **AI news digests.** Feed real article URLs into an LLM summarizer or a RAG index.
- **Localized news.** The German or Brazilian front page, in its own language.

### Example inputs

#### Latest news for a keyword, past 24 hours

```json
{ "queries": ["openai"], "timeRange": "1d", "maxArticlesPerQuery": 50 }
```

#### Deep history: every day of September

```json
{ "queries": ["\"electric vehicles\""], "dateFrom": "2026-09-01", "dateTo": "2026-09-25", "maxArticlesPerQuery": 0 }
```

#### Only one publisher

```json
{ "queries": ["interest rates site:reuters.com"], "timeRange": "7d" }
```

#### German front page and tech section

```json
{ "topics": ["TOP", "TECHNOLOGY"], "language": "de", "country": "DE", "maxArticlesPerQuery": 30 }
```

### Input

| Field | What it does |
|---|---|
| `queries` | Search terms, one per line. Google News operators work. |
| `topics` | Headline sections: `TOP`, `WORLD`, `NATION`, `BUSINESS`, `TECHNOLOGY`, `ENTERTAINMENT`, `SPORTS`, `SCIENCE`, `HEALTH`. |
| `language`, `country` | The Google News edition. Default `en` and `US`. |
| `timeRange` | `1h`, `1d`, `7d`, `30d`, `1y`. Ignored when a date range is set. |
| `dateFrom`, `dateTo` | Search one day at a time across this range (up to 366 days). |
| `maxArticlesPerQuery` | New articles per query or topic, counted after duplicates are removed. `0` means no limit. Default 100. |
| `resolveUrls` | Resolve publisher URLs. Default on. |
| `deduplicate` | Skip articles already saved in this run. Default on. |

### Output

A real row from a test run on 2026-09-26 (related coverage shortened to two entries):

```json
{
  "title": "Roads flood in New Jersey, inundate some homes, as major signs of nor’easter take shape",
  "url": "https://apnews.com/article/noreaster-storm-cape-cod-new-england-4057bec1ced5c4883d36c30c778df410",
  "googleNewsUrl": "https://news.google.com/rss/articles/CBMinAFBVV95cUxQdGRqdGEzMGFJUjh3UDU4cDJaaGdaWnNMYVpJS1Fid19ueFRYeXRMTGtJODdIeGE4RlYwZnpzU2JWYTdaQnk5LW5KRGluRjVjX2U5cS11UzY4Q0FDQXdWRVJENnluUHBMT1hKa0NRUDlUNU9pVVR0MXN0MFRNLV94WjduYUVkODlBNTdoQUtHUDNvNHp1REV3NWZXQ2c?oc=5",
  "articleId": "CBMinAFBVV95cUxQdGRqdGEzMGFJUjh3UDU4cDJaaGdaWnNMYVpJS1Fid19ueFRYeXRMTGtJODdIeGE4RlYwZnpzU2JWYTdaQnk5LW5KRGluRjVjX2U5cS11UzY4Q0FDQXdWRVJENnluUHBMT1hKa0NRUDlUNU9pVVR0MXN0MFRNLV94WjduYUVkODlBNTdoQUtHUDNvNHp1REV3NWZXQ2c",
  "source": "AP News",
  "sourceUrl": "https://apnews.com",
  "sourceDomain": "apnews.com",
  "publishedAt": "2026-09-26T12:11:00Z",
  "relatedCoverage": [
    {
      "title": "A major nor’easter is battering the East Coast",
      "source": "The Washington Post",
      "googleNewsUrl": "https://news.google.com/rss/articles/CBMilAFBVV95cUxPU08zV3pjZVhQWHlHT1dLSHFoQTJTTEg2d003R1RON2Vpak42STdlc3E3Vzd1WnVYVWVZTnNFY2VOVjZXNWVjWWFuNEdtNjV2bm5ybWVxdmYyQnpfb3RFbU9BRDlVT2hJQWpMNk1QVTI1S2Q1bWlFMWpXT3ZSLWtDWUR3NVZpWUFNZ1RUb29kZHNEZE5E?oc=5"
    },
    {
      "title": "Live updates: Powerful nor'easter batters East Coast with coastal flooding, high winds and power outages",
      "source": "Yahoo",
      "googleNewsUrl": "https://news.google.com/rss/articles/CBMi8wFBVV95cUxOMXN4ZTNmdFBUOWZadG9mODBsSHUyZlRpdkFEQ09tcUlBZ1RST1RKVzdSZ3NFUTNZNXhXSzB2MlNLbFJhSU45QWlobmZJZmd3Vlc5OEQtM29EOTJ0LXVCZ0R2SlBLNXBLOVkzUm5JdkFpX0RqOUZ6TXpmNDZWSi1pSDVuUUVMa2ltXzJ5N0hhX3ZJVjFMbm1IZHlRb1lRTGhpV01pSVUtN3dESW4zZUd5UkE0eWtEMDQteDNuRV94REpyb1BKbFlaODRncVR4T2NOZHJDa0pBZnp6QW5rMUx4OWNCcVBxR252SGFCaWhtN0J3UDA?oc=5"
    }
  ],
  "query": null,
  "topic": "TOP",
  "language": "en",
  "country": "US"
}
```

### Pricing

Pay per event:

- **$0.005** per run start
- **$0.002** per article on the Free plan (lower on higher Apify plans)

Apify platform usage is billed separately to your account and is tiny: 30 articles with resolved URLs used $0.0003.

| Articles | Actor cost |
|---|---|
| 30 | about $0.065 |
| 100 | about $0.21 |
| 1,000 | about $2.01 |

Set a **Maximum cost per run** in the run options and the actor stops cleanly when it reaches it.

### Automate it with n8n

Each workflow uses n8n's official **Apify** node, operation **Run actor and get dataset**, actor `labrat011/google-news-scraper`. Paste the input into **Input JSON**.

#### 1. Brand mentions to Slack every hour

```
Schedule Trigger (every hour)
  > Apify: Run actor and get dataset   { "queries": ["\"Your Brand\""], "timeRange": "1h" }
  > Remove Duplicates (n8n, across executions, on url)
  > Slack: "{{ $json.source }}: {{ $json.title }} {{ $json.url }}"
```

#### 2. Morning industry digest by email, written by AI

```
Schedule Trigger (weekdays 07:00)
  > Apify: Run actor and get dataset   { "queries": ["semiconductors", "chip export"], "timeRange": "1d", "maxArticlesPerQuery": 30 }
  > Aggregate: titles, sources and urls into one list
  > OpenAI / Anthropic: "Write a 5-bullet briefing. Cite the source name and link for each bullet."
  > Gmail: send to the team
```

#### 3. News history to Google Sheets for research

```
Manual Trigger
  > Apify: Run actor and get dataset   { "queries": ["NVDA"], "dateFrom": "2026-01-01", "dateTo": "2026-09-25", "maxArticlesPerQuery": 0, "resolveUrls": false }
  > Google Sheets: Append rows (publishedAt, source, title, url)
```

Turn off `resolveUrls` for big historical pulls when you only need headlines and dates; it is faster.

#### 4. Competitor press tracker in Airtable

```
Schedule Trigger (daily)
  > Apify: Run actor and get dataset   { "queries": ["CompetitorA", "CompetitorB", "CompetitorC"], "timeRange": "1d" }
  > Airtable: Upsert by articleId (query as the competitor column)
```

#### 5. Feed a RAG knowledge base

```
Schedule Trigger (daily)
  > Apify: Run actor and get dataset   (your topic, past 24 hours)
  > HTTP Request: GET {{ $json.url }} (the real publisher page)
  > HTML Extract / Readability: article text
  > Embeddings + Vector Store insert (Pinecone, Qdrant, Supabase)
```

This only works because the URLs are real publisher links.

### For AI agents

- **Actor:** `labrat011/google-news-scraper`
- **Smallest input:** `{ "queries": ["openai"] }`
- **One row = one article.** Key fields: `title`, `url` (publisher link), `source`, `sourceDomain`, `publishedAt` (ISO 8601 UTC), `query` or `topic`.
- **Billing:** `apify-actor-start` once per run, `article` once per row. Set `maxArticlesPerQuery` or a maximum cost per run to cap spend.
- **Run it:** `POST https://api.apify.com/v2/acts/labrat011~google-news-scraper/run-sync-get-dataset-items` with the input as the JSON body, or call it from the Apify MCP server.
- **Done signal:** the run's status message starts `Saved N articles in M requests.` and may add how many links stayed unresolved and which queries or topics had `No results for: ...`.
- **Bad input fails fast** with a readable status message, for example an unknown topic or a reversed date range.

### FAQ

**Why do some rows keep a news.google.com URL?** Google occasionally will not resolve an article. The run summary tells you how many; the row still has title, source and date.

**How far back can I go?** Google News search reaches back years, but coverage thins out for older dates. The date range is limited to 366 days per run.

**Why are a few articles dated just outside my date range?** Google draws its day boundaries in its own time zone, not UTC, while `publishedAt` is always UTC. Expect a few percent of rows to sit a few hours past the last day. Filter on `publishedAt` if you need a hard cut.

**Do I need a proxy?** No.

### Support

Open an issue on the actor's Issues tab with the run ID and it will be looked at.

# Changelog

This Actor's version history is a separate document: https://apify.com/labrat011/google-news-scraper/changelog.md

# Actor input Schema

## `queries` (type: `array`):

One per line. Google News operators work: "exact phrase", site:reuters.com, -excluded, OR.

## `topics` (type: `array`):

Headline sections to read. Top stories is the front page.

## `language` (type: `string`):

Language code: en, de, fr, es, pt-BR, ja...

## `country` (type: `string`):

Two-letter country code: US, GB, DE, IN, BR...

## `timeRange` (type: `string`):

Only recent articles. Ignored when a date range is set.

## `dateFrom` (type: `string`):

Search one day at a time from this date. Each day returns up to about 100 articles, so a date range goes far past Google's usual cap. Format 2026-09-01. Up to 366 days.

## `dateTo` (type: `string`):

Last day of the range, inclusive. Defaults to today.

## `maxArticlesPerQuery` (type: `integer`):

0 means no limit. Without a date range Google returns about 100 per query.

## `resolveUrls` (type: `boolean`):

Turn news.google.com links into the publisher's own URL. Adds about a quarter second per article. The Google link is always kept in googleNewsUrl.

## `deduplicate` (type: `boolean`):

Skip articles already saved by an earlier query or day in this run.

## `proxyConfiguration` (type: `object`):

Not needed normally.

## Actor input object example

```json
{
  "queries": [
    "artificial intelligence"
  ],
  "language": "en",
  "country": "US",
  "timeRange": "",
  "maxArticlesPerQuery": 20,
  "resolveUrls": true,
  "deduplicate": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `articles` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "artificial intelligence"
    ],
    "language": "en",
    "country": "US",
    "maxArticlesPerQuery": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("labrat011/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["artificial intelligence"],
    "language": "en",
    "country": "US",
    "maxArticlesPerQuery": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("labrat011/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "artificial intelligence"
  ],
  "language": "en",
  "country": "US",
  "maxArticlesPerQuery": 20
}' |
apify call labrat011/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,labrat011/google-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TGHoMTjtg0yEek5bO/builds/FWsgqd24Yy5rkUUqP/openapi.json
