# Google News Scraper | Real Article URLs, Only New (`akatra/google-news-scraper`) Actor

Scrape Google News by keyword or topic in 50 country and language editions: headline, publisher, publication time and the real article URL instead of a Google redirect. Go past the 100-result limit with date ranges. Turn on 'Only new articles' to monitor news on a schedule.

- **URL**: https://apify.com/akatra/google-news-scraper.md
- **Developed by:** [Akatra](https://apify.com/akatra) (community)
- **Categories:** News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.10 / 1,000 news articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper: real article URLs, only new articles, 50 editions

Collect news from Google News by keyword or by topic, in 50 country and language editions. For every article you get the headline, the publisher, the publication time and **the publisher's real article URL**, not a `news.google.com` redirect link that you would have to open one by one.

Google shows at most 100 results per search. Ask for more and the scraper **splits your time range into single days** and searches each one, so you can pull hundreds or thousands of articles on a topic.

Run it on a schedule with **Only new articles** turned on and you get just what was published since your last run.

**Pay per result** · **No login, no API key** · **Public data only** · **Tested daily** · **LLM-ready JSON for AI agents (MCP)**

### 🛠️ What does Google News Scraper do?

- 🔗 **Real publisher URLs**, not `news.google.com` redirect links you would have to open one by one.
- 🌍 **50 country and language editions.** Search by keyword or follow a topic.
- 📅 **More than Google's 100 results per search.** The time range is split into single days and each day is searched.
- 🆕 **Only new articles mode.** Schedule it and get just what was published since your last run.
- 🛡️ **Built to get through blocking.** Requests go out with a real browser fingerprint through rotating proxies, and a blocked request is retried from a fresh IP. A result that could not be completed is never charged.

#### What people use it for

- **Media monitoring.** Track your brand, competitors, products or people in the news, hourly or daily, and send only new articles to Slack, email, Google Sheets or your own app.
- **Market and trend research.** Pull every article on a topic for a date range and see how coverage changes over time and across countries.
- **News feeds for AI and RAG.** Feed fresh, deduplicated article links with real URLs into your pipeline, ready for a content extractor.
- **PR and SEO reporting.** See which publishers covered a story and when.

### 📊 What data can you extract?

Every result is one JSON object (one row in CSV or Excel) with these fields:

| Field | Meaning |
|---|---|
| `title` | Headline, without the " - Publisher" suffix Google adds |
| `publisher`, `publisherUrl` | News outlet name and its website |
| `publishedAt` | Publication time (ISO, UTC) |
| `url` | The publisher's own article URL. If a link could not be resolved, this holds the Google News link and `urlResolved` is `false` |
| `googleNewsUrl`, `articleId` | The Google News link and Google's ID for the article |
| `related` | In topic feeds, other outlets' coverage of the same story: headline, publisher and Google News link |
| `search` | The query that found the article, or `topic:NAME` for a topic feed |
| `edition` | Country and language edition, for example `US:en` |

### 🚀 How to use Google News Scraper

1. Create a free [Apify](https://apify.com) account and open this Actor.
2. Add one or more **search queries**, for example `electric vehicles`. Google operators work: `"exact phrase"`, `site:reuters.com`, `intitle:tesla`, `OR`, `-exclude`.
3. Or pick one or more **topics** (Top stories, World, Business, Technology and so on) to collect the current headlines of that section.
4. Choose the **country and language** edition.
5. Set **Published within**, or a **From date** and **To date**, to limit the time window.
6. Raise **Max articles per query** above 100 if you want deep coverage of the time window.
7. Turn on **Only new articles** and add a schedule to monitor a topic.
8. Run it and download the results as JSON, CSV, Excel, or connect them through the API.

### 📥 Input example (JSON)

```json
{
  "queries": ["electric vehicles", "\"solid state battery\" site:reuters.com"],
  "topics": ["TECHNOLOGY"],
  "edition": "US:en",
  "maxItemsPerQuery": 300,
  "timeRange": "7d",
  "resolveUrls": true,
  "onlyNew": false
}
```

### 📤 Sample output (JSON)

![Sample output table of Google News Scraper](https://api.apify.com/v2/key-value-stores/GkJnrNz6YWKHQYnZr/records/google-news-scraper-output.png)

```json
{
  "articleId": "CBMijAFBVV95cUxQUmlUS1drVC1NdUdkWFF3TFNOZ0FSd2V2eUtlb3B2Q0wx...",
  "title": "One year after the end of the EV tax credit, is the future still electric?",
  "publisher": "NPR",
  "publisherUrl": "https://www.npr.org",
  "publishedAt": "2026-09-30T09:00:00+00:00",
  "url": "https://www.npr.org/2026/09/30/nx-s1-5979915/ev-hybrid-sales-2026-federal-tax-credit",
  "urlResolved": true,
  "googleNewsUrl": "https://news.google.com/rss/articles/CBMijAFBVV95cUxQUmlUS1drVC1NdUdkWFF3TFNOZ0FSd2V2eUtlb3B2Q0wx...",
  "related": [],
  "search": "electric vehicles",
  "edition": "US:en",
  "scrapedAt": "2026-10-01T22:00:00+00:00"
}
```

### 🤖 Use it with AI agents, MCP and the API

The output is clean JSON with stable field names, ready for LLM pipelines, RAG and agent tools without extra parsing.

- **AI agents and MCP (Claude, ChatGPT, Cursor and others).** Add the Apify MCP server with this Actor as a tool: `https://mcp.apify.com?tools=akatra/google-news-scraper`. The agent can then run it and read the results by itself.
- **API.** Start a run and get the results in a single call:

```bash
curl -X POST "https://api.apify.com/v2/acts/akatra~google-news-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"queries": ["electric vehicles", "\"solid state battery\" site:reuters.com"], "topics": ["TECHNOLOGY"], "edition": "US:en", "maxItemsPerQuery": 300, "timeRange": "7d", "resolveUrls": true, "onlyNew": false}'
```

- **No-code tools and SDKs.** Works with Make, n8n, Zapier and LangChain through Apify's integrations, and with the Apify clients for Python and JavaScript. Schedule it and send the results to a webhook, Google Sheets or your own database.

### 💰 Pricing: pay per result

You pay per article, nothing for empty runs beyond a tiny start fee. See the price on the Pricing tab. Real article URLs are included at no extra charge. There is no subscription or rental fee. New to Apify? The free plan includes $5 of platform credit every month, enough for about 1,600 articles with this Actor, and no credit card is needed.

### 🚦 Run status and error messages

Every run ends with a plain status message, so automated workflows (API, Make, n8n, AI agents) can tell what happened without reading the log:

| Situation | What you get |
|---|---|
| Finished normally | `Saved N articles.` |
| Nothing matched | The run succeeds with `No articles saved. Try a broader query, a longer time range, or turn off 'Only new articles'.` You pay only the start fee. |
| Missing or unsupported input | The run fails at once and the log names the field to fix. Nothing is scraped. |
| The site kept blocking some pages | `Some pages stayed blocked, so the results are incomplete; run again to get the rest.` |
| Your maximum charge is reached | Stops cleanly with `Stopped at your maximum charge limit.` Everything saved so far stays in the dataset. |

### ℹ️ Good to know

- One search returns up to 100 articles, which is Google's own limit. Above 100 the scraper searches day by day, newest day first, and stops when it reaches your maximum. With no time window set, it goes back 30 days.
- A single very busy day can still have more than 100 matching articles. To get them all, split the topic into narrower queries, for example one per `site:` or per keyword.
- Google ranks search results by relevance, so articles inside one day are not in strict time order. The output is sorted newest first within each search.
- An article is delivered once per run, even when it matches several of your queries. If you set a maximum charge for a run, the scraper stops as soon as that limit is reached.
- Day-by-day search covers at most the 366 most recent days of your date range.
- Topic feeds return the current headlines of that section, usually 30 to 70 stories, and ignore the time filters.
- **Only new articles** remembers what it delivered per query and edition. The first run is your baseline. From the second run on you get articles published since the previous run (with a two-day safety margin), so older articles that did not fit into an earlier run are not sold to you as new. Changing the query or the edition starts a fresh memory.
- The scraper returns headlines and links, not the article text. Pass the `url` values to a content extractor if you need full text.
- No personal data: the output has no author names or contact details.

### 📝 Changelog

- **2026-10-02** 'Only new articles' now returns articles published since the previous run; unfinished searches are retried on the next run. README restructured: data table, API and MCP examples, run status messages.
- **2026-10-02** Sturdier parsing, protection against endless paging, maximum charge limit respected, after two independent audits.
- **2026-10-02** First release: keyword and topic feeds, real article URLs, day-by-day deep search, only-new mode.

### ⚖️ Legal and privacy

This scraper is an independent tool and is not affiliated with, endorsed by, or sponsored by Google. "Google News" is a trademark of its owner and is used here only to show which website the tool works with. Headlines and articles belong to their publishers. Use the data in line with the laws that apply to you and the source site's terms.

The source site's name and logo are trademarks of their owner and appear here only to identify the website this tool works with. The Actor reads only pages that are public without logging in. How you store and use the data is your responsibility: follow the laws that apply to you, such as GDPR and CCPA.

### 🔗 More scrapers from Akatra

- [Indeed Jobs Scraper](https://apify.com/akatra/indeed-jobs-scraper): 💼 job ads from Indeed in 30 countries, with yearly salary and only-new mode.
- [Google Trends Scraper](https://apify.com/akatra/google-trends-scraper): 📈 interest over time, by region, and rising related queries.
- [Google Ads Transparency Scraper](https://apify.com/akatra/google-ads-transparency-scraper): 🔎 every ad a company runs on Google, with regions and dates.
- [YouTube Transcript Scraper](https://apify.com/akatra/youtube-transcript-scraper): 🎬 video transcripts as clean text and timestamped lines.

### 💬 Support

Found a bug or need a field that is missing? Open an issue on the Issues tab and we usually reply within a day.

# Actor input Schema

## `queries` (type: `array`):

Keywords to search in Google News, one per line. Google search operators work: "exact phrase", site:reuters.com, intitle:tesla, OR, -exclude.

## `topics` (type: `array`):

Optional. Also collect the current headlines of these Google News sections for the chosen edition. Use this with or without search queries.

## `edition` (type: `string`):

Which Google News edition to read. It decides the language and the country the results are ranked for.

## `maxItemsPerQuery` (type: `integer`):

Stop after this many articles for each query. Google returns at most 100 per search; when you ask for more, the scraper splits the time range into single days and searches each day.

## `timeRange` (type: `string`):

Only articles from this time window. Ignored when you set the dates below.

## `dateFrom` (type: `string`):

Optional. Only articles published on or after this date.

## `dateTo` (type: `string`):

Optional. Only articles published on or before this date.

## `resolveUrls` (type: `boolean`):

Google News links point to a Google redirect. Keep this on to get the publisher's own article URL. Turn off for a faster run that returns the Google News link only.

## `onlyNew` (type: `boolean`):

The first run is your baseline. Later runs with the same query and edition return only articles published since your previous run that you have not received yet. Ideal for an hourly or daily schedule.

## `proxyConfiguration` (type: `object`):

Apify Proxy is recommended. The scraper spreads requests over several IPs automatically.

## Actor input object example

```json
{
  "queries": [
    "electric vehicles"
  ],
  "edition": "US:en",
  "maxItemsPerQuery": 100,
  "timeRange": "",
  "resolveUrls": true,
  "onlyNew": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `articles` (type: `string`):

All scraped articles, shown in the overview table.

## `articlesFull` (type: `string`):

All scraped articles with every field.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "electric vehicles"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("akatra/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["electric vehicles"],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("akatra/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "electric vehicles"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call akatra/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,akatra/google-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wGRJ8rT2lRlpS5hXj/builds/whgffhKrwuen5nl2A/openapi.json
