# Google News Scraper: Real Article Links & Top Stories (`precious_bathmat/google-news-scraper`) Actor

Google News articles for any keyword, topic or place in 35 countries, with the publisher's real link, stories grouped across outlets, and a report of top outlets and a daily timeline. Past the 100-result cap. No proxy.

- **URL**: https://apify.com/precious\_bathmat/google-news-scraper.md
- **Developed by:** [Mariam Ahmed](https://apify.com/precious_bathmat) (community)
- **Categories:** News, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper: Real Article Links & Top Stories

Scrape **Google News** for any keyword, topic section or city in **35 countries** and get every article with the **publisher's real URL** (not a news.google.com redirect), the outlet, the publish time and, if you want, the **image and summary**. Articles about the same event are **grouped into stories**, and each search comes with a **news report**: top outlets, the biggest stories, and articles per day.

Most Google News scrapers stop at **100 results per search**, because that is all Google returns. This one searches the time window **one day at a time**, so a single search can return **thousands of articles**: 2,000 for "nvidia" over a month in one test run.

No API key, no login, no proxy: it reads Google News's public RSS feeds.

### What does this Google News scraper do?

```
query "nvidia" · United States · past month · 2,000 articles in 4 minutes

 2,000 real publisher links        408 outlets        1,315 stories
 top outlet   Yahoo Finance 25%    busiest day   23 Sep (110 articles)

 biggest stories   16 outlets  China weighs allowing ByteDance, Alibaba to buy new Nvidia chips
                   16 outlets  Nvidia-backed Nscale files for IPO
                   13 outlets  Nvidia unveils security platform to stop AI agents from going rogue
```

1. **Search, topic or local feeds**: keywords with Google News operators (`"exact phrase"`, `OR`, `-word`, `site:reuters.com`, `intitle:`), sections such as Business or Technology, or local news for a city.
2. **Past the 100 cap**: longer windows are split into daily searches automatically.
3. **Real links**: every Google News link is turned into the article's own URL.
4. **Stories**: headlines that report the same event are grouped, so you can see which stories were covered by 30 outlets and which by one.
5. **Report**: one summary per search with top outlets, top stories, a daily timeline and the newest headlines.

### Who is it for?

- **PR and communications teams**: track coverage of a brand, a launch or a crisis, and see which outlets picked it up
- **Investors and analysts**: every headline about a company or sector, with the stories that moved the most outlets
- **Media monitoring and market research**: daily or weekly news feeds per country, straight into a sheet or dashboard
- **AI and data teams**: clean, deduplicated news datasets with real URLs, ready for summarising or sentiment analysis
- **Journalists and researchers**: who reported a story first, and how coverage spread over the following days

### What data do you get?

**Every article** (one row each, newest first):

| Field | What it is |
|---|---|
| `title`, `outlet`, `outletDomain` | The headline and who published it |
| `articleUrl` | The publisher's own article URL |
| `googleNewsUrl` | The original Google News link |
| `publishedAt`, `ageHours` | When it was published, and how long ago |
| `storyId`, `storySize` | Which story it belongs to, and how many outlets covered that story |
| `countries` | The Google News editions it appeared in |
| `image`, `description`, `language`, `siteName` | From the publisher page, when "Add image and summary" is on |

**One news report per search, topic or place** (`NEWS_REPORT`): article and outlet counts, `topOutlets`, `topStories` (with who reported first), `articlesPerDay`, `busiestDay`, `medianAgeHours` and the ten `newest` headlines.

### Examples from live runs, 29 September 2026

| Input | Articles | Real links | Outlets | Time |
|---|---|---|---|---|
| **"nvidia"**, US, past month, up to 2,000 | **2,000** | 2,000 | 408 | 239 s |
| **"Federal Reserve"**, US + UK, past week, up to 300 each | 535 | 535 | 263 | 35 s |
| **"artificial intelligence"**, US, past week (default) | 100 | 100 | 72 | 10 s |
| **"Elektroauto"**, Germany, 1 to 10 September | 183 | 183 | 75 | 14 s |
| **Top stories + Technology + Nairobi**, DE, JP, KE, BR, with image and summary | 745 | 745 | 196 in top stories | 121 s |

Times are from single runs; big runs spend most of their time looking up real links (about 8 a second at the default 1 GB of memory). In the Federal Reserve run the biggest story was the Fed's stablecoin proposals under the GENIUS Act, picked up by **29 outlets**. Governor Barr's call for more rate hikes came next with 11. With image and summary switched on, **85% of articles** came back with the publisher's image.

### How to use it

1. Enter one or more **search queries**, or pick **topic sections**, or add **places** for local news.
2. Choose the **countries** (Google News editions) and a **time range** or exact dates.
3. Set **max articles per query**. Up to 100 needs one search; more is fetched day by day.
4. Run it, then download the articles as **JSON, CSV or Excel**, or open the news report.

### Pricing

**$0.001 per article**, with its real link, story grouping and the report included.
**+$0.001 per article** for image and summary, charged only when the publisher's page could be read.

100 articles cost **$0.10**; the 2,000-article Nvidia month above cost **$2.00**. Apify's free monthly credit covers several thousand articles.

### Good to know

- **Google's own limits.** Google News returns at most 100 articles per search and only for the recent past. Daily splitting gets around the cap, but a quiet subject simply has fewer articles. A one-hour or one-day window cannot be split further, so it gives up to 100 per country.
- **Newest first.** When the maximum is reached before the whole window is read, you get the most recent days.
- **Dates are in UTC.** Google reads its own date filters in US Pacific time; this Actor corrects for that, so a date range means exactly those UTC days.
- **Topic and local feeds** show what Google News lists right now. They also list extra outlets under each story without a publish time; with image and summary on, the time is read from the publisher page, and otherwise `storyLeadPublishedAt` gives the lead article's time.
- **Some publishers refuse automated visits** (Reuters, WSJ and AP among them). Those rows still have their headline, outlet, date and real link; only the image and summary are missing, and they are not charged for them.
- **Story grouping reads headlines**, not full articles. It is tuned to keep unrelated news apart, so an event occasionally appears as two stories rather than being merged with a different one.
- **No personal data.** No author names or bylines are collected.

### Input

| Field | Meaning |
|---|---|
| **Search queries** | Keywords, one per line; Google News operators work |
| **Topic sections** | Top stories, World, National, Business, Technology, Entertainment, Sports, Science, Health |
| **Local news for places** | Cities or regions, e.g. London, Texas, Nairobi |
| **Countries** | 35 Google News editions; duplicates across editions are removed |
| **Published within** | Past hour, day, week, month, year, or any time |
| **From date / To date** | Exact UTC dates, instead of the time range |
| **Max articles per query and country** | 1 to 5,000 |
| **Real publisher links** | On by default |
| **Add image and summary** | Reads each publisher page; off by default |

### Integrations

Export to JSON, CSV, Excel or Google Sheets, or connect through the Apify API, webhooks, Make, Zapier and n8n. **Schedule it** hourly or daily to build a running news feed for your brand, competitors or market.

# Actor input Schema

## `queries` (type: `array`):

What to search Google News for, one per line, e.g. "nvidia", "electric vehicles", ""interest rates" site:reuters.com". Google News search operators work: quotes, OR, -word, site:, intitle:.

## `topics` (type: `array`):

Google News sections to read as they stand right now. Optional; use with or instead of queries.

## `locations` (type: `array`):

Cities, regions or countries for Google News local headlines, e.g. "London", "Texas", "Nairobi". Optional.

## `countries` (type: `array`):

Which Google News editions to read. Each edition has its own language and outlets. The same article found in several editions is returned once, listing all of them.

## `timeRange` (type: `string`):

For search queries. Ignored when a date range is set below.

## `dateFrom` (type: `string`):

Optional exact start date for search queries (YYYY-MM-DD).

## `dateTo` (type: `string`):

Optional exact end date, included (YYYY-MM-DD). Leave empty for today.

## `maxArticlesPerQuery` (type: `integer`):

Google News gives at most 100 per search. Above 100, the time window is searched one day at a time, so a week can give about 700 articles and a month about 3,000 for a busy subject.

## `resolveUrls` (type: `boolean`):

Turns every news.google.com link into the publisher's own article URL. Switch off only for the fastest possible run.

## `fetchDetails` (type: `boolean`):

Opens each article on the publisher's site for its main image, summary and language. Needs real publisher links. Charged only for pages that could be read.

## Actor input object example

```json
{
  "queries": [
    "artificial intelligence"
  ],
  "countries": [
    "US"
  ],
  "timeRange": "week",
  "maxArticlesPerQuery": 100,
  "resolveUrls": true,
  "fetchDetails": false
}
```

# Actor output Schema

## `articles` (type: `string`):

One row per article, newest first, with the publisher's own link and how many outlets carried the same story.

## `report` (type: `string`):

Per query: top outlets, the biggest stories, articles per day and the newest headlines.

## `summary` (type: `string`):

How many links resolved, any problems, and the limits of the data.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "artificial intelligence"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("precious_bathmat/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": ["artificial intelligence"] }

# Run the Actor and wait for it to finish
run = client.actor("precious_bathmat/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "artificial intelligence"
  ]
}' |
apify call precious_bathmat/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,precious_bathmat/google-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JoZgH8bHpmd9TRY6V/builds/GC0XbCAWtSrmYsSyA/openapi.json
