# Google News Scraper - Real URLs, Images, Any Language (`unbrowseai/google-news-scraper`) Actor

Scrape Google News by keyword, topic or section in 50+ editions: headline, publisher, real article URL (not the Google redirect), publish time, image, description and related coverage. Time filters and date ranges, over 100 results per search. Half the usual price.

- **URL**: https://apify.com/unbrowseai/google-news-scraper.md
- **Developed by:** [Unbrowse AI](https://apify.com/unbrowseai) (community)
- **Categories:** News, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper – Real Article URLs, Images and Related Coverage

Turn Google News into clean, structured data. Search by keyword, follow a topic such as Technology or Business, or pull the front page of any of 50+ country and language editions. Every article comes back with the headline, the publisher and its website, the **real article URL on the publisher's site** (not the `news.google.com` redirect), the exact publish time in ISO format, a preview image, a short description, and, for clustered stories, the other outlets covering the same event.

### Why this one

- **Publisher links, not redirects.** Google News feeds only give you encoded Google links. This scraper resolves them to the article's own URL by default, in fast batches, so you can open, deduplicate or crawl the articles directly.
- **Beyond 100 results.** Google caps a search at about 100 articles. Ask for more and keyword searches are split into one-day windows, going further back until your limit is reached.
- **Related coverage.** Top stories and topic sections group several outlets under one story. You get them all, each with its own resolved link.
- **About half the price** of the most used Google News scrapers.

### What people use it for

- **Media monitoring and PR** – track every mention of a brand, product, person or competitor.
- **Market and investment research** – follow tickers, sectors and events with timestamps you can chart.
- **AI and LLM pipelines** – feed fresh, source-attributed headlines and URLs into summarisers, RAG indexes and agents.
- **SEO and content teams** – see which outlets rank in Google News for your topics.
- **Academic and policy research** – build dated news datasets by country and language.

### How to use it

1. Add **search keywords** (Google News operators work: `-word`, `"exact phrase"`, `OR`, `site:reuters.com`, `intitle:`), pick **topics**, or paste **Google News URLs** of topic or publication pages.
2. Choose the **region and language** edition, for example United States (English), Germany (German) or Brazil (Portuguese).
3. Set a **time period** (last hour to last year) or exact **from/to dates**.
4. Set **max articles** per keyword or topic.
5. Run, then export JSON, CSV or Excel, or read the dataset through the API. Schedule it to get a feed of new articles every hour or day.

### Output example

```json
{
  "title": "5 arrested in U.K. near base used by U.S. on suspicion of planning terror act",
  "source": "The Washington Post",
  "sourceUrl": "https://www.washingtonpost.com",
  "url": "https://www.washingtonpost.com/world/2026/09/27/5-arrested-uk-near-base-used-by-us-suspicion-planning-terror-act/",
  "urlResolved": true,
  "googleNewsUrl": "https://news.google.com/rss/articles/CBMiswFBVV95cUxN…?oc=5",
  "publishedAt": "2026-09-28T00:03:02.000Z",
  "publishedTimestamp": 1790553782000,
  "description": "President Trump said U.S. and British authorities were working together on the investigation.",
  "image": "https://www.washingtonpost.com/wp-apps/imrs.php?src=…&w=1440",
  "author": null,
  "relatedArticles": [
    { "title": "5 Arrested on Suspicion of Terrorism Near RAF Fairford Air Base in UK", "source": "The New York Times", "url": "https://www.nytimes.com/2026/09/27/world/europe/raf-fairford-airbase-arrests.html", "googleNewsUrl": "https://news.google.com/rss/articles/…" }
  ],
  "position": 1,
  "query": null,
  "topic": "TOP_STORIES",
  "language": "en",
  "country": "US",
  "articleId": "CBMiswFBVV95cUxN…",
  "scrapedAt": "2026-09-28T04:10:00.000Z"
}
```

`image`, `description` and `author` come from the article page itself. Some publishers refuse automated visits; for those articles these fields are `null`, while the headline, publisher, time and link are always filled. Turn off **Add image and description** for the fastest runs.

### Pricing

One price per article:

| Apify plan | Price per 1,000 articles |
|---|---|
| Free | $2.00 |
| Starter | $2.00 |
| Scale | $1.88 |
| Business and Enterprise | $0.50 |

Searches that fail or find nothing are free. A tiny fee applies per run start. Set a maximum cost per run and the scraper stops cleanly at that limit.

### FAQ

**How fresh is the data?** Articles are read live from Google News at run time. Use the "Last hour" period and a schedule for near real-time monitoring.

**Why do some articles have no image?** The image and description are read from the publisher's page, and a few large publishers block automated visits. Everything Google News itself shows is always present.

**Can I get full article text?** Not in this scraper. Feed the resolved `url` values into a web page content extractor.

**Is it legal?** It collects publicly available headlines and links. You are responsible for how you use the data, including copyright, the publishers' terms and Google's terms.

Missing a field or edition? Open an issue on the Issues tab.

# Actor input Schema

## `keywords` (type: `array`):

Search terms, one per line. Google News operators work: -word to exclude, "exact phrase", OR, site:reuters.com, intitle:word.

## `topics` (type: `array`):

Google News sections to collect. TOP\_STORIES is the front page with story clusters.

## `topicUrls` (type: `array`):

Optional. Any news.google.com topic, search or publication page URL (e.g. a niche topic you follow). Its locale parameters are kept.

## `maxItems` (type: `integer`):

Google shows up to 100 articles per search; above that, keyword searches are split into one-day windows to go further back. 0 means up to 5,000.

## `timeframe` (type: `string`):

How far back keyword searches look. Ignored when From/To dates are set. Topics return the current Google News selection.

## `dateFrom` (type: `string`):

Optional. Only articles published on or after this date (YYYY-MM-DD). Keyword searches only.

## `dateTo` (type: `string`):

Optional. Only articles published before this date (YYYY-MM-DD). Keyword searches only.

## `regionLanguage` (type: `string`):

Google News edition: which country's news and in which language.

## `decodeUrls` (type: `boolean`):

Replace Google redirect links with the real article URL on the publisher site. Fast (batched); on by default.

## `extractDetails` (type: `boolean`):

Read each article page header for its preview image, description and author. Sites that block automated visits return null for these fields.

## `includeRelated` (type: `boolean`):

For clustered stories (top stories and topics), add the other outlets covering the same story, with their resolved URLs.

## Actor input object example

```json
{
  "keywords": [
    "nvidia"
  ],
  "topics": [],
  "topicUrls": [],
  "maxItems": 10,
  "timeframe": "7d",
  "regionLanguage": "US:en",
  "decodeUrls": true,
  "extractDetails": true,
  "includeRelated": true
}
```

# Actor output Schema

## `articles` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "nvidia"
    ],
    "maxItems": 10,
    "timeframe": "7d"
};

// Run the Actor and wait for it to finish
const run = await client.actor("unbrowseai/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["nvidia"],
    "maxItems": 10,
    "timeframe": "7d",
}

# Run the Actor and wait for it to finish
run = client.actor("unbrowseai/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "nvidia"
  ],
  "maxItems": 10,
  "timeframe": "7d"
}' |
apify call unbrowseai/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,unbrowseai/google-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Tv8Td3bDJogCn7l7G/builds/GSU7DFsOq1DTYBBCM/openapi.json
