# Google News Scraper - Real Article URLs + Full Text (`datafetch_labs/google-news-scraper`) Actor

Search Google News or pull topic headlines for any country and language. Returns real publisher URLs (not news.google.com redirects), source, date and related coverage, plus optional full article text, author and image. Time filters and monitor mode.

- **URL**: https://apify.com/datafetch\_labs/google-news-scraper.md
- **Developed by:** [DataFetch Labs](https://apify.com/datafetch_labs) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 article scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper: Real Article URLs + Full Text

**Search Google News or pull topic headlines for any country and language, and get the real publisher URL for every article**, not a `news.google.com` redirect. You also get the source, publish date and related coverage from other outlets. Optionally the Actor visits each article and extracts the **full text, author, description and image**.

- ✅ **Real article URLs**: every `news.google.com/rss/articles/…` link is resolved to the publisher's page.
- ✅ **Search with Google operators** (`OR`, `site:`, `intitle:`, quotes) and **time ranges** (past hour, day, week, month, year).
- ✅ **Topic headlines**: Top stories, World, Business, Technology, Sports, Science, Health, Entertainment.
- ✅ **Any country and language**: US/en, GB/en, DE/de, FR/fr, IN/hi, BR/pt, EG/ar and more.
- ✅ **Full article text** (optional), with author, image and description.
- ✅ **Monitor mode** for news alerts: returns only articles you haven't seen yet.
- 💲 **$1.50 per 1,000 articles**, plus $2.50 per 1,000 articles with full text extracted.

### Use cases

- **Media monitoring and PR**: track mentions of your brand, competitors or executives.
- **Finance and trading**: follow news about companies, tickers and sectors, hour by hour.
- **AI and RAG pipelines**: fresh news with full text for LLM summaries, briefings and agents.
- **Research and journalism**: build datasets of coverage on any topic across countries.
- **Newsletters**: auto-collect the day's top stories in your niche.

### Input example

```json
{
    "queries": ["artificial intelligence", "site:reuters.com tesla"],
    "topics": ["TECHNOLOGY"],
    "language": "en",
    "country": "US",
    "timeRange": "1d",
    "maxArticlesPerQuery": 50,
    "resolveUrls": true,
    "extractText": true
}
```

### Output example

```json
{
    "query": "openai",
    "topic": null,
    "title": "Introducing GPT-6 Sol and Luna",
    "source": "OpenAI",
    "sourceUrl": "https://openai.com",
    "publishedAt": "2026-09-23T16:30:34.000Z",
    "url": "https://openai.com/index/introducing-gpt-6-sol-and-luna/",
    "googleNewsUrl": "https://news.google.com/rss/articles/CBMiZ0FV...",
    "relatedCoverage": [
        { "title": "OpenAI launches two new models", "source": "The Verge", "googleNewsUrl": "https://news.google.com/rss/articles/..." }
    ],
    "articleTitle": "Introducing GPT-6 Sol and Luna",
    "description": "Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work...",
    "imageUrl": "https://images.ctfassets.net/...",
    "author": null,
    "text": "More ways to bring frontier intelligence into the work you do every day...",
    "wordCount": 1397,
    "language": "en",
    "country": "US",
    "scrapedAt": "2026-09-28T22:50:57.970Z"
}
```

Fields from `articleTitle` onward appear only when *Extract full article text* is on.

### Pricing

- **$0.0015 per article ($1.50 per 1,000)**: title, source, date, real URL, related coverage.
- **$0.0025 per article with full text ($2.50 per 1,000)**: charged only when text was actually extracted.

### Notes

- Google News returns up to ~100 articles per search. Use several narrower queries or time ranges to cover more.
- Paywalled or bot-protected publishers (e.g. some major newspapers) may not return full text. Those articles still include the real URL and metadata, and you aren't charged for text.

# Actor input Schema

## `queries` (type: `array`):

Google News searches. Supports Google operators: "tesla OR rivian", "site:reuters.com ai", "intitle:earnings".

## `topics` (type: `array`):

Also (or instead) get headlines from these Google News sections.

## `language` (type: `string`):

Language code: en, es, de, fr, pt, ar, hi, ja...

## `country` (type: `string`):

Two-letter country code: US, GB, IN, DE, FR, BR, EG, AU...

## `timeRange` (type: `string`):

Only for search queries.

## `maxArticlesPerQuery` (type: `integer`):

Google News returns up to 100 articles per search. 0 = all.

## `resolveUrls` (type: `boolean`):

Convert news.google.com links into the publisher's own article URL.

## `extractText` (type: `boolean`):

Visit each article and extract text, author, description and image (extra charge per article with text). Paywalled sites may not return text.

## `dedupe` (type: `boolean`):

Skip articles already returned by another query in the same run.

## `onlyNewArticles` (type: `boolean`):

Return only articles not returned in earlier runs. Schedule it for news alerts.

## `monitorStateName` (type: `string`):

Use a different name for each separate monitor.

## `maxConcurrency` (type: `integer`):

Articles processed at once.

## Actor input object example

```json
{
  "queries": [
    "artificial intelligence"
  ],
  "language": "en",
  "country": "US",
  "timeRange": "any",
  "maxArticlesPerQuery": 20,
  "resolveUrls": true,
  "extractText": false,
  "dedupe": true,
  "onlyNewArticles": false,
  "monitorStateName": "google-news-monitor",
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "artificial intelligence"
    ],
    "maxArticlesPerQuery": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("datafetch_labs/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["artificial intelligence"],
    "maxArticlesPerQuery": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("datafetch_labs/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "artificial intelligence"
  ],
  "maxArticlesPerQuery": 20
}' |
apify call datafetch_labs/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datafetch_labs/google-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xKhfRic5e7ABcyMdg/builds/5I0i8jyHVxqXKK09p/openapi.json
