# Google News Scraper - Real Article URLs & Full Text (`pantomath/google-news-scraper-real-article-urls-full-text`) Actor

Google News by keyword, operators or topic in any country. Decodes real publisher URLs, optional full article text, day-by-day deep history and only-new monitoring. $2 per 1,000 articles.

- **URL**: https://apify.com/pantomath/google-news-scraper-real-article-urls-full-text.md
- **Developed by:** [abdallah alramahi](https://apify.com/pantomath) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 article scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper — Real Article URLs & Full Text

**Get Google News articles with the publisher's real URL (not Google's redirect), optional full text, any country & language, and unlimited history by splitting date ranges by day. $2 per 1,000 articles.**

Most Google News scrapers return `news.google.com/rss/articles/...` redirect links that break in your pipeline. This Actor **decodes every link into the real article URL** — and can extract the full article text for you.

### What you get

- **Real article URL** (decoded), Google News URL, title, source name & domain, publish date (UTC)
- **Related coverage** from the same story cluster
- **Full text (optional):** clean article text, author, main image, language, site name
- Works with **keywords and Google operators**: `"exact phrase"`, `OR`, `-exclude`, `site:reuters.com`, `intitle:`
- **Topics:** top headlines for World, Business, Technology, Sports, Science, Health…
- **Any edition:** `en-US`, `en-GB`, `en-IN`, `de-DE`, `fr-FR`, `es-ES`, `pt-BR`, `ja-JP`…
- **Deep history:** set a date range and the Actor queries **day by day** to go beyond Google's ~100-results-per-query cap
- **Monitoring mode:** skips articles you already received — schedule hourly for alerts

### Use cases

- **PR & brand monitoring** — every mention of your brand or competitors, with real links
- **Market & investment research** — news flow per company, sector or country
- **SEO & content** — who covers what, which outlets rank in Google News
- **AI / LLM pipelines & RAG** — fresh, full-text news for summarization and agents
- **Media analysis** — coverage volume over time, by source and edition

### Pricing (pay per event)

| Event | Price |
|---|---|
| Article (with decoded URL) | $0.002 |
| Full text extracted (optional, only when successful) | +$0.003 |

1,000 articles ≈ **$2** (or $5 with full text). No start fee.

### Example output

```json
{
  "query": "openai",
  "edition": "en-US",
  "title": "OpenAI safety leader quits, warning AI company’s culture is ‘broken’",
  "source": "theguardian.com",
  "sourceUrl": "https://www.theguardian.com",
  "publishedAt": "2026-10-03T19:41:00+00:00",
  "url": "https://www.theguardian.com/technology/2026/oct/03/openai-safety-leader-quits-warning-ai-companys-culture-is-broken",
  "googleNewsUrl": "https://news.google.com/rss/articles/CBMitgFB..."
}
```

With **Extract full article text** enabled you also get `text`, `author`, `imageUrl`, `language`, `siteName`.

### Tips

- **Brand alerts:** queries = your brand names, time range = last 24 hours, monitoring mode on, schedule hourly, add a Slack/email integration.
- **One outlet only:** `site:bloomberg.com nvidia`
- **Exact phrase:** `"interest rates"`
- **History for a month:** set *Date from* and *Date to* — up to ~100 articles per day per query.

### FAQ

**Why are some `url` values null?** Rarely Google doesn't expose the target; you still get the Google News link.

**Why is `text` empty for some articles?** Paywalls and bot protection on some publishers. You're only charged the full-text fee when text is extracted.

**Need a feature?** Open an issue — I usually reply within 24 hours.

# Actor input Schema

## `queries` (type: `array`):

One per line. Google search operators work: "exact phrase", OR, -exclude, site:reuters.com, intitle:tesla.

## `topics` (type: `array`):

Top headlines for: WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE, HEALTH.

## `editions` (type: `array`):

Google News editions, e.g. en-US, en-GB, en-IN, de-DE, fr-FR, es-ES, pt-BR, ja-JP. Each query runs in each edition.

## `timeRange` (type: `string`):

Only articles from the last hour/day/week/month/year. Ignored when a custom date range is set.

## `dateFrom` (type: `string`):

Custom range start. The range is split into single days (up to ~100 articles per day), so you can collect far more than Google's 100-result limit.

## `dateTo` (type: `string`):

Custom range end (defaults to today).

## `maxArticlesPerQuery` (type: `integer`):

Without a date range Google returns up to ~100 articles per query.

## `decodeUrls` (type: `boolean`):

Converts Google News redirect links (news.google.com/rss/articles/...) into the publisher's real URL.

## `extractFullText` (type: `boolean`):

Downloads each article and extracts clean text, author, main image and language. Some publishers block bots or use paywalls.

## `onlyNewSinceLastRun` (type: `boolean`):

Remembers articles already returned for each query & edition and skips them. Schedule hourly for brand/PR alerts.

## `proxyConfiguration` (type: `object`):

Usually not needed.

## Actor input object example

```json
{
  "queries": [
    "openai",
    "site:reuters.com electric vehicles"
  ],
  "editions": [
    "en-US"
  ],
  "timeRange": "any",
  "maxArticlesPerQuery": 100,
  "decodeUrls": true,
  "extractFullText": false,
  "onlyNewSinceLastRun": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "openai",
        "site:reuters.com electric vehicles"
    ],
    "editions": [
        "en-US"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pantomath/google-news-scraper-real-article-urls-full-text").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": [
        "openai",
        "site:reuters.com electric vehicles",
    ],
    "editions": ["en-US"],
}

# Run the Actor and wait for it to finish
run = client.actor("pantomath/google-news-scraper-real-article-urls-full-text").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "openai",
    "site:reuters.com electric vehicles"
  ],
  "editions": [
    "en-US"
  ]
}' |
apify call pantomath/google-news-scraper-real-article-urls-full-text --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pantomath/google-news-scraper-real-article-urls-full-text"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gzB1eN3hNFcaxn2xN/builds/VRloPGV7Cpge7TdEr/openapi.json
