# Google News Scraper - Articles, Real URLs & Full Text (`tenfoldfleet/google-news-scraper`) Actor

Get news articles from Google News by keyword, topic or top stories in any language and country: title, publisher, date, the real article URL (decoded from news.google.com links) and optional full text. Date filters, dedupe. $2 per 1,000 articles, +$3 per 1,000 full texts; failed queries are free.

- **URL**: https://apify.com/tenfoldfleet/google-news-scraper.md
- **Developed by:** [Tenfold Fleet](https://apify.com/tenfoldfleet) (community)
- **Categories:** News, AI, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Google News Scraper do?

**Google News Scraper** extracts news articles from [Google News](https://news.google.com) by **keyword search**, **topic section** (World, Business, Technology, Science, Sports, Health and more) or **top stories**, in any **language and country**. For every article you get the **title, publisher, publish date, Google News link and the real article URL**, decoded from the `news.google.com/rss/articles/...` redirect links. Turn on full text and it also returns the **main article text**, author and image from the publisher's page.

It works like a **Google News API**: run it from the Apify Console, call it over the REST API, schedule it, or plug it into Make, Zapier, n8n, LangChain or an AI agent through MCP. It reads Google News RSS over plain HTTP, so it is fast and cheap. No browser, no login.

### Why use it?

- **Media monitoring and brand tracking**: get every new article about your company, competitors or executives each hour or day.
- **Market and financial research**: collect news about stocks, sectors and commodities for sentiment analysis.
- **AI and LLM pipelines**: feed fresh news with full article text into RAG, summarization or newsletters.
- **PR and SEO**: see which publishers cover a topic and link to the real article URLs.
- **News archives and datasets**: back-fill months of coverage day by day with the date-split option.

### What data can it extract?

| Field | Description |
|---|---|
| `query` | Your search query, `topic:BUSINESS` or `top-stories` |
| `position` | Rank of the article in the feed |
| `title` | Headline (publisher suffix removed) |
| `publisher`, `publisherUrl` | News site name and homepage |
| `publishedAt` | Publish time, ISO 8601 (UTC) |
| `googleNewsUrl` | The news.google.com article link |
| `articleUrl` | **Real publisher URL**, decoded (null if Google won't resolve it) |
| `snippet` | Article description (filled from the publisher page when full text is on) |
| `text`, `textLength` | Main article text (with *Fetch full article text*) |
| `textError` | Why there is no full text: HTTP error, paywall teaser only (not charged), no article found |
| `author`, `imageUrl` | Author and lead image (with *Fetch full article text*) |
| `language`, `country`, `scrapedAt` | Edition and scrape time |

### How to scrape Google News

1. Enter one or more **search queries** (Google operators work: `"exact phrase"`, `OR`, `-word`, `site:reuters.com`, `intitle:`), and/or pick **topic sections**. Leave both empty to get today's top stories.
2. Set **language** (`en-US`, `de`, `fr`, `es-419`, `ja`...) and **country** (`US`, `DE`, `IN`...).
3. Optionally limit by time: **Published within** (last hour, day, week...) or a **date range**.
4. Keep **Decode real article URLs** on, and turn on **Fetch full article text** if you need the article body.
5. Click **Start**, then download the results as JSON, CSV, Excel or HTML, or read them via the API.

### How much does it cost to scrape Google News?

Pay only for results:

- **$2.00 per 1,000 articles** (`article` event), real URLs included. Rows whose real URL Google won't give are free.
- **+$3.00 per 1,000 full texts** (`article-text` event), charged only when the full text was extracted. Paywall teasers (the text is still returned), video and blocked pages are free.
- Failed queries and empty results are free. Set a maximum cost per run in the Console and the scraper never charges past it.

### Input

```json
{
    "queries": ["artificial intelligence", "site:reuters.com tesla"],
    "topics": ["BUSINESS", "TECHNOLOGY"],
    "language": "en-US",
    "country": "US",
    "timeWindow": "1d",
    "maxArticlesPerQuery": 100,
    "decodeUrls": true,
    "fetchFullText": false
}
```

Date range instead of `timeWindow`: `"dateFrom": "2026-09-01", "dateTo": "2026-09-30"` (both days included). Add `"splitByDay": true` and raise `maxArticlesPerQuery` to get more than 100 articles (see Tips).

### Output

```json
{
    "query": "site:reuters.com nvidia",
    "topic": null,
    "position": 1,
    "title": "DeepSeek partners with Huawei to develop chip programming tools, reducing reliance on Nvidia",
    "publisher": "Reuters",
    "publisherUrl": "https://www.reuters.com",
    "publishedAt": "2026-09-30T03:03:00.000Z",
    "googleNewsUrl": "https://news.google.com/rss/articles/CBMizgFBVV95cUxP...?oc=5",
    "articleUrl": "https://www.reuters.com/world/asia-pacific/deepseek-partners-with-huawei-develop-chip-programming-tools-reducing-reliance-2026-09-30/",
    "snippet": null,
    "language": "en-US",
    "country": "US",
    "scrapedAt": "2026-10-01T02:46:41.809Z",
    "error": null
}
```

Download the dataset as JSON, CSV, Excel, XML or HTML table, or fetch it from the API.

### Tips

- **Getting more than 100 articles.** Google News RSS returns at most about **100 articles per query**. To get more, split the time range: Google's `after:` and `before:` operators restrict results to a date window, so each day (or week) returns its own ~100 articles. Set `dateFrom` and `dateTo`, turn on **Split date range by day** and raise **Max articles per query** (for example 3000 for a month): the scraper runs one search per day, newest first, and stops at `maxArticlesPerQuery`. With the default of 100 you only get the newest day. You can also split manually by adding queries like `tesla after:2026-09-01 before:2026-09-08`.
- **Same story, many publishers.** Articles are de-duplicated by URL across all your queries in a run, so you never pay twice for the same article.
- **Monitoring.** Schedule the actor hourly with `timeWindow: "1h"` or daily with `"1d"` and send new rows to Slack, Google Sheets or a webhook.
- **Topics** follow the edition you choose: `NATION` is US news for `US`, German news for `DE`, and so on. Date filters apply to search queries only.
- **Real URLs.** Google now hides the publisher URL behind an encoded id. The scraper resolves it the same way the Google News site does. In the rare case Google refuses, `articleUrl` is `null` and the row is still returned, free of charge.
- **Rate limits.** If Google answers with HTTP 429 or a CAPTCHA page, the scraper backs off and switches to Apify's datacenter proxy for the rest of the run. It never solves CAPTCHAs.
- **Large runs.** Hundreds of queries or a long day-by-day range can take a while. Near the run timeout the scraper stops starting new searches and marks the rest as skipped (free), so the run still succeeds. Raise the timeout or split the input for more.

### FAQ

#### Is there an official Google News API?

No. Google retired the Google News API years ago. This actor gives you the same data through Google News RSS feeds with a clean JSON output.

#### Why is `articleUrl` sometimes null?

Google could not resolve that article id (rare, usually during heavy rate limiting). The row still has the title, publisher, date and Google News link, and it is not charged.

#### Can I get the full article text?

Yes. Turn on **Fetch full article text**. The text is extracted from the publisher's page. Paywalled sites (for example Bloomberg, WSJ or FT) only show a teaser or nothing. A teaser is still returned in `text`, with `textError` saying so, and like pages without text it is not charged.

#### Which languages and countries are supported?

Every Google News edition: set `language` (hl) and `country` (gl). The edition id (ceid) is built automatically; set `ceid` yourself only to force a specific edition. If Google has no edition for your pair (for example `de` in `US`), the run stops with a free error row instead of returning another country's news.

#### Is it legal to scrape Google News?

The actor reads public RSS feeds and public article pages. It does not log in or collect personal data. You are responsible for how you use the data, including copyright of article texts and Google's terms. Use full article text for analysis, not for republishing. Consult a lawyer if unsure.

### Disclaimer

This actor is not affiliated with, endorsed or sponsored by Google LLC. Google News is a trademark of Google LLC.

### More tools from Tenfold Fleet

| Actor | Price |
|---|---|
| [ATS Jobs Scraper - Greenhouse, Lever, Ashby & 5 More](https://apify.com/tenfoldfleet/ats-jobs-scraper) | $2 per 1,000 (job posting) |
| [Website Contact Scraper - Emails, Phones & Socials](https://apify.com/tenfoldfleet/company-contact-finder) | $8 per 1,000 (website with contacts) |
| [Google Ads Transparency Center Scraper - Competitor Ads by Site](https://apify.com/tenfoldfleet/google-ads-transparency-scraper) | $1 per 1,000 (ad scraped) |
| [Keyword Search Volume, CPC & Difficulty Checker (Bulk)](https://apify.com/tenfoldfleet/keyword-search-volume) | $5 per 1,000 (keyword with data) |
| [Remote Jobs Aggregator - RemoteOK, WWR, Himalayas & 4 more](https://apify.com/tenfoldfleet/remote-jobs-aggregator) | $1.2 per 1,000 (unique remote job) |
| [SEO Audit Tool - Technical On-Page SEO Checker & Score](https://apify.com/tenfoldfleet/seo-audit-tool) | $20 per 1,000 (page audited) |
| [Shopify Products Scraper - Prices, Variants & Stock](https://apify.com/tenfoldfleet/shopify-products-scraper) | $0.8 per 1,000 (product) |
| [Shopify App & Theme Detector - Store Analyzer & Contacts](https://apify.com/tenfoldfleet/shopify-store-analyzer) | $8 per 1,000 (shopify store analyzed) |
| [Sitemap URL Extractor - All URLs from sitemap.xml & robots.txt](https://apify.com/tenfoldfleet/sitemap-url-extractor) | $0.4 per 1,000 (url extracted) |
| [Technology Detector - Wappalyzer & BuiltWith Alternative](https://apify.com/tenfoldfleet/tech-stack-detector) | $10 per 1,000 (website analyzed) |
| [URL to Markdown - Web Page to LLM-Ready Text for AI Agents](https://apify.com/tenfoldfleet/url-to-markdown) | $1 per 1,000 (page converted) |
| [Workday Jobs Scraper - myworkdayjobs.com API with Salary](https://apify.com/tenfoldfleet/workday-jobs-scraper) | $1.5 per 1,000 (job posting) |
| [YouTube Comments Scraper - Replies, Likes & Sort by Newest](https://apify.com/tenfoldfleet/youtube-comments-scraper) | $0.6 per 1,000 (comment scraped) |
| [YouTube Transcript Scraper - Captions & Subtitles API](https://apify.com/tenfoldfleet/youtube-transcript-scraper) | $3 per 1,000 (transcript extracted) |

# Actor input Schema

## `queries` (type: `array`):

Keywords to search on Google News, one per line. Google search operators work: "exact phrase", OR, -exclude, site:reuters.com, intitle:. Leave empty (and no topics) to get today's top stories.

## `topics` (type: `array`):

Optional Google News sections to scrape in addition to the queries.

## `maxArticlesPerQuery` (type: `integer`):

Google News RSS returns up to ~100 articles per feed. To get more, set a date range, turn on "Split date range by day" and raise this limit (it is the total across all days).

## `language` (type: `string`):

Google News interface language, e.g. en-US, en-GB, de, fr, es-419, pt-BR, ja, zh-TW.

## `country` (type: `string`):

Two-letter country code of the Google News edition, e.g. US, GB (not UK), DE, FR, IN, BR, JP.

## `ceid` (type: `string`):

Advanced. Built automatically from country and language (US:en, DE:de). Set it only to force a specific edition such as BR:pt-419.

## `timeWindow` (type: `string`):

Only articles from the last hour/day/week/... (Google's when: operator). Applies to search queries, not topics.

## `dateFrom` (type: `string`):

Only articles published on or after this date (YYYY-MM-DD, Google's after: operator). Applies to search queries.

## `dateTo` (type: `string`):

Only articles published on or before this date (YYYY-MM-DD, Google's before: operator). Applies to search queries.

## `splitByDay` (type: `boolean`):

Needs both dates. Runs one search per day in the range (newest first), so each day can return up to ~100 articles. Also raise "Max articles per query" (e.g. 3000 for a month): it caps the total across all days. Use it for archives and backfills.

## `decodeUrls` (type: `boolean`):

Resolve news.google.com/rss/articles/... links to the publisher's real URL (articleUrl). If Google won't resolve one, articleUrl is null and the article is still returned, free of charge.

## `fetchFullText` (type: `boolean`):

Download each publisher page and extract the main article text, author and image. Slower. Charged as an extra event only when the full text was extracted: paywall teasers (still returned), blocked and video pages are free.

## `proxyConfiguration` (type: `object`):

Requests go direct first. If Google rate-limits (429 or CAPTCHA page), the scraper backs off and retries through Apify Proxy with a fresh IP. Publisher pages that answer 401/403/429 to the full-text fetch are retried once through it too. Only Apify datacenter proxy (or your own proxy URLs) is used; residential groups are ignored.

## Actor input object example

```json
{
  "queries": [
    "artificial intelligence",
    "site:reuters.com tesla"
  ],
  "maxArticlesPerQuery": 100,
  "language": "en-US",
  "country": "US",
  "splitByDay": false,
  "decodeUrls": true,
  "fetchFullText": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "artificial intelligence",
        "site:reuters.com tesla"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tenfoldfleet/google-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "artificial intelligence",
        "site:reuters.com tesla",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tenfoldfleet/google-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "artificial intelligence",
    "site:reuters.com tesla"
  ]
}' |
apify call tenfoldfleet/google-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tenfoldfleet/google-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZjdXAakwm4tsh6AXR/builds/7WnCSEUc2pdIcboOU/openapi.json
