# Google News Scraper (search, headlines, alerts, no login) (`datahamster/google-news-feed`) Actor

Search Google News by keyword or pull a headlines/topic feed and get one row per article: title, source, publisher URL, published date and link. Monitor mode alerts on new articles for saved keywords. No login, no cookies — reads Google News public RSS feeds directly.

- **URL**: https://apify.com/datahamster/google-news-feed.md
- **Developed by:** [Viktor Dubnytskiy](https://apify.com/datahamster) (community)
- **Categories:** Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 result items

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper: search, headlines and keyword alerts

Read Google News' own public RSS feeds — keyword search, topic/headlines feeds, or a `source:`/`site:` search for one publication — and get one flat row per article. No login, no cookies, no scraping of the Google News web app.

### What you get

Real rows from the example dataset (`queries: ["web scraping"]`):

| title | source | publishedAt |
|---|---|---|
| Extracting Data at Scale: Modern Approaches to Web Scraping | InfoWorld | 2026-09-24 |
| Best web scraping tools for 2026 | TechRadar | 2026-09-20 |

Full row: `id` (Google's own article guid), `url`, `googleUrl` (the Google News redirect link), `title`, `source`, `sourceUrl`, `publishedAt`, `snippet`, `query`, `language`, `country`, `resolved`, `scrapedAt`.

Google News RSS carries no author names or bylines — nothing person-level is collected or added.

### Use cases

- **Keyword alert feed** — put a brand, product or topic in `queries` and run it in monitor mode on a schedule; each run returns only the articles that are new since the previous run.
- **Competitor / industry watch** — track a `source:` query for a competitor's own site, or a topic feed, to see what is being published without opening Google News yourself.
- **One-off research pull** — a single run across several keywords for a report or a dataset, with `maxItemsPerQuery` capping how deep each query goes.

### Try it in 10 seconds

**One-off list** — hit **Start**/**Try it**, the input already works: `queries: ["web scraping"]`, `maxItemsPerQuery: 20`, `maxItems: 20`, nothing required.

**Monitor mode** — save the task, then set:

```json
{
  "mode": "monitor",
  "monitorStateId": "web-scraping-watch",
  "queries": ["web scraping"],
  "webhookUrl": "https://your-endpoint.example.com/hook"
}
```

and put it on a schedule (Apify → Schedules). Each monitor run charges one `monitor-check` event ($0.005) and returns only articles that are new since the previous run of that task, each billed as one `change` event ($0.0005) — a run with nothing new charges only the check. A news article never "changes" once Google has given it an id — monitor mode tracks each article by that id alone, so a `change` event always means a genuinely new article, never the same article re-billed because its resolved URL differed between runs.

### Related actors

- [Google Trends Scraper: Interest, Related Queries, Spikes](https://apify.com/datahamster/google-trends) — quantify whether a headline is actually driving search interest.
- [Reddit Scraper (subreddit posts, search, comments, no login)](https://apify.com/datahamster/reddit-posts) — see how a story is landing in community discussion, not just headlines.
- [Google Ads Transparency Scraper (advertiser, domain, alerts)](https://apify.com/datahamster/google-ads-transparency) — track advertiser activity alongside the news cycle.
- [Substack Scraper (posts, publications, likes, full text)](https://apify.com/datahamster/substack-posts) — for long-form takes instead of headline monitoring.

### How it works

1. Each entry in `queries` becomes either a Google News search URL (`news.google.com/rss/search?q=...`) or, if you paste a full `news.google.com` feed URL yourself, that feed is fetched as-is — this is how headline, topic and publication feeds are supported without a separate input field.
2. The feed's `<item>` entries are parsed into rows: title (with the trailing " - Source" Google appends stripped out), publisher, publish date and the Google redirect link.
3. With `resolveUrls: true`, each article's Google redirect is fetched once to try to reach the publisher's own URL. Google resolves most of these with client-side JavaScript, so a plain request often stays on the Google link — those rows come back with `resolved: false` and `url` equal to `googleUrl`, never as a missing or broken link.
4. A query that comes back as a non-feed page (consent wall, rate limit) on every proxy tier, with nothing pushed yet, ends the run as a block, not a silent empty dataset — checked even for a single-query run.

### Input

| Field | Meaning | Default |
|---|---|---|
| `queries` | Keywords, or a full Google News feed URL | `["web scraping"]` |
| `language` | Google News UI language (`hl`) | `"en"` |
| `country` | Google News edition (`gl`/`ceid`) | `"US"` |
| `maxItemsPerQuery` | Cap per query/feed (Google itself caps a feed at 100) | `100` |
| `resolveUrls` | Try to resolve the publisher URL per article | `false` |
| `sinceHours` | Only articles published within this many hours | none |
| `maxItems` | Stop after this many rows total | `200` |
| `mode` | `scrape` or `monitor` | `scrape` |
| `monitorStateId`, `webhookUrl`, `telegramBotToken`, `telegramChatId` | Monitor-mode state key and alert targets | empty |

### Pricing

| Event | Price |
|---|---|
| result | $0.0005 per article ($0.50 per 1,000) |
| monitor-check | $0.005 per monitor run |
| change | $0.0005 per new article |

Charged only for articles actually pushed. No proxy is required for the default configuration (tier `none`); a proxy tier is used automatically, and billed by Apify, only if Google ever answers with a rate limit or a consent wall.

**Found it useful?** A short review on the Store page helps other people find this actor and tells us what to improve. If a headline or alert looks wrong, open an issue on the actor page — issues are answered within a day.

### Why this actor

- No login, no cookies, no headless browser — reads Google's own public RSS feeds directly.
- Monitor mode with webhook/Telegram change alerts for a recurring keyword watch, not just a one-off pull.
- Accepts any Google News feed URL, not only plain search — headline, topic and publication (`source:`) feeds all go through the same field.
- A walled run is reported as a block with a reason, never a quietly empty dataset.

### Limits

- Google News RSS has no article summary/excerpt — `snippet` is the feed's own description with HTML tags stripped, which usually just repeats the headline and source.
- No author names or bylines — Google News RSS does not carry them, and none are added.
- A feed caps out at 100 items regardless of `maxItemsPerQuery`; older articles for a busy keyword are not reachable through this feed.
- `resolveUrls` is best-effort: Google resolves most article redirects with client-side JavaScript, so many rows stay on the Google link (`resolved: false`).

### FAQ

**Does it need a Google account or an API key?** No. Every request reads Google News' own public RSS feed, the same one a feed reader would use.

**What happens when a query has no articles?** No rows are pushed and no result events are charged. A real empty result (verified on the platform) is different from a wall: the `RUN_SUMMARY` record in the run's key-value store carries `emptyReason`, so a genuine zero-article query never looks like a block, and a block never looks like a genuine zero.

**Can I pull a specific publication's articles?** Yes — put `source:reuters.com` (or any Google News search operator) in `queries`; it goes through the same search endpoint as a plain keyword.

**What does monitor mode save me?** It keeps state per `monitorStateId` (or per saved task) across runs, so a schedule returns only new articles instead of the whole feed again — one `monitor-check` plus one `change` event per new article, not a full `result` per article every time.

### Changelog

- 0.1: initial release — search/headline/topic/publication feeds, `resolveUrls`, `sinceHours` filter, monitor mode.

***

If this actor saved you time, a short review on its Store page genuinely helps other people find it. Found a bug or need a field that is missing? Open a ticket on the **Issues** tab.

# Actor input Schema

## `maxItems` (type: `integer`):

Stop after this many results (you are charged only for pushed items)

## `mode` (type: `string`):

scrape = full results; monitor = only new/changed items since the previous run of this task

## `monitorStateId` (type: `string`):

Optional state id when not running as a saved task (monitor mode)

## `webhookUrl` (type: `string`):

POST a change summary here in monitor mode

## `telegramBotToken` (type: `string`):

Optional: bot token for monitor-mode change summaries

## `telegramChatId` (type: `string`):

Optional: chat id that receives monitor-mode summaries

## `queries` (type: `array`):

Keywords to search Google News for, e.g. "apify" or "web scraping". You can also paste a full Google News feed URL (headlines, topic or an already-built search URL) and it is fetched as-is, e.g. "https://news.google.com/rss?hl=en-US\&gl=US\&ceid=US:en" for top headlines.

## `language` (type: `string`):

Google News UI language code used to build the search URL (`hl`), e.g. "en".

## `country` (type: `string`):

Google News edition/country code used to build the search URL (`gl`, `ceid`), e.g. "US".

## `maxItemsPerQuery` (type: `integer`):

Stop reading a single query/feed after this many articles (Google News caps a feed at 100 items). Example: 20.

## `resolveUrls` (type: `boolean`):

true = follow each Google redirect link and try to return the publisher's own URL in `url` (one extra request per article; Google resolves most of these client-side with JavaScript, so many stay on the Google link with `resolved: false`). false = `url` is always the Google News link. Example: false.

## `sinceHours` (type: `integer`):

Only return articles published within this many hours. Leave empty for no time filter. Example: 24.

## Actor input object example

```json
{
  "maxItems": 200,
  "mode": "scrape",
  "queries": [
    "web scraping"
  ],
  "language": "en",
  "country": "US",
  "maxItemsPerQuery": 20,
  "resolveUrls": false
}
```

# Actor output Schema

## `results` (type: `string`):

All pushed rows (dataset, JSON)

## `resultsTable` (type: `string`):

Dataset in the Console viewer

## `runSummary` (type: `string`):

RUN\_SUMMARY record

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "web scraping"
    ],
    "language": "en",
    "country": "US",
    "maxItemsPerQuery": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("datahamster/google-news-feed").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["web scraping"],
    "language": "en",
    "country": "US",
    "maxItemsPerQuery": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("datahamster/google-news-feed").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "web scraping"
  ],
  "language": "en",
  "country": "US",
  "maxItemsPerQuery": 20
}' |
apify call datahamster/google-news-feed --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datahamster/google-news-feed"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/eLYBa5ULSamxyELvI/builds/UV8MwFdeUIxU8V5oY/openapi.json
