# Google News Scraper (`tacps126/google-news`) Actor

Scrape Google News: search results, topic sections and top stories in any language and country — with the publisher's real article URL, publisher, date and related coverage. Site and time filters, new-article monitoring, instant API.

- **URL**: https://apify.com/tacps126/google-news.md
- **Developed by:** [Tapaswai Ashok Choudhary](https://apify.com/tacps126) (community)
- **Categories:** News, SEO tools
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.60 / 1,000 articles

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google News Scraper

Get news articles from **Google News** — search results, topic sections (Business, Technology, Sports…) and top stories, in any language and country — as a clean, spreadsheet-ready dataset. Every article comes with its **real URL on the publisher's site** (not a Google redirect), the publisher, the publication time and related coverage from other outlets. Export to JSON, CSV or Excel, send to Google Sheets or Slack, or call it as an API.

**Built for** PR and communications teams tracking coverage, marketers and brand monitors, analysts and researchers following a topic, investors watching companies, and developers feeding news into apps and AI agents.

### Why this scraper

- **Real article links.** Google News hides articles behind redirect links; this Actor returns the publisher's own URL for each article, plus its domain — ready to open, share or crawl further.
- **Fast.** Dozens of articles with resolved links in a few seconds. No browser is started — the Actor reads Google News' own feeds directly.
- **Costs almost nothing to run.** A small native program that runs in 256 MB, so platform usage for a typical run is a fraction of a cent. You pay for articles, not machine time.
- **Precise searches.** Use Google News operators (`"exact phrase"`, `-exclude`, `OR`), limit to specific publishers, and pick a time window from the past hour to the past year.
- **Any edition.** Choose the language and country — English for the US, UK or India, German, French, Spanish, Hindi, Japanese and more.
- **News alerts built in.** Schedule it with *Only new articles* and each run returns only articles you haven't received yet.
- **Handles rate limits for you.** When Google says "slow down" (HTTP 429), the Actor waits as long as asked and carries on; timeouts and dropped connections are retried automatically.
- **Never loses or double-bills work.** Results are saved as each query finishes; if a run is moved to another server it resumes without charging twice, and it stops cleanly at your maximum charge.

### How to use it

1. Add **Search queries** (e.g. `artificial intelligence`, `"Tesla" earnings`) and/or pick **Topics**.
2. Optionally limit to **sites**, choose **Published within**, and set **Language** and **Country**.
3. Press **Start**. Download results from the **Output** tab, or connect an integration.

#### Example input

```json
{
  "queries": ["artificial intelligence", "\"OpenAI\" -stock"],
  "sites": ["reuters.com", "bbc.co.uk"],
  "timeframe": "7d",
  "topics": ["TECHNOLOGY"],
  "language": "en",
  "country": "US",
  "maxItems": 200
}
```

### What you get

| Field | Description |
|-------|-------------|
| `title` | Headline |
| `publisher` / `publisherUrl` | News outlet and its homepage |
| `url` / `domain` | The article on the publisher's site, and its domain |
| `googleNewsUrl` | The Google News link for the article |
| `publishedAt` | Publication time (ISO 8601, UTC) |
| `relatedArticles` | Other outlets' coverage of the same story (topic and top-story feeds): title, publisher, link |
| `query` / `topic` | Which search or topic found it |
| `position` | Rank in that feed |
| `language` / `country` | Edition it came from |
| `urlResolved` / `scrapedAt` | Whether the publisher URL was found, and when the article was collected |

Example:

```json
{
  "title": "OpenAI shelves new AI model release over safety concerns",
  "publisher": "reuters.com",
  "url": "https://www.reuters.com/business/openai-shelves-new-ai-model-after-...",
  "domain": "reuters.com",
  "publishedAt": "2026-09-28T22:34:00+00:00",
  "query": "openai",
  "position": 2,
  "language": "en",
  "country": "US",
  "urlResolved": true
}
```

### News alerts on a schedule

Turn on **Only new articles** and schedule the Actor (hourly, daily). Each run returns only articles earlier runs haven't delivered — you're charged only for new ones. Connect Slack, email, Google Sheets, Zapier, Make or a webhook to get them pushed to you.

### Use it as an instant API

The Actor also runs in **Standby mode** — an always-ready endpoint that returns articles straight in the response, handy for apps and AI agents. Find the URL on the **Standby** tab:

```bash
curl "https://<standby-url>/?q=artificial%20intelligence&timeframe=1d&maxItems=20" \
  -H "Authorization: Bearer <YOUR_APIFY_TOKEN>"
```

Any input field works as a query parameter (`q` is shorthand for one search), or `POST` the full input as JSON. The response is `{"success": true, "count": 20, "items": [ … ]}`.

### Pricing

Pay per result: about **$2 per 1,000 articles** (lower on higher Apify plans) plus a tiny start fee. Duplicates and (with *Only new articles*) already-delivered articles are free. See the **Pricing** tab for exact rates, and set a maximum charge per run to cap spend.

On Apify's **free plan** each run returns up to 500 articles — plenty to try everything. Any paid Apify plan removes the limit.

### FAQ

**How many articles can I get per search?** Google News shows up to about 100 articles for one search or topic. For more, split your search (by keywords, sites or time window) into several queries.

**Can I get the full article text?** This Actor returns the article's metadata and real URL. To fetch full text, pass the URLs to a website content crawler.

**Is it legal?** The Actor collects publicly listed headlines and links. You're responsible for how you use them and for respecting publishers' copyrights.

### Support

Found a problem or need a field that isn't there yet? Open an issue on the **Issues** tab with your input and what you expected — requests genuinely decide what gets built next.

# Actor input Schema

## `queries` (type: `array`):

What to search for, one per line. Google News operators work: `"exact phrase"`, `-exclude`, `OR`, `intitle:word`.

## `topics` (type: `array`):

Topic sections to include, as shown on Google News. Can be combined with search queries.

## `sites` (type: `array`):

Limit searches to these publishers' domains, e.g. `reuters.com`, `bbc.co.uk`. With no search query, returns everything recent from these sites.

## `excludeSites` (type: `array`):

Exclude these publishers' domains from searches, e.g. `example-tabloid.com`.

## `excludeKeywords` (type: `array`):

Skip articles whose headline contains any of these words (case-insensitive), e.g. `stock`, `opinion`.

## `timeframe` (type: `string`):

Only articles published in this window (applies to searches).

## `dateFrom` (type: `string`):

Only articles published on or after this date. Use instead of “Published within” for a fixed period.

## `dateTo` (type: `string`):

Only articles published before this date.

## `language` (type: `string`):

Edition language as a 2-letter code, e.g. `en`, `de`, `fr`, `es`, `hi`, `ja`.

## `country` (type: `string`):

Edition country as a 2-letter code, e.g. `US`, `GB`, `IN`, `DE`, `AU`.

## `resolveUrls` (type: `boolean`):

Return each article's real URL on the publisher's site instead of a Google News redirect link. Recommended.

## `includeRelated` (type: `boolean`):

Add other outlets' coverage of the same story (topic and top-story feeds).

## `maxItems` (type: `integer`):

Maximum number of articles to return in total.

## `maxItemsPerQuery` (type: `integer`):

Optional cap per search or topic (Google News shows up to about 100 each).

## `onlyNew` (type: `boolean`):

Return only articles that no earlier run with the same input has delivered. Ideal for news alerts — you pay only for new articles.

## `monitorName` (type: `string`):

Optional name for the “only new” history, so several monitors can run side by side.

## `proxyConfiguration` (type: `object`):

Not needed for most runs. Turn on for very frequent or very large runs.

## `dedupe` (type: `boolean`):

Drop the same article appearing in several queries or topics.

## `tenantId` (type: `string`):

Optional label used to keep separate histories for different clients or teams.

## `debug` (type: `boolean`):

Save extra diagnostics with the run (for support requests).

## Actor input object example

```json
{
  "queries": [
    "artificial intelligence"
  ],
  "timeframe": "",
  "language": "en",
  "country": "US",
  "resolveUrls": true,
  "includeRelated": true,
  "maxItems": 200,
  "onlyNew": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "dedupe": true,
  "debug": false
}
```

# Actor output Schema

## `articles` (type: `string`):

Every article found, in one consistent format.

## `summary` (type: `string`):

Counts, warnings, and whether the run stopped early.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "artificial intelligence"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tacps126/google-news").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": ["artificial intelligence"] }

# Run the Actor and wait for it to finish
run = client.actor("tacps126/google-news").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "artificial intelligence"
  ]
}' |
apify call tacps126/google-news --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tacps126/google-news"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5PgHMPt14b1gnAPTw/builds/sAiwY6qMZTcdZvFsh/openapi.json
