# Google News Scraper (`scrapeai/google-news-actor`) Actor

Scrapes news articles from Google News search results

- **URL**: https://apify.com/scrapeai/google-news-actor.md
- **Developed by:** [ScrapeAI](https://apify.com/scrapeai) (community)
- **Categories:** News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.49 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Google News Scraper do?

Scrapes news articles from Google News search results It requests the live source at <https://news.google.com/> and turns the returned page or response into structured dataset records. This actor is useful when you need to monitor news coverage by query, topic, date, and publisher. Results are parsed from the configured Google source, and each record includes run metadata when the source exposes a usable result.

### Why use this actor?

Use this actor when a repeatable Apify run is more useful than manually checking Google. Inputs are exposed as a schema, results are written to an Apify Dataset, and the checked sample in `input.json` can be copied into the Console or API. The actor keeps result fields predictable while retaining source and run context in `metadata` where available.

### What data can it collect?

The dataset fields depend on the selected Google result type. Typical records include the visible title, URL, summary, rank, source details, and actor-specific fields such as prices, ratings, dates, language, coordinates, or travel information. See the Output table below for the fields declared by this actor's dataset schema. A field can be null or absent when Google does not expose it for a particular result.

### How to scrape Google News Scraper?

1. Open the actor in Apify and create a new run.
2. Enter the search, location, URL, language, date, or other values shown in the Input table.
3. Keep the prefilled proxy configuration for normal runs; use an appropriate Apify Proxy group when the source requires it.
4. Start the run and open the Dataset tab to export JSON, CSV, Excel, or RSS-compatible records.
5. For automation, send the same JSON to the Apify API or schedule the actor from the Console.

A complete checked example is available in `input.json`:

```json
{
  "query": "javascript 2024",
  "country": "us",
  "language": "en",
  "maxResults": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

### How much does it cost?

The total cost varies with run duration, compute usage, proxy traffic, result volume, and the Apify plan. Google pages can require retries or proxy requests, so a large pagination or expansion setting costs more than a single small query. Check the current estimate shown by Apify for the selected actor, proxy, and run before scaling it.

### Input

All fields are optional unless marked required by the schema. Use the Apify Console form or pass JSON matching `input_schema.json`.

| Field | Type | Description | Prefill / default | Required |
| --- | --- | --- | --- | --- |
| Search Query | `string` | Search query to scrape results for | technology | No |
| Country | `string` | Country code for Google search (e.g. us, uk, de) | us | No |
| Language | `string` | Language code for search results (e.g. en, de, fr) | en | No |
| Max Results | `integer` | Maximum number of results to return | 100 | No |
| Proxy Configuration | `object` | Proxy configuration. Use Apify Proxy with RESIDENTIAL group for best results. | {"useApifyProxy":true,"apifyProxyGroups":\["RESIDENTIAL"]} | No |

### Output

Each successful result is pushed to the Apify Dataset. The following fields are declared in `.actor/dataset_schema.json`:

| Field | Type | Description |
| --- | --- | --- |
| Title | `string` | Result or item title. |
| Url | `string` | Canonical or destination URL. |
| Source | `string` | Source or publisher name. |
| Date | `string \| null` | Publication or result date. |
| Description | `string \| null` | Short description or visible summary. |
| Thumbnail | `string \| null` | Thumbnail image URL. |
| Category | `string \| null` | Category extracted from the source result when available. |
| Rank | `integer` | Position in the returned results. |
| Query | `string` | Input query associated with the record. |
| Metadata | `object` | Run metadata such as source URL, crawl time, actor ID, and run ID. |

### Tips / Advanced options

- Start with one focused query and a conservative result limit, then increase the limit or page/expansion settings after checking the output.
- Keep the country, language, location, date, and device settings aligned with the search context you want to measure.
- Use a stable proxy configuration for repeatable runs; changing geography can change the visible Google results.
- Treat missing, null, or empty fields as normal because Google result layouts vary by query and over time.
- Store the input and run ID with downstream records so a later run can be compared with an earlier snapshot.

### FAQ, Disclaimers, and Support

**Why did a run return fewer records?** Google may show fewer results, a consent page, a CAPTCHA, a block page, or no matching items for the query. Try a narrower query, a supported location, or an appropriate proxy and review the run log.

**Does this provide guaranteed or permanent data?** No. Google page structure, availability, ranking, prices, and policies can change. Results are a point-in-time snapshot and should be validated before high-impact decisions.

**Can I use the data commercially?** Check Google's terms, the source page's terms, applicable privacy rules, and your Apify plan before collecting or redistributing data. Do not use the actor to bypass access controls.

**Where can I get help?** Review the run log and input schema first, then contact the actor owner through the Apify Console with the actor version, input JSON, run ID, and a short description of the problem.

# Changelog

This Actor's version history is a separate document: https://apify.com/scrapeai/google-news-actor/changelog.md

# Actor input Schema

## `query` (type: `string`):

Search query to scrape results for

## `country` (type: `string`):

Country code for Google search (e.g. us, uk, de)

## `language` (type: `string`):

Language code for search results (e.g. en, de, fr)

## `maxResults` (type: `integer`):

Maximum number of results to return

## `proxyConfiguration` (type: `object`):

Proxy configuration. Use Apify Proxy with RESIDENTIAL group for best results.

## Actor input object example

```json
{
  "query": "technology",
  "country": "us",
  "language": "en",
  "maxResults": 100,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset containing live actor results.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "technology",
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapeai/google-news-actor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "technology",
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapeai/google-news-actor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "technology",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call scrapeai/google-news-actor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapeai/google-news-actor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7lV8LLDIKcaBfMuGn/builds/4KdUzWK4dNwNkCdEi/openapi.json
