Google News Scraper avatar

Google News Scraper

Pricing

from $2.49 / 1,000 results

Go to Apify Store
Google News Scraper

Google News Scraper

Scrapes news articles from Google News search results

Pricing

from $2.49 / 1,000 results

Rating

0.0

(0)

Developer

ScrapeAI

ScrapeAI

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Categories

Share

What does Google News Scraper do?

Scrapes news articles from Google News search results It requests the live source at https://news.google.com/ and turns the returned page or response into structured dataset records. This actor is useful when you need to monitor news coverage by query, topic, date, and publisher. Results are parsed from the configured Google source, and each record includes run metadata when the source exposes a usable result.

Why use this actor?

Use this actor when a repeatable Apify run is more useful than manually checking Google. Inputs are exposed as a schema, results are written to an Apify Dataset, and the checked sample in input.json can be copied into the Console or API. The actor keeps result fields predictable while retaining source and run context in metadata where available.

What data can it collect?

The dataset fields depend on the selected Google result type. Typical records include the visible title, URL, summary, rank, source details, and actor-specific fields such as prices, ratings, dates, language, coordinates, or travel information. See the Output table below for the fields declared by this actor's dataset schema. A field can be null or absent when Google does not expose it for a particular result.

How to scrape Google News Scraper?

  1. Open the actor in Apify and create a new run.
  2. Enter the search, location, URL, language, date, or other values shown in the Input table.
  3. Keep the prefilled proxy configuration for normal runs; use an appropriate Apify Proxy group when the source requires it.
  4. Start the run and open the Dataset tab to export JSON, CSV, Excel, or RSS-compatible records.
  5. For automation, send the same JSON to the Apify API or schedule the actor from the Console.

A complete checked example is available in input.json:

{
"query": "javascript 2024",
"country": "us",
"language": "en",
"maxResults": 5,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": [
"RESIDENTIAL"
]
}
}

How much does it cost?

The total cost varies with run duration, compute usage, proxy traffic, result volume, and the Apify plan. Google pages can require retries or proxy requests, so a large pagination or expansion setting costs more than a single small query. Check the current estimate shown by Apify for the selected actor, proxy, and run before scaling it.

Input

All fields are optional unless marked required by the schema. Use the Apify Console form or pass JSON matching input_schema.json.

FieldTypeDescriptionPrefill / defaultRequired
Search QuerystringSearch query to scrape results fortechnologyNo
CountrystringCountry code for Google search (e.g. us, uk, de)usNo
LanguagestringLanguage code for search results (e.g. en, de, fr)enNo
Max ResultsintegerMaximum number of results to return100No
Proxy ConfigurationobjectProxy configuration. Use Apify Proxy with RESIDENTIAL group for best results.{"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]}No

Output

Each successful result is pushed to the Apify Dataset. The following fields are declared in .actor/dataset_schema.json:

FieldTypeDescription
TitlestringResult or item title.
UrlstringCanonical or destination URL.
SourcestringSource or publisher name.
Datestring | nullPublication or result date.
Descriptionstring | nullShort description or visible summary.
Thumbnailstring | nullThumbnail image URL.
Categorystring | nullCategory extracted from the source result when available.
RankintegerPosition in the returned results.
QuerystringInput query associated with the record.
MetadataobjectRun metadata such as source URL, crawl time, actor ID, and run ID.

Tips / Advanced options

  • Start with one focused query and a conservative result limit, then increase the limit or page/expansion settings after checking the output.
  • Keep the country, language, location, date, and device settings aligned with the search context you want to measure.
  • Use a stable proxy configuration for repeatable runs; changing geography can change the visible Google results.
  • Treat missing, null, or empty fields as normal because Google result layouts vary by query and over time.
  • Store the input and run ID with downstream records so a later run can be compared with an earlier snapshot.

FAQ, Disclaimers, and Support

Why did a run return fewer records? Google may show fewer results, a consent page, a CAPTCHA, a block page, or no matching items for the query. Try a narrower query, a supported location, or an appropriate proxy and review the run log.

Does this provide guaranteed or permanent data? No. Google page structure, availability, ranking, prices, and policies can change. Results are a point-in-time snapshot and should be validated before high-impact decisions.

Can I use the data commercially? Check Google's terms, the source page's terms, applicable privacy rules, and your Apify plan before collecting or redistributing data. Do not use the actor to bypass access controls.

Where can I get help? Review the run log and input schema first, then contact the actor owner through the Apify Console with the actor version, input JSON, run ID, and a short description of the problem.