Google News Scraper - Full Article Text, Alerts & RSS avatar

Google News Scraper - Full Article Text, Alerts & RSS

Pricing

from $3.50 / 1,000 results

Go to Apify Store
Google News Scraper - Full Article Text, Alerts & RSS

Google News Scraper - Full Article Text, Alerts & RSS

Scrape Google News by keyword: headline, publisher, real article link, date and full article text, with only new articles on each scheduled run. Add RSS or Atom feeds to the same table. No API key. Export to CSV, JSON or Excel, schedule runs or use the API.

Pricing

from $3.50 / 1,000 results

Rating

0.0

(0)

Developer

Automly

Automly

Maintained by Community

Actor stats

0

Bookmarked

5

Total users

2

Monthly active users

2 days ago

Last modified

Categories

Share

What is Google News Scraper?

Google News Scraper is a tool that lets you scrape news articles from Google News by keyword: headline, the publisher's real link, publisher, publication date, summary, image and, if you want it, the full article text. On repeat runs it gives you only the stories you haven't seen yet. Add the GDELT global news index and any RSS or Atom feed to the same table. Type a keyword, click Start, and download the articles as Excel, CSV or JSON.

  • 🔗 Real publisher links: the publisher's own URL, not a news.google.com/rss/articles/CBMi… redirect
  • 📄 Full article text: the article body, not just a headline and two lines of summary
  • 🔔 Only new articles: a scheduled run reports only the stories earlier runs didn't deliver
  • 🌍 146 languages and 235 countries: pick the Google News edition from a dropdown
  • 🔓 No login: no API key, account, login or cookies for any of the sources
  • 💸 Free to try: the $5 of free usage every Apify account gets each month covers up to about 1,400 articles

What can Google News Scraper do?

  • Scrape Google News by keyword and get the publisher's own link
  • Get the full text of news articles, not only the summary
  • Get more than 100 Google News results for one query by adding a date range
  • Google Alerts alternative: monitor a keyword and get only new articles on every run
  • Filter news by publisher domain and by whole words, so car never matches cargo
  • Get news in a language or country of your choice
  • Read the GDELT news index and any RSS or Atom feed into the same table, with BBC, CNN, NPR, The New York Times, The Guardian, Google News and Hacker News ready to pick
  • Remove duplicate stories that appear in several sources
  • Export to Excel, CSV, JSON, HTML or XML, or send the articles to Google Sheets, Slack, Make, Zapier and more

What data can you extract from Google News?

📰 Headline🔗 Publisher URL🧹 Clean canonical URL
🏢 Publisher name and domain📅 Publication date📝 Summary
✍️ Author (RSS)🖼️ Image🏷️ Categories (RSS)
🌐 Language and country (GDELT)🔎 The query that found it📄 Full article text and length (optional)

How to scrape Google News

  1. Create a free Apify account (no credit card needed).
  2. Open Google News Scraper and type a keyword into Search queries, say tesla recall.
  3. Choose your Sources (google-news, gdelt, rss), and add Preset feeds or your own feed URLs if you want them.
  4. Set Maximum results, turn on Include full text if you need the article body, and click Start.
  5. When the run finishes, download the articles as Excel, CSV, JSON, HTML or XML.

Three switches matter most: Include full text adds the article body, Resolve Google News URLs swaps Google's redirect links for publisher links (full text does this for you), and Skip already seen articles makes each scheduled run return only what the last one missed.

How much does it cost to scrape Google News?

You pay $3.50 per 1,000 articles, plus a start fee of $0.001 per run at the default settings. You also pay the Apify platform usage of the run, separately from these prices.

For example, 1,000 articles cost $3.50 plus the $0.001 start fee, plus the platform usage of that run.

Every Apify account gets $5 of free usage each month, enough for up to about 1,400 articles (platform usage comes out of the same $5). See the Pricing tab for details.

⬇️ Input

SettingWhat it does
Search queriesSearch terms, one per line, up to 100. Sent to every search source you enabled
RSS/Atom feed URLsFeed URLs to read directly, up to 200
Preset feedsWell-known feeds by name: bbc, cnn, npr, nyt, guardian, google-news-top, hacker-news
SourcesAny of google-news, gdelt, rss (default google-news)
Maximum resultsHow many articles to collect in total, 1 to 50,000 (default 100)
Published within the last N hoursOnly articles from the last N hours, 1 to 8,760
From date / To dateYYYY-MM-DD bounds. A date range wins over the last-N-hours setting
Include keywordsKeep articles whose headline or summary contains one of these. Whole words, any case
Exclude keywordsDrop articles containing any of these. Exclusions win over includes
Include domainsKeep only these publishers, like reuters.com. Subdomains count, and www. is ignored
Exclude domainsDrop these publishers and anything under them
LanguagePick a language. It sets the Google News edition and drops rows from sources that report a different language
CountryPick a country. It sets the Google News edition and drops rows from sources that report a different country
Resolve Google News URLsTurn Google News redirect links into real publisher links
Include full textGet each article's body text. Turns on link resolving for you. Some publishers don't share their article text, so a few rows keep their summary and leave the body empty
Skip already seen articlesReturn only articles that earlier runs didn't deliver
Seen articles store nameNames the list of delivered articles. Use a separate one per monitored query

Connection settings: the default works for most runs.

You need at least one search query, feed URL or preset feed. Adding a feed or a preset switches the rss source on by itself.

Example: Google News and GDELT plus two preset feeds, with real publisher links.

{
"queries": ["tesla recall"],
"sources": ["google-news", "gdelt"],
"presetFeeds": ["bbc", "guardian"],
"maxResults": 50,
"includeKeywords": ["tesla"],
"resolveGoogleUrls": true
}

⬆️ Output

You get one row per article, nineteen fields. You can view them as a table in Apify Console or download them as Excel, CSV, JSON, HTML or XML. This is a real row from a run on 12 September 2026:

{
"title": "Lawmaker Demand Feds Take Action After 43 Tesla Drivers Filmed Asleep At The Wheel",
"url": "https://insideevs.com:443/news/807893/tesla-sleeping-congress-nhtsa-demand/",
"canonicalUrl": "https://insideevs.com/news/807893/tesla-sleeping-congress-nhtsa-demand",
"source": "gdelt",
"sourceName": "insideevs.com",
"sourceDomain": "insideevs.com",
"publishedAt": "2026-09-11T23:15:00+00:00",
"summary": null,
"author": null,
"imageUrl": "https://cdn.motor1.com/images/mgl/W87NGL/s1/tesla-fsd-with-a-dead-battery.jpg",
"categories": [],
"language": "English",
"country": "United States",
"query": "tesla recall",
"guid": null,
"resolved": true,
"fullText": null,
"fullTextChars": null,
"scrapedAt": "2026-09-12T12:52:24.373597+00:00"
}

canonicalUrl is the field duplicates are matched on. It lowercases the host, drops www., the fragment and a redundant port, and strips the usual tracking parameters such as utm_*, at_*, fbclid and gclid. That's why the same story reaching you from two different feeds still counts once.

Switch on Resolve Google News URLs and a Google News row carries the publisher's own link instead:

{
"title": "Driver crashes Tesla into scaffolding in midtown Manhattan, killing passenger: Police - ABC News",
"url": "https://abcnews.com/US/tesla-crashes-scaffolding-midtown-manhattan-killing-passenger/story?id=136300943",
"source": "google-news",
"sourceDomain": "abcnews.com",
"resolved": true
}

Switch on Include full text and each row gains the body:

{
"title": "Tesla and others begin record vehicle recall in China - Reuters",
"url": "https://www.reuters.com/world/tesla-fix-software-millions-china-made-imported-evs-china-2026-08",
"sourceDomain": "reuters.com",
"fullTextChars": 4258,
"fullText": "BEIJING, Aug 21 (Reuters) - Tesla and eight other automakers said on Friday they will recall a total of about 4.3 million vehicles in China over concerns that doors may be difficult to open…"
}

The three sources don't publish the same things, so some columns are fuller than others:

SourceYou getYou don't
rssheadline, link, date, summary, guid, and author, image or categories when the feed includes themlanguage, country
google-newsheadline, link, date, summary, guid, publisher domainauthor, categories, language, country
gdeltheadline, link, date, image, language, countrysummary, author, guid

fullText and fullTextChars only appear when you ask for full text.

How can I use Google News data?

  • Brand and competitor monitoring: watch for mentions, with one row per story instead of the same piece three times
  • Media monitoring and PR reporting: the publisher, the date and a link that stays stable
  • Research datasets: article text across a stretch of dates
  • News alerts: schedule it, get only what's new, and send it to Slack or a webhook

How many Google News articles can you get per query?

Google News gives you about 100 articles for a single query, and GDELT at most 250. To get more, ask for more than that and give a date range: each day of the range can then add up to about another hundred articles, so a week of history can return several hundred rather than one hundred. Ask for 500 with no date range and you'll still hit the ceiling.

When you pick several sources, you get a fair share from each. Ask for 60 across three sources and you'll get roughly 20 from each, instead of 60 from just one.

Monitoring: only new articles on every run

Switch on Skip already seen articles and every run returns only the articles no earlier run delivered, which is what makes polling on a schedule worthwhile. In testing, a first run stored 60 articles and an identical second run stored 34. Give each query you monitor its own Seen articles store name so they keep separate histories.

Use it as a Google News API

One request returns the rows directly, from any language:

curl -X POST "https://api.apify.com/v2/acts/automly~google-news-scraper-api/run-sync-get-dataset-items" \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
-H 'Content-Type: application/json' \
-d '{
"queries": ["tesla recall"],
"sources": ["google-news", "gdelt"],
"maxResults": 50,
"resolveGoogleUrls": true
}'

You can also use the Python and JavaScript clients, or schedule the same run in Apify Console.

What it doesn't do

No sentiment scoring, no entity extraction, no translation, no paywall bypass.

Scrape more news and social data

ActorWhat it gets
News API - Multi-Source Headlines & ArticlesToday's headlines from BBC, CNN, NPR, NYT, The Guardian, Google News and Hacker News in one table
Hacker News Trends APIHacker News stories and comments by keyword, points and date
Reddit Keyword + Comment Thread MonitorNew Reddit posts and comments that mention your keywords
Reddit Scraper - Posts, Comments, Users, Subreddits & SearchReddit posts, comments and keyword search results
LinkedIn Post Search Scraper (No Login)Public LinkedIn posts by keyword with author, date and likes

🤖 Use with AI agents (MCP)

This actor works as a tool for AI agents. Connect any MCP-compatible AI assistant or agent to the Apify MCP server at mcp.apify.com (sign in with your Apify account), and it can find this actor, fill in the input and read the results by itself. Ask in plain words, for example "Collect the news about the Tesla recall from the last 24 hours", and the agent runs automly/google-news-scraper-api with an input like {"queries": ["tesla recall"], "sinceHours": 24, "maxResults": 100}. Every input and output field has a plain description, so the agent knows what to type and what each column means. Agent runs cost the same as runs you start yourself.

❓FAQ

How much does it cost to scrape Google News?

$3.50 per 1,000 articles plus a $0.001 start fee, and the run's platform usage is paid separately. See How much does it cost to scrape Google News? above.

Do I need an API key for Google News or GDELT?

No. You need no API key, account, login or cookies for any of the sources.

Can I get more than 100 articles from one Google News query?

Yes. Set Maximum results above 100 and give it a From date and To date. Each day of the range can add up to about another hundred articles.

Because Google News publishes encoded redirect links, and following one lands you back inside Google rather than on the article. Switch on Resolve Google News URLs and every link becomes the publisher's own. That also lets the same story be matched against GDELT and RSS copies when duplicates are removed.

How do I get the full article text and not just the summary?

Switch on Include full text. Each article's main text is pulled out (a Reuters piece came back at 4,258 characters), and Google News links are resolved along the way so the body can be reached.

Expect a few rows to come back without text. On one 80-article run, 71 had text and 9 didn't, because those publishers did not share the article text. Those rows keep their summary, and the body is left empty rather than filled with navigation text. The run log tells you the split, so you're not left guessing why a column has gaps.

Is the same article returned twice if two sources carry it?

Not once Google News links are resolved. Duplicates are matched on canonicalUrl, so one story counts once across all three sources. The exception is a Google News row you left unresolved: it has no publisher link to compare against, so it can't be matched with a GDELT or RSS copy of the same piece. Switching on Resolve Google News URLs brings those rows in too.

Do the keyword filters search the full article text?

No. Include keywords and Exclude keywords check the headline and summary.

Can I use it instead of Google Alerts?

For keyword monitoring, yes. Schedule it, switch on Skip already seen articles, and connect the run to Slack, email or a webhook through Apify integrations. Each run then delivers only the stories earlier runs didn't, with the publisher link, the date and optionally the full text.

What output formats are supported?

JSON, CSV, Excel, XML and JSONL, from the Storage tab or the Apify API, plus Apify integrations for Google Sheets, Slack, webhooks and S3.

Can I connect Google News Scraper to other tools or AI agents?

Yes. Start runs and download results with the Apify API or the Python and JavaScript clients, send results to Make, Zapier, n8n, Google Sheets or Slack with Apify integrations and webhooks. AI assistants and agents can run it through the Apify MCP server, see the Use with AI agents (MCP) section above.

Google News Scraper only collects publicly available article details without logging in, and full text comes from the publisher's own public page. Articles can name people, and personal data may be protected by laws such as GDPR, so only scrape for a legitimate reason, respect the publishers' copyright, and ask a lawyer if you are unsure. You can read more in Is web scraping legal?

Google News Scraper is an independent tool. It is not affiliated with, endorsed by or sponsored by Google.

Something isn't working. What should I do?

Open the Issues tab and tell us your input and what you expected.

⭐ Your feedback

Have an idea or found a problem? Tell us on the Issues tab. If Google News Scraper saved you time, a short review on the Reviews tab helps other people find it.