Google News Scraper - Real Publisher URLs, Not Redirects avatar

Google News Scraper - Real Publisher URLs, Not Redirects

Pricing

from $2.85 / 1,000 articles

Go to Apify Store
Google News Scraper - Real Publisher URLs, Not Redirects

Google News Scraper - Real Publisher URLs, Not Redirects

Track Google News by keyword or topic, one row per article: title, publisher, date, snippet and the publisher's real link, decoded from Google's redirect, with links that would not decode marked. Full text, and a summary on your own OpenAI key, are optional. $3.00 per 1,000 articles.

Pricing

from $2.85 / 1,000 articles

Rating

5.0

(1)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

6

Total users

2

Monthly active users

2 days ago

Last modified

Share

Give it a keyword or pick a Google News topic feed, and get one row per article: headline, publisher, date, snippet and a link. Optionally the full body text, and a short summary with a sentiment label if you bring your own OpenAI key.

Here is the part that decides whether you need this. Every link in a Google News feed is a redirect blob, https://news.google.com/rss/articles/CBMi..., and the publisher's actual address has to be decoded out of it. That decoding is best effort. When a link will not decode, the row still arrives with the Google link in url and urlResolved: false, so you always know which one you are holding.

InputA keyword search, or one of eight Google News topic feeds
OutputOne row per article: headline, publisher, date, snippet and the publisher's own link
Ceiling100 articles per run, and Google's feeds hold about 100 anyway
Account neededNone. An OpenAI key only if you want the summaries
Price$3.00 per 1,000 articles, flat on every plan. The free plan's $5 a month covers about 1,600

๐Ÿ” What Google News Scraper does

It reads a Google News feed and writes a row per article, with the original Google link kept beside the decoded one in case you want it.

Search or topic. A query takes Google operators, so tesla OR rivian, site:reuters.com and intitle:layoffs all work. A topic instead gives you the edition feed for World, Nation, Business, Technology, Entertainment, Sports, Science or Health.

Full text, if you ask. With fetchArticleText on, each resolved article is downloaded and its body pulled out, up to 20,000 characters. Plenty of publishers will not give that up to a plain request, and those rows come back with articleText: null and are not charged.

Summaries, on your key. aiSummary adds a one or two sentence summary and a sentiment label, using the OpenAI key you supply. OpenAI bills you for those calls directly. This actor adds nothing for them.

๐Ÿ“‹ What data you get from each Google News article

What you getField
Headlinetitle
Publisher name and homepagesource, sourceUrl
When it was publishedpublishedAt
The snippet from the feedsnippet
The publisher's own link, and whether it decodedurl, urlResolved
The original Google redirect linkgoogleUrl
The article body, with full text onarticleText
A short summary and a sentiment label, with AI summary onaiSummary, sentiment

โ–ถ๏ธ How to scrape Google News

  1. Open Google News Scraper and click Try for free.
  2. Put keywords into Search query, or leave it empty and pick a Topic instead.
  3. Set Max articles, and Freshness if you only want recent pieces.
  4. Tick Fetch full article text only if you need bodies, then click Start.
  5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

๐Ÿ’ฐ How much does it cost to scrape Google News?

$3.00 per 1,000 articles, which is $0.003 each. The same rate on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 1,600 articles.

Resolving links, fetching article text and the AI summary all cost nothing extra here. The OpenAI calls behind aiSummary are billed to you by OpenAI, on your own key.

What is not charged: a row where you asked for the body or the summary and it could not be had, every diagnostic row, and a search that finds nothing.

๐Ÿ“ฅ What you give it

{
"query": "openai funding",
"freshness": "7d",
"language": "en-US",
"country": "US",
"maxItems": 30,
"resolveUrls": true
}
FieldDefaultWhat it is
querynone, the box starts at artificial intelligenceKeywords, with Google operators. Leave empty if you are using a topic.
topicnoneA topic feed instead of a search: World, Nation, Business, Technology, Entertainment, Sports, Science, Health.
freshnessany time1h, 1d, 7d, 30d or 1y. It narrows a query search. A topic feed returns whatever the feed holds.
languageen-USGoogle's hl. en-GB, fr, de, es-419 and the rest.
countryUSGoogle's gl. GB, CA, DE, IN and so on.
maxItems50Articles to return, up to 100.
resolveUrlsonDecode the Google redirect into the publisher's own link.
fetchArticleTextoffDownload each article and pull out the body. Needs resolveUrls on.
aiSummaryoffSummary and sentiment per article. Needs your own key in openaiApiKey.
openaiApiKeynoneRead only when aiSummary is on. Stored as a secret.
notionConnectornoneOptional. Write each article into your Notion when the run finishes.
notionParentIdnoneOptional. The Notion data source to write into.
proxyConfigurationoffOptional, and a normal run does not want it.

Send neither a query nor a topic and you get one uncharged BAD_INPUT row telling you which is missing. Same if aiSummary is on without a key.

๐Ÿ“ค What you get back

A real row from a run on 16 September 2026, with the long Google link cut short:

{
"ok": true,
"title": "Kapolri Dorong Pekerja Tingkatkan Kompetensi dan Penguasaan Teknologi - Tribratanews Polda Metro Jaya",
"publishedAt": "2026-09-16T07:04:36.000Z",
"source": "Tribratanews Polda Metro Jaya",
"sourceUrl": "https://tribratanews.metro.polri.go.id",
"snippet": "Kapolri Dorong Pekerja Tingkatkan Kompetensi dan Penguasaan Teknologi Tribratanews Polda Metro Jaya",
"url": "https://tribratanews.metro.polri.go.id/kapolri-dorong-pekerja-tingkatkan-kompetensi-dan-penguasaan-teknologi/",
"urlResolved": true,
"articleText": null,
"googleUrl": "https://news.google.com/rss/articles/CBMirgFBVV95cUxNbUp3aDlPMFNNNG5Qa3A3QlV0WFVKVjJHLU5GWExSal9qdEcz..."
}
FieldHow to read it
urlThe publisher's own link once decoded, or the Google link when it would not decode.
urlResolvedtrue means url is the publisher's address. false means it is still the Google redirect. Check it per row.
source, sourceUrlPublisher name and homepage, as Google's feed gives them. null when the feed leaves them out.
articleTextOnly with fetchArticleText on, and only when the body could be pulled out. Capped at 20,000 characters.
aiSummary, sentimentOnly with aiSummary on and the call came back. When the call fails, they stay empty and the row is marked ok: false.

๐Ÿงพ Reading the output

RowHow to spot itCharged
A complete articleok: trueyes
An article missing the text or summary you asked forok: false, no errorCodeno
A diagnosticok: false and an errorCodeno

There is no sample row on this one. An empty input goes straight to BAD_INPUT.

CodeWhat it means
BAD_INPUTNo query and no topic, or aiSummary on with no key.
NO_RESULTSThe feed came back empty. Widen freshness, or try the keyword without operators.
NOT_FOUNDThat feed address does not exist. Usually a bad language or country pairing.
BLOCKEDThat feed could not be read this run. Re-run it.
RATE_LIMITEDToo many requests in a short window. Wait and re-run.
SERVER_ERRORGoogle answered with an error of its own.
NETWORKGoogle was unreachable or the response never finished.

The Overview table hides urlResolved, ok, articleText and aiSummary, which are exactly the fields you want to check. Switch the dataset to JSON to see them.

๐Ÿ’ก What people use it for

  • Watching a brand, a competitor's product name or a ticker across publishers, on a daily schedule.
  • Building a reading list for one topic feed in one country, with real links you can open.
  • Pulling headline, date and publisher into a sheet for a weekly roundup, skipping the body entirely.
  • Feeding resolved article URLs into your own pipeline, since googleUrl is useless to most tools.

๐Ÿšง What it does not do

  • About 100 articles per feed. That is Google's cap, not ours. Split a bigger job by keyword, by site:, or by freshness window.
  • Full text is a gamble per article. Paywalls and pages that build themselves in the browser give a plain request nothing. Those rows are flagged and not charged, so do not budget on getting the body for every article.
  • Some links never decode. You get the Google link and urlResolved: false rather than a dropped row.
  • freshness narrows a search, not a topic feed. Pick a topic and you get what that feed holds.
  • No archive. Google News feeds carry what is current. There is no way to ask for last March.
  • The Notion connector receives every row, including the incomplete ones. Filter on ok if you only want complete articles in there.
  • One query or one topic per run. Use a schedule or several runs for a list of keywords.

๐Ÿงญ Which Google scraper do you need?

If you wantUse
The web results page for a keyword: organic, ads, People Also AskGoogle Search Results Scraper
What a search box suggests as people typeKeyword Autocomplete Scraper
Interest in a keyword over time and by regionGoogle Trends Scraper
News articles with the publisher's real linkThis one
Image results and the pages they sit onGoogle Images Scraper
The ads a brand runs on GoogleGoogle Ads Transparency Scraper
World news in 65+ languages, with each outlet's countryGDELT News Scraper

โ“ Questions people ask

Do I need a Google API key to scrape Google News?

No. Nothing to apply for, nothing to log into.

Do I need an OpenAI key?

Only for aiSummary. Leave that off and the actor never touches OpenAI.

Why is url sometimes still a news.google.com address?

Because that particular link would not decode. urlResolved tells you, per row, and googleUrl always holds the original.

Why did I get fewer articles than maxItems?

The feed ran out. A narrow keyword with a tight freshness window often returns a handful.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/google-news-scraper. Either way the run happens on your Apify account at the same price.

These are public news feeds, and the rows point back at the publishers. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

๐Ÿ†˜ If something breaks

Open the Issues tab on the actor page. Send the query or topic and the run ID. The errorCode on the diagnostic row usually names the problem on its own.