Google News API Scraper – RSS, Real URLs, News Monitoring avatar

Google News API Scraper – RSS, Real URLs, News Monitoring

Pricing

from $1.00 / 1,000 articles

Go to Apify Store
Google News API Scraper – RSS, Real URLs, News Monitoring

Google News API Scraper – RSS, Real URLs, News Monitoring

Search Google News RSS by keyword, topic or publisher and get structured articles: headline (raw and with the " - Publisher" suffix stripped), publisher name and domain, date, feed rank, resolved publisher URL, related coverage of the same story, and optional full article text. Pay per article.

Pricing

from $1.00 / 1,000 articles

Rating

0.0

(0)

Developer

Fetch Smith

Fetch Smith

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

14 hours ago

Last modified

Categories

Share

Google News Scraper

Search Google News and get clean, structured articles as JSON, CSV or Excel: headline (raw and with the - Publisher suffix stripped), publisher name and domain, publish time, feed rank and the real publisher URL (Google's encoded redirect links are resolved for you). Optionally extract the full article text — body, author, image, keywords and section — from the publisher's page. Works for any language and country. Pay only per article returned.

Use cases

  • Media monitoring and brand mentions for any keyword, in any market
  • Feeding news into AI agents, RAG pipelines and newsletters
  • Competitor and industry tracking (site: and when:7d operators supported)
  • Building datasets of headlines by topic, region or publisher

Input

FieldTypeDescription
queriesarraySearch terms. Operators work: "exact phrase", site:reuters.com, when:7d, before:2026-01-01, after:2026-06-01
rssUrlsarrayOptional Google News RSS feed URLs (topics, sections, publications)
topicsarrayOptional: browse Google News' built-in sections without knowing an RSS URL — WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SCIENCE, SPORTS, HEALTH, plus the narrower sub-sections POLITICS, ECONOMY, REAL_ESTATE, JOBS, EDUCATION, AUTOS, MOVIES, MUSIC, CELEBRITIES, ARTS, SOCCER, BASKETBALL
excludeWordsarrayWords/phrases to drop from every query, e.g. ["iphone"] on a query apple removes iPhone coverage. Same effect as typing -word yourself, just a manageable list. Applies to queries only
siteFilterarrayRestrict every query to these publisher domains, e.g. ["nytimes.com", "reuters.com"] (OR'd, filtered by Google itself — no extra requests, no under-filled results). Applies to queries only
excludeSitesarrayDrop results from these publisher domains, e.g. ["pinterest.com"]. Applies to queries only
timePeriodstringOnly articles published within this window: 1h, 6h, 12h, 1d, 7d, 30d, 90d, 1y. Filtered by Google itself, so it costs no extra requests and you are never charged for articles outside the window. Applies to queries only
publishedAfterstringYYYY-MM-DD — only articles published on or after this date. Overrides timePeriod. Applies to queries only
publishedBeforestringYYYY-MM-DD — only articles published before this date. Overrides timePeriod. Applies to queries only
languagestringhl code such as en-US, de, fr, pt-BR, ar, ja (default en-US)
countrystringgl code such as US, GB, DE, IN (default US)
maxItemsPerQueryintegerUp to 100 (Google's feed limit). Default 25 — decoding the real publisher URL is a 2-request round trip per article, and Google's per-IP rate limit on that endpoint has gotten slower recently, so 50+ articles can now take 4-5 minutes
decodeUrlsbooleanResolve the publisher URL for each article (default true)
fetchArticleBodybooleanOpen each publisher page and extract the full article text, author, image, keywords and section (default false)
articleBodyMaxCharsintegerTruncate articleBody to this length (default 20000)
extractTickersbooleanPull stock tickers into a tickers array from the title (and article body, if fetchArticleBody is on). Rule-based, no extra request, costs nothing extra (default false)
maxResultsintegerTotal cap across queries
proxyConfigurationobjectApify Proxy for Google/publisher requests (default: on). Rotating IPs is the real fix for Google's URL-decode rate limit — see FAQ

Output (one item per article)

{
"title": "Apify raises new funding to scale web data platform - TechCrunch",
"titleClean": "Apify raises new funding to scale web data platform",
"url": "https://techcrunch.com/2026/09/02/apify-funding/",
"googleNewsUrl": "https://news.google.com/rss/articles/CBMi...",
"source": "TechCrunch",
"sourceUrl": "https://techcrunch.com",
"sourceDomain": "techcrunch.com",
"publishedAt": "2026-09-02T07:00:00.000Z",
"snippet": "Apify raises new funding to scale web data platform",
"relatedArticles": [
{ "title": "Apify closes funding round", "source": "Reuters", "googleNewsUrl": "https://news.google.com/rss/articles/CBMi..." }
],
"position": 1,
"query": "apify",
"topic": null,
"language": "en-US",
"country": "US"
}

A few of these are worth knowing about:

  • titleClean — Google News always appends - Publisher to the headline. title keeps it exactly as Google sends it; titleClean is the same headline with that suffix removed, which is what you usually want in a dashboard or digest.
  • sourceDomain — the publisher's bare domain (www. stripped), so you can group or filter by outlet without parsing URLs yourself.
  • position — the article's 1-based rank inside its own feed, so Google's relevance/recency ordering survives export and re-sorting.
  • relatedArticles — when Google groups several outlets covering the same story, the other outlets land here with their own title, source and Google News URL. Costs no extra requests. Most items have an empty array; expect it on roughly 1% of results for a typical search.
  • snippet — kept for backward compatibility, but be aware Google News RSS ships no real article summary: this field always repeats the headline. For an actual summary, turn on fetchArticleBody and read articleDescription.

With fetchArticleBody: true each item also carries:

{
"articleBody": "Apify, the web scraping and automation platform, said on Tuesday...",
"articleWordCount": 812,
"articleDeclaredWordCount": null,
"articleBodyComplete": true,
"articleBodyIncompleteReason": null,
"articleBodyTruncated": false,
"articleBodySource": "jsonld",
"articleAuthor": "Jane Doe",
"articleImage": "https://techcrunch.com/wp-content/uploads/2026/09/apify.jpg",
"articleKeywords": ["funding", "web scraping"],
"articleSection": "Startups",
"articleDescription": "The platform raised a new round to...",
"articlePublishedAt": "2026-09-02T07:00:00Z",
"articleModifiedAt": "2026-09-02T09:14:00Z",
"articleFetchStatus": "ok"
}

articleFetchStatus always tells you where the text came from or why it is missing: ok, blocked (publisher refused the request, e.g. hard paywall), no-body (page had no readable article text), error (request failed) or no-url (Google's redirect could not be resolved). Body extraction reads the page's Article JSON-LD first and falls back to the article paragraphs; publishers that serve different HTML to different clients are retried under several request fingerprints, and the one that works is reused for the rest of that publisher's articles.

With extractTickers: true each item also carries:

{ "tickers": ["TSLA"] }

Rule-based and free (no extra request): it catches a cashtag ($TSLA), an exchange prefix/suffix (NASDAQ:AAPL, AAPL:NASDAQ), or a capitalized name immediately followed by (TICKER) — the most common real convention in financial headlines, e.g. "Tesla, Inc. (TSLA)". It deliberately does not match a bare capitalized word (TSLA Stock Rises) — that would flood results with false positives from ordinary acronyms (WSJ, IPO, EV, SEC, UN...), which are also excluded by name when they appear in the (XXX) position. The tradeoff: plain-text mentions with no notation at all are missed. tickers is always [] when nothing matches, never omitted.

Pricing

result — charged per article returned. Failed feeds and duplicates are free, and full article text costs nothing extra. HTTP-only, no browser, so runs finish in seconds.

Tips

  • Combine queries with when:1d to get only fresh news for daily runs.
  • Use excludeWords to drop off-topic coverage that shares a keyword with your query (e.g. exclude iphone when tracking apple as a company, not a product line).
  • Google News' RSS search only understands the hour, day and year units in a time filter. when:1m and when:12m come back as an empty feed, not an error — which looks exactly like "no articles matched". That is why timePeriod offers 30d, 90d and 1y rather than a "last month" option. If you type when:/after:/before: directly into a query, that query keeps your operator and timePeriod is not applied on top of it.
  • publishedAfter/publishedBefore are Google's own after:/before: operators, and Google evaluates the day boundary in its locale, not in UTC. Expect the edges to be loose by a few hours — a before:2026-09-05 query can return an article stamped 2026-09-05T01:51Z (observed live). If you need a hard UTC cut-off, filter the publishedAt field yourself after the run.
  • Combining publishedAfter/publishedBefore with excludeSites trips a Google bug: the feed starts leaking articles months to years outside the window (measured live: 7 in 100, including a 2011 article, on a June-2026 window). when:-style timePeriod is unaffected, and so are excludeWords and siteFilter. This Actor drops those leaked articles for you — anything more than a day outside your window is discarded before it is decoded or charged, and the run status message tells you how many. The one-day allowance is what keeps the legitimate timezone-edge articles above from being thrown away with them. Same protection applies if you type after:/before: directly into a query yourself instead of using the publishedAfter/publishedBefore fields (measured live: 3 in 100 on a query -site:pinterest.com shape) — the Actor reads the dates out of your query text and polices that window too.
  • topics covers more than the eight sections Google shows in its own nav. Twelve narrower sub-sections are served by the same endpoint but are not linked anywhere in the Google News UI: POLITICS, ECONOMY, REAL_ESTATE, JOBS, EDUCATION, AUTOS, MOVIES, MUSIC, CELEBRITIES, ARTS, SOCCER, BASKETBALL. Use them when BUSINESS or ENTERTAINMENT is too broad to monitor — ECONOMY and REAL_ESTATE are far tighter feeds than BUSINESS, and SOCCER/BASKETBALL beat SPORTS for a single-league watch. All 20 are verified live; an unrecognised code returns an empty feed rather than an error, so the Actor warns you instead of silently returning nothing.
  • For any section not in the list above, pass its feed URL in rssUrls directly, e.g. https://news.google.com/rss/headlines/section/topic/TECHNOLOGY?hl=en-US&gl=US&ceid=US:en.
  • Turn off decodeUrls for the fastest runs if you only need headlines and sources.
  • fetchArticleBody adds one request per article, so it is slower — but it costs no extra: you are still charged once per article returned, body or no body.
  • Hard-paywalled publishers come back as blocked. Metered ones are the tricky case — they return a teaser that looks like a real article; filter on articleBodyComplete !== false rather than guessing from articleWordCount.

FAQ

I set excludeWords / siteFilter / timePeriod and some rows ignored it — why? Those are Google search operators, so they only exist on the search endpoint. topics and rssUrls are fixed feeds: Google serves them whole, and there is nothing to attach an operator to. A run that mixes both — say queries: ["nba"] plus topics: ["SPORTS"] — filters the query rows and returns the topic rows untouched. The run log now says so explicitly, naming how many of your feeds are fixed, so you can see it in the run rather than infer it from the output. To filter a section, search it instead: a query like nba site:espn.com when:1d covers the same ground with every operator honoured. Why is url null on some articles? Google occasionally rate-limits the redirect-resolving endpoint per IP; googleNewsUrl still works, and the run status message tells you how many articles were affected. Every run routes through rotating Apify Proxy IPs by default specifically to avoid this — if you turned proxyConfiguration off, turn it back on first. Does fetchArticleBody cost more? No — you pay once per article returned whether or not the body was fetched. Why does articleFetchStatus say blocked or no-body? The publisher likely paywalls the article or serves it without readable paragraph text; both are reported explicitly instead of a silently empty articleBody. How do I know the articleBody is the WHOLE article and not a paywall teaser? Read articleBodyComplete. A metered paywall is the dangerous case: it serves the first few hundred words as ordinary paragraphs, so the body clears every length check and articleFetchStatus is a perfectly honest ok — nothing about the row looks wrong. articleBodyComplete is three-state on purpose: true = we observed the whole article arrive, false = we observed that it did not (articleBodyIncompleteReason names which observation), null = the page carried no completeness signal, so we are not guessing either way. Never read null as a problem, and never read false as a guess. The reasons are:

  • paywalled-section-missing / paywalled-section-empty — the page's own structured data names the paywalled region by CSS selector (Google's paywall markup), we looked, and that region was either absent from the HTML we received or held no article text and what we did extract is teaser-sized (≤220 words). Both of those had to be true. You got the free part only.
  • short-vs-declared-wordcount — the publisher declared a wordCount (also surfaced as articleDeclaredWordCount) and we extracted under 60% of it. Both numbers ride on the row so you can apply your own threshold.
  • paywall-declared-unverifiable — the page is gated but nothing we can check came back conclusive, so completeness is genuinely unknown. This is a null, not a false.

Two things this deliberately does not do, both because we measured them failing:

  • It does not treat isAccessibleForFree: false as proof of a teaser. theatlantic.com carries that flag on every article yet served us a complete 3,479-word body. Metered paywalls gate the Nth read, so the flag describes the publisher's intent, never what your request received — we use it only to find out where to look.
  • It does not treat an empty paywalled region as proof either. scmp.com points its selector at .piano-metering__paywall-container, the client-side paywall overlay, which is correctly empty on a free read — trusting that alone flagged two complete SCMP articles (849 and 1,123 words) as truncated. That is why false needs a second, independent observation and null is a first-class answer here. Why did a run return 0 articles with status SUCCEEDED? The status message distinguishes "Google returned nothing for this query" from "every result was a duplicate of another feed" from "the request failed" — check it before assuming your query is wrong. What's the difference between siteFilter and typing site: into queries? None functionally — siteFilter just OR's multiple domains together (site:a.com OR site:b.com) and applies them to every query in your list, so you don't have to hand-append the operator to each one. What happens if Google's RSS feed has a transient blip mid-run? Every feed and article-URL-decoding request is retried up to 3 times on a connection-level failure (measured at roughly 1 fresh request in 4 for HTTP/2 faults across this fleet, 2026-09-21) before that feed is given up on and named in the status message — a single blip no longer silently empties a query's results into erroredFeeds.

Engineering write-ups behind this Actor:

Only publicly available data is collected. Questions or feature requests: support@fetchsmith.com. Also available as a hosted API at https://fetchsmith.com/tools/google-news-scraper

Source code: https://github.com/Fetchsmith/fetchsmith/tree/main/actors/google-news-scraper