Google News Scraper with Real Article URLs & Full Text
Pricing
from $2.50 / 1,000 articles
Google News Scraper with Real Article URLs & Full Text
Google News scraper: search by keyword, topic or top stories. Get real publisher article URLs (decoded Google links), source, date and full article text.
Pricing
from $2.50 / 1,000 articles
Rating
0.0
(0)
Developer
Marlon Lee
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
17 hours ago
Last modified
Categories
Share
Google News Scraper
Export Google News articles by keyword, topic or top stories, with the real publisher URL instead of the news.google.com/rss/articles/... redirect, plus title, source, publish date and related coverage. Turn on text extraction to also get the article body, author and main image. PR teams, traders, researchers and AI builders use it for media monitoring and news datasets in any Google News language or country edition.
What you get
Each article is one row with:
- The feed it came from (your query, a topic or top stories)
- Title, source name and source website, and publish date
- The Google News link and article ID
- The decoded publisher URL (on by default)
- Related articles from other outlets covering the same story, for topic and top-story feeds
- Optional full text: the article headline, body text, author and main image, read from the publisher page
The Actor removes duplicate articles across all your queries and topics in a run.
Use cases
- Media monitoring and PR: track mentions of your brand, competitors or executives.
- Market signals: watch tickers and companies with
when:1hqueries on a schedule. - AI, LLM and RAG pipelines: feed clean article text with canonical URLs into summarizers or embeddings.
- Research and journalism: build news datasets by topic, country and date range.
- SEO and content teams: see which publishers rank in Google News for your keywords.
How to use the Google News scraper
- Open the Actor in Apify Console.
- Enter search queries (Google News operators work), pick topic sections, turn on top stories, or combine them.
- Set the language and country edition, for example
deandDE, orpt-BRandBR. - Keep Resolve real article URLs on, and turn on Extract full article text if you need the body. Set Max articles per query/topic.
- Click Start, then download the results as JSON, CSV, Excel or HTML, or read them through the Apify API.
Input example
{"queries": ["openai when:7d", "\"interest rates\" site:reuters.com"],"topics": ["TECHNOLOGY", "BUSINESS"],"topStories": true,"language": "en-US","country": "US","maxItems": 100,"decodeUrls": true,"extractArticleText": false}
Search operators: "exact phrase", OR, -exclude, site:reuters.com, intitle:, when:1h, when:1d, when:7d, after:2026-01-01, before:2026-02-01. Topics: WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SPORTS, SCIENCE, HEALTH. For editions where the ID isn't COUNTRY:language, set ceid, for example BR:pt-419.
Output example
A real item from a test run with extractArticleText: true. The article text is shortened.
{"feed": "query:openai when:7d","title": "OpenAI Targets $30 Billion in New Funding at $1.4 Trillion Value","source": "Bloomberg","sourceUrl": "https://www.bloomberg.com","publishedAt": "2026-09-29T17:35:13.000Z","googleNewsUrl": "https://news.google.com/rss/articles/CBMiswFBVV95cUxObmVY...?oc=5","articleId": "CBMiswFBVV95cUxObmVY...","description": "OpenAI Targets $30 Billion in New Funding at $1.4 Trillion Value Bloomberg","relatedArticles": [],"decodedUrl": "https://www.bloomberg.com/news/articles/2026-09-29/openai-targets-30-billion-in-new-funding-at-1-4-trillion-value","articleTitle": "OpenAI Targets $30 Billion in New Funding at $1.4 Trillion Value","articleText": "OpenAI aims to raise at least $30 billion from investors in a new round of funding, according to people familiar with the matter...","author": "Rebecca Torrence, Edward Ludlow, Shirin Ghaffary","image": "https://assets.bwbx.io/images/users/iqjWHBFdfxIU/iBQYMNSiiUWc/v1/1200x800.jpg"}
Items from topic and top-story feeds fill relatedArticles with { title, source, googleNewsUrl } for other outlets covering the same story.
Pricing
You pay per article saved, with no monthly rental. Each article counts as one of two results:
article: an article with its metadata and the decoded publisher URLarticle-with-text: an article whose full text was extracted
If text extraction fails for an article (a paywall or a bot wall), it's charged as an article, not an article-with-text. Set a maximum spend per run and the Actor stops once it reaches that limit.
FAQ
Is it legal to scrape Google News? The Actor collects headlines and links that Google News publishes in its public feeds, plus publicly accessible publisher pages. Publishers own the copyright in their articles, so check your rights before you republish full text or use it for AI training. You're responsible for how you use the data.
How does URL decoding work?
Google News wraps each link in an encoded news.google.com/rss/articles/CBMi... redirect. For each article, the Actor reads the Google News article page to get a signature, then asks Google's own batchexecute endpoint for the real URL, the same approach the open-source googlenewsdecoder library uses. If decoding still fails after retries, you get the article with decodedUrl: null.
Will Google block it? Which proxy should I use? The default is Apify residential proxy in the US, the most reliable choice for Google. If you turn off URL decoding, the Actor only reads Google News feeds, and a cheaper datacenter proxy is usually enough.
How many articles can I get per query?
Google News feeds return up to about 100 articles per query or topic. To get more, split your search with date operators, for example after:2026-09-01 before:2026-09-08, or use narrower queries.
Why is description so short?
Google News feeds include only the headline and the source, plus related headlines for topic stories. Turn on text extraction to get the article content.
Why is articleText sometimes null?
Some publishers (WSJ, FT and other paywalled or bot-protected sites) don't serve the article to scrapers. You still get the item with its decoded URL.
What are the limits?
- Up to about 100 articles per query or topic feed, a Google limit.
- Text extraction uses a lightweight readability-style heuristic, not a browser. JavaScript-only or paywalled pages return no text.
- Text extraction needs URL decoding on.
- Google may change its link encoding. If decoding breaks, items still arrive with
decodedUrl: nulluntil the Actor is updated.