Google News AI Scraper
Pricing
from $0.05 / ai data enrichment
Google News AI Scraper
Search Google News by keyword, optionally extract full article text and AI-generated summaries via OpenRouter, and never re-scrape the same article twice across runs via built-in cross-run deduplication.
Pricing
from $0.05 / ai data enrichment
Rating
0.0
(0)
Developer
ActorFlow
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
17 hours ago
Last modified
Categories
Share
Search Google News by keyword and get structured article data — title, source, publish date, and link — with optional full-article-text extraction and AI-generated summaries or paraphrases. Runs against news.google.com, the public Google News search index.
Target website: news.google.com
✨ Features
- News article extraction — title, publisher, publish date, link, and snippet for each search result
- Full-text extraction — optionally visit each article and extract the cleaned, full article body
- AI enrichment — optionally summarize, paraphrase, extract keywords, run sentiment analysis, or apply a custom instruction to each article via an OpenRouter model
- Cross-run deduplication caching — set a project name and the actor automatically skips articles it already returned in a previous run under the same Apify account
- Proxy support — optional Apify Proxy configuration
- Query filtering — language, country, date range (or absolute date window), site filter, and excluded words
🔧 Input Configuration
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
projectName | String | — | — | Set to persist deduplication across runs; articles already scraped for this project are skipped. |
queries | Array of strings | ✅ | ["artificial intelligence"] | Search queries to run on Google News. |
maxResultsPerQuery | Integer | — | 3 | Max articles returned per query, after deduplication. |
language | String | — | en | 2-letter BCP-47 language code. |
country | String | — | US | ISO 3166-1 alpha-2 country code. |
dateRange | Select | — | any | 1h / 1d / 7d / 1m / 1y / any. Ignored if dateFrom/dateTo is set. |
dateFrom | String | — | — | Only articles on/after this date (YYYY-MM-DD). |
dateTo | String | — | — | Only articles on/before this date (YYYY-MM-DD). |
siteFilter | String | — | — | Restrict results to one domain, e.g. reuters.com. |
excludeWords | Array of strings | — | — | Words that must not appear in results. |
includeImages | Boolean | — | false | Fetch the article's main image during extraction. Visits the article page with a browser. Not currently included in the output. |
extractFullText | Boolean | — | false | Extract the full, cleaned article text via a browser visit. Automatically enabled when aiEnabled is on. |
aiEnabled | Boolean | — | false | Enable AI enrichment via OpenRouter. |
aiApiKey | String (secret) | — | — | Your OpenRouter API key. Optional — if empty, this Actor's own key is used and a per-article platform fee applies. |
aiModel | Select | — | openai/gpt-4o-mini | OpenRouter model used for AI enrichment. |
aiFeatures | Array (select) | — | ["summarize"] | summarize / paraphrase / keywords / sentiment / custom. |
aiCustomInstructions | String | — | — | Extra instructions used when custom is selected. |
proxyConfiguration | Proxy object | — | { "useApifyProxy": false } | Optional Apify Proxy configuration. |
📦 Output
One dataset view (Overview) showing title, source, publish date, query, AI output, and link.
Sample output:
{"query": "artificial intelligence","title": "US efforts to secure AI supply chains doomed to fail: Chinese academic - South China Morning Post","link": "https://www.scmp.com/news/china/diplomacy/article/3364973/us-efforts-secure-ai-supply-chains-are-doomed-fail-says-chinese-academic","source": "South China Morning Post","publishedAt": "2026-08-23T11:00:09.000Z","description": "US efforts to secure AI supply chains doomed to fail: Chinese academic South China Morning Post","fullText": "The latest US artificial intelligence (AI) and supply chain initiatives are doomed to fail, according to a prominent political scientist who also argued that efforts to restrict US companies’ access to low-cost Chinese models would backfire... (line truncated to 2000 chars)","ai": {"summary": "A prominent Chinese political scientist has expressed concerns that the US' AI and supply chain initiatives are doomed to fail due to efforts to restrict US companies' access to low-cost Chinese AI models. He suggested that the US and China collaborate on driving global economic development to stimulate external demand for long-term growth.","paraphrase": "A prominent Chinese political scientist, Zheng Yongnian, has expressed concerns that the United States' latest artificial intelligence (AI) and supply chain initiatives are doomed to fail... (line truncated to 2000 chars)","keywords": ["US AI initiatives","Supply chain rivalry","China-US economic relations","Artificial intelligence","Global economic development","Commercial reality","Market competition","AI-driven growth model","Systemic risks","Global economy"],"sentiment": "negative","custom": null},"scrapedAt": "2026-08-30T08:03:58.202Z"}
💡 Uses of This Data
- Ongoing brand or competitor mention monitoring
- Building a topic-specific news feed or newsletter
- Feeding AI-summarized news into an internal dashboard
- Market and sentiment research over recent coverage
- Tracking coverage of a specific publisher or domain
🚀 How to Use
- Sign up for a free Apify account — includes $5 monthly credit.
- Open the actor page and click Try for free.
- Fill in Search queries (required), and optionally a Project name to enable caching.
- Click Start and wait for the run to complete.
- Download results from the Output tab in JSON, CSV, or Excel format.
You can also run this actor via the Apify API or integrate it directly into your workflows using Zapier, Make, or n8n.
⚠️ Limitations & Known Issues
- Maximum ~100 results per query — Google News RSS search feeds do not expose pagination, so the actor cannot retrieve more than the feed's limit no matter how high
maxResultsPerQueryis set. - Paywalled articles — Full-text extraction may return an empty or partial body for paywalled publishers; the article is still returned with its search-result snippet.
- Rate limiting — Very high query volumes may be rate-limited by Google News; space out large runs if you see missing results.
📝 Notes
- API integration — This actor can be called as an API from any automation platform (Zapier, Make, n8n, custom scripts).
- Caching is per project name — runs with different project names (or no project name) do not share deduplication state.
- AI enrichment pricing — Bring your own OpenRouter API key and there's no extra fee (you're billed by OpenRouter directly). Leave the key empty and this Actor uses its own key instead, charging a small per-article platform fee.
- AI accuracy — AI-generated summaries, paraphrases, keywords, and sentiment are based on the model's interpretation and may contain errors or hallucinations. Please fact-check any AI-enriched content before relying on it for critical decisions.
🔗 Other Actors
| Scraper | Description |
|---|---|
| 🚗 Cheapest Cars & Bids Scraper | Scrapes the cheapest vehicle listings from Cars & Bids, making it easy to discover the best-value auction deals. |
| 🔍 Incidecoder Scraper | Extracts cosmetic ingredient and product information from Incidecoder for research and analysis. |
| 🎟️ Church Finder Scraper [💰Free] | Scrape church listings and profiles from Church Finder. Extracts church name, denomination, address, phone number, service times, ratings and reviews from city search pages or individual church URLs. |
| 🔴 BidNet Direct [Closed Links] Bid Opportunities Scraper | Scrapes closed🔴 government bid and RFP solicitations from BidNet Direct, including title, issuing organization, location, publication date, closing date, and days since closing |
| ✅ BidNet Direct [Open Links] Bid Opportunities Scraper | Scrapes open✅ government bid and RFP solicitations from BidNet Direct, including title, publication date, closing date, and days remaining until closing. |
⚖️ Legality of this actor
This actor only collects article metadata and text that is already publicly visible on Google News and the linked publisher pages — no login, paywall bypass, or private content is accessed. Scraping publicly available data is generally considered legal (see hiQ Labs v. LinkedIn). You are responsible for complying with the target sites' Terms of Service and applicable laws (e.g. GDPR/CCPA) if you process personal data found in scraped content.
💡 Use Case
Teams that need an ongoing feed of news mentions for a topic, brand, or competitor — without re-processing articles they've already collected in a previous run.
🏭 Industry
Media & Entertainment, Marketing & Advertising, Market Research, Financial Services
📤 Output
Structured JSON dataset (exportable to CSV/Excel) with one record per article: query, title, link, source, publish date, description, full text, and AI-generated fields.
🌐 Domain
News search and article extraction (Google News)
🏷️ Label
google-news, news-scraper, ai-summarization, openrouter, deduplication, rss
💬 Support & Contact
If you encounter any issues or have questions, please open an issue
You can also find more of our actors on the Actor Flow .