# Changelog of Google News Scraper (Pay Per Event) (`data_xplorer/google-news-scraper-fast`) Actor

- **URL**: https://apify.com/data\_xplorer/google-news-scraper-fast/changelog.md
- **Full Actor documentation**: https://apify.com/data\_xplorer/google-news-scraper-fast.md

## Changelog

### 2026-09-17 / v2.8

- 🔁 **Smarter URL decoding**: automatically adapts (proxy + session rotation) when decoding slows down or gets blocked, with a safeguard that stops retrying once it's clearly not working, avoiding wasted time on the run
- ⚡ **Optimized fetch strategy for keyword search**: streamlined for more consistent results

### 2026-09-05 / v2.7

- ✅ **HTML scraping restored**: Direct HTML extraction is back as the primary strategy, returning up to 102 articles per query. RSS feed is now the fallback when HTML is unavailable
- 📡 **RSS fallback with enrichment**: When HTML scraping fails, falls back to RSS (up to 100 articles) with automatic og:image and description extraction from article pages

### 2026-09-05 / v2.6

- 🔀 **Direct-first fetch strategy**: HTML requests now attempt a direct connection (no proxy) first, since Apify container IPs are not flagged by Google. Proxy is only used as fallback (3 attempts with rotation). This resolves 429s caused by flagged datacenter/residential proxy IPs
- 🔥 **Proxy warm-up**: Before each proxied request, visits `news.google.com` homepage to establish a fresh Google session on that IP — mimics real browser navigation
- 🖥️ **Updated Chrome impersonation**: `chrome120` → `chrome131` (kept up to date with current Chrome release)

## 2026-08-03 / v2.5

- 📅 **Multi-window pagination**: When `maxArticles > 100`, keyword searches are automatically split into date-range sub-queries (`after:/before:` operators) to bypass Google News's ~100 articles per request limit — up to ~700 articles for `7d`, ~3 000 for `30d`, ~5 200 for `1y`
- 🔢 Default `maxArticles` changed from 20 to 100
- ⚡ Reduced first fetch attempt timeout from 20s to 10s for faster proxy rotation on connection stalls

### 2026-07-03 / v2.4

- 🗓️ Added "Last month" (30d) time period option for keyword searches

### 2026-03-07 / v2.3.2

- 📰 Added topic scraping via RSS for predefined Google News categories (World, Business, Technology, Sports…)
- 🔗 Added support for custom topic and section URLs pasted directly from Google News
- 🖼️ Added extractImages input to control whether article images are included in the output
- 🏷️ Added sourceType field in article metadata to distinguish keywords, topics, and custom URLs

### 2026-02-26 / v2.2.16

- 🔧 Fixed 407 Proxy Authentication error on HTTPS
- 🔧 Fixed proxy type
- 🌍 Added GDPR consent page bypass
- ♻️ Replaced synchronous RSS fetch (blocking event loop) with async `_fetch_url()` via curl\_cffi

### 2026-02-14 / v2.1.5

- 🛠️ Fixed an issue where descriptions were not extracted if "Decode URLs" was disabled.
- 💰 Enhanced Pay Per Event logic to be compatible with newer Apify Python SDK versions

### 2026-02-13 / v2.0.2

- 💰 Implemented logic to respect `ACTOR_MAX_PAID_DATASET_ITEMS`

### 2026-01-06 / v1.9.3

- 🔢 Query limit increased: Maximum number of queries raised to 50

### 2025-12-05 / v1.8.4

- 🛡️ Proxy simplification: Google News requests now always use Apify Datacenter proxy when enabled

### 2025-12-05 / v1.7.6

- 🧭 RSS fallback as detector only: RSS is used to detect the presence of results, never to populate output.
- 🔁 Proxy rotation & retries: Rotation now happens on internal fetch retries, external scrape retries, and on RSS-positive escalation.

### 2025-12-04 / v1.6.12

- 🛡️ Proxy: Fixed Apify proxy usage in Python SDK
- 📰 Parser: Updated Google News search parser to current markup

### 2025-11-24 / v1.5.3

- 🔐 Limited permissions compatibility
- 📤 Output schema integrated: added `.actor/output_schema.json`
- 🧹 Logs cleanup: removed account plan detection and related logs; clearer proxy status messages; reduced noisy warnings.

### 2025-09-13 / v1.3.3

- 💸 Pay‑per‑result compliance: enforce ACTOR\_MAX\_PAID\_DATASET\_ITEMS
- 🧠 Memory safeguards
- ⚙️ DefaultRunMemoryMbytes=1024 and defaultRunTimeoutSecs=120
- 🛑 Forcefully exit the process right after the SDK’s clean exit to prevent residual “RUNNING” time and extra costs.

### 2025-06-25 / v1.2.36

- 🌐 Update documentation

### 2025-03-26 / v1.2.3

- 🚀 Added timeout and retry mechanism for page fetching
- 📊 Optimized logging to reduce noise and improve readability
- 🛠️ Fixed errors related to timeouts and dependencies

### 2025-03-18 / v1.1.6

- ⚡ Apify Proxy Default 'Datacenter' to improve results extraction

### 2025-03-15 / v1.0.6

- 🛡️ Improved Description Extraction: Added multi-step fallback system with three methods to extract article descriptions
- ☁️ Implemented cloudscraper as a third fallback option to handle Cloudflare-protected websites
- ⚡ Optimized timeouts and error handling for better stability
- 🔧 Fixed Actor.fail() method error for proper error reporting

### 2025-03-03 / v0.9.13

- 🚀 Optimized extraction performance with parallel processing for faster scraping
- ✨ **New feature**: Article description extraction with advanced anti-blocking technology for major news sites

### 2025-03-02 / v0.6.15

- 💰 Implemented cost optimization measures to reduce resource consumption
- 🌐 Simplified proxy options (datacenter, residential, none) with datacenter as default
- ⏱️ Added global timeout of 120 seconds to prevent runaway executions
- 🔢 Limited maximum keywords to 5 to optimize performance and costs
- 🧠 Set memory limit to 4GB for more efficient resource utilization

### 2025-02-27 / v0.6.7

- 🌐 Default time period set to "Last hour" (1h)
- 🕒 Updated time period options to match Google News format (1h, 1d, 7d, 1y, all)

### 2025-02-24 / v0.4.8

- ⚖️ Added limitation of 10 articles per keyword for free users
- 💰 Unlimited articles for paid users
- 🔄 Improved proxy handling with better support for DATACENTER and RESIDENTIAL groups

### 2025-02-22 / v0.2.40

- 🌐 Added URL decoder option to get original article URLs from Google News redirections

### 2025-02-22 / v0.1.9

- 🎯 First stable release with comprehensive Google News data extraction
- 🏗️ Structured data format with standardized fields
