# News & RSS Scraper: Google News to JSON (`axiorasolutions/news-feed-scraper`) Actor

Scrape news and blog articles from Google News searches, Google News topics, and any RSS 2.0, RSS 1.0, Atom or JSON Feed URL. One identical row per article: title, link, publisher, publish date, author, summary, categories, images and optional full article text for RAG pipelines.

- **URL**: https://apify.com/axiorasolutions/news-feed-scraper.md
- **Developed by:** [Axiora Solutions](https://apify.com/axiorasolutions) (community)
- **Categories:** News, AI, Marketing
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## News & RSS Feed Scraper — Google News, RSS, Atom, JSON Feed

Google News and RSS feed scraper that turns any source into one clean, analysis-ready article list. Give it Google News searches and topics, RSS 2.0, RSS 1.0/RDF, Atom 1.0 or JSON Feed URLs — or just a bare website domain — and it unifies every one of them into a single identical schema. No API key or login is needed, and the fastest way to try it is to leave the prefilled sources and click **Start**.

### What you get

- `title`, `titleWithoutPublisher` and `url` for every article, with Google News publisher names recovered.
- `publishedAt` in ISO 8601, plus `updatedAt` and `ageHours` for monitoring runs.
- `contentText` and `contentChars` — the article body when the feed already ships it, at zero extra cost.
- Optional `articleText` from the page, with `articleTextStatus` and `articleTextSelector` for auditing quality.
- `author`, `summary`, `categories`, `enclosures`, `imageUrl`, `publisherName` and `publisherDomain`.
- Stable `articleUid` and `contentHash` keys for incremental indexing and deduplication.

### Quick start

1. Open the Actor and leave the prefilled `gnews:web scraping` and `https://techcrunch.com/feed/` in **Sources**.
2. Add any mix of `gnews:`, `gnews-topic:` or feed/website URLs to the same **Sources** list.
3. Set **Max articles per source** and any filters, then click **Start**.
4. Read the **Articles**, **Text for AI** and **By source** dataset tabs.

```json
{
  "sources": ["gnews:\"artificial intelligence\" when:2d", "https://techcrunch.com/feed/"],
  "maxItemsPerSource": 50,
  "publishedAfter": "48 hours",
  "deduplicateBy": "title"
}
```

### Example output

One representative dataset row, with the body already present in the feed:

```json
{
  "ok": true,
  "errorCode": null,
  "articleUid": "a41d7c9b2f0e5538",
  "sourceLabel": "https://techcrunch.com/feed/",
  "sourceType": "rss-2.0",
  "feedUrl": "https://techcrunch.com/feed/",
  "feedTitle": "TechCrunch",
  "feedLanguage": "en-US",
  "title": "Startup raises $40M to automate data pipelines",
  "titleWithoutPublisher": "Startup raises $40M to automate data pipelines",
  "url": "https://techcrunch.com/2026/10/01/startup-raises-40m/",
  "urlResolution": "direct",
  "publisherName": null,
  "publisherDomain": "techcrunch.com",
  "guid": "https://techcrunch.com/?p=2884412",
  "publishedAt": "2026-10-01T14:22:00.000Z",
  "ageHours": 21.6,
  "author": "Jane Doe",
  "summary": "The round was led by...",
  "contentText": "The round was led by... full article body from the feed ...",
  "contentChars": 4820,
  "articleText": "The round was led by... extracted from the page ...",
  "articleTextChars": 5102,
  "articleTextStatus": "ok",
  "articleTextSelector": "article",
  "canonicalUrl": "https://techcrunch.com/2026/10/01/startup-raises-40m/",
  "categories": ["Startups", "Funding"],
  "imageUrl": "https://techcrunch.com/wp-content/uploads/round.jpg",
  "contentHash": "3fb7a1c08d2e4956",
  "scrapedAt": "2026-10-02T12:00:00.000Z"
}
```

### One input list, four kinds of source

Everything goes in the **Sources** array:

| Entry | What it does |
|---|---|
| `gnews:web scraping` | Google News search, up to 100 articles |
| `gnews:"series A" when:7d` | Google News search with its own operators (`when:`, `site:`, quotes) |
| `gnews:top` | Google News top stories for your language and country |
| `gnews-topic:TECHNOLOGY` | A Google News topic: `WORLD`, `NATION`, `BUSINESS`, `TECHNOLOGY`, `ENTERTAINMENT`, `SPORTS`, `SCIENCE`, `HEALTH` |
| `https://techcrunch.com/feed/` | Any RSS, Atom or JSON feed, read directly |
| `blog.cloudflare.com` | A plain website — the Actor **finds the feed for you** |

**Feed auto-discovery** reads `<link rel="alternate">` declarations first, then tries ten conventional paths (`/feed`, `/rss.xml`, `/atom.xml`, `/index.xml`, `/feed.json`, …). You paste a domain; you get articles.

### What makes this feed scraper different

- 🧩 **Four formats, one schema** — `publishedAt` is always ISO 8601 whether the feed used RFC 822, RFC 3339 or a JSON Feed timestamp. `categories`, `enclosures` and `author` are normalised the same way.
- 🪪 **Honest Google News URL handling** — Google News items link to a consent-gated redirect page that **cannot** be resolved without a browser. Most scrapers hand you that URL and say nothing. This Actor sets `urlResolution: "google-news-redirect"` and gives you `publisherName` and a clean `titleWithoutPublisher` so attribution still works.
- 🧠 **Feed content is often the whole article** — many publishers ship the full body inside the feed. `contentText` captures it with **zero extra requests**, so check `contentChars` before paying the time cost of full-text extraction.
- 📄 **Real article extraction when you need it** — optional full text uses a readability-style pass over `article`, `[itemprop=articleBody]`, `.entry-content` and similar containers, and reports `articleTextSelector` so you can audit quality. `articleTextStatus` always explains a missing body instead of leaving a silent `null`.
- 🔁 **Deduplication you control** — by canonical URL (tracking parameters stripped), by normalised title (collapses the same story syndicated across twenty outlets), by feed GUID, or off.
- 🎯 **Filters cut your bill** — `publishedAfter`, `keywords` and `excludeKeywords` run before anything is written, so filtered articles are never charged.
- 🛟 **One dead feed never fails the run** — it becomes an `ok: false` row with a stable `errorCode`.

Running on Apify adds scheduling, webhooks, monitoring, API and SDK access, and one-click export to JSON, CSV, Excel, Google Sheets and 20+ integrations.

### How to use it

1. Put your searches, topics, feeds or domains in **Sources**.
2. Set **Google News language** and **Google News country** if you want a non-US edition (`de-DE` + `DE`, `en-GB` + `GB`, `pt-BR` + `BR`).
3. Set **Max articles per source** and **Max articles for the whole run**.
4. Add **Published after** (`24 hours`, `7 days`) for monitoring runs.
5. Turn on **Fetch full article text** only if `contentChars` from a test run shows the feeds are truncated.
6. Click **Start**, then use the **Articles**, **Text for AI** and **By source** dataset tabs.

#### How do I monitor brand mentions?

Use `gnews:"your brand"` plus `gnews:"your brand" when:1d`, set **Published after** to `24 hours`, deduplicate by **Normalised title**, and schedule the Actor hourly or daily. Diff on `articleUid` downstream to get only new stories.

### How much does it cost to scrape news feeds?

Pricing is **pay per event** with exactly one event:

| Event | What triggers it | Billed |
|---|---|---|
| Article | One article written to the dataset | per article |
| Actor start | Once per run, platform fee | per run |

**Full-text extraction is included in the per-article price.** Turning it on costs you run time, not money. Articles removed by filters or de-duplication are **not** billed, because they are never written. A source that fails is not billed.

500 articles is 500 billed events — that is the whole calculation. Compute, bandwidth and storage are included; there is no separate platform-usage charge on top.

Set **Max cost per run** in the run options for a hard ceiling. Higher Apify plans get progressively lower per-article pricing through Apify Store tier discounts.

Evaluating? Set **Max articles for the whole run** to `20` with one source.

### Example input

```json
{
  "sources": [
    "gnews:\"artificial intelligence\" when:2d",
    "gnews-topic:TECHNOLOGY",
    "https://techcrunch.com/feed/",
    "blog.cloudflare.com"
  ],
  "googleNewsLanguage": "en-US",
  "googleNewsCountry": "US",
  "maxItemsPerSource": 50,
  "maxItemsTotal": 300,
  "publishedAfter": "48 hours",
  "keywords": ["funding", "launch"],
  "excludeKeywords": ["sponsored"],
  "deduplicateBy": "title",
  "includeArticleText": true,
  "articleTextMaxChars": 12000
}
```

#### A Google News output row

Note `urlResolution` and the recovered publisher:

```json
{
  "ok": true,
  "articleUid": "b72e1f904ac35d10",
  "sourceLabel": "gnews:web scraping",
  "sourceType": "google-news",
  "title": "She Code Africa and Apify announce hackathon winners - TechCabal",
  "titleWithoutPublisher": "She Code Africa and Apify announce hackathon winners",
  "url": "https://news.google.com/rss/articles/CBMipgFBVV95cUxNOFlQWk1F",
  "urlResolution": "google-news-redirect",
  "publisherName": "TechCabal",
  "publisherDomain": null,
  "publishedAt": "2026-10-01T09:11:00.000Z",
  "articleTextStatus": "skippedGoogleNewsRedirect"
}
```

### Use cases

- **Brand and competitor monitoring** — Google News searches plus de-duplication by title gives one row per story instead of twenty.
- **Deal and funding signals** — `gnews:"series A" when:7d` on a schedule, filtered by your sector keywords.
- **RAG and LLM ingestion** — `contentText` and `articleText` are clean plain text with stable `articleUid` and `contentHash` keys for incremental indexing.
- **Newsroom and research dashboards** — mix twenty publisher feeds and five Google News topics in one dataset.
- **Podcast catalogues** — podcast RSS feeds expose audio in `enclosures` with type and byte length.
- **Content gap analysis** — pull competitor blog feeds by domain and compare `categories` and publishing cadence.

### Related Actors by Axiora Solutions

| Actor | Use it for |
|---|---|
| **Substack Newsletter Archive Scraper** | Newsletter archives with engagement metrics, which RSS does not expose |
| **ATS Job Scraper** | Hiring signals for the companies appearing in your news feed |
| **SEC EDGAR API** | The filings behind the financial headlines |

### Frequently asked questions

#### Why is the Google News URL not the publisher's URL?

Because Google does not let it be. A `news.google.com/rss/articles/...` link responds with a consent redirect and then a JavaScript application; there is no `Location` header and no readable link in the HTML. Resolving it requires a real browser, which would multiply the cost of every run. This Actor is explicit about it: `urlResolution` is set to `google-news-redirect`, `publisherName` is recovered from the title, and full-text extraction reports `skippedGoogleNewsRedirect` instead of silently returning nothing.

**If you need canonical publisher URLs and article bodies, add the publishers' own feeds as sources.** That path returns `urlResolution: "direct"` and works with full-text extraction.

#### Which feed formats are supported?

RSS 2.0, RSS 1.0 / RDF, Atom 1.0 and JSON Feed 1.x. Namespaces are stripped, so `dc:creator`, `content:encoded` and `media:*` extensions are picked up as well.

#### Do I need full-text extraction?

Often not. Run once without it and look at `contentChars` — many publishers put the whole article in the feed. Only enable it when `contentChars` is small compared with the real article.

#### How many articles does Google News return?

Up to 100 per search, around 70 per topic and around 40 for top stories. That is Google's limit. Use several narrower searches, or add `when:1d` and schedule the Actor more often, to cover more ground.

#### Can I use Google News search operators?

Yes. `when:7d`, `site:example.com`, quoted phrases and `-exclusions` all pass straight through. `gnews:"series A" when:7d -crypto` is a valid source.

#### Is scraping RSS feeds legal?

Feeds exist to be consumed by software; that is their entire purpose. This Actor sends ordinary HTTP requests and identifies itself. You remain responsible for how you use and republish the content, including the publisher's copyright and licensing terms.

#### Can I run this on a schedule?

Yes, and it is the intended pattern. Set **Published after** to a short window, pick a de-duplication strategy, and diff on `articleUid`.

#### Something looks wrong — how do I report it?

Open the **Issues** tab on this Actor page with the source entry and the field you expected. Feed-format edge cases are treated as bugs and get fixed in the normaliser.

***

Runnable examples and how-to guides for these Actors: [github.com/batow133/axiora-apify-actors](https://github.com/batow133/axiora-apify-actors)

# Actor input Schema

## `sources` (type: `array`):

Mix any of these in one run. gnews:artificial intelligence runs a Google News search. gnews-topic:TECHNOLOGY pulls a Google News topic. gnews:top pulls top stories. A feed URL is read directly. A plain website URL is auto-resolved to its RSS, Atom or JSON feed.

## `googleNewsLanguage` (type: `string`):

Language and region code for Google News sources, for example en-US, en-GB, de-DE, fr-FR, es-ES, pt-BR, ja-JP. Ignored by plain feed URLs.

## `googleNewsCountry` (type: `string`):

Two-letter edition country for Google News sources, for example US, GB, DE, FR, IN, AU. Ignored by plain feed URLs.

## `maxItemsPerSource` (type: `integer`):

Stop after this many articles from each source. Google News returns at most 100 per search.

## `maxItemsTotal` (type: `integer`):

Hard ceiling across every source. The run stops cleanly when it is reached.

## `publishedAfter` (type: `string`):

Keep only articles published on or after this date. Accepts 2026-01-31 or a relative value such as 24 hours or 7 days. Articles with no date are kept.

## `keywords` (type: `array`):

Keep only articles whose title or summary contains one of these words, case-insensitive. Leave empty to keep everything the sources return.

## `excludeKeywords` (type: `array`):

Drop articles whose title or summary contains any of these words, case-insensitive.

## `deduplicateBy` (type: `string`):

Google News returns the same story from many outlets. url keeps one row per link, title collapses re-publications of the same headline, guid trusts the feed's own ID, none keeps everything.

## `includeArticleText` (type: `boolean`):

Fetch each article page and extract the main body as clean plain text, for embedding or summarisation. Adds one request per article. Google News links are redirect pages that cannot be read without a browser, so they are skipped and reported.

## `articleTextMaxChars` (type: `integer`):

Truncate extracted article text at this many characters.

## `requestTimeoutSecs` (type: `integer`):

Give up on a single feed or article request after this many seconds.

## `proxyConfiguration` (type: `object`):

Optional. Most feeds answer direct requests. Turn on datacenter proxy rotation if a publisher blocks you, especially when fetching full article text at volume.

## Actor input object example

```json
{
  "sources": [
    "gnews:\"series A\" when:7d",
    "gnews-topic:TECHNOLOGY",
    "https://techcrunch.com/feed/",
    "blog.cloudflare.com"
  ],
  "googleNewsLanguage": "en-US",
  "googleNewsCountry": "US",
  "maxItemsPerSource": 50,
  "maxItemsTotal": 500,
  "publishedAfter": "24 hours",
  "keywords": [
    "funding",
    "acquisition"
  ],
  "excludeKeywords": [
    "sponsored",
    "horoscope"
  ],
  "deduplicateBy": "url",
  "includeArticleText": false,
  "articleTextMaxChars": 12000,
  "requestTimeoutSecs": 25,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `articles` (type: `string`):

Every article collected, deduplicated by the strategy you chose.

## `runSummary` (type: `string`):

Per-source feed URL, detected format, items available and written, filter effects, full-text extraction statistics, billing and network totals.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sources": [
        "gnews:web scraping",
        "https://techcrunch.com/feed/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiorasolutions/news-feed-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "sources": [
        "gnews:web scraping",
        "https://techcrunch.com/feed/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("axiorasolutions/news-feed-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sources": [
    "gnews:web scraping",
    "https://techcrunch.com/feed/"
  ]
}' |
apify call axiorasolutions/news-feed-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiorasolutions/news-feed-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Odb8KyYGCx1ZKdRBl/builds/wVZcEbCQxs0TPyeRR/openapi.json
