# Multi-Source Business News Aggregator (`axery/business-news-aggregator`) Actor

Scrape and normalize business/finance headlines from Forbes, CNBC, Fortune, NYTimes and the Financial Times into one schema, with a cross-outlet keyword filter. No login, no API key.

- **URL**: https://apify.com/axery/business-news-aggregator.md
- **Developed by:** [Axery](https://apify.com/axery) (community)
- **Categories:** News, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.25 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Multi-Source Business News Aggregator

Scrapes and normalizes business/finance headlines from Forbes, CNBC, Fortune, The New York Times and the Financial Times into one schema, with an optional cross-outlet keyword filter. No login, no API key.

Useful for media monitoring, competitive/market intelligence, and building a single feed out of sources that would otherwise each need their own scraper.

### What makes this different

**One outlet-agnostic schema, not five separate actors.** Each outlet formats its feed differently — CDATA here, plain text there, a different date format, a different category scheme — and a scraper built against just one of them breaks silently on the others. This Actor reads all five defensively and normalizes them to identical fields, so filtering or sorting works the same way regardless of which outlet a row came from.

**The keyword filter runs across all selected outlets in one pass.** Tracking a topic ("tariff", "interest rate", "AI") across five outlets normally means five separate scrapes and a manual merge afterward. Here it's one input field, checked against every outlet's title and description before the row is even counted toward your item limit.

**Dates you can actually sort and compare.** Every outlet's publish timestamp is normalized to UTC ISO-8601 regardless of the RFC-822 or outlet-specific format it originally shipped in — so a Forbes article and an FT article land on a shared timeline instead of five incompatible date strings.

**Article IDs are namespaced by outlet.** The same headline occasionally gets syndicated across outlets; namespacing prevents a coincidental ID collision from quietly merging two different articles' rows.

### Input

| Field | Type | Notes |
|---|---|---|
| `outlets` | array | Any of `forbes`, `cnbc`, `fortune`, `nytimes`, `ft`. |
| `keyword` | string | Optional. Case-insensitive, checked across title + description. |
| `maxItemsPerOutlet` | integer | Cap per outlet, applied after the keyword filter. `0` = everything the feed serves. |
| `proxyConfiguration` | object | Defaults to Residential. |

#### One limit worth knowing up front

Each outlet's feed is a **fixed window of its most recent (or, for Forbes specifically, currently-trending) items** — there's no pagination parameter to reach further back. To build a longer history on a topic, run this on a schedule with the same keyword and let the dataset accumulate.

### Output

```json
{
  "article_id": "cnbc:108353620",
  "source": "cnbc",
  "title": "Trump targets Iran's trade lifelines — here are the countries most exposed",
  "description": "Washington's threat of \"economic D-Day\" collides with a small group of governments...",
  "published_at": "2026-08-25T06:24:36Z",
  "url": "https://www.cnbc.com/2026/08/25/us-iran-secondary-sanctions-china-india-uae-hormuz-trade-.html"
}
```

Each run also writes a `RUN_COVERAGE` record to the key-value store with what was requested, what came back, and any per-outlet failures.

### Local development

```bash
pip install -r requirements.txt
python test_local.py all --max 10 --out sample_output.json
python test_local.py cnbc fortune --keyword ai
```

`sample_output.json` is real output from a live run across all five outlets.

# Actor input Schema

## `outlets` (type: `array`):

Which outlets to pull from: forbes, cnbc, fortune, nytimes, ft. Each is fetched independently into the same normalized dataset.

## `keyword` (type: `string`):

Keep only articles whose title or description contains this text (case-insensitive), checked across every selected outlet at once.

## `maxItemsPerOutlet` (type: `integer`):

Cap per outlet, applied after the keyword filter. `0` returns everything each feed serves.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings. Defaults to Residential - these feeds are public, but Apify's own container IP range can still be rate-limited or blocked by services that treat cloud IPs as suspicious regardless of any WAF.

## Actor input object example

```json
{
  "outlets": [
    "cnbc",
    "fortune"
  ],
  "keyword": "tariff",
  "maxItemsPerOutlet": 0,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `articles` (type: `string`):

One row per article: title, description, author, publish date, categories, thumbnail.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "outlets": [
        "forbes",
        "cnbc",
        "fortune",
        "nytimes",
        "ft"
    ],
    "maxItemsPerOutlet": 0
};

// Run the Actor and wait for it to finish
const run = await client.actor("axery/business-news-aggregator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "outlets": [
        "forbes",
        "cnbc",
        "fortune",
        "nytimes",
        "ft",
    ],
    "maxItemsPerOutlet": 0,
}

# Run the Actor and wait for it to finish
run = client.actor("axery/business-news-aggregator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "outlets": [
    "forbes",
    "cnbc",
    "fortune",
    "nytimes",
    "ft"
  ],
  "maxItemsPerOutlet": 0
}' |
apify call axery/business-news-aggregator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axery/business-news-aggregator"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TyvXfNgNpsubOJMQu/builds/n4ZNdH3V8IQHjfe3y/openapi.json
