# Company News Scraper — Press, Funding & Announcements Feed (`inovaflow/company-news-scraper`) Actor

Company news and press releases for your accounts: give names or domains (or a topic) and get one deduplicated, typed row per story — funding, M\&A, launches, partnerships, hires, layoffs — merged across every outlet, with the company newsroom, source URLs and a new-since-last-run delta.

- **URL**: https://apify.com/inovaflow/company-news-scraper.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** News, Lead generation, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 news events

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Company news for your accounts, as a clean feed: one deduplicated, typed row per story — funding, acquisitions, launches, partnerships, executive moves, layoffs — with every source that reported it.**

If you sell to, invest in or keep an eye on a list of companies, you know the ritual. Before every call, every planning cycle, every Monday: search each company, wade through five outlets reprinting the same press release, skip the share-price noise, and try to remember whether you already saw that story last week. Your AI agent does the same thing, only faster and just as confused: it gets five near-identical articles, a stock ticker page and a K-drama star who happens to share the company's name.

We needed this for our own account agents, so we built it — and now we're sharing it. Give it company names or domains; it searches the news, reads each company's own newsroom and press releases, keeps only headlines that are really about that company, merges the same story from every outlet into **one row**, labels what kind of event it is, and remembers what it already told you so a scheduled run returns **only what's new**. No summaries, no AI rewriting: titles, dates, URLs and snippets come straight from the source, so your agent reasons over facts.

### Who it's for

- **Account-based sales & SDR teams** — *"what happened at my 50 target accounts this week?"* before writing the next touch.
- **AI agents and GTM automations** — a raw, typed, MCP-callable news feed to poll on a schedule, with a stable id per story.
- **Investors & analysts** — *"any funding, M\&A or leadership changes at my portfolio companies?"*
- **Customer success & partnerships** — spot expansions, layoffs, security incidents or new partners at your customers early.
- **Market & competitor watchers** — follow a topic (*"SOC 2"*, *"agentic commerce"*) or a set of competitors in one feed.

### What you get

One row per **news event** (a story), newest first:

```json
{
  "companyName": "HubSpot",
  "companyDomain": "hubspot.com",
  "eventType": "award",
  "eventTypes": ["award"],
  "title": "HubSpot named a Leader in the 2026 Gartner® Magic Quadrant™ for B2B Marketing Automation Platforms",
  "url": "https://www.hubspot.com/company-news/gartner-magic-quadrant-2026",
  "publisher": "HubSpot",
  "publishedAt": "2026-09-24T15:36:38.000Z",
  "snippet": "Marketing Hub recognized as a Leader for the sixth consecutive year",
  "snippetSource": "company-feed",
  "sources": [{ "url": "https://www.hubspot.com/company-news/gartner-magic-quadrant-2026", "publisher": "HubSpot", "publishedAt": "2026-09-24T15:36:38.000Z", "title": "HubSpot named a Leader in the 2026 Gartner® Magic Quadrant™ …" }],
  "sourceCount": 1,
  "isFirstParty": true,
  "attribution": "first-party",
  "isNew": false,
  "canonicalId": "cn_4f1c2a9b7d3e0a61",
  "firstSeenAt": "2026-09-26T08:12:40.000Z"
}
```

- **One story = one row.** The same announcement from TechCrunch, Yahoo Finance and three wire re-posts is one event with five `sources` — and one charge. Same-day share-price chatter about a company is one event, not twenty.
- **Typed events** — `funding`, `acquisition`, `product`, `partnership`, `expansion`, `leadership`, `layoffs`, `earnings`, `legal`, `security`, `award`, `stock` (share price, analyst ratings, insider trades) and `other`. Typed by transparent keyword rules, all matching types listed in `eventTypes`.
- **The company's own voice** — its newsroom, press feed or announcement blog is read too and merged with the coverage (`isFirstParty: true`).
- **Only news that is really about the company.** The headline has to name it. For names that are also everyday words (Stripe, Gong, Ramp, Notion) the headline needs extra evidence — "Stripe's", "Stripe launches", "at Stripe", a link to stripe.com — so football "stripes" and actors called Gong never reach you. Each row says how it was matched in `attribution`.
- **Verbatim, never generated** — `snippet` is the article's own description or lead paragraph (or the company's feed summary), copied as-is; `snippetSource` says which.
- **New since last run** — a stable `canonicalId` per story and a per-watch memory. Flip on *Only stories that are new since my last run* and a daily schedule becomes a clean feed of what changed; a story that gets picked up by more outlets later is not "new" again.

### Input

Everything is optional — run it empty to get the latest company announcements worldwide.

- **Companies** — names (`Datadog`), domains (`hubspot.com`) or both (`Gong (gong.io)`). A domain is the most precise, and it unlocks the company's newsroom.
- **Topics** — with companies: narrow their news (`stablecoin`). Without companies: follow the topic itself (`SOC 2`, `data center`).
- **Published within** — 1–90 days (default 7). **Event types** — keep only the kinds you act on (leave out *Stock & analyst chatter* to keep business events only).
- **Sources** — newsroom (on), second news index (on), global news database (off, slower), open articles for the publisher URL and lead text (on).
- **Delta & limits** — only-new mode, Watch ID, caps per company and in total, and whether to include unconfirmed matches.

### How to use it

1. **In the Console:** add your companies, press **Start**, open the *News events* view. Export to CSV/Excel/JSON or connect Google Sheets, Slack or your CRM through Apify integrations.
2. **On a schedule:** save the input as a Task with *Only stories that are new since my last run* on, and schedule it daily. Each run delivers only the stories you haven't seen.
3. **From an AI agent (MCP):** add the Actor through the Apify MCP server and ask *"latest news on datadoghq.com and hubspot.com this week, funding and product launches only"* — the agent gets typed JSON back.
4. **From code:** call it via the Apify API with `{"companies": ["hubspot.com"], "daysBack": 7}` and read the dataset.

And that's it.

### Pricing

Pay per event: **$0.005 per delivered news event** (plus Apify's small start fee). Duplicate reports of the same story, filtered-out events, unconfirmed matches and empty runs are never charged.

A daily watch of 20 accounts typically delivers 20–60 new stories a day — **about $0.10–$0.30 a day**, and you pay once per story, not once per outlet that reprinted it.

> **Tip:** turn on *Only stories that are new since my last run* for scheduled watches, and leave out *Stock & analyst chatter* if you only care about business events — you pay only for rows you keep.

### Notes

- Sources are public news pages and the companies' own public newsrooms; no login, no API key, no personal data. Please use the data in line with the publishers' terms.
- The Actor delivers raw facts; it does not summarize or rewrite articles.
- Missing a source, a newsroom that isn't picked up, or a company that is matched wrongly? Open an issue in the **Issues** tab — we fix those quickly. Custom feeds and integrations: reach us through Apify.

# Actor input Schema

## `companies` (type: `array`):

The accounts to watch, one per line: a name (`Datadog`), a domain (`hubspot.com`) or both (`Gong (gong.io)`). A domain gives the most precise matching and adds the company's own newsroom and press releases. Leave empty to search by topic, or to get the latest company announcements worldwide.

## `topics` (type: `array`):

With companies: only their news that mentions one of these (`stablecoin`, `AI agents`). Without companies: news about the topic itself (`SOC 2`, `data center`, `layoffs in fintech`). Quotes and OR work like in a search box.

## `daysBack` (type: `integer`):

How far back to look. 7 days suits a weekly sweep; use 1–2 for a daily schedule together with the new-since-last-run option.

## `eventTypes` (type: `array`):

Only deliver these kinds of news. Empty = everything. Share-price chatter, analyst ratings and insider trades are "Stock & analyst chatter" — leave it out to keep only business events.

## `includeNewsroom` (type: `boolean`):

Reads the company's own newsroom, press feed or announcement blog (found from its domain) — the first-party version of its news, merged with the press coverage of the same story.

## `useSecondNewsIndex` (type: `boolean`):

Adds a second news search. It finds outlets the first one misses and ships each article's own lead text, which is often the only snippet a paywalled article has.

## `useGlobalNewsIndex` (type: `boolean`):

Adds a global event database that covers many regional and international outlets. It is slow (about 10–15 seconds per company) and rate-limited, so it is off by default.

## `readArticles` (type: `boolean`):

Opens the lead article of every delivered story to record its real publisher URL, its own description or first paragraph (copied as-is) and its language, and to confirm stories about companies whose name is also an everyday word. Turn off for the fastest possible run.

## `onlyNewSinceLastRun` (type: `boolean`):

Deliver only stories this watch has not reported before — ideal for a schedule. A story that gets picked up by more outlets later is not new again. The first run of a watch returns everything as the baseline.

## `watchId` (type: `string`):

Optional name for this watch, e.g. `top-accounts`. Runs sharing a Watch ID share the memory of what was already reported. Leave empty to derive it from the companies and topics.

## `maxEventsPerCompany` (type: `integer`):

Cap per company (or per topic). When it bites, business events and the company's own posts are kept ahead of share-price chatter; the rows come out newest first.

## `maxItems` (type: `integer`):

Cap on delivered rows for the whole run. Also caps what you pay.

## `maxCompanies` (type: `integer`):

Only the first N companies of the list are covered.

## `includeUnconfirmedMatches` (type: `boolean`):

For a company whose name is also an everyday word (Stripe, Gong, Ramp), a headline needs extra evidence that it is about the company. Turn on to also receive the headlines that name it without that evidence, flagged `attribution: "unconfirmed"`.

## `newsLanguage` (type: `string`):

Language edition of the news search, e.g. `en-US`, `de-DE`, `fr-FR`. Articles in other languages are dropped from the second index.

## `newsCountry` (type: `string`):

Two-letter country edition, e.g. `US`, `GB`, `DE`. Pair it with the matching language for strong local coverage.

## `resolveDomainsWithCompanyEnrichment` (type: `boolean`):

For companies given by name whose domain cannot be confirmed from their homepage or their news, run our Company Enrichment Actor once to find it (billed separately by that Actor). Off by default — most runs never need it.

## `maxConcurrency` (type: `integer`):

Concurrent requests to news sources and article pages. Lower it if a source starts rate-limiting.

## `secondIndexPages` (type: `integer`):

Each page adds about ten articles from the second news index.

## `proxyConfiguration` (type: `object`):

Optional. Every source is public; the Actor already falls back to Apify Proxy if a source rate-limits the run. Set one only to route every request through a specific proxy.

## Actor input object example

```json
{
  "companies": [
    "hubspot.com",
    "datadoghq.com",
    "Stripe"
  ],
  "topics": [
    "AI agents",
    "stablecoin"
  ],
  "daysBack": 7,
  "eventTypes": [
    "funding",
    "acquisition",
    "leadership",
    "product"
  ],
  "includeNewsroom": true,
  "useSecondNewsIndex": true,
  "useGlobalNewsIndex": false,
  "readArticles": true,
  "onlyNewSinceLastRun": false,
  "watchId": "top-accounts",
  "maxEventsPerCompany": 10,
  "maxItems": 300,
  "maxCompanies": 100,
  "includeUnconfirmedMatches": false,
  "newsLanguage": "de-DE",
  "newsCountry": "DE",
  "resolveDomainsWithCompanyEnrichment": false,
  "maxConcurrency": 6,
  "secondIndexPages": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `events` (type: `string`):

One row per story: company, event type, headline, publisher URL, verbatim snippet, publication time and every outlet that reported it.

## `sources` (type: `string`):

Every outlet merged into each story.

## `newEvents` (type: `string`):

Stories this watch had not reported before.

## `summary` (type: `string`):

Counts by company and event type, the attribution funnel, per-source and newsroom reports, the watch id and the time budget.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "hubspot.com",
        "datadoghq.com"
    ],
    "daysBack": 7,
    "maxEventsPerCompany": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/company-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "hubspot.com",
        "datadoghq.com",
    ],
    "daysBack": 7,
    "maxEventsPerCompany": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/company-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "hubspot.com",
    "datadoghq.com"
  ],
  "daysBack": 7,
  "maxEventsPerCompany": 10
}' |
apify call inovaflow/company-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/company-news-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4hQ5xnvWRpNoyTb0U/builds/LLE33h6rm9Opws73S/openapi.json
