# Hiring Intent Scraper – Companies Hiring & Hiring Signals (`inovaflow/hiring-intent-scraper`) Actor

Companies hiring by role as company-level buying signals: open-role counts, momentum vs your last run, departments, seniority mix and role→intent tags (scaling outbound, building data team…). Multi-board coverage merged and deduped per company. Dataset-only, MCP-ready.

- **URL**: https://apify.com/inovaflow/hiring-intent-scraper.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** Lead generation, Jobs, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $20.00 / 1,000 company signals

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

**Companies hiring by role, packaged as company-level buying signals — not job postings.** Tell it which roles mean intent for you (SDRs, data engineers, RevOps…), and get one row per company: how many matching roles are open, how that number moved since your last run, which departments and cities are growing, and what the hiring pattern says the company is about to buy.

If you sell to companies that are *about to need you*, job postings are the earliest public signal there is: a company hiring three SDRs is building an outbound motion; a company hiring its first data engineer is about to buy a warehouse; a VP of Security posting means a security budget just got approved. Signal-based outreach converts several times better than cold lists — but every job scraper hands you the seeker's view: thousands of postings with apply links, the same role syndicated across five boards, no company roll-up, no memory of last week. This Actor is the seller's view.

### Hiring intent data, one row per company

Each run searches several public job boards and employer career pages for the roles you watch, merges every posting to its company (the same req on three boards counts once), and writes a **company-level signal row**:

| Field | What it tells you |
| --- | --- |
| `company`, `domain`, `website`, `linkedinUrl`, `careersUrl` | Who — ready for enrichment or CRM matching |
| `signalScore` (0–100) | One sortable number: volume × recency × cross-board presence × momentum × intent breadth |
| `matchingOpenRoles`, `matchedRoleTitles`, `matchedQueries` | How many of *your* roles are open, and which |
| `totalOpenRoles`, `departments`, `seniorityMix` | The whole hiring picture when a career page is readable — every department, from intern to VP |
| `momentum`, `momentumTrend`, `previousMatchingOpenRoles` | Δ open roles vs your previous run: `growing`, `steady`, `shrinking`, `new`, `gone` |
| `newRolesSinceLastRun`, `closedRolesSinceLastRun`, `newLocationsSinceLastRun` | Exactly what opened and closed, and where a company started hiring for the first time |
| `rolesPosted7d`, `rolesPosted30d`, `newestPostedAt` | Velocity — useful even on the very first run |
| `signalTags`, `intents[]` | The role→intent layer: `scaling-outbound`, `building-data-team`, `investing-in-security`, `opening-berlin`, `hiring-surge-ahead`… each with evidence titles and a strength |
| `locations`, `countries`, `remoteRoles`, `boards`, `isUrgentHiring` | Where, how many remote, on which boards, and whether the company flagged the hire as urgent |
| `roles[]` | Up to 25 underlying roles (title, location, seniority, department, posted date, link) for drill-down |
| `industry`, `companySize` | Firmographics when a board reports them — also usable as discovery filters |

The **Companies hiring** dataset view is the signal table; **Momentum** shows what changed; **Intent & departments** shows the sell-into layer.

### Companies hiring: discover or watch

- **Discovery** — leave the watchlist empty. Enter the roles and locations, optionally narrow by industry words or headcount band, and get every company hiring those roles in that market, ranked by signal score.
- **Watchlist** — paste your target accounts (company names, domains, LinkedIn company URLs or career-board handles). Every account gets a row, including the ones *not* hiring, so an agent can tell "no signal" from "not checked". Career pages of watchlisted companies are read directly for the complete open-role picture.

### Hiring signals that move: momentum across runs

Run the same watch again (same roles + locations + watchlist, or the same **Watch ID**) and each row carries its change since last time. Momentum is measured, not sampled: every company seen before is re-read directly on every run (its career page and a company-scoped board search), so a company is reported as `momentumTrend: gone` only when its own checks come back empty — once, and it is tagged `resumed-hiring` if it comes back. `momentum` is the exact change in matching roles; `momentumTrend` calls it growing or shrinking only when the change is meaningful (at least two roles and 20 %, or a 50 % swing), so a +1 on a company with 30 open roles stays `steady`. Put it on a daily or weekly **Schedule** and the dataset becomes a live hiring-momentum feed.

### Role → intent mapping (what hiring for X means)

The intent layer is a curated vocabulary of role families, each mapped to a buying signal and to what typically gets bought:

| Hiring for… | Tag | Sell into |
| --- | --- | --- |
| SDR / BDR / outbound | `scaling-outbound` | sales engagement, dialers, B2B data, intent data, outbound agencies |
| Account executives | `expanding-sales-team` | CRM seats, enablement, conversation intelligence |
| RevOps / SalesOps / MarketingOps | `building-revops-stack` | CRM consulting, revenue intelligence, forecasting, CPQ |
| Data / analytics engineers | `building-data-team` | warehouses, ELT, orchestration, data quality, BI |
| ML / AI / LLM engineers | `building-ai-capability` | LLM APIs, vector DBs, GPU compute, ML platforms |
| DevOps / SRE / platform | `scaling-infrastructure` | cloud, Kubernetes, observability, incident management |
| Security | `investing-in-security` | security tooling, SOC/MDR, compliance automation |
| Customer success / support | `scaling-customer-success` / `scaling-support` | CS platforms, help desks, AI support |
| Demand gen / content / product marketing | `ramping-demand-gen` / `investing-in-content-seo` / `launching-products` | marketing automation, agencies, SEO, ABM |
| Recruiters / talent acquisition | `hiring-surge-ahead` | ATS, sourcing tools, employer branding |
| Finance / legal / people ops | `finance-build-out` / `compliance-build-out` / `scaling-people-ops` | ERP, FP\&A, CLM, HRIS, payroll |
| VP / Head of / Director / C-level | `leadership-hire` | new function or reorg — executive-level budgets |
| Roles in a city the company never hired in | `opening-<city>` | local vendors, office services, regional partners |

Twenty-plus role families cover sales, marketing, CS, data, AI, engineering, infrastructure, security, product, design, people, finance, legal, operations, IT and clinical roles. Abbreviations are understood (`SDR`, `AE`, `SRE`, `CSM`, `MLE`…) and matched to their family, so `SDR` also catches "Sales Development Representative" — while the loosely related listings boards return for a keyword (retail sales associates for an SDR search) are gated out by title.

### Who uses it

- **Outbound / GTM teams** — a signal feed for account prioritisation: who is scaling the function you sell into, this week.
- **Sales-intelligence & data teams** — a company-level hiring dataset to join with enrichment and CRM data instead of deduping postings yourself.
- **Agencies and consultancies** — spot companies building a capability before they go to market for help.
- **Investors and analysts** — hiring momentum by company, department and city as a growth proxy.
- **AI agents** — a keyword-discoverable tool that runs unattended and returns a clean, typed dataset.

### Set it up in a minute

1. **Roles that signal intent** — one per line (`Sales Development Representative`, `Data Engineer`, `Revenue Operations`).
2. **Locations** — cities, regions or countries; or leave empty with **Remote roles only**.
3. Optionally paste a **watchlist**, set **industry** / **size** filters, or pick the **posting window** (default: last 30 days).
4. Start. Re-run (or schedule) to get momentum.

Sources: several public job boards plus employer career pages, all built in — nothing is outsourced to third-party scrapers and no nested Actor is run. Each board can be toggled; the defaults are sized for a run of a few minutes on a national search, well under a minute on a city or a watchlist.

### Ask your AI assistant (MCP)

Connect [mcp.apify.com](https://mcp.apify.com) to Claude, Cursor or any MCP client and ask: *"Which companies in Berlin are hiring data engineers right now?"*, *"Give me hiring intent signals for SDR roles in the US this week"*, *"Is anyone on my watchlist scaling outbound?"* — the assistant runs this Actor with the right input and reads the company rows back. The input has no required secrets, every field has a working default, and the output is one dataset with typed, stable field names, so agents can call it unattended and poll the same watch for momentum.

Via API:

```bash
curl -X POST "https://api.apify.com/v2/acts/inovaflow~hiring-intent-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"roles": ["SDR", "Account Executive"], "locations": ["United States"], "postedWithin": "week", "watchId": "outbound-us"}'
```

### Pricing

Pay per **company signal** delivered — a company row with at least one matching open role. Rows for watchlisted companies with nothing open, rows for companies that stopped hiring, and empty runs are free, and so are duplicates (a company is one row no matter how many boards list it). A one-role, one-country run typically returns 50–150 companies. The platform's small per-run start fee applies. Cap what you pay with **Maximum companies to report**.

### FAQ

**How is this different from a job scraper?** A job scraper returns postings for job seekers. This returns companies for sellers: merged across boards, deduplicated by company, with momentum, intent tags and firmographics — the dataset an outbound tool consumes directly.

**Does it use my LinkedIn account or any login?** No. Only publicly visible listings and public career pages are read; no credentials are needed.

**What does `matchingOpenRolesAnywhere` mean?** For companies whose career page was readable, the number of matching roles company-wide, regardless of your searched locations — "hiring 2 SDRs in Austin, 12 worldwide".

**Can I get the raw postings too?** Turn on **Also export the underlying postings** — they go to a separate named dataset (ID and URL in the run's `OUTPUT` record) so the company table stays clean.

**Why did a company show `new` and then `gone`?** Discovery samples very broad searches, so a company can first surface on any run (`new`). Once seen, it is re-read directly every run, and `gone` means those direct reads found none of your roles. Raise **Max postings per role × location × board** to widen discovery.

**Does it work outside the US?** Yes — locations map to the right country editions of each board; board strength varies by country, which is why several boards are on by default.

**Can nested runs bill a different Apify account?** Optionally: set the **MCP connector** field (Advanced) and the board searches run as our published Actors under that account. Leave it empty to run everything built in, on this run.

### Notes

Reads publicly visible listings and public career pages only. Signals are derived from public hiring activity and role titles; treat them as a prioritisation input, not a statement about any company's plans. Momentum state lives in a named key-value store in your own Apify storage (`hiring-intent-state-<watchId>`); delete it to reset a watch.

# Actor input Schema

## `roles` (type: `array`):

Job titles or keywords, one per line. Abbreviations are understood (SDR, BDR, AE, SRE, CSM…) and matched to their role family, so `SDR` also catches "Sales Development Representative". Boards return loosely related roles too — only postings whose TITLE matches the role family are counted.

## `locations` (type: `array`):

Cities, regions or countries to search in, one per line — e.g. `United States`, `Berlin, Germany`, `London`, `Sydney`. Each role is searched in each location. Leave empty with Remote only for a location-free search.

## `postedWithin` (type: `string`):

Only count postings published inside this window. A shorter window = a sharper "hiring right now" signal; a longer one = fuller coverage of every company hiring the role.

## `seniority` (type: `array`):

Keep only postings at these levels (derived from the title). Empty = any level. Executive-level hires (VP, Head of, Director, Chief) are the strongest "new function is being built" signal.

## `remoteOnly` (type: `boolean`):

Count only remote postings.

## `companies` (type: `array`):

Target accounts, one per line: company names (`Gong`), domains (`gong.io`), LinkedIn company URLs, or career-board handles (`greenhouse:stripe`, `lever:palantir`, `ashby:ramp`). When set, only these companies are reported, and their career pages are read directly for the full open-role picture. Leave empty to discover companies from the role searches.

## `industries` (type: `array`):

Keep only companies whose reported industry contains one of these words, e.g. `software`, `saas`, `financial`, `health`. Companies with no reported industry are kept. Applies to discovery runs only.

## `companySizes` (type: `array`):

Keep only companies in these headcount bands when a board reports the size. Companies with an unknown size are kept. Applies to discovery runs only.

## `hideStaffingAgencies` (type: `boolean`):

Drop recruiting / staffing agencies — an agency posting a role for a client is not the client's buying signal. Turn off if agencies ARE your buyers.

## `useLinkedIn` (type: `boolean`):

Public LinkedIn job listings — no login, no cookies. Uses residential proxy (included in Apify Proxy) for reliability.

## `useIndeed` (type: `boolean`):

Indeed job search in 60+ countries — also the source of company industry, headcount and website for the discovery filters and domain matching.

## `useCareerPages` (type: `boolean`):

Read employer career pages directly (free, no proxy). For watchlisted companies this gives the complete open-role picture; for discovered companies the top ones are probed automatically to add the careers URL and total open roles.

## `useGoogleJobs` (type: `boolean`):

Adds Google for Jobs, which aggregates 20+ boards and employer sites. Off by default because each result page uses Apify's Google SERP proxy (a few cents per role × location). Turn on for maximum coverage of smaller boards.

## `minMatchingRoles` (type: `integer`):

Companies with fewer matching open roles than this are not reported (they are still remembered for momentum). Raise to 2–3 to keep only companies hiring the role in volume.

## `maxCompanies` (type: `integer`):

Cap on delivered company rows (highest signal score first). Also caps what you pay.

## `includePostings` (type: `boolean`):

Write every underlying job posting (title, company, location, posted date, board, URL) to a separate named dataset for drill-down. Its ID and URL are in the run's OUTPUT record. Off by default — company rows already carry up to 25 matching roles each.

## `watchId` (type: `string`):

Optional name for this watch, e.g. `outbound-us-q4`. Runs sharing a Watch ID share momentum state, so an agent can poll one watch with slightly different inputs. Leave empty to derive it from roles + locations + watchlist.

## `maxJobsPerSearch` (type: `integer`):

How many postings to read per role × location on each board before aggregating (newest first). Broad roles in big markets have 300+ postings a month, so a bigger budget = more companies and steadier momentum; the run reports the date its coverage stopped at. Board pages hold 10 (LinkedIn), 100 (Indeed) or 10 (Google) postings.

## `maxCareerPageProbes` (type: `integer`):

In discovery runs, how many of the top companies get their Greenhouse / Lever / Ashby board probed for the careers URL and total open roles. Free, but each probe is up to three requests.

## `maxConcurrency` (type: `integer`):

Concurrent board requests. Lower it if a board starts rate-limiting.

## `proxyConfiguration` (type: `object`):

Proxy used for the LinkedIn source (Indeed and career pages use Apify's datacenter proxy / direct requests). Apify residential proxy is recommended — LinkedIn rate-limits datacenter IPs.

## `apifyMcpConnector` (type: `string`):

Optional. Connect an Apify MCP connector (Console → Settings → Integrations → MCP Connectors, server URL https://mcp.apify.com) and the board searches run as our published Actors under THAT account — for agencies billing a client, or to keep this run's own footprint minimal. Leave empty (the default) to run everything built in, on this run, with no nested Actor.

## Actor input object example

```json
{
  "roles": [
    "SDR",
    "Data Engineer",
    "Revenue Operations"
  ],
  "locations": [
    "New York",
    "Austin, Texas"
  ],
  "postedWithin": "month",
  "seniority": [],
  "remoteOnly": false,
  "companies": [
    "gong.io",
    "greenhouse:stripe",
    "https://www.linkedin.com/company/notion"
  ],
  "industries": [
    "software",
    "information technology"
  ],
  "companySizes": [],
  "hideStaffingAgencies": true,
  "useLinkedIn": true,
  "useIndeed": true,
  "useCareerPages": true,
  "useGoogleJobs": false,
  "minMatchingRoles": 1,
  "maxCompanies": 200,
  "includePostings": false,
  "watchId": "outbound-us",
  "maxJobsPerSearch": 200,
  "maxCareerPageProbes": 100,
  "maxConcurrency": 6,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `companies` (type: `string`):

One row per company: matching open roles, signal score, momentum vs the previous run, intents, departments, locations, boards.

## `momentum` (type: `string`):

Growing / shrinking / new / stopped-hiring companies with the roles and cities that opened or closed since your last run.

## `summary` (type: `string`):

Counts, per-source report, top signals, intent-tag totals, the watch id and (if enabled) the postings dataset id.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "roles": [
        "Sales Development Representative",
        "Account Executive"
    ],
    "locations": [
        "United States"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/hiring-intent-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "roles": [
        "Sales Development Representative",
        "Account Executive",
    ],
    "locations": ["United States"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/hiring-intent-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "roles": [
    "Sales Development Representative",
    "Account Executive"
  ],
  "locations": [
    "United States"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call inovaflow/hiring-intent-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/hiring-intent-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aUl1cZrnOf8MLF5pw/builds/EGxQeGlW0l323vw6e/openapi.json
