# Hiring Signal Changefeed (Greenhouse, Lever, Ashby) (`changefeeds/ats-hiring-signal-changefeed`) Actor

Track which companies open or close roles: each run returns job changes versus the previous run from the companies' public job boards (Greenhouse, Lever, Ashby). Reads only job-board metadata, never people.

- **URL**: https://apify.com/changefeeds/ats-hiring-signal-changefeed.md
- **Developed by:** [Changefeeds Tools](https://apify.com/changefeeds) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hiring Signal Changefeed (Greenhouse, Lever, Ashby)

Give it a list of company job boards. Each run returns their **open roles
compared with the previous run of the same list**: which roles were added, which
were removed, and which changed their title, department or location. Plus a
per-board summary: open-role count, the delta since last time, and how the
roles distribute across departments.

Schedule it daily or weekly and the dataset becomes a change feed for
sales prospecting, RevOps account intel, or recruiting-market tracking — using
only each company's official public job-board JSON API (Greenhouse, Lever,
Ashby). No API keys, no logins, no scraping of people: only job metadata.

### What you get

One `board` row per board (`type: "board"`):

- `status`: `ok`, `not_found` (no such board), `error` (the API failed) or
  `invalid` (the input entry is not a recognizable board URL or shorthand)
- `open_roles`, `open_roles_delta` versus the previous run
- `added`, `removed`, `changed`, `baseline` counts, and `by_department`
- `is_baseline` (first run of the list), `previous_checked_at`

One `job` row per job with a change (`type: "job"`), with a `change_type`:

- `baseline` — first run, only when you set `includeBaselineJobs`: every open role
- `added` — a job id the board did not have last run
- `removed` — a job id that was on the board last run and is gone now (only
  ever reported when the board fetch fully succeeded)
- `changed` — the title, department or location of a known job id moved; the
  row carries `previous` (the old title/department/location)

Each row carries `id`, `title`, `department`, `team`, `location`, `remote`,
`employment_type`, `url`, `published_at`, `compensation_summary` (Ashby), and
`description` (Greenhouse/Lever, only when you set `includeDescriptions`).

The key-value store record `OUTPUT` holds a per-board summary plus totals
(`changes`, `job_rows_returned`, `first_run`, `charged_events`,
`charge_limit_reached`). If you set `webhookUrl`, the same summary is POSTed
there as JSON when the run ends.

### Input

| Field | Default | Notes |
|---|---|---|
| `boards` | required | Board URLs or shorthands: `greenhouse:<token>`, `https://boards.greenhouse.io/<token>`, `https://job-boards.greenhouse.io/<token>`, `lever:<site>`, `https://jobs.lever.co/<site>`, `ashby:<org>`, `https://jobs.ashbyhq.com/<org>`. Up to 1,000 per run. Invalid entries become an error row; the run continues. |
| `includeBaselineJobs` | false | Off: a board's first run records its current roles silently and returns only the board summary row (open roles, by department), so you start watching 100 companies for 100 board checks. On: the first run also returns every open role as a charged `baseline` row. |
| `includeDescriptions` | false | Also return each job's description text. Makes rows bigger and Greenhouse responses heavier. |
| `snapshotKey` | derived | Which saved state to compare against. By default it is derived from the sorted board list, so the same list always compares with itself. Set it explicitly to keep history when you edit the list. Distinct explicit keys always get distinct storage records (the key is sanitised and suffixed with a hash of the full original), so `watch/a` and `watch?a` never share history. |
| `webhookUrl` | none | Receives the run summary as a JSON POST. |

```json
{
  "boards": ["greenhouse:gitlab", "lever:palantir", "ashby:ramp"],
  "includeDescriptions": false
}
```

If a run finds nothing new, the dataset holds one `no_changes` row (not charged), so a quiet run is never mistaken for a broken one.

### Sample output

```json
{
  "type": "board",
  "board": "greenhouse:gitlab",
  "ats": "greenhouse",
  "token": "gitlab",
  "change_type": "board",
  "status": "ok",
  "open_roles": 199,
  "open_roles_delta": -3,
  "added": 1,
  "removed": 4,
  "changed": 2,
  "baseline": 0,
  "by_department": { "Engineering": 62, "Sales": 40 },
  "is_baseline": false,
  "previous_checked_at": "2026-09-27T06:00:05.000Z"
}
```

```json
{
  "type": "job",
  "board": "greenhouse:gitlab",
  "ats": "greenhouse",
  "company": "gitlab",
  "change_type": "added",
  "previous": null,
  "id": "9000000001",
  "title": "Brand New Role",
  "department": "Engineering",
  "team": null,
  "location": "Remote, Earth",
  "remote": null,
  "employment_type": null,
  "url": "https://job-boards.greenhouse.io/gitlab/jobs/9000000001",
  "published_at": "2026-09-28T04:00:00-04:00",
  "compensation_summary": null,
  "description": null
}
```

```json
{
  "type": "job",
  "board": "greenhouse:gitlab",
  "change_type": "changed",
  "id": "8556658002",
  "title": "AI Engineer II",
  "previous": { "title": "AI Engineer", "department": "Engineering", "location": "Remote, Bangalore" }
}
```

### Pricing

Pay per event, nothing else:

- **$0.003 per board checked** (a board that was fetched successfully and
  compared). Missing, invalid or failed boards are not charged.
- **$0.003 per job change returned**: `added`, `removed` or `changed` rows
  (and `baseline` rows, only if you turn on `includeBaselineJobs`). The board
  summary row is covered by the board check.

Worked examples:

- **Starting a watch** of 100 companies: the first run costs 100 × $0.003 =
  **$0.30** (baseline roles are recorded, not returned).
- **A daily run** over those 100 boards: $0.30 plus $0.003 per actual change.
  If each company adds, removes or edits about three roles a week, that is
  \~43 changes a day, about $0.43 a day or **~$13 a month** for 100 companies
  watched daily; fast-hiring companies cost more. Weekly runs cost a seventh of the board checks.
- **With `includeBaselineJobs` on**, the first run also returns every open role:
  20 boards averaging 150 roles cost 20 × $0.003 + 3,000 × $0.003 = $9.06
  once.

If you set a maximum total charge for the run, the actor returns only as many
rows as fit, saves state for what it returned, stops fetching further boards,
and says so in `OUTPUT` (`stopped_reason: "max_total_charge_reached"`). It
never charges for a row it did not deliver, and rows it did deliver are
remembered, so the next run does not bill them again. (Exception: if the run
is killed between delivering a board's rows and saving that board's state,
for example by a timeout or an abort, the next run reports that board's
changes again.)

### How the change feed works

- State lives in the named key-value store `hiringwatch-snapshots`, one record
  per board and snapshot key. The record holds the last checked time, the
  open-role count and, per job id, a SHA-256 hash of the title, department and
  location plus those three fields.
- The first run of a snapshot key is the baseline: every open role is
  remembered (and returned as a `baseline` row only with `includeBaselineJobs`).
- Later runs fetch the whole board again and diff job ids against the
  snapshot: new ids are `added`, missing ids are `removed`, and a known id
  whose hash changed is `changed` (with the old values in `previous`).
- A board is fetched with a single request to its official public API, so a
  successful fetch covers the whole board: removals are only ever reported
  from a complete fetch, never guessed from an error.
- Failed or missing boards get a free row with a status and never touch their
  snapshot — the next run compares against the last good state.
- A snapshot save applies only that run's own changes to what is currently
  stored, so a slower overlapping run cannot restore a job title or a job a
  newer run already replaced. (Overlapping runs are normally refused anyway;
  see the one-run-at-a-time limit below.)

### Limits, stated plainly

- **Job-board metadata only.** The actor reads each company's public job-board
  API and never visits profiles, applications or anything about people. It
  cannot see jobs a company withholds from the public board.
- **Only Greenhouse, Lever and Ashby.** Workday, SmartRecruiters, Teamtailor
  and other ATSes are not supported.
- **Boards are public pages.** A company can relabel or retire its board at any
  time; a retired board comes back `not_found` until you remove it.
- **`changed` means title, department or location.** A changed job description,
  salary band or published date alone does not emit a row (Greenhouse's
  `updated_at` also moves on unrelated edits, so it is deliberately not
  watched).
- **A board that suddenly lists no roles is checked twice.** If a board that
  had roles comes back empty, the actor waits 5 s and fetches it again before
  reporting every role as removed; a board that is still empty is taken at its
  word.
- **`published_at`** is Greenhouse's `first_published` (falling back to
  `updated_at`), Lever's `createdAt`, Ashby's `publishedAt`.
- **Removed rows carry only remembered fields.** The job is gone from the
  board, so its `url` is null and only the last-seen title, department and
  location can be shown.
- **Descriptions change without notice.** When `includeDescriptions` is on, the
  description reflects this run's fetch; description-only edits are not
  detected as changes.
- **No pagination, one request per board.** A full board listing can be large
  (some boards list thousands of roles); each response is fetched in one
  request per run.
- **At most 10,000 open jobs per board.** A larger board is reported as an
  uncharged `error` row rather than tracked partially.
- **A board with an unreadable posting is skipped.** If a board lists a posting
  without an id (or the same id twice), the board gets a free `error` row for
  that run instead of a guessed diff.
- **Ashby postings marked unlisted are ignored**, like on the public board.
- **It stays polite:** at most 2 boards fetched at a time, at least 1 s between
  two requests to the same board, HTTP 429 honoured (Retry-After, capped at
  60 s, max 3 retries; a longer Retry-After skips that board for this run), a
  descriptive User-Agent, 20 s timeouts.
- **One run at a time per snapshot key.** If a scheduled run starts while the
  previous one with the same `snapshotKey` is still going, the new run fails
  immediately, before fetching or charging anything. A run that crashed
  without cleaning up blocks its key for at most 30 minutes. (The lock is best
  effort: two runs started within the same couple of seconds can, rarely, both
  proceed; their saved history is merged, not overwritten.)
- **If saved history can't be read, the board is skipped**, with a free
  `error` row, rather than re-sent as a paid baseline. If it can't be *saved*
  after 3 attempts, the run is marked failed and names the boards, because the
  next run would report those jobs again.

### Local development

```bash
pnpm --filter @mmnm/hiringwatch test        # unit tests, no network
pnpm --filter @mmnm/hiringwatch build
```

`node dist/main.js` runs the actor locally with Apify's local storage
(`CRAWLEE_STORAGE_DIR=./storage` by default).

# Actor input Schema

## `boards` (type: `array`):

Company job boards to watch. Each entry is either a board URL — https://boards.greenhouse.io/<token>, https://job-boards.greenhouse.io/<token>, https://jobs.lever.co/<site>, https://jobs.ashbyhq.com/<org> — or a shorthand: greenhouse:<token>, lever:<site>, ashby:<org>. Boards that do not exist are reported as not\_found. Up to 1,000 per run.

## `includeDescriptions` (type: `boolean`):

Also return each job's description text (Greenhouse and Lever; Ashby's summary). Makes rows bigger and Greenhouse responses heavier.

## `includeBaselineJobs` (type: `boolean`):

Off (default): a board's first run only records its current jobs as the baseline and returns one summary row per board (open roles, by department), so starting to watch 100 companies costs 100 board checks, not thousands of job rows. Later runs return what changed. On: the first run also returns every open job as a baseline row (each charged as a job row).

## `snapshotKey` (type: `string`):

Name of the saved state used to compute changes between runs. Leave empty to derive it from the sorted board list, so the same list always compares against its own last run. Set it explicitly to keep history when you edit the list.

## `webhookUrl` (type: `string`):

Optional. When the run finishes, the OUTPUT summary is POSTed here as JSON.

## Actor input object example

```json
{
  "boards": [
    "greenhouse:gitlab",
    "lever:palantir",
    "ashby:ramp"
  ],
  "includeDescriptions": false,
  "includeBaselineJobs": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "boards": [
        "greenhouse:gitlab",
        "lever:palantir",
        "ashby:ramp"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("changefeeds/ats-hiring-signal-changefeed").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "boards": [
        "greenhouse:gitlab",
        "lever:palantir",
        "ashby:ramp",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("changefeeds/ats-hiring-signal-changefeed").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "boards": [
    "greenhouse:gitlab",
    "lever:palantir",
    "ashby:ramp"
  ]
}' |
apify call changefeeds/ats-hiring-signal-changefeed --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,changefeeds/ats-hiring-signal-changefeed"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/tVyFHXoGeoPo9rANU/builds/g9hZgjtMD8VajzcLa/openapi.json
