# Greenhouse, Lever & Ashby Job Board Scraper with Change Feed (`magenta_courser/niche-job-feed-scraper`) Actor

Enter company job boards on Greenhouse, Lever or Ashby, get every open job as a flat row: title, location, department, published salary, plain-text description, URL, plus new/changed flags since your last run. Official public APIs, no proxy. Pay per job row.

- **URL**: https://apify.com/magenta_courser/niche-job-feed-scraper.md
- **Developed by:** [SUNGHWAN CHO](https://apify.com/magenta_courser) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.80 / 1,000 job rows

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Greenhouse, Lever & Ashby Job Board Scraper with Change Feed

Enter company job boards on Greenhouse, Lever or Ashby and get every open job as one flat JSON row: title, location, department, published salary, plain-text description and URL, plus `isNew` / `isChanged` flags since your last run.

Built for niche job sites, recruiting tools and market researchers who **sync the same list of company boards every day**. It reads the three platforms' official public job-board APIs with one request per board. No browser, no proxy, no login.

### What you get

- **One schema for three ATS platforms.** The same 26 fields for Greenhouse, Lever and Ashby; a value a platform does not publish is `null`.
- **Change feed.** Give the run a history name and each row says whether the job is new or changed. Turn on `onlyChanges` to store only those.
- **Salaries as numbers, only when published.** `salaryMin`, `salaryMax`, `salaryCurrency`. Nothing is estimated from the title, and equity is never counted as salary.
- **Plain-text descriptions.** Greenhouse's escaped HTML is converted to text; e-mail addresses are masked.
- **Honest status per board.** Every board ends as `completed`, `partial`, `failed`, `blocked` or `budgetLimited` in the free `RUN_SUMMARY` record. A failed board is never reported as "no open jobs".
- **Pay per job row.** Failed boards are free.

### What this is not

- **Not a company database or job search engine.** You supply the boards. The Actor does not discover which companies use Greenhouse, Lever or Ashby, and `keyword` only filters the boards you listed.
- **Only Greenhouse, Lever and Ashby.** No Workday, SmartRecruiters, LinkedIn or Indeed.
- **No "job closed" signal.** A job missing from a later run is not reported as closed.

### Quick start

1. Open a company's careers page and copy the board URL, for example `https://boards.greenhouse.io/stripe`, `https://jobs.lever.co/spotify` or `https://jobs.ashbyhq.com/Ashby`.
2. Paste the URLs into **Company job boards** and run.
3. For a daily feed, set a **History name** (for example `my-job-feed`), turn on **Save only new or changed jobs** and schedule the Actor.

### Input

Only `boards` is required.

| Field | Type | Default | Description |
|---|---|---|---|
| `boards` | array | required | Up to 100 boards per run. Each item is a board URL, `"source:slug"`, or an object (see below). |
| `maxJobsPerBoard` | integer | 100 | Store at most this many jobs from each board (1 to 1000). A board with more is `partial`. |
| `keyword` | string | none | Keep only jobs whose title, location or department contains this text. |
| `historyStoreName` | string | none | Name of a key-value store in your account that remembers saved jobs. Turns on change tracking. |
| `onlyChanges` | boolean | `false` | Store only new or changed jobs. Needs `historyStoreName`. |
| `maxItems` | integer | 1000 | Maximum rows stored per run (up to 10,000). |
| `maxRequests` | integer | 300 | Maximum HTTP requests. Each board needs one. |
| `maxRunSeconds` | integer | 600 | Time budget. Boards not reached are `budgetLimited`. |
| `requestIntervalMs` | integer | 1000 | Pause between two requests (minimum 500). |

A board object:

| Key | Description |
|---|---|
| `source` | `greenhouse`, `lever` or `ashby`. Required. |
| `slug` | The company part of the board URL (`stripe` in `boards.greenhouse.io/stripe`). Required. Ashby slugs are case-sensitive. |
| `company` | Optional display name written to the `company` field. |
| `region` | `eu` for Lever boards on `jobs.eu.lever.co`. |

#### Example 1: snapshot of three boards

```json
{
  "boards": [
    "https://boards.greenhouse.io/stripe",
    "https://jobs.lever.co/spotify",
    "https://jobs.ashbyhq.com/Ashby"
  ],
  "maxJobsPerBoard": 20
}
```

#### Example 2: daily feed of new and changed jobs

```json
{
  "boards": [
    { "source": "greenhouse", "slug": "airbnb" },
    { "source": "greenhouse", "slug": "figma" },
    { "source": "ashby", "slug": "notion", "company": "Notion" },
    { "source": "lever", "slug": "palantir", "company": "Palantir" }
  ],
  "maxJobsPerBoard": 1000,
  "historyStoreName": "my-job-feed",
  "onlyChanges": true
}
```

#### Example 3: only data roles

```json
{
  "boards": ["greenhouse:databricks", "ashby:openai"],
  "keyword": "data",
  "maxJobsPerBoard": 200
}
```

### Output

One row per job. All fields are always present; a value the board does not publish is `null`.

Real rows from a run on 2026-10-02 (descriptions shortened here):

```json
{
  "recordKey": "greenhouse|global|airbnb|8184174",
  "source": "greenhouse",
  "board": "airbnb",
  "company": "Airbnb",
  "jobId": "8184174",
  "title": "Account Manager",
  "location": "London, United Kingdom",
  "country": null,
  "department": "Business Development",
  "employmentType": null,
  "workplaceType": null,
  "salaryMin": 46000,
  "salaryMax": 54000,
  "salaryCurrency": "GBP",
  "salaryInterval": null,
  "salaryText": "United Kingdom Annual Pay Range",
  "description": "Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 b...",
  "publishedAt": "2026-09-09T04:35:19-04:00",
  "updatedAt": "2026-09-29T20:00:54-04:00",
  "url": "https://careers.airbnb.com/positions/8184174?gh_jid=8184174",
  "availability": "listed",
  "schemaVersion": 1,
  "observedAt": "2026-10-02T07:00:16.968Z",
  "isNew": true,
  "isChanged": null,
  "previousObservedAt": null
}
```

```json
{
  "recordKey": "ashby|global|Ashby|7458d4e9-da2e-47bd-98cb-adfda43d42b2",
  "source": "ashby",
  "board": "Ashby",
  "company": "Ashby",
  "jobId": "7458d4e9-da2e-47bd-98cb-adfda43d42b2",
  "title": "Engineering Manager - EU",
  "location": "Remote - European Union",
  "country": "European Union",
  "department": "Engineering",
  "employmentType": "FullTime",
  "workplaceType": "Remote",
  "salaryMin": 110000,
  "salaryMax": 185000,
  "salaryCurrency": "EUR",
  "salaryInterval": "1 YEAR",
  "salaryText": "€110K - €185K",
  "description": "Hi 👋 I’m Colin https://uk.linkedin.com/in/colinhoweuk, Director of Engineering, Europe. How do you feel about software engineers writing product specs, making p...",
  "publishedAt": "2024-03-04T14:29:08.532+00:00",
  "updatedAt": null,
  "url": "https://jobs.ashbyhq.com/Ashby/7458d4e9-da2e-47bd-98cb-adfda43d42b2",
  "availability": "listed",
  "schemaVersion": 1,
  "observedAt": "2026-10-02T07:00:16.968Z",
  "isNew": true,
  "isChanged": null,
  "previousObservedAt": null
}
```

| Field | Type | Greenhouse | Lever | Ashby | Description |
|---|---|---|---|---|---|
| `recordKey` | string | yes | yes | yes | Stable key: source, region, board, job id. Join runs on it. |
| `source`, `board` | string | yes | yes | yes | Platform and board slug. |
| `company` | string | published name | your label or slug | your label or slug | The `company` you gave wins. |
| `jobId`, `title`, `url` | string | yes | yes | yes | |
| `location` | string | null | yes | yes | yes | As published. |
| `country` | string | null | `null` | ISO code | address country | |
| `department` | string | null | yes (joined with `\|`) | department or team | yes | |
| `employmentType` | string | null | `null` | yes | yes | As published, not normalized. |
| `workplaceType` | string | null | `null` | yes | yes | As published (`Remote`, `hybrid`, ...). |
| `salaryMin`, `salaryMax` | number | null | single pay range only | when published | single salary component only | |
| `salaryCurrency` | string | null | yes | yes | yes | |
| `salaryInterval` | string | null | `null` | yes | yes | For example `1 YEAR`. |
| `salaryText` | string | null | range title | salary description | summary text | |
| `description` | string | null | plain text | plain text | plain text | E-mail addresses become `[email]`. |
| `publishedAt` | string | null | first published | created | published | ISO 8601. |
| `updatedAt` | string | null | yes | `null` | `null` | |
| `availability` | string | `listed` | `listed` | `listed` | The job was on the board at `observedAt`. |
| `schemaVersion` | integer | | | | `1`. |
| `observedAt` | string | | | | Run observation time, ISO 8601 UTC. |
| `isNew` | boolean | | | | Not in the history before (always `true` without a history name). |
| `isChanged` | boolean | null | | | | Any field differs from the last saved observation. `null` when new. |
| `previousObservedAt` | string | null | | | | Time of the last saved observation. |

Dataset views: **Jobs** (core columns), **Salaries** and **Change feed**.

#### Salary coverage we saw

In the 3,970-job test run: Ashby 890 of 1,219 rows had a salary range (73%), Greenhouse 550 of 2,351 (23%), Lever 0 of 400 (the two Lever companies tested publish none). When a Greenhouse job lists several pay ranges (for example one per US pay zone), the salary fields stay `null` instead of picking one.

#### Run summary (free)

The key-value record `RUN_SUMMARY` has one entry per board. Shortened real example from the first test run (14 boards, 50 jobs per board):

```json
{
  "status": "partial",
  "stopReason": "someTargetsIncomplete",
  "targetsRequested": 14,
  "targetsCompleted": 1,
  "targetsPartial": 10,
  "targetsFailed": 3,
  "rowsSaved": 530,
  "requests": 14,
  "targets": [
    { "target": "greenhouse:stripe", "status": "partial", "reason": "Board observation cap; no expiry inference", "observed": 50, "saved": 50, "unchanged": 0, "requests": 1 },
    { "target": "ashby:linear", "status": "completed", "reason": null, "observed": 30, "saved": 30, "unchanged": 0, "requests": 1 },
    { "target": "lever:plaid", "status": "failed", "reason": "HTTP 404", "observed": 0, "saved": 0, "unchanged": 0, "requests": 1 }
  ]
}
```

| Board status | Meaning | Jobs charged |
|---|---|---|
| `completed` | The whole board was read and stored. A board with no open jobs is `completed` with 0 rows. | The rows stored |
| `partial` | The board has more jobs than `maxJobsPerBoard` (or `maxItems` was reached). | The rows stored |
| `failed` | Unknown board (HTTP 404), response larger than 16 MB, or malformed data. | None |
| `blocked` | The API answered HTTP 403 or 429. | None |
| `budgetLimited` | Not reached within `maxRequests`, `maxRunSeconds` or your maximum cost. | The rows stored before the limit |

The run itself fails only when every board failed.

### Change feed

- Without `historyStoreName` every run is an independent snapshot: all rows have `isNew: true`.
- With it, the Actor keeps a hash of each **saved** job in a key-value store of that name in your account. Next run, `isNew` marks jobs not seen before and `isChanged` marks jobs where any field differs (title, location, salary, description, Greenhouse `updatedAt`, ...).
- `onlyChanges: true` stores only new or changed jobs. A quiet day can give zero rows; `RUN_SUMMARY.rowsUnchanged` shows how many jobs were identical.
- Jobs cut off by `maxJobsPerBoard` are not in the history. For a full feed set `maxJobsPerBoard` to 1000.
- Only rows that were stored move the baseline. If a dataset write fails, the history is not advanced and the run fails.
- Use one history name per board list, and do not run two runs with the same name at the same time. Up to 10,000 jobs are remembered per history.

### Pricing

Pay per event: **$0.001 for each job row stored** (event name `job`), plus the standard Actor start event ($0.00005 per run). That is $1 per 1,000 jobs.

| What you run | Job rows | Cost |
|---|---|---|
| 3 boards, 20 jobs each, once | 60 | $0.06 |
| 1 board with 400 jobs, once | 400 | $0.40 |
| 14 boards, all jobs, first sync | 3,970 | $3.97 |
| The same 14 boards daily with `onlyChanges`, if 2% of jobs are new or changed per day | about 80 per day | about $0.08 per day |

**Not charged:** boards that end as `failed` or `blocked`; unchanged jobs skipped by `onlyChanges`; the run summary.

**Good to know:**

- If you set a maximum cost per run, rows are trimmed to the budget before they are stored, so stored rows, reported rows and charged rows are always the same number. Boards that no longer fit are not fetched.
- `maxJobsPerBoard` and `maxItems` cap what a run can cost.

### Reliability: what we measured

Cloud runs on 2026-10-02, 256 MB, no proxy:

| Run | Boards | Result | Time |
|---|---|---|---|
| 14 boards, 50 jobs per board | 6 Greenhouse, 3 Lever, 5 Ashby | 530 rows. 11 boards read; 1 unknown slug (HTTP 404); 2 boards failed because their response was above the old 8 MB limit | 42 s |
| 14 boards, all jobs (limit raised to 16 MB) | 6 Greenhouse, 3 Lever, 5 Ashby | 3,970 rows, 14 of 14 boards read (one valid board had 0 open jobs), no duplicates | 95 s |
| The same 14 boards again with `onlyChanges` | 6 Greenhouse, 3 Lever, 5 Ashby | 0 rows stored, 3,970 reported unchanged, 14 of 14 boards read | 56 s |
| The two largest boards only | OpenAI (Ashby), Databricks (Greenhouse) | 1,705 rows, peak memory 120 MB of 256 MB | 6 s |

The largest boards tested were OpenAI on Ashby (833 jobs, 14.6 MB) and Databricks on Greenhouse (872 jobs, 10.3 MB). No HTTP 403 or 429 was seen. This is a same-day sample of 14 boards.

### Limits and honest notes

- **Boards above 16 MB fail explicitly.** Greenhouse and Ashby return the whole board in one response; the Actor runs in 256 MB. Such a board ends as `failed` with the reason and is free.
- **Lever boards are read in one page** of up to `maxJobsPerBoard + 1` postings.
- **No proxy, no retries.** One request per board, at least 500 ms apart. If an API answers 403 or 429 the board is `blocked`.
- **Salary only when the company publishes a single comparable range.** Several ranges or equity-only compensation give `null`.
- **`employmentType`, `workplaceType` and `salaryInterval` are passed through as published**, so values differ by platform.
- **Descriptions are public text.** E-mail addresses and Korean mobile numbers are masked; this is not complete anonymization of free text.
- The public APIs are meant for showing a company's jobs. Check each company's terms before republishing descriptions.

Sources: [Greenhouse Job Board API](https://developers.greenhouse.io/job-board.html), [Lever Postings API](https://github.com/lever/postings-api), [Ashby Job Postings API](https://developers.ashbyhq.com/docs/public-job-posting-api).

### Use from the API

```bash
curl -X POST "https://api.apify.com/v2/acts/magenta_courser~niche-job-feed-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"boards":["https://boards.greenhouse.io/stripe"],"maxJobsPerBoard":20}'
```

Schedule it in Apify Console (Schedules) or call it from n8n, Make or Zapier with the same JSON.

### Use with AI agents (MCP)

The Actor works as a tool through the Apify MCP server. One required input (`boards`, a list of board URLs), flat output fields, numbers as numbers and `null` for unknown values. A good instruction: "Get the open jobs of these boards with niche-job-feed-scraper and list the ones with `salaryMin` above 150000 USD."

### FAQ

**How do I find a company's slug?** Open its careers page and look at the job links: `boards.greenhouse.io/<slug>`, `job-boards.greenhouse.io/<slug>`, `jobs.lever.co/<slug>` or `jobs.ashbyhq.com/<slug>`. You can paste the URL as it is.

**Why HTTP 404?** The company has no public board with that slug on that platform, or a Lever EU board needs `"region": "eu"`.

**Why is salary `null`?** The company does not publish it through the board API, or publishes several ranges.

**Can it tell me when a job closes?** Not in this version. Compare `recordKey` lists between full snapshots yourself, and only for boards that ended as `completed`.

**Can I search all companies for a keyword?** No. `keyword` filters the boards you listed.

# Actor input Schema

## `boards` (type: `array`):

Public job boards to read, up to 100 per run. Each item is a board URL (https://boards.greenhouse.io/stripe, https://jobs.lever.co/spotify, https://jobs.ashbyhq.com/Ashby) or an object: {"source": "greenhouse | lever | ashby", "slug": "company part of the board URL", "company": "optional display name", "region": "eu for Lever EU boards"}.

## `maxJobsPerBoard` (type: `integer`):

Save at most this many jobs from each board. A board with more jobs is reported as partial.

## `keyword` (type: `string`):

Optional. Keep only jobs whose title, location or department contains this text (case-insensitive). It filters the boards you listed; it is not a job search across companies.

## `historyStoreName` (type: `string`):

Name of a key-value store in your account that remembers the jobs saved before, for example "my-job-feed". With it, each row gets isNew and isChanged. Leave empty for an independent snapshot. Use one name per board list and do not run two runs with the same name at the same time.

## `onlyChanges` (type: `boolean`):

Needs a history name. Jobs identical to the last saved observation are skipped, so a run can save zero rows.

## `maxItems` (type: `integer`):

Stop saving after this many job rows in total.

## `maxRequests` (type: `integer`):

Each board needs one request. Boards that do not fit are reported as budgetLimited.

## `maxRunSeconds` (type: `integer`):

Boards not reached within this time are reported as budgetLimited.

## `requestIntervalMs` (type: `integer`):

Minimum time between two requests.

## Actor input object example

```json
{
  "boards": [
    {
      "source": "greenhouse",
      "slug": "stripe"
    },
    {
      "source": "lever",
      "slug": "spotify"
    },
    {
      "source": "ashby",
      "slug": "Ashby"
    }
  ],
  "maxJobsPerBoard": 20,
  "onlyChanges": false,
  "maxItems": 1000,
  "maxRequests": 300,
  "maxRunSeconds": 600,
  "requestIntervalMs": 1000
}
```

# Actor output Schema

## `jobs` (type: `string`):

One row per open job with all fields, including the plain-text description.

## `overview` (type: `string`):

The same rows with the core columns only.

## `runSummary` (type: `string`):

Free key-value record: status and reason for every board (completed, partial, failed, blocked, budgetLimited), rows saved and unchanged, requests, and what was charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "boards": [
        {
            "source": "greenhouse",
            "slug": "stripe"
        },
        {
            "source": "lever",
            "slug": "spotify"
        },
        {
            "source": "ashby",
            "slug": "Ashby"
        }
    ],
    "maxJobsPerBoard": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("magenta_courser/niche-job-feed-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "boards": [
        {
            "source": "greenhouse",
            "slug": "stripe",
        },
        {
            "source": "lever",
            "slug": "spotify",
        },
        {
            "source": "ashby",
            "slug": "Ashby",
        },
    ],
    "maxJobsPerBoard": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("magenta_courser/niche-job-feed-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "boards": [
    {
      "source": "greenhouse",
      "slug": "stripe"
    },
    {
      "source": "lever",
      "slug": "spotify"
    },
    {
      "source": "ashby",
      "slug": "Ashby"
    }
  ],
  "maxJobsPerBoard": 20
}' |
apify call magenta_courser/niche-job-feed-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,magenta_courser/niche-job-feed-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JX5bCBsXUclQ4KApW/builds/nMHIjGLzdR5Ge8FMC/openapi.json
