# ATS Jobs API — Greenhouse, Lever & Ashby Job Scraper (`kittiwake/ats-jobs-api`) Actor

Open jobs from Greenhouse, Lever and Ashby boards through their official public APIs, with a per-company hiring summary by department and location, and a watch mode that reports jobs opened and closed since the last run.

- **URL**: https://apify.com/kittiwake/ats-jobs-api.md
- **Developed by:** [Kittiwake Data](https://apify.com/kittiwake) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.05 / 1,000 job returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ATS Jobs API — Greenhouse, Lever & Ashby Job Scraper

**A Greenhouse jobs API, a Lever jobs scraper and Ashby jobs in one Actor — plus the hiring signals
on top.** Give it a list of companies. For each one it reads every open job from the company's
Greenhouse, Lever or Ashby job board through that ATS's **official public posting API**, and returns
clean job rows and a per-company summary: open roles by department and by location, and how many
jobs opened and closed since the last run.

Schedule it weekly in **watch mode** and you get only what changed — the jobs a company just opened
and the ones it just closed. That is the hiring-surge signal: a team that opens ten sales roles in a
month is buying, growing, or both.

### What it is for

- **B2B sales and account-based marketing** — hiring signals for your target accounts: who is
  opening roles in the department you sell to, and where.
- **Recruiters and talent teams** — a clean feed of new openings at the companies you track.
- **Investors and analysts** — headcount intent by department and location, week over week.
- **Job boards and aggregators** — structured job rows from three ATSs through one input format.

### What makes it different

| | a raw job scraper | this Actor |
|---|---|---|
| input | one ATS, one URL format | Greenhouse, Lever and Ashby URLs, `ats:token`, or just a company domain |
| output | job rows | job rows **and** a per-company summary: open jobs by department and top 10 locations |
| change | re-download everything | `mode: "watch"` returns only jobs opened or closed since the last run |
| personal data | whatever the page has | no person fields at all; descriptions (which can name a hiring manager) are opt-in |
| source | HTML scraping | each ATS's official public job-posting API |

### Quick start

**All open jobs at three companies:**

```json
{
  "companies": [
    "https://boards.greenhouse.io/airbnb",
    "https://jobs.lever.co/palantir",
    "https://jobs.ashbyhq.com/ashby"
  ]
}
```

**Weekly hiring-signal watch, from domains and tokens:**

```json
{
  "companies": ["palantir.com", "greenhouse:airbnb", "ashby:notion"],
  "mode": "watch"
}
```

The first watch run of a board returns its open jobs as a **baseline** (`changeType: "baseline"`);
every run after that returns only `opened` and `closed` rows, plus the summaries.

### Input

| field | type | default | what it does |
|---|---|---|---|
| `companies` | string\[] | **required** | One per company: a board URL (`boards.greenhouse.io/…`, `job-boards.greenhouse.io/…`, `jobs.lever.co/…`, `jobs.ashbyhq.com/…`), a token with its ATS (`greenhouse:airbnb`, `lever:palantir`, `ashby:ashby`), or a company domain (`palantir.com`) |
| `mode` | `scan` | `watch` | `scan` | `scan`: every open job. `watch`: only jobs opened or closed since the last run |
| `includeDescription` | boolean | `false` | Add each job's description as plain text |
| `maxJobsPerCompany` | integer | `50` | Cap on job rows per company in scan mode and on a first watch run; watch-mode opened/closed rows are never capped; the summary still counts every job |
| `stateStoreName` | string | `ats-jobs-api-state` | Named key-value store that remembers each board's open jobs between runs |
| `requestDelayMs` | integer | `1000` | Pause between requests. Lever calls are always at least 1 s apart |

**Domains.** For a bare domain the Actor checks its `robots.txt`, then reads `/careers` (and `/jobs`
only if `/careers` has no board link) **once**, and looks for a Greenhouse, Lever or Ashby board link.
If none is found the company is reported as `ats-not-detected` and **not charged**. A bare token
without an ATS (`airbnb`) is not guessed — the same token can be a different company on each ATS —
so write `greenhouse:airbnb`.

### Output

Two row types in one dataset, with a **Jobs** view and a **Company summaries** view.

**A job row** (`type: "job"`):

```json
{
  "type": "job",
  "company": "palantir",
  "ats": "lever",
  "jobId": "6ed76ce8-4156-4b60-b120-403538bd66cd",
  "title": "Administrative Business Partner",
  "department": null,
  "team": "Administrative",
  "location": "Singapore, Singapore",
  "workplaceType": "hybrid",
  "remote": false,
  "employmentType": "full-time",
  "publishedAt": "2026-08-11T17:38:11.368Z",
  "updatedAt": null,
  "url": "https://jobs.lever.co/palantir/6ed76ce8-4156-4b60-b120-403538bd66cd",
  "changeType": "baseline",
  "checkedAt": "2026-09-26T07:20:46.388Z"
}
```

`changeType` is `baseline` (first run of that board), `opened`, `closed` (watch mode), or `null` (a
job already there last run, in scan mode). `workplaceType` is `remote`, `hybrid`, `onsite` or `null`;
`employmentType` is normalised to `full-time`, `part-time`, `contract`, `intern` or `temporary` where
the ATS says so. Greenhouse publishes no employment type or workplace field, so those are `null` for
Greenhouse jobs (`remote` is `true` only when the location says "Remote"). `description` appears only
with `includeDescription: true`.

**A company summary row** (`type: "company-summary"`), one per company per run:

```json
{
  "type": "company-summary",
  "company": "ashby",
  "ats": "ashby",
  "input": "https://jobs.ashbyhq.com/ashby",
  "status": "ok",
  "openJobs": 66,
  "byDepartment": { "Engineering": 24, "Customer Success": 23, "Sales": 13, "Marketing": 3, "Design": 2, "People & Talent": 1 },
  "byLocation": { "Remote - US": 36, "United Kingdom": 8, "Remote - European Union": 7, "Remote - Canada": 5, "Australia": 4 },
  "opened": null,
  "closed": null,
  "source": "https://api.ashbyhq.com/posting-api/job-board/ashby",
  "message": null,
  "checkedAt": "2026-09-26T07:20:47.638Z"
}
```

`status` is `ok`, `not-found` (the ATS has no public board by that name), `ats-not-detected` or
`fetch-failed`. `byLocation` holds the top 10 locations. `byDepartment` uses the team where a board
publishes no department (common on Lever). `opened` and `closed` are `null` on a board's first run.
A run summary is saved as `RUN_SUMMARY` in the run's key-value store.

### What you pay for

| event | price | when |
|---|---|---|
| **Job returned** | $0.0015 | one job row written — every open job in scan mode, and a board's baseline on its first watch run |
| Job opened or closed | $0.003 | one `opened` or `closed` row in watch mode. Never for a baseline |
| Company summary | $0.005 | one company's board read and summarised successfully |

Each row is charged once — a watch-mode change is never also charged as a job. Charges are made only
after the board was read, the rows written and the state saved. **Nothing is charged** for a board
that does not exist, a company with no board detected, or a failed request. The run stops at the
spending limit you set and writes no more job rows than that limit covers.

Apify Store discounts apply by subscription plan: the prices above are the Free-plan prices;
Starter pays 10% less, Scale 20% less and Business 30% less on every event.

**Example:** a weekly watch of 50 companies that together open or close 40 jobs a week costs
50 × $0.005 + 40 × $0.003 = **$0.37 a week** after the first run.

### Sources and legal basis

This Actor calls only each ATS's **official, documented, public job-posting API**, with GET
requests and one identifying User-Agent. It never touches an application or apply endpoint.

| ATS | endpoint | what the publisher says | robots.txt |
|---|---|---|---|
| Greenhouse | `boards-api.greenhouse.io/v1/boards/{token}/jobs` | [Job Board API docs](https://developers.greenhouse.io/job-board.html): *"Job Board data is publicly available, so authentication is not required for any GET endpoints"* | disallows only `/embed/` |
| Lever | `api.lever.co/v0/postings/{company}?mode=json` | [postings-api README](https://github.com/lever/postings-api): *"all job postings in the `published` state are publicly viewable"* | `Allow: /`, `Crawl-delay: 1` — honoured: Lever calls are at least 1 s apart |
| Ashby | `api.ashbyhq.com/posting-api/job-board/{name}` | publishes a [Job Postings API](https://developers.ashbyhq.com/docs/public-job-posting-api) for public job boards; no key needed | none |

**SmartRecruiters is deliberately not supported:** its API host's robots.txt disallows every crawler
except LinkedInBot. Workday, iCIMS and others without a documented public posting API are out of
scope too.

A company domain's careers page is fetched only if its `robots.txt` allows it; an unreadable
`robots.txt` is treated as "no". Private, local and IP-address hosts are refused. A 403 or 429 from
any host is not retried, and that host is not asked again in the same run.

### Personal data

No row carries a person field — no recruiter, hiring manager, email or phone. Greenhouse's
`metadata` and `data_compliance` blocks are never read. Job descriptions are free text written by
the company and **can** name a hiring manager, so they are off by default; switch
`includeDescription` on only if you need them, and handle them under your own data-protection terms.

### FAQ

**Which ATSs are supported?** Greenhouse, Lever and Ashby. Paste any board URL, job URL, `ats:token`
or company domain.

**How do I find a company's board token?** It is the first path segment of its board URL:
`boards.greenhouse.io/`**`airbnb`**, `jobs.lever.co/`**`palantir`**, `jobs.ashbyhq.com/`**`ashby`**.
Or give the company domain and let the Actor find it.

**How do I get only new jobs?** Use `mode: "watch"` on a schedule. Keep the same `stateStoreName`
between runs; use a different one per list if you run several.

**Why are some fields null?** The ATS does not publish them. Nothing is guessed or inferred beyond
the `remote` flag for Greenhouse locations that say "Remote".

**Can I export to CSV or Excel?** Yes. The dataset downloads as JSON, CSV, Excel or XML, or read it
over the Apify API. Use the Jobs view for a flat table.

**Is this a LinkedIn or Indeed scraper?** No. It reads companies' own job boards, which is where
LinkedIn and Indeed listings usually come from.

### Support

Use the **Issues** tab on this Actor. Include the run id and the input you used.

# Actor input Schema

## `companies` (type: `array`):

One item per company. A job-board URL (https://boards.greenhouse.io/airbnb, https://jobs.lever.co/palantir, https://jobs.ashbyhq.com/ashby), a board token with its ATS (greenhouse:airbnb, lever:palantir, ashby:ashby), or a company domain (palantir.com) — for a domain the Actor reads its /careers page, then /jobs, once, and looks for a Greenhouse, Lever or Ashby board link. Items with no board found are reported as ats-not-detected and not charged.

## `mode` (type: `string`):

scan: every open job plus a summary per company. watch: only jobs opened or closed since the last run, plus the summary. The first run of a board in watch mode returns its open jobs as a baseline.

## `includeDescription` (type: `boolean`):

Add the job description as plain text. Off by default: descriptions are free text and sometimes name a hiring manager or recruiter. No other field names a person.

## `maxJobsPerCompany` (type: `integer`):

Cap on job rows written per company in scan mode and on a board's first watch run. Watch-mode opened/closed rows are never capped. The summary still counts every open job.

## `stateStoreName` (type: `string`):

Named key-value store that keeps each board's open jobs between runs, so the next run can tell what opened and closed. Use a different name per watchlist if you run several.

## `requestDelayMs` (type: `integer`):

Pause between requests. Lever calls are always at least 1 second apart, as Lever's robots.txt asks.

## Actor input object example

```json
{
  "companies": [
    "https://boards.greenhouse.io/airbnb",
    "https://jobs.lever.co/palantir",
    "https://jobs.ashbyhq.com/ashby"
  ],
  "mode": "scan",
  "includeDescription": false,
  "maxJobsPerCompany": 50,
  "stateStoreName": "ats-jobs-api-state",
  "requestDelayMs": 1000
}
```

# Actor output Schema

## `jobs` (type: `string`):

No description

## `companies` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "https://boards.greenhouse.io/airbnb",
        "https://jobs.lever.co/palantir",
        "https://jobs.ashbyhq.com/ashby"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("kittiwake/ats-jobs-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "companies": [
        "https://boards.greenhouse.io/airbnb",
        "https://jobs.lever.co/palantir",
        "https://jobs.ashbyhq.com/ashby",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("kittiwake/ats-jobs-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "https://boards.greenhouse.io/airbnb",
    "https://jobs.lever.co/palantir",
    "https://jobs.ashbyhq.com/ashby"
  ]
}' |
apify call kittiwake/ats-jobs-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,kittiwake/ats-jobs-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dfcSsqwB8K3dCcI4M/builds/rUOxRHttwVDy7zVmd/openapi.json
