# ATS Jobs Scraper: Greenhouse, Lever, Workday, Ashby & more (`changefeeds/ats-jobs-scraper`) Actor

Scrape open jobs from company career pages on 11 ATS platforms (Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee, BambooHR, Personio, Teamtailor, Breezy HR) via their public job boards. Auto-detects the ATS from a company website; one normalized schema.

- **URL**: https://apify.com/changefeeds/ats-jobs-scraper.md
- **Developed by:** [Changefeeds Tools](https://apify.com/changefeeds) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ATS Jobs Scraper: Greenhouse, Lever, Workday, Ashby & more

Scrape the open jobs on company career pages, from **11 applicant tracking
systems in one actor**: Greenhouse, Lever, Ashby, Workday, Workable,
SmartRecruiters, Recruitee, BambooHR, Personio, Teamtailor and Breezy HR.

Give it careers-page URLs, `platform:token` references, or just company
websites. The actor works out which ATS each company uses, reads the
company's **public job board** (the same JSON or XML feed the careers page
itself loads), and returns every job in **one normalized schema**: title,
department, team, locations, remote, employment type, structured salary where
the platform publishes one, posting date, apply URL and, if you want it, the
full description as plain text.

Use it as a Greenhouse jobs scraper, a Workday jobs scraper, a Lever or Ashby
jobs scraper, or a general career page scraper for a list of target companies:
job aggregators, recruiting and sourcing tools, sales prospecting on hiring
signals, labour-market research.

No API keys, no logins, no browser. Only public job-board data.

### Supported platforms

| Platform | Careers-page URL shapes you can paste | Shorthand | Public endpoint used |
|---|---|---|---|
| Greenhouse | `https://boards.greenhouse.io/<token>`, `https://job-boards.greenhouse.io/<token>`, `…/embed/job_board?for=<token>` | `greenhouse:<token>` | `boards-api.greenhouse.io/v1/boards/<token>/jobs` |
| Lever | `https://jobs.lever.co/<site>` | `lever:<site>` | `api.lever.co/v0/postings/<site>` |
| Ashby | `https://jobs.ashbyhq.com/<org>` | `ashby:<org>` | `api.ashbyhq.com/posting-api/job-board/<org>` |
| Workday | `https://<tenant>.wd5.myworkdayjobs.com/<site>` (any `wdN`, optional `/en-US/` locale), `https://wdN.myworkdaysite.com/recruiting/<tenant>/<site>` | `workday:<tenant>.wd5/<site>` | `…/wday/cxs/<tenant>/<site>/jobs` (paginated) |
| Workable | `https://apply.workable.com/<account>` | `workable:<account>` | `apply.workable.com/api/v1/widget/accounts/<account>` |
| SmartRecruiters | `https://careers.smartrecruiters.com/<company>`, `https://jobs.smartrecruiters.com/<company>` | `smartrecruiters:<company>` | `api.smartrecruiters.com/v1/companies/<company>/postings` (paginated) |
| Recruitee | `https://<company>.recruitee.com` | `recruitee:<company>` | `<company>.recruitee.com/api/offers/` |
| BambooHR | `https://<company>.bamboohr.com/careers` | `bamboohr:<company>` | `<company>.bamboohr.com/careers/list` |
| Personio | `https://<company>.jobs.personio.de` (or `.com`) | `personio:<company>` | `<company>.jobs.personio.de/xml` |
| Teamtailor | `https://<company>.teamtailor.com` | `teamtailor:<company>` | `<company>.teamtailor.com/jobs.rss` |
| Breezy HR | `https://<company>.breezy.hr` | `breezy:<company>` | `<company>.breezy.hr/json` |

Paste any page of a board (a job URL works too); the actor reduces it to the
board. For anything else, give the company's website and let auto-detection
find the board.

#### Auto-detection from a company website

For an entry like `ramp.com` or `https://www.figma.com`, the actor loads, at
most 4 pages in total and without retries: the exact page you gave (when the
URL has a path, e.g. `https://acme.com/about/open-roles`), the homepage, the
first careers link on it (same domain or a subdomain), `/careers` and `/jobs`. It stops at
the first page that links to, embeds, or redirects to a supported job board;
if several boards are referenced, the most-referenced one wins. The page it
was found on is reported in `OUTPUT.companies[].detected_on`.

If nothing is found, the company gets a free status row with
`status: "ats_not_found"`. Custom career sites that render jobs from their own
backend (GitLab's is one) cannot be detected this way; if you know the
company's board, pass its URL or `platform:token` directly.

### Input

| Field | Default | Notes |
|---|---|---|
| `companies` | required | Careers-page URLs, `platform:token` references or company websites, up to 1,000 per run. Duplicates are skipped. |
| `keywords` | none | Keep jobs whose **title** contains any of these (case-insensitive). |
| `locations` | none | Keep jobs with a location containing any of these (case-insensitive substring), e.g. `Berlin`, `United Kingdom`, `Remote`. |
| `remoteOnly` | false | Keep only remote jobs (see `remote` below). |
| `postedWithinDays` | none | Keep only jobs posted in the last N days. Jobs without a posting date are left out when this is set. |
| `includeDescription` | false | Add `description`: plain text, HTML stripped, capped at 20,000 characters. |
| `maxJobsPerCompany` | 1000 | At most this many jobs (after filters) per company. |

Examples:

```json
{
  "companies": ["https://boards.greenhouse.io/gitlab", "lever:spotify"],
  "maxJobsPerCompany": 50
}
```

```json
{
  "companies": [
    "figma.com",
    "https://jobs.ashbyhq.com/ramp",
    "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
    "smartrecruiters:BoschGroup"
  ],
  "keywords": ["engineer", "data scientist"],
  "remoteOnly": true,
  "postedWithinDays": 30,
  "maxJobsPerCompany": 200
}
```

### Output

One row per job (`type: "job"`). A real row from a live run on 2026-09-30:

```json
{
  "type": "job",
  "company": "ashby:ramp",
  "platform": "ashby",
  "board_token": "ramp",
  "job_id": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "title": "Security Engineer, Cloud",
  "department": "Engineering",
  "team": "Backend",
  "location": "New York, NY (HQ)",
  "locations": ["New York, NY (HQ)", "Remote (Canada)", "Remote (US)", "Miami, FL"],
  "remote": true,
  "employment_type": "FullTime",
  "salary_min": 211400,
  "salary_max": 290600,
  "salary_currency": "USD",
  "salary_period": "year",
  "posted_at": "2026-04-07T17:12:35.753Z",
  "updated_at": null,
  "url": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "scraped_at": "2026-09-30T03:53:18.482Z"
}
```

- `company` is your input entry, exactly as you typed it.
- `description` is present only when `includeDescription` is on.
- `salary_*` are filled **only when the platform publishes structured
  compensation** (Ashby, Lever `salaryRange`, Recruitee `salary`). Free-text
  salary strings are never parsed, so on other platforms they stay `null`.
- `remote` is the platform's own flag where it has one (Lever, Ashby,
  Workable, SmartRecruiters, Recruitee, Teamtailor, Breezy, BambooHR, Workday
  when listed). Otherwise it is `true` when a location or the title says
  "Remote", and `null` (unknown) when nothing says either way.
- `posted_at` / `updated_at` are ISO 8601 (UTC) or `null`. BambooHR's list has
  no dates; the job detail does, and it is fetched when you ask for
  descriptions or set `postedWithinDays`. **Workday** lists only "Posted N days
  ago": the actor turns that into a date (accurate to about a day, since
  Workday counts days in its own time zone) and leaves "30+ days ago" as
  `null`; with `includeDescription` on, the exact start date from the job page
  is used instead.
- `locations` lists every location the platform gives. Workday lists
  multi-location jobs as "5 Locations"; `locations` is then empty unless
  descriptions (which include the job page) are requested.

Companies that yield no job rows get **one free status row** each
(`type: "company"`):

| `status` | Meaning | Charged |
|---|---|---|
| `ok_no_jobs` | Board read; it has no open jobs, or none matched your filters (`jobs_found` says how many it has) | company-checked only |
| `not_found` | No such board on that platform | no |
| `ats_not_found` | Website checked, no supported job board found | no |
| `error` | The board or website failed to load (message in `error`) | no |
| `invalid` | Not a recognizable entry, or a website whose detected board is already in the run | no |
| `skipped_budget` | Your maximum total charge was reached before this company | no |

Every run leaves at least one row.

The key-value store record `OUTPUT` holds the run summary: per-company
`platform`, `board_token`, `status`, `jobs_found` (as the platform reports
it), `jobs_matched`, `jobs_returned`, `error`, `detected_on`; totals;
`charged_events`; `stopped_reason` (`completed`, `max_total_charge_reached`
or `emit_failure`) and `incomplete`.

### Pricing

Pay per event, nothing else:

- **Company checked: $0.001** per company whose job board was read
  successfully, including boards with no matching jobs.
- **Job returned: $0.0015** per job row in the dataset.

Status rows for boards that are not found, websites without a detectable ATS,
errors, invalid entries and skipped companies are free.

Worked examples:

- 20 companies, 50 jobs each returned: 20 × $0.001 + 1,000 × $0.0015 =
  **$1.52**.
- 200 companies with a keyword filter that leaves about 10 jobs each:
  200 × $0.001 + 2,000 × $0.0015 = **$3.20**.
- 100 companies checked where nothing matches your filters: 100 × $0.001 =
  **$0.10**.

Filters are applied before rows are written, so filtering in the actor is
cheaper than filtering afterwards.

**Maximum total charge.** If you set one, the actor returns only as many rows
as fit, then stops cleanly: remaining companies get a free `skipped_budget`
row, `OUTPUT.stopped_reason` is `max_total_charge_reached` and
`OUTPUT.incomplete` is `true`. A company's job rows are delivered before its
company-checked charge, so the budget running out never charges you for a
board that delivered nothing; the last company can be cut short (its
`OUTPUT` entry says "Stopped after N of M jobs").

### Tested with

Verified live on 2026-09-30 (job counts are what each board listed that day):

| Platform | Input | Jobs on board |
|---|---|---|
| Greenhouse | `https://boards.greenhouse.io/gitlab` | 201 |
| Lever | `lever:spotify` | 79 |
| Ashby | `https://jobs.ashbyhq.com/ramp` | 155 |
| Workday | `https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite` | 2,000 (Workday's reported total) |
| Workable | `https://apply.workable.com/skroutz` | 11 |
| SmartRecruiters | `https://careers.smartrecruiters.com/BoschGroup` | 4,825 |
| Recruitee | `https://bunq.recruitee.com` | 16 |
| BambooHR | `https://gusto.bamboohr.com/careers` | 5 |
| Personio | `https://pulsegroup.jobs.personio.de` | 14 |
| Teamtailor | `https://polestar.teamtailor.com` | 30 |
| Breezy HR | `https://euler.breezy.hr` | 19 |
| Auto-detect | `figma.com` → Greenhouse, `notion.so` → Ashby, `ramp.com` → Ashby, `polestar.com` → Teamtailor | |

### Limits, stated plainly

- **Public job boards only.** The actor reads what each company publishes on
  its ATS job board. It does not scrape LinkedIn, Indeed, Glassdoor or other
  job sites, never logs in, and cannot see internal or unlisted postings
  (Ashby postings marked unlisted are skipped).
- **Only the 11 platforms above.** iCIMS, Taleo, SuccessFactors, Jobvite,
  Eightfold and custom career sites are not supported. Greenhouse's EU-hosted
  boards (`job-boards.eu.greenhouse.io`) are not supported yet.
- **Auto-detection can miss.** It looks at 4 pages at most and finds boards
  that are linked, embedded or redirected to. Career sites on a custom domain
  that load jobs from their own backend, or behind bot protection, come back
  `ats_not_found`.
- **Workday and SmartRecruiters are paginated.** Big boards take many
  requests; the actor stops paging once it has `maxJobsPerCompany` matching
  jobs. `jobs_found` is the total the platform reports.
- **Some filters need the job page.** Where a platform's list lacks or
  summarises a field, jobs that fail only on that field are checked against
  their job detail before being dropped: Workday multi-location jobs ("5
  Locations"), Workday jobs without a remote type or older than "30+ days"
  (for `locations`, `remoteOnly`, `postedWithinDays`), and BambooHR jobs (the
  list has no country and no date). That is one extra request per job, only
  as many as needed to fill `maxJobsPerCompany`, and at most 500 per company;
  jobs left unchecked past that limit are counted in
  `OUTPUT.companies[].jobs_unresolved`. `remoteOnly` on a large Workday board
  is the slowest case.
- **Descriptions cost requests on some platforms.** Workday, SmartRecruiters
  and BambooHR need one request per job for the description, so large runs
  with `includeDescription` are slower. Breezy's public feed has no
  description; `description` is `null` there.
- **Field coverage differs by platform.** For example Greenhouse has no remote
  flag or employment type, Workday has no department, Personio and Teamtailor
  have no salary. Missing fields are `null`, never guessed.
- **Polite by design.** At most 2 requests at a time per host, a short gap
  between requests, HTTP 429 and 5xx retried with backoff (Retry-After
  honoured up to 60 s), 20 s timeouts, and a User-Agent that identifies the
  actor.

### FAQ

**How do I find a company's board token?** Open the company's careers page
and click a job: the URL usually shows the platform and token
(`jobs.lever.co/<token>/…`, `boards.greenhouse.io/<token>/jobs/…`). Paste that
URL as is. Or just give the company's website.

**Can I scrape a Workday careers site?** Yes. Paste the careers-site URL,
e.g. `https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite`. The
`wdN` number and the site id come from that URL.

**Why did a company return fewer jobs than its careers page shows?** Check
your filters and `maxJobsPerCompany`, then `OUTPUT.companies[]`:
`jobs_found` is the board's own count and `jobs_matched` the count after
filters. A careers page can also combine several boards; each board is a
separate entry.

**Does it return salary?** Only when the ATS publishes structured pay data
(numbers, currency and period). Many postings mention pay only in the
description text, which is returned as-is with `includeDescription` but not
parsed into the salary fields.

**Is this allowed?** The actor reads public job-board endpoints that
companies publish so their jobs can be found, at a polite rate. You are
responsible for how you use the data.

### Local development

```bash
pnpm --filter @mmnm/atsjobs test     # unit tests on recorded fixtures, no network
pnpm --filter @mmnm/atsjobs build
```

`node dist/main.js` runs the actor locally with Apify's local storage
(`CRAWLEE_STORAGE_DIR=./storage`, input in
`storage/key_value_stores/default/INPUT.json`).

# Actor input Schema

## `companies` (type: `array`):

One entry per company, in any of three forms. (1) A careers-page URL: https://boards.greenhouse.io/<token>, https://jobs.lever.co/<site>, https://jobs.ashbyhq.com/<org>, https://apply.workable.com/<account>, https://careers.smartrecruiters.com/<company>, https://<company>.recruitee.com, https://<tenant>.wd5.myworkdayjobs.com/<site>, https://<company>.bamboohr.com/careers, https://<company>.jobs.personio.de, https://<company>.teamtailor.com, https://<company>.breezy.hr. (2) A platform:token reference, e.g. greenhouse:stripe, lever:spotify, workday:nvidia.wd5/NVIDIAExternalCareerSite. (3) A company website or domain, e.g. ramp.com: the actor looks for a supported job board on its homepage, careers link, /careers and /jobs (at most 4 requests). Up to 1,000 per run.

## `keywords` (type: `array`):

Optional. Keep only jobs whose title contains at least one of these words or phrases (case-insensitive), e.g. engineer, product manager.

## `locations` (type: `array`):

Optional. Keep only jobs with a location containing at least one of these (case-insensitive substring), e.g. London, Germany, Remote.

## `remoteOnly` (type: `boolean`):

Keep only jobs that are remote: the platform's own remote flag where it has one, otherwise a location or title that says "Remote".

## `postedWithinDays` (type: `integer`):

Optional. Keep only jobs posted in the last N days. Jobs whose platform gives no posting date are left out when this is set (see the README for which platforms have dates).

## `includeDescription` (type: `boolean`):

Also return each job's description as plain text (HTML stripped, capped at 20,000 characters). On Workday, SmartRecruiters and BambooHR this needs one extra request per job, so runs are slower.

## `maxJobsPerCompany` (type: `integer`):

At most this many jobs (after filters) are returned per company.

## Actor input object example

```json
{
  "companies": [
    "https://boards.greenhouse.io/gitlab",
    "lever:spotify"
  ],
  "remoteOnly": false,
  "includeDescription": false,
  "maxJobsPerCompany": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "https://boards.greenhouse.io/gitlab",
        "lever:spotify"
    ],
    "maxJobsPerCompany": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("changefeeds/ats-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "https://boards.greenhouse.io/gitlab",
        "lever:spotify",
    ],
    "maxJobsPerCompany": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("changefeeds/ats-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "https://boards.greenhouse.io/gitlab",
    "lever:spotify"
  ],
  "maxJobsPerCompany": 50
}' |
apify call changefeeds/ats-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,changefeeds/ats-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ccVBx12BPpoStzUjO/builds/YpFKcAcrTwFtSdIIo/openapi.json
