# ATS Job Scraper: Greenhouse, Ashby & Lever (`axiorasolutions/ats-job-scraper`) Actor

Scrape every open job from a company's own ATS board. Returns one row per job: title, department, locations, workplace type, employment type, posted date, apply URL, salary with provenance and clean description text. Batch many boards per run. Greenhouse, Ashby, Lever, SmartRecruiters.

- **URL**: https://apify.com/axiorasolutions/ats-job-scraper.md
- **Developed by:** [Axiora Solutions](https://apify.com/axiorasolutions) (community)
- **Categories:** Jobs, Lead generation, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.20 / 1,000 job postings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ATS Job Scraper — Greenhouse, Ashby, Lever, SmartRecruiters

ATS job scraper that pulls every open role directly from a company's own applicant tracking system and returns one clean row per job — title, department, all locations, workplace type, posted date, apply URL and salary. Point it at one or more Greenhouse, Ashby, Lever or SmartRecruiters boards and you get a normalised dataset in seconds, with **no API key and no login**. Four providers, one identical schema, so you stop writing a parser per vendor.

### What you get

- `title`, `department` and `team` for every open role.
- `locations`, `locationPrimary` and `workplaceType` (`REMOTE`, `HYBRID` or `ON_SITE`).
- `postedAt`, `updatedAt` and `ageDays` so you can spot fresh postings.
- `salary` with min/max, currency, interval and `evidence` provenance.
- `applyUrl`, `jobUrl` and `requisitionId` for one-click applications.
- `descriptionText`, `descriptionHash`, `jobUid` and `recordHash` for change detection.

### Quick start

1. Open the ATS Job Scraper on Apify Store and find the **ATS boards** field.
2. Add one entry per company: a `provider:token` pair like `greenhouse:databricks`, a careers URL like `https://jobs.ashbyhq.com/ramp`, or a bare token that is auto-detected across providers.
3. Click **Start**. The prefilled `greenhouse:airtable` and `ashby:Linear` boards return real postings from two different ATS platforms in a few seconds.
4. Open the **Jobs**, **Compensation** or **Descriptions** tab to browse the dataset, then export to JSON, CSV, Excel or Google Sheets.

Minimal input:

```json
{
  "boards": ["greenhouse:airtable", "ashby:Linear"]
}
```

### Example output

One row per job. Fields are stable across all four providers:

```json
{
  "ok": true,
  "errorCode": null,
  "jobUid": "greenhouse:databricks:7712345",
  "source": "greenhouse",
  "sourceLabel": "Greenhouse",
  "boardToken": "databricks",
  "companyName": "Databricks",
  "boardUrl": "https://boards.greenhouse.io/databricks",
  "jobId": "7712345",
  "title": "Senior Data Engineer",
  "department": "Engineering",
  "departments": ["Engineering"],
  "team": null,
  "employmentType": "Full time",
  "locationPrimary": "San Francisco, CA",
  "locations": ["San Francisco, CA", "Remote - US"],
  "isRemote": false,
  "remoteFlagFromAts": null,
  "workplaceType": null,
  "postedAt": "2026-09-18T09:14:02.000Z",
  "updatedAt": "2026-09-30T11:02:00.000Z",
  "ageDays": 14,
  "applyUrl": "https://boards.greenhouse.io/databricks/jobs/7712345",
  "jobUrl": "https://boards.greenhouse.io/databricks/jobs/7712345",
  "requisitionId": "REQ-4821",
  "salary": {
    "text": "$180,000 - $230,000 per year",
    "minAmount": 180000,
    "maxAmount": 230000,
    "currency": "USD",
    "interval": "YEAR",
    "evidence": "ats-field"
  },
  "descriptionText": "About the role\nWe are looking for...",
  "descriptionChars": 4211,
  "descriptionHash": "1f3a9c4b7d2e5061",
  "atsFields": { "function": "Engineering" },
  "recordHash": "9bd2c1ea77405f38",
  "scrapedAt": "2026-10-02T12:00:00.000Z"
}
```

### What this ATS job scraper does

- 🏢 **Four providers, one contract** — Greenhouse, Ashby, Lever and SmartRecruiters are normalised into the same fields, so you stop writing a parser per vendor.
- 📦 **Batch mode by design** — the main input is an **array**, so hundreds of employers are one run instead of one run per company. Fewer runs means lower cost and far less orchestration.
- 🔍 **Auto-detection** — paste a bare board token and the Actor probes every provider to find it. Paste a careers URL and it extracts the token for you.
- 🏠 **Hybrid is not remote** — `workplaceType` is read from the ATS field and normalised to `REMOTE`, `HYBRID` or `ON_SITE`. Some providers set their own `isRemote` boolean to `true` for Hybrid roles; this Actor does not, so **Remote jobs only** returns jobs that are actually remote.
- 💰 **Salary with provenance** — every row carries `salary.evidence`: `ats-field` when the ATS published a compensation field, `description-text` when a figure was quoted from the job text, `none` when nothing was published. **Numbers are parsed from published figures, never estimated.**
- 🧭 **Change detection built in** — `jobUid`, `descriptionHash` and `recordHash` are deterministic, so a scheduled run tells you exactly which roles are new, edited or gone.
- 🧹 **Server-side filters** — title keywords, excluded keywords, location keywords, remote-only and `postedAfter` narrow the dataset *and* the number of billed results.
- 🛟 **One bad board never fails the run** — an unreachable or empty board becomes an `ok: false` row with a stable `errorCode` and a remediation hint. Everything else still lands.

Because it runs on Apify you also get scheduling, run history, webhooks, monitoring, API and SDK access, and one-click exports to JSON, CSV, Excel, Google Sheets and over 20 integrations — without building any of it.

#### Failure rows and the run summary

A board that cannot be read produces a traceable failure row instead of killing the run:

```json
{
  "ok": false,
  "errorCode": "NOT_FOUND",
  "source": "greenhouse",
  "requestedInput": "greenhouse:does-not-exist",
  "error": {
    "code": "NOT_FOUND",
    "message": "HTTP 404 from https://boards-api.greenhouse.io/v1/boards/does-not-exist/jobs",
    "httpStatus": 404,
    "hint": "The target does not exist or is not public. Verify the identifier and that the resource is publicly reachable."
  }
}
```

The **run summary** is written to the key-value store under `OUTPUT`: per-board counts, how many jobs each filter removed, billed event totals and request/transfer totals.

### How to use it

1. Open the Actor and find the **ATS boards** field.
2. Add one entry per company. All of these work:
   - `greenhouse:databricks` — explicit provider and token
   - `https://jobs.ashbyhq.com/ramp` — a careers URL
   - `spotify` — a bare token, auto-detected across providers
3. Set **Max jobs per board** and **Max jobs for the whole run** to bound the dataset.
4. Optional: add **Title keywords**, **Location keywords**, **Remote jobs only** or **Posted after** to filter.
5. Click **Start**, then open the **Jobs**, **Compensation** or **Descriptions** tab on the dataset.
6. To keep it fresh, add a **Schedule** and a webhook, then diff on `recordHash`.

#### Where do I find a board token?

Open the company's careers page. If the URL looks like `boards.greenhouse.io/figma`, `jobs.ashbyhq.com/ramp`, `jobs.lever.co/spotify` or `jobs.smartrecruiters.com/Continental`, the last path segment is the token. Paste the whole URL and the Actor handles the rest.

### How much does it cost to scrape ATS job boards?

Pricing is **pay per event**, so your bill is a straight multiplication you can predict before you start:

| Event | What triggers it | Billed |
|---|---|---|
| Job posting | One job written to the dataset | per job |
| Actor start | Once per run, platform fee | per run |

There is **no separate platform-usage charge on top** — compute, bandwidth and storage are already covered by the per-job price. Filtered-out jobs are **not** billed, because they are never written.

So 40 companies averaging 50 open roles is 2,000 rows, and 2,000 × the per-job price is your total. Set **Max cost per run** in the run options for a hard ceiling; the Actor stops cleanly at the limit and tells you in the log instead of silently truncating. Higher Apify plans get progressively lower per-job pricing through Apify Store tier discounts.

Want to spend nothing while you evaluate? Set **Max jobs per board** to `5` and run a single board.

### Example input

```json
{
  "boards": ["greenhouse:databricks", "https://jobs.ashbyhq.com/ramp", "lever:spotify"],
  "maxJobsPerBoard": 100,
  "maxJobsTotal": 500,
  "includeDescription": true,
  "descriptionMaxChars": 6000,
  "titleKeywords": ["engineer", "data"],
  "remoteOnly": false,
  "postedAfter": "30 days"
}
```

### Use cases

- **Recruiting intelligence** — track which competitors are hiring, where and for what, week over week.
- **Sales triggers** — a new Head of Data role is a buying signal; `postedAfter` plus a schedule turns that into a feed.
- **Talent market research** — salary evidence and location mix across a named peer set.
- **Job boards and aggregators** — a legitimate, first-party source instead of scraping aggregator HTML.
- **LLM and RAG pipelines** — `descriptionText` is already clean plain text with a stable hash.

### Related Actors by Axiora Solutions

| Actor | Use it for |
|---|---|
| **Domain Contact Enricher** | Turn the companies you found here into contactable records: emails, phones, socials, tech stack |
| **News & RSS Feed Scraper** | Watch the same companies for funding, launch and leadership news |
| **SEC EDGAR API** | Pull filings and financial facts for the public companies in your list |

### Frequently asked questions

#### Which ATS platforms are supported?

Greenhouse, Ashby, Lever and SmartRecruiters — all through their own public, unauthenticated job-board endpoints. Boards embedded in a company's own website via an iframe or script tag are **not** supported; use the underlying ATS URL instead.

#### Do I need an API key or a login?

No. Every source is a public job board. Nothing in the input is a credential, which also means autonomous agents can call this Actor without human help.

#### Why is `companyName` sometimes null?

Ashby and Lever do not publish a company name on their public board endpoints. `boardToken` and `companySlug` are always present and are reliable join keys.

#### Why is a salary missing when the posting shows one?

`salary.evidence` tells you exactly why. Many employers put compensation in an image, a regional addendum, or wording this Actor will not guess from. We would rather return `none` than a number we cannot defend.

#### Does it get descriptions from SmartRecruiters?

Yes, but SmartRecruiters only serves descriptions from a per-posting endpoint, so one extra request per job is made when **Include job description text** is on. Turn it off for a much faster run on large SmartRecruiters boards.

#### What happens if a board is empty or blocked?

You get an `ok: false` row with `errorCode` (`NOT_FOUND`, `ACCESS_DENIED`, `RATE_LIMITED`, `EMPTY_RESULT`, `UNSUPPORTED_SOURCE`, `TIMEOUT`) and a hint. The run still succeeds and every other board is still written.

#### Can I run this on a schedule?

Yes. Add an Apify **Schedule**, then diff consecutive runs on `recordHash` (any change) or `jobUid` (new and removed roles).

#### Is scraping public job boards legal?

These are public endpoints that employers publish so their jobs can be distributed. You are responsible for how you use the output, including any applicable data-protection obligations. Job postings are company data, not personal data, and this Actor does not collect recruiter names or contact details.

#### Something is broken — how do I report it?

Open the **Issues** tab on this Actor page with the board reference and the `errorCode` you saw. ATS vendors change their endpoints occasionally; issues with a concrete reproduction get fixed fastest.

***

Runnable examples and how-to guides for these Actors: [github.com/batow133/axiora-apify-actors](https://github.com/batow133/axiora-apify-actors)

# Actor input Schema

## `boards` (type: `array`):

Add one entry per company. Accepted formats: a careers URL (https://boards.greenhouse.io/figma, https://jobs.ashbyhq.com/ramp, https://jobs.lever.co/spotify, https://jobs.smartrecruiters.com/Continental), a provider:token pair (greenhouse:figma), or a bare token, which is auto-detected across all four providers.

## `maxJobsPerBoard` (type: `integer`):

Stop after this many jobs from each board. Keeps a run over a large employer predictable. Use a high value to take everything.

## `maxJobsTotal` (type: `integer`):

Hard ceiling across every board. The run stops cleanly when it is reached.

## `includeDescription` (type: `boolean`):

Add the full job description as clean plain text. Turn off for a smaller, faster dataset when you only need titles and locations.

## `descriptionMaxChars` (type: `integer`):

Truncate description text at this length. Lower values keep exports small and are usually enough for classification or embedding.

## `titleKeywords` (type: `array`):

Keep only jobs whose title contains at least one of these words, case-insensitive. Leave empty to keep every job.

## `excludeTitleKeywords` (type: `array`):

Drop jobs whose title contains any of these words, case-insensitive. Useful for removing intern, contract or sales roles.

## `locationKeywords` (type: `array`):

Keep only jobs whose location text contains one of these words, case-insensitive, for example a city, country or region.

## `remoteOnly` (type: `boolean`):

Keep only fully remote jobs. Where the ATS publishes an explicit workplace type, that is used, so Hybrid roles are excluded rather than counted as remote. Boards with no workplace field fall back to the location text.

## `postedAfter` (type: `string`):

Keep only jobs first published on or after this date. Accepts 2026-01-31 or a relative value such as 14 days. Jobs with no publish date from the provider are always kept.

## `requestTimeoutSecs` (type: `integer`):

Give up on a single ATS request after this many seconds. Raise it for very large employers whose boards are slow to generate.

## Actor input object example

```json
{
  "boards": [
    "greenhouse:databricks",
    "https://jobs.ashbyhq.com/ramp",
    "lever:spotify"
  ],
  "maxJobsPerBoard": 200,
  "maxJobsTotal": 2000,
  "includeDescription": true,
  "descriptionMaxChars": 6000,
  "titleKeywords": [
    "engineer",
    "data",
    "security"
  ],
  "excludeTitleKeywords": [
    "intern",
    "contract"
  ],
  "locationKeywords": [
    "London",
    "Remote",
    "Germany"
  ],
  "remoteOnly": false,
  "postedAfter": "30 days",
  "requestTimeoutSecs": 45
}
```

# Actor output Schema

## `jobs` (type: `string`):

Every job posting collected in this run, one row per job. Deduplicated on jobUid (provider:boardToken:jobId).

## `runSummary` (type: `string`):

Per-board counts, how many jobs each filter removed, billed event totals, request and transfer totals.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "boards": [
        "greenhouse:airtable",
        "ashby:Linear"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiorasolutions/ats-job-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "boards": [
        "greenhouse:airtable",
        "ashby:Linear",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("axiorasolutions/ats-job-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "boards": [
    "greenhouse:airtable",
    "ashby:Linear"
  ]
}' |
apify call axiorasolutions/ats-job-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiorasolutions/ats-job-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/a574TkZHiTHwucai6/builds/8b0wGZ2l6zQW3yCu8/openapi.json
