# ATS Jobs Scraper - Greenhouse, Lever, Ashby, Workday & More (`clearfetch/ats-jobs-scraper`) Actor

Every open job from company career sites on Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee and Personio. Paste a careers page, domain or board link. One schema with salary, remote flag and post date; filters; new-jobs-only mode. No login or proxy.

- **URL**: https://apify.com/clearfetch/ats-jobs-scraper.md
- **Developed by:** [Nada Hanad](https://apify.com/clearfetch) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ATS Jobs Scraper - Greenhouse, Lever, Ashby, Workday & More

Get every open job from company career sites, straight from the applicant tracking system behind them:
Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee and Personio. Paste a careers page, a
plain company domain or a job board link, and the Actor finds the board, reads it through the board's public API
and returns every job in one flat schema, with salary ranges, remote flags and post dates. **$1.00 per 1,000
jobs.** No login, no proxy, no browser.

### What data you get

One row per job, the same columns whichever board it came from:

- `company`, `title`, `url` (the posting), `applyUrl`, `jobId` (stable, so it works as a key across runs)
- `location`, `locations` (every location listed), `country`, `remote`, `workplaceType` (remote, hybrid, onsite)
- `employmentType` (full-time, part-time, contract, internship, temporary), `department`, `team`
- `postedAt`, `updatedAt`
- Pay: `salaryMin`, `salaryMax`, `salaryCurrency`, `salaryPeriod` (year, month, hour), `salaryText`, and
  `salarySource`, which says whether the board publishes the range as data (`ats`: Greenhouse pay transparency,
  Lever, Ashby, Recruitee) or it was read from the job text (`description`)
- `description` as plain text, and optionally `descriptionHtml` and the board's untouched record in `raw`
- `ats`, `boardSlug`, `detectedBy` (how the board was found) and `isNew` for scheduled monitoring

A `SUMMARY` record in the run's key-value store lists every company with how its board was found, how many jobs
the board has, how many passed your filters and how many were written.

### How to use

1. Add companies, one per line: `stripe.com`, `https://jobs.lever.co/palantir`, a careers page, or a pair such
   as `greenhouse:airbnb`.
2. Optionally filter by title words, location, remote, or jobs posted in the last N days. For monitoring, turn
   on **Only jobs not seen in earlier runs** and put the Actor on a schedule.
3. Run it, then download JSON, CSV or Excel, or pull the rows through the API.

### Input

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `companies` | array | — | Board links, careers pages, domains or `ats:slug` pairs. Also `urls`, `startUrls`, `domains`. |
| `titleKeywords` | array | `[]` | Keep jobs whose title contains any of these. Also searched on Workday's side. |
| `excludeTitleKeywords` | array | `[]` | Leave out jobs whose title contains any of these. |
| `locationKeywords` | array | `[]` | Keep jobs whose location, country or workplace type contains any of these. |
| `remoteOnly` | boolean | `false` | Only remote jobs. |
| `postedWithinDays` | integer | `0` | Only jobs published in the last N days (0 = all). |
| `onlyNew` | boolean | `false` | Write only jobs not written in earlier runs, per board. The first run returns everything. |
| `stateStoreName` | string | `ats-jobs-scraper-state` | Where `onlyNew` remembers what it has seen. |
| `maxJobsPerCompany` | integer | `0` | Stop after this many jobs per board (0 = all). |
| `maxItems` | integer | `0` | Stop the run after this many jobs (0 = no limit). |
| `includeDescription` | boolean | `true` | Description as plain text. |
| `includeDescriptionHtml` | boolean | `false` | Description as HTML too. |
| `workdayDetails` | boolean | `true` | Read each Workday job's own record for its exact date, locations and description. |
| `guessBoards` | boolean | `true` | When a site links to no board, try its domain name as a board id and keep it only if the company name fits. |
| `includeRaw` | boolean | `false` | The board's own record for each job. |
| `maxConcurrency` | integer | `10` | Boards read at the same time (1-50). |
| `timeoutSecs` | integer | `30` | Per request. |
| `proxyConfiguration` | object | off | Not needed; available if you want your own IPs. |

Example:

```json
{
    "companies": ["stripe.com", "https://jobs.lever.co/palantir", "greenhouse:figma", "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"],
    "titleKeywords": ["engineer"],
    "locationKeywords": ["remote", "united states"],
    "postedWithinDays": 14
}
```

### Output example

A real row from a run on 2026-09-29 (description shortened here):

```json
{
    "ok": true,
    "ats": "ashby",
    "company": "Ramp",
    "boardSlug": "ramp",
    "jobId": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
    "title": "Security Engineer, Cloud",
    "url": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
    "applyUrl": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245/application",
    "location": "New York, NY (HQ)",
    "locations": ["New York, NY (HQ)", "Remote (Canada)", "Remote (US)", "Miami, FL"],
    "country": "USA",
    "remote": false,
    "workplaceType": "hybrid",
    "employmentType": "full-time",
    "department": "Engineering",
    "team": "Backend",
    "postedAt": "2026-04-07T17:12:35.753Z",
    "postedAtApproximate": false,
    "updatedAt": null,
    "salaryMin": 211400,
    "salaryMax": 290600,
    "salaryCurrency": "USD",
    "salaryPeriod": "year",
    "salaryText": "$211.4K – $290.6K • Offers Equity",
    "salarySource": "ats",
    "description": "About Ramp\nRamp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging ris …",
    "isNew": null,
    "detectedBy": "url",
    "inputUrl": "https://jobs.ashbyhq.com/ramp",
    "scrapedAt": "2026-09-29T12:18:10.134Z"
}
```

An input that cannot be read comes back as one row with `ok: false` and a plain reason, such as
`no lever job board found for "acme"` or `no supported job board is linked from this site or its careers page`.
Those rows are free.

### Pricing

- **$0.001 per job**, which is $1.00 per 1,000. A company with 200 open jobs costs $0.20.
- You pay only for jobs written to your dataset. Inputs that fail, jobs your filters leave out and, with
  `onlyNew`, jobs already seen in earlier runs are never charged.
- Set a maximum cost on the run and it stops cleanly when it gets there.

Paid Apify plans pay less: 10% off on Bronze, 20% on Silver and 30% on Gold and higher tiers.

### Use cases

- **Sales prospecting on hiring signals**: which target accounts are hiring for the role your product serves,
  with the department and location, every morning.
- **Recruiting and talent intelligence**: competitors' open roles, salary bands and where they hire.
- **Job boards and newsletters**: fresh postings from a curated list of companies, straight from the source,
  with only-new mode for daily updates.
- **Market and compensation research**: pay ranges by title, level and location across hundreds of companies.
- **Investors**: headcount growth and hiring focus across a portfolio or a watchlist.

### FAQ

**Which job boards are supported?** Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee and
Personio. Together they run the careers pages of most tech companies and a large share of enterprises.

**What if I only know the company's website?** Give the domain or the careers page. The Actor looks for a board
link or embed on the home page, then on the careers page it links to. If there is none, it tries the domain
name as a board id and keeps a match only when the board's company name fits; those rows say
`detectedBy: "slug-guess"`. A site that runs its own careers system comes back as a free error row.

**Do I need a proxy?** No. These are the public APIs the boards publish for career sites, so they answer
ordinary requests. The proxy option is there only if you want to use your own IPs.

**How fresh is the data?** Live: every run reads the boards at that moment, nothing is cached.

**Why is a Workday date marked approximate?** Workday's list says "Posted 3 Days Ago". With **Read each Workday
job** on (the default), the exact date comes from each job's record and `postedAtApproximate` is false.

**Why does a salary sometimes come from the description?** Many boards do not publish pay as data. When a
posting states a range in its text ("$60,000 - $97,000/year"), it is read from there and marked
`salarySource: "description"`. The reader is deliberately strict and ignores numbers that are not clearly pay.

**How does only-new mode work?** For each board it remembers the ids of jobs it has written, in a key-value
store in your account. The next run writes only ids it has not seen. Jobs that close are dropped from memory,
so a re-posted role counts as new again.

**Is it legal?** It reads job postings that companies publish so that anyone can see them, through the APIs the
boards provide for that. You are responsible for how you use the data.

### Integrations

Run it from the Apify API or a client library, schedule it in Apify Console, or connect it to n8n, Make,
Zapier or any MCP client through Apify's integrations. Results are available as JSON, CSV, Excel and through
the dataset API.

### More tools from clearfetch

- [Website Contact Extractor](https://apify.com/clearfetch/website-contact-extractor): emails, phone numbers and social profiles from company websites
- [Tech Stack Detector](https://apify.com/clearfetch/tech-stack-detector): the CMS, frameworks, analytics and hosting behind any website
- [Google Trends Scraper](https://apify.com/clearfetch/google-trends-scraper): interest over time, by region and related queries, plus today's trending searches

### Changelog

- **1.0.0** (2026-09) — first release: eight job boards, board detection from careers pages and domains, one
  schema with salary, remote flag and post date, Greenhouse pay transparency, title, location, remote and date
  filters, only-new mode for schedules, per-company summary.

# Actor input Schema

## `companies` (type: `array`):

One per line. Any of: a job board link (boards.greenhouse.io/figma, jobs.lever.co/palantir, jobs.ashbyhq.com/ramp, nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite, apply.workable.com/..., careers.smartrecruiters.com/..., \*.recruitee.com, \*.jobs.personio.de), a company's careers page or plain domain (the board is found on the page), or an "ats:slug" pair such as greenhouse:airbnb. Also accepts "urls", "startUrls" and "domains".

## `titleKeywords` (type: `array`):

Keep jobs whose title contains any of these words (case-insensitive), e.g. engineer, designer. Empty keeps all. On Workday boards these are also searched on Workday's side, which is much faster on large boards.

## `excludeTitleKeywords` (type: `array`):

Leave out jobs whose title contains any of these words, e.g. intern, senior.

## `locationKeywords` (type: `array`):

Keep jobs whose location, country or workplace type contains any of these, e.g. Berlin, Germany, remote. Empty keeps all.

## `remoteOnly` (type: `boolean`):

Keep only jobs the board marks as remote, or whose location says remote.

## `postedWithinDays` (type: `integer`):

Keep jobs published in the last N days. 0 keeps all. Jobs whose board gives no date are left out when this is set.

## `onlyNew` (type: `boolean`):

For scheduled runs: write only jobs this Actor has not written before for the same board, so each run returns just the new postings. The first run returns everything. Seen jobs are kept in a named key-value store in your account.

## `stateStoreName` (type: `string`):

The key-value store that remembers seen jobs for onlyNew. Use a different name per schedule if two schedules watch the same boards with different filters.

## `maxJobsPerCompany` (type: `integer`):

Stop after this many jobs from one board (after filters). 0 means all.

## `maxItems` (type: `integer`):

Stop the run after this many jobs. 0 means no limit. You can also cap the run's cost in the run options.

## `includeDescription` (type: `boolean`):

Adds the job description as plain text. On SmartRecruiters this reads each job's own record.

## `includeDescriptionHtml` (type: `boolean`):

Adds the job description as HTML as well.

## `workdayDetails` (type: `boolean`):

Workday lists give only the title and "Posted 3 Days Ago". On, each job's record is read for its exact date, all locations and description. Off is faster but those fields stay approximate or empty, and location filters cannot be applied to multi-location jobs.

## `guessBoards` (type: `boolean`):

When a company site links to no board, try the domain name as a board id on Greenhouse, Lever, Ashby, Workable and Recruitee, and keep a match only if the board's company name fits. Rows found this way have detectedBy slug-guess.

## `includeRaw` (type: `boolean`):

Adds each board's own record for the job under raw, for fields this Actor does not map.

## `maxConcurrency` (type: `integer`):

How many boards are read at the same time.

## `timeoutSecs` (type: `integer`):

Per request.

## `proxyConfiguration` (type: `object`):

Not needed: these are public job board APIs. Available if you want your own IPs.

## Actor input object example

```json
{
  "companies": [
    "https://jobs.lever.co/palantir",
    "https://job-boards.greenhouse.io/figma",
    "https://jobs.ashbyhq.com/ramp",
    "huggingface.co"
  ],
  "titleKeywords": [],
  "excludeTitleKeywords": [],
  "locationKeywords": [],
  "remoteOnly": false,
  "postedWithinDays": 0,
  "onlyNew": false,
  "stateStoreName": "ats-jobs-scraper-state",
  "maxJobsPerCompany": 10,
  "maxItems": 0,
  "includeDescription": true,
  "includeDescriptionHtml": false,
  "workdayDetails": true,
  "guessBoards": true,
  "includeRaw": false,
  "maxConcurrency": 10,
  "timeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One row per job: company, title, links, locations, remote flag, employment type, department, posted date, salary range and description. Inputs or boards that could not be read appear with ok=false and a reason, and are not charged.

## `summary` (type: `string`):

For each board: how it was found, jobs on the board, jobs that passed the filters and jobs written.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "https://jobs.lever.co/palantir",
        "https://job-boards.greenhouse.io/figma",
        "https://jobs.ashbyhq.com/ramp",
        "huggingface.co"
    ],
    "maxJobsPerCompany": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("clearfetch/ats-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "https://jobs.lever.co/palantir",
        "https://job-boards.greenhouse.io/figma",
        "https://jobs.ashbyhq.com/ramp",
        "huggingface.co",
    ],
    "maxJobsPerCompany": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("clearfetch/ats-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "https://jobs.lever.co/palantir",
    "https://job-boards.greenhouse.io/figma",
    "https://jobs.ashbyhq.com/ramp",
    "huggingface.co"
  ],
  "maxJobsPerCompany": 10
}' |
apify call clearfetch/ats-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,clearfetch/ats-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JX7zHSakDcg8NpuoL/builds/bKOfFDGlHFeggF6ZJ/openapi.json
