# ATS Job Scraper: Workday, Greenhouse, Lever & More (`digital_influx/ats-jobs-scraper`) Actor

Extract every open job from company career pages on 12 applicant tracking systems: Workday, Greenhouse, Lever, Ashby, Workable, Personio, BambooHR and more. Filter by title, description keywords, remote, employment type, salary or visa. Get location, salary, seniority and full description.

- **URL**: https://apify.com/digital\_influx/ats-jobs-scraper.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** Jobs, Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 job saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ATS Job Scraper: Workday, Greenhouse, Lever & More

Get **every open job from any company's careers page** in one clean, consistent table. Paste careers page URLs (or just company names) and get title, location, country, remote flag, **salary**, **seniority**, **visa sponsorship**, **years of experience**, department, employment type, dates and the full description.

It reads the **public job-board APIs** that these applicant tracking systems serve to their own careers pages. That means:

- **No proxies, no headless browser, no login, no CAPTCHAs.** Runs are fast and cheap: in our test, companies on 11 of the 12 platforms, including NVIDIA and Salesforce on Workday, were read in about 12 seconds.
- **Stable.** Public APIs don't break when a careers page is redesigned.
- **One schema for twelve platforms.** Workday, Greenhouse, Lever, Ashby, Workable, Recruitee, Personio, Rippling, Teamtailor, Breezy, BambooHR and Gem all come out with the same columns, so you can mix companies freely in one run instead of paying for a separate scraper per platform.
- **Salary even when it is only in the text.** Many employers write the pay range inside the description (US pay-transparency ranges, "€50.000 – 60.000", "$45/hr"). The Actor reads it into numbers. In our test on 469 real jobs, 382 came out with a salary, 262 of them from the description alone.

### What you can use it for

- **Job boards and newsletters:** aggregate fresh roles from hundreds of companies you choose, filtered by keyword, location, seniority and remote.
- **Recruiters and talent teams:** monitor competitors' hiring and find companies hiring for the roles you place, at the level you place.
- **Sales intelligence:** hiring is a buying signal. A company opening 5 data engineering roles is shopping for data tools.
- **Market research:** compare salary ranges, seniority mix, remote share and hiring velocity by company.
- **Candidates and job sites for international talent:** keep only jobs that say they sponsor visas.
- **AI agents:** the output is plain JSON with plain-text descriptions, ready for an LLM. Agents can call this Actor through the Apify MCP server.

### Input

| Field | What it does |
|---|---|
| **Companies** (required) | One per line. A careers URL, `platform:slug`, or just the board name. |
| Title keywords | Keep jobs whose title has any of these words. Whole-word match: `intern` does not match `internal`, `java` does not match `javascript`. |
| Exclude title keywords | Drop titles with these words, e.g. `manager`. |
| Locations | Keep jobs whose location, secondary locations or country code contain the text: `Berlin`, `Germany`, `DE`. A two-letter code matches as a whole word or the country code, so `US` does not match "Dusseldorf". `remote` also matches jobs flagged remote. |
| Remote only | Only jobs with a fully remote option. |
| Seniority | Keep only these levels, read from the title: intern, entry, not stated (usually mid-level), senior, lead, director, executive. |
| Posted within (days) | Only jobs published in the last N days. |
| Workplace type | Remote, hybrid, on-site, or not stated. Greenhouse jobs rarely state it (85% of the 195,000 Greenhouse jobs in our index), unless the location says Remote, Hybrid or On-site: add *Not stated* to keep them. |
| Employment type | Full-time, part-time, contract, temporary, internship, or not stated. Each platform writes it its own way ("FullTime", "Permanent", "fulltime\_fixed\_term", "Short Term"); the Actor maps them to these values. Greenhouse lists none, so its jobs are *Not stated* unless the title says so, as in "Intern" or "(Contract)". |
| Description keywords / Exclude description keywords | Whole-word match on the full description: `Kubernetes`, `GDPR`; exclude `clearance`. |
| Only jobs that state a salary | A pay range from the platform or written in the description. |
| Only jobs that offer visa sponsorship | The description says the company sponsors visas. Jobs that say nothing are dropped. |
| Job description | Plain text (default), HTML, both, or none. |
| Max jobs per company / in total | Newest first. |
| Only new jobs (for scheduled runs) | Save only jobs that earlier runs with the same companies and filters did not save. |
| Memory name | Optional: share one memory between tasks, or start over with a new name. |

Accepted company formats:

```
https://boards.greenhouse.io/airbnb
https://job-boards.greenhouse.io/figma
https://jobs.lever.co/spotify
https://jobs.eu.lever.co/yourcompany
https://jobs.ashbyhq.com/ramp
https://apply.workable.com/huggingface
https://bunq.recruitee.com
https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite
https://wd3.myworkdaysite.com/recruiting/acme/External
https://ottonova.jobs.personio.de
https://ats.rippling.com/rippling/jobs
https://career.teamtailor.com
https://yourcompany.breezy.hr
https://yourcompany.bamboohr.com/careers
https://jobs.gem.com/11x-ai
lever:palantir
notion            <- just a name: the Actor tries each platform and uses the one that has the board
                     (not for Workday, Personio or BambooHR: paste their careers URL)
```

Example input:

```json
{
  "companies": ["https://boards.greenhouse.io/stripe", "jobs.lever.co/palantir", "notion"],
  "keywords": ["engineer", "developer"],
  "locations": ["remote", "London"],
  "seniority": ["senior", "lead"],
  "postedWithinDays": 30
}
```

### Output

One item per job. Every platform fills the same fields, and anything a platform doesn't publish is `null`, never missing:

```json
{
  "id": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "platform": "ashby",
  "companySlug": "ramp",
  "company": "Ramp",
  "title": "Security Engineer, Cloud",
  "department": "Engineering",
  "team": "Backend",
  "location": "New York, NY (HQ)",
  "locations": ["New York, NY (HQ)", "Remote (Canada)", "Remote (US)", "Miami, FL"],
  "country": "US",
  "remote": true,
  "workplaceType": "hybrid",
  "employmentType": "Full time",
  "experienceLevel": null,
  "seniority": null,
  "postedAt": "2026-04-07T17:12:35.753Z",
  "updatedAt": null,
  "url": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "applyUrl": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245/application",
  "salary": { "min": 211400, "max": 290600, "currency": "USD", "period": "year", "text": "$211.4K - $290.6K", "source": "platform" },
  "yearsOfExperience": null,
  "visaSponsorship": null,
  "descriptionText": "ABOUT RAMP\n\nRamp is building the smart infrastructure for finance teams...",
  "descriptionHtml": null,
  "scrapedAt": "2026-09-23T20:00:00.000Z"
}
```

Field notes:

- `country` is an ISO 3166-1 alpha-2 code (`US`, `GB`, `DE`). If a platform doesn't send it, it is taken from the end of the location ("London, United Kingdom" → `GB`, "Austin, TX" → `US`), and it stays `null` rather than being guessed: "Chicago, IL" stays `null`, because IL is also Israel.
- `remote: true` means the job offers a fully remote option, even when it also lists offices. `workplaceType` is what the employer declared: `remote`, `hybrid` or `onsite`.
- `salary` comes from the platform when it publishes one (`"source": "platform"`), otherwise from the pay range written in the description (`"source": "description"`). A number is taken only when it is clearly pay: it needs a pay word, a period or a currency code next to it, and funding rounds, revenue or deal sizes are skipped. When in doubt the field stays `null`.
- Ashby jobs also carry `payRanges` (each salary range the company set, per tier: `{ "title": "SF/NYC", "min": 150000, "max": 200000, "currency": "USD", "period": "year" }`), `compensationSummary` (the pay line Ashby shows, with commission and bonus) and `offersEquity`. With several tiers in one currency, `salary` spans them all.
- `seniority` is read from the title: `intern`, `entry`, `senior`, `lead` (lead, staff, principal), `director` (director, head of) or `executive` (VP, C-level). `null` means the title names no level. `experienceLevel` is the level the employer set on the platform, when it has one.
- `yearsOfExperience` is the minimum asked for in the description ("5+ years of experience" → 5, "3-5 years" → 3).
- `visaSponsorship` is `true` when the description offers visa sponsorship, `false` when it says there is none, and `null` when it does not say.
- The run's **OUTPUT** record in the key-value store lists, for every company, which platform was used, how many jobs were open, how many matched, and why a company failed (not found, invalid input, unsupported platform, HTTP error).

### Only new jobs

Turn on **Only new jobs** and schedule the task (daily or hourly). Each run saves only the jobs that earlier runs with the same companies and filters did not save, so you get just the new postings and pay only for those. The memory is kept for 180 days in a key-value store named `ats-jobs-state` in your own Apify account. With a per-company limit, the new jobs are looked for among the newest ones.

Each set of companies and filters has its own memory. To share one between two tasks, give both the same **Memory name**; to start over, give a new one.

### Pricing

Pay per event: you pay **only for jobs saved to the dataset**, not for jobs that were read and filtered out, and not for jobs skipped by *Only new jobs*. Filters cost nothing, so a precise filter is the cheapest run. Set a maximum cost for a run and the Actor stops cleanly when it is reached.

### Good to know

- Only **public** job boards are available: the same jobs anyone sees on the company's careers page.
- **SmartRecruiters is not supported:** the robots.txt of its job API allows only LinkedIn's crawler, and this Actor follows robots.txt. SmartRecruiters links are reported in OUTPUT as unsupported and cost nothing.
- **Lever** asks crawlers for one request per second, and the Actor keeps to it. A company with hundreds of Lever jobs takes a few more seconds.
- **Workday** lists jobs newest first, 20 per page. With keywords the Actor uses Workday's own search, and with a job cap it stops reading pages once it has enough candidates, so big employers stay fast and cheap. Description, salary, visa, workplace and employment type filters are decided on each job page: the Actor reads the whole list, then the job pages newest first, and stops once enough jobs pass your cap. Without a cap it reads every candidate job page, which takes a while on big employers. Workday itself lists at most about 2,000 jobs per search.
- **Rippling** lists a job once per location; the Actor merges them into one job with all its locations, and reads each job page for the date and description.
- **Teamtailor** is supported on company.teamtailor.com addresses. Career sites on a custom domain are not detected yet.
- **Gem** boards (from Gem's documented Job Board API) carry no company name: it comes from the board name, so `11x-ai` becomes "11x Ai".
- **Breezy** feeds have no job description, so `descriptionText` is empty for Breezy jobs, and so are the fields read from it. The description filters drop Breezy jobs for that reason.
- **BambooHR** lists no date or description; the Actor reads each job page for them, and the company name from the careers page.
- **Rippling and BambooHR** list jobs without a date, so the Actor reads each job page before applying the per-company cap: you get the newest jobs by their exact date, and the date filter is exact too.
- iCIMS, Taleo and other platforms are not supported yet.
- **Only Workday?** Our **Workday Jobs Scraper** takes just Workday career sites, with the same columns and filters.
- **Don't have a list of companies?** Our **Job Search API** searches 20,000+ company career sites on Greenhouse, Ashby, Lever, Personio, Recruitee, Teamtailor and Gem by job title, location, seniority and date, and can return only the jobs that are new since your last run.

### Support

Found a company whose board doesn't load, or a field that looks wrong? Open an issue on the Actor's Issues tab with the company URL. Issues get an answer within a few days.

# Actor input Schema

## `companies` (type: `array`):

One per line. Paste a careers page URL (boards.greenhouse.io/airbnb, jobs.lever.co/spotify, jobs.ashbyhq.com/ramp, apply.workable.com/huggingface, bunq.recruitee.com, nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite, ottonova.jobs.personio.de, ats.rippling.com/rippling/jobs, career.teamtailor.com, company.breezy.hr, company.bamboohr.com/careers, jobs.gem.com/11x-ai), write platform:slug (lever:spotify), or just the board name (ramp) and the Actor finds the platform for you. Workday, Personio and BambooHR need the full careers URL.

## `keywords` (type: `array`):

Keep jobs whose title contains any of these words (whole-word match, so 'intern' does not match 'internal'). Leave empty for all jobs.

## `excludeKeywords` (type: `array`):

Drop jobs whose title contains any of these words, e.g. senior, manager.

## `locations` (type: `array`):

Keep jobs whose location, any secondary location or country code contains one of these texts (e.g. Berlin, Germany, DE). The word 'remote' also matches jobs flagged remote.

## `remoteOnly` (type: `boolean`):

Keep only jobs that offer a fully remote option.

## `seniority` (type: `array`):

Keep only these levels, read from the job title: "Senior Data Engineer" is senior, "Head of Sales" is director. "Not stated" keeps titles without a level word, which are mostly mid-level roles. Empty = all levels.

## `postedWithinDays` (type: `integer`):

Keep only jobs published in the last N days. Jobs without a publish date are dropped when this is set.

## `workplaceTypes` (type: `array`):

Keep only these workplace types, as the company set them on its job board. Greenhouse jobs rarely state it (85% of the 195,000 in our index), unless the location says Remote, Hybrid or On-site: add "Not stated" to keep them. Empty = all.

## `employmentTypes` (type: `array`):

Keep only these employment types. Most platforms give the type; Greenhouse lists none, so its jobs count as "Not stated" unless the title says intern, part-time or (contract). Empty = all.

## `descriptionKeywords` (type: `array`):

Keep jobs whose description contains any of these words (whole-word match), e.g. Python, Kubernetes, GDPR. Breezy boards publish no description, so their jobs are dropped by description filters. On Workday, Rippling and BambooHR the text of every candidate job is fetched, so those runs take longer.

## `excludeDescriptionKeywords` (type: `array`):

Drop jobs whose description contains any of these words, e.g. clearance, relocation.

## `onlyWithSalary` (type: `boolean`):

Keep only jobs with a salary given by the job board or written in the description (for example a US pay-transparency range).

## `onlyVisaSponsorship` (type: `boolean`):

Keep only jobs whose description says the company sponsors visas. Jobs that say nothing about it are dropped.

## `descriptionFormat` (type: `string`):

Plain text is best for spreadsheets and AI agents. 'None' makes runs faster and results smaller.

## `maxJobsPerCompany` (type: `integer`):

Newest jobs first. 0 or empty = no limit.

## `maxJobs` (type: `integer`):

Stop after saving this many jobs across all companies. 0 or empty = no limit. You also never pay more than the spending limit you set for the run.

## `maxConcurrency` (type: `integer`):

How many companies are read at the same time (1 to 10).

## `onlyNew` (type: `boolean`):

Save only jobs that earlier runs with these same companies and filters did not save, so a daily schedule gives you just the new postings and you pay only for those. With a per-company limit, the new jobs are looked for among the newest ones. The memory is kept for 180 days in a key-value store named "ats-jobs-state" in your account.

## `stateKey` (type: `string`):

By default each set of companies and filters has its own memory. Give a name to share one memory between tasks, or a new name to start over. Letters, digits, dot, dash and underscore.

## Actor input object example

```json
{
  "companies": [
    "https://boards.greenhouse.io/airbnb",
    "https://jobs.lever.co/spotify",
    "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
    "ramp"
  ],
  "remoteOnly": false,
  "onlyWithSalary": false,
  "onlyVisaSponsorship": false,
  "descriptionFormat": "text",
  "maxJobsPerCompany": 20,
  "maxConcurrency": 4,
  "onlyNew": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "https://boards.greenhouse.io/airbnb",
        "https://jobs.lever.co/spotify",
        "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
        "ramp"
    ],
    "maxJobsPerCompany": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/ats-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "https://boards.greenhouse.io/airbnb",
        "https://jobs.lever.co/spotify",
        "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
        "ramp",
    ],
    "maxJobsPerCompany": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/ats-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "https://boards.greenhouse.io/airbnb",
    "https://jobs.lever.co/spotify",
    "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
    "ramp"
  ],
  "maxJobsPerCompany": 20
}' |
apify call digital_influx/ats-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/ats-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/abgKfkx3I2YI6tg4w/builds/dMYrBxExBQkftI50h/openapi.json
