ATS Jobs Scraper: Greenhouse, Lever, Workday, Ashby & more avatar

ATS Jobs Scraper: Greenhouse, Lever, Workday, Ashby & more

Pricing

Pay per event

Go to Apify Store
ATS Jobs Scraper: Greenhouse, Lever, Workday, Ashby & more

ATS Jobs Scraper: Greenhouse, Lever, Workday, Ashby & more

Scrape open jobs from company career pages on 11 ATS platforms (Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee, BambooHR, Personio, Teamtailor, Breezy HR) via their public job boards. Auto-detects the ATS from a company website; one normalized schema.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Changefeeds Tools

Changefeeds Tools

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Scrape the open jobs on company career pages, from 11 applicant tracking systems in one actor: Greenhouse, Lever, Ashby, Workday, Workable, SmartRecruiters, Recruitee, BambooHR, Personio, Teamtailor and Breezy HR.

Give it careers-page URLs, platform:token references, or just company websites. The actor works out which ATS each company uses, reads the company's public job board (the same JSON or XML feed the careers page itself loads), and returns every job in one normalized schema: title, department, team, locations, remote, employment type, structured salary where the platform publishes one, posting date, apply URL and, if you want it, the full description as plain text.

Use it as a Greenhouse jobs scraper, a Workday jobs scraper, a Lever or Ashby jobs scraper, or a general career page scraper for a list of target companies: job aggregators, recruiting and sourcing tools, sales prospecting on hiring signals, labour-market research.

No API keys, no logins, no browser. Only public job-board data.

Supported platforms

PlatformCareers-page URL shapes you can pasteShorthandPublic endpoint used
Greenhousehttps://boards.greenhouse.io/<token>, https://job-boards.greenhouse.io/<token>, …/embed/job_board?for=<token>greenhouse:<token>boards-api.greenhouse.io/v1/boards/<token>/jobs
Leverhttps://jobs.lever.co/<site>lever:<site>api.lever.co/v0/postings/<site>
Ashbyhttps://jobs.ashbyhq.com/<org>ashby:<org>api.ashbyhq.com/posting-api/job-board/<org>
Workdayhttps://<tenant>.wd5.myworkdayjobs.com/<site> (any wdN, optional /en-US/ locale), https://wdN.myworkdaysite.com/recruiting/<tenant>/<site>workday:<tenant>.wd5/<site>…/wday/cxs/<tenant>/<site>/jobs (paginated)
Workablehttps://apply.workable.com/<account>workable:<account>apply.workable.com/api/v1/widget/accounts/<account>
SmartRecruitershttps://careers.smartrecruiters.com/<company>, https://jobs.smartrecruiters.com/<company>smartrecruiters:<company>api.smartrecruiters.com/v1/companies/<company>/postings (paginated)
Recruiteehttps://<company>.recruitee.comrecruitee:<company><company>.recruitee.com/api/offers/
BambooHRhttps://<company>.bamboohr.com/careersbamboohr:<company><company>.bamboohr.com/careers/list
Personiohttps://<company>.jobs.personio.de (or .com)personio:<company><company>.jobs.personio.de/xml
Teamtailorhttps://<company>.teamtailor.comteamtailor:<company><company>.teamtailor.com/jobs.rss
Breezy HRhttps://<company>.breezy.hrbreezy:<company><company>.breezy.hr/json

Paste any page of a board (a job URL works too); the actor reduces it to the board. For anything else, give the company's website and let auto-detection find the board.

Auto-detection from a company website

For an entry like ramp.com or https://www.figma.com, the actor loads, at most 4 pages in total and without retries: the exact page you gave (when the URL has a path, e.g. https://acme.com/about/open-roles), the homepage, the first careers link on it (same domain or a subdomain), /careers and /jobs. It stops at the first page that links to, embeds, or redirects to a supported job board; if several boards are referenced, the most-referenced one wins. The page it was found on is reported in OUTPUT.companies[].detected_on.

If nothing is found, the company gets a free status row with status: "ats_not_found". Custom career sites that render jobs from their own backend (GitLab's is one) cannot be detected this way; if you know the company's board, pass its URL or platform:token directly.

Input

FieldDefaultNotes
companiesrequiredCareers-page URLs, platform:token references or company websites, up to 1,000 per run. Duplicates are skipped.
keywordsnoneKeep jobs whose title contains any of these (case-insensitive).
locationsnoneKeep jobs with a location containing any of these (case-insensitive substring), e.g. Berlin, United Kingdom, Remote.
remoteOnlyfalseKeep only remote jobs (see remote below).
postedWithinDaysnoneKeep only jobs posted in the last N days. Jobs without a posting date are left out when this is set.
includeDescriptionfalseAdd description: plain text, HTML stripped, capped at 20,000 characters.
maxJobsPerCompany1000At most this many jobs (after filters) per company.

Examples:

{
"companies": ["https://boards.greenhouse.io/gitlab", "lever:spotify"],
"maxJobsPerCompany": 50
}
{
"companies": [
"figma.com",
"https://jobs.ashbyhq.com/ramp",
"https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
"smartrecruiters:BoschGroup"
],
"keywords": ["engineer", "data scientist"],
"remoteOnly": true,
"postedWithinDays": 30,
"maxJobsPerCompany": 200
}

Output

One row per job (type: "job"). A real row from a live run on 2026-09-30:

{
"type": "job",
"company": "ashby:ramp",
"platform": "ashby",
"board_token": "ramp",
"job_id": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
"title": "Security Engineer, Cloud",
"department": "Engineering",
"team": "Backend",
"location": "New York, NY (HQ)",
"locations": ["New York, NY (HQ)", "Remote (Canada)", "Remote (US)", "Miami, FL"],
"remote": true,
"employment_type": "FullTime",
"salary_min": 211400,
"salary_max": 290600,
"salary_currency": "USD",
"salary_period": "year",
"posted_at": "2026-04-07T17:12:35.753Z",
"updated_at": null,
"url": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
"scraped_at": "2026-09-30T03:53:18.482Z"
}
  • company is your input entry, exactly as you typed it.
  • description is present only when includeDescription is on.
  • salary_* are filled only when the platform publishes structured compensation (Ashby, Lever salaryRange, Recruitee salary). Free-text salary strings are never parsed, so on other platforms they stay null.
  • remote is the platform's own flag where it has one (Lever, Ashby, Workable, SmartRecruiters, Recruitee, Teamtailor, Breezy, BambooHR, Workday when listed). Otherwise it is true when a location or the title says "Remote", and null (unknown) when nothing says either way.
  • posted_at / updated_at are ISO 8601 (UTC) or null. BambooHR's list has no dates; the job detail does, and it is fetched when you ask for descriptions or set postedWithinDays. Workday lists only "Posted N days ago": the actor turns that into a date (accurate to about a day, since Workday counts days in its own time zone) and leaves "30+ days ago" as null; with includeDescription on, the exact start date from the job page is used instead.
  • locations lists every location the platform gives. Workday lists multi-location jobs as "5 Locations"; locations is then empty unless descriptions (which include the job page) are requested.

Companies that yield no job rows get one free status row each (type: "company"):

statusMeaningCharged
ok_no_jobsBoard read; it has no open jobs, or none matched your filters (jobs_found says how many it has)company-checked only
not_foundNo such board on that platformno
ats_not_foundWebsite checked, no supported job board foundno
errorThe board or website failed to load (message in error)no
invalidNot a recognizable entry, or a website whose detected board is already in the runno
skipped_budgetYour maximum total charge was reached before this companyno

Every run leaves at least one row.

The key-value store record OUTPUT holds the run summary: per-company platform, board_token, status, jobs_found (as the platform reports it), jobs_matched, jobs_returned, error, detected_on; totals; charged_events; stopped_reason (completed, max_total_charge_reached or emit_failure) and incomplete.

Pricing

Pay per event, nothing else:

  • Company checked: $0.001 per company whose job board was read successfully, including boards with no matching jobs.
  • Job returned: $0.0015 per job row in the dataset.

Status rows for boards that are not found, websites without a detectable ATS, errors, invalid entries and skipped companies are free.

Worked examples:

  • 20 companies, 50 jobs each returned: 20 × $0.001 + 1,000 × $0.0015 = $1.52.
  • 200 companies with a keyword filter that leaves about 10 jobs each: 200 × $0.001 + 2,000 × $0.0015 = $3.20.
  • 100 companies checked where nothing matches your filters: 100 × $0.001 = $0.10.

Filters are applied before rows are written, so filtering in the actor is cheaper than filtering afterwards.

Maximum total charge. If you set one, the actor returns only as many rows as fit, then stops cleanly: remaining companies get a free skipped_budget row, OUTPUT.stopped_reason is max_total_charge_reached and OUTPUT.incomplete is true. A company's job rows are delivered before its company-checked charge, so the budget running out never charges you for a board that delivered nothing; the last company can be cut short (its OUTPUT entry says "Stopped after N of M jobs").

Tested with

Verified live on 2026-09-30 (job counts are what each board listed that day):

PlatformInputJobs on board
Greenhousehttps://boards.greenhouse.io/gitlab201
Leverlever:spotify79
Ashbyhttps://jobs.ashbyhq.com/ramp155
Workdayhttps://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite2,000 (Workday's reported total)
Workablehttps://apply.workable.com/skroutz11
SmartRecruitershttps://careers.smartrecruiters.com/BoschGroup4,825
Recruiteehttps://bunq.recruitee.com16
BambooHRhttps://gusto.bamboohr.com/careers5
Personiohttps://pulsegroup.jobs.personio.de14
Teamtailorhttps://polestar.teamtailor.com30
Breezy HRhttps://euler.breezy.hr19
Auto-detectfigma.com → Greenhouse, notion.so → Ashby, ramp.com → Ashby, polestar.com → Teamtailor

Limits, stated plainly

  • Public job boards only. The actor reads what each company publishes on its ATS job board. It does not scrape LinkedIn, Indeed, Glassdoor or other job sites, never logs in, and cannot see internal or unlisted postings (Ashby postings marked unlisted are skipped).
  • Only the 11 platforms above. iCIMS, Taleo, SuccessFactors, Jobvite, Eightfold and custom career sites are not supported. Greenhouse's EU-hosted boards (job-boards.eu.greenhouse.io) are not supported yet.
  • Auto-detection can miss. It looks at 4 pages at most and finds boards that are linked, embedded or redirected to. Career sites on a custom domain that load jobs from their own backend, or behind bot protection, come back ats_not_found.
  • Workday and SmartRecruiters are paginated. Big boards take many requests; the actor stops paging once it has maxJobsPerCompany matching jobs. jobs_found is the total the platform reports.
  • Some filters need the job page. Where a platform's list lacks or summarises a field, jobs that fail only on that field are checked against their job detail before being dropped: Workday multi-location jobs ("5 Locations"), Workday jobs without a remote type or older than "30+ days" (for locations, remoteOnly, postedWithinDays), and BambooHR jobs (the list has no country and no date). That is one extra request per job, only as many as needed to fill maxJobsPerCompany, and at most 500 per company; jobs left unchecked past that limit are counted in OUTPUT.companies[].jobs_unresolved. remoteOnly on a large Workday board is the slowest case.
  • Descriptions cost requests on some platforms. Workday, SmartRecruiters and BambooHR need one request per job for the description, so large runs with includeDescription are slower. Breezy's public feed has no description; description is null there.
  • Field coverage differs by platform. For example Greenhouse has no remote flag or employment type, Workday has no department, Personio and Teamtailor have no salary. Missing fields are null, never guessed.
  • Polite by design. At most 2 requests at a time per host, a short gap between requests, HTTP 429 and 5xx retried with backoff (Retry-After honoured up to 60 s), 20 s timeouts, and a User-Agent that identifies the actor.

FAQ

How do I find a company's board token? Open the company's careers page and click a job: the URL usually shows the platform and token (jobs.lever.co/<token>/…, boards.greenhouse.io/<token>/jobs/…). Paste that URL as is. Or just give the company's website.

Can I scrape a Workday careers site? Yes. Paste the careers-site URL, e.g. https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite. The wdN number and the site id come from that URL.

Why did a company return fewer jobs than its careers page shows? Check your filters and maxJobsPerCompany, then OUTPUT.companies[]: jobs_found is the board's own count and jobs_matched the count after filters. A careers page can also combine several boards; each board is a separate entry.

Does it return salary? Only when the ATS publishes structured pay data (numbers, currency and period). Many postings mention pay only in the description text, which is returned as-is with includeDescription but not parsed into the salary fields.

Is this allowed? The actor reads public job-board endpoints that companies publish so their jobs can be found, at a polite rate. You are responsible for how you use the data.

Local development

pnpm --filter @mmnm/atsjobs test # unit tests on recorded fixtures, no network
pnpm --filter @mmnm/atsjobs build

node dist/main.js runs the actor locally with Apify's local storage (CRAWLEE_STORAGE_DIR=./storage, input in storage/key_value_stores/default/INPUT.json).