Job Board Scraper: Greenhouse, Lever, Ashby & More avatar

Job Board Scraper: Greenhouse, Lever, Ashby & More

Pricing

$1.50 / 1,000 job scrapeds

Go to Apify Store
Job Board Scraper: Greenhouse, Lever, Ashby & More

Job Board Scraper: Greenhouse, Lever, Ashby & More

Extract normalized public jobs from Greenhouse, Lever, Ashby, Workable and SmartRecruiters using board slugs or supported careers pages. Filter by title, location and posted date. Export job details for recruiting, job boards and market research.

Pricing

$1.50 / 1,000 job scrapeds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

ATS Job Aggregator

Collect public vacancies from Greenhouse, Lever, Ashby, Workable, and SmartRecruiters in one consistent dataset. Supply board identifiers or company careers pages, then export jobs as JSON, CSV, Excel, or another Apify dataset format. The actor supports recruiting research, job alerts, and comparisons across employers without requiring an ATS account or private API credentials.

Quick start

Use Python 3.12. From this directory, create an isolated environment and install the dependencies:

python -m venv .venv
& .venv/Scripts/python.exe -m pip install -r requirements.txt
apify validate-schema .actor/input_schema.json
& .venv/Scripts/python.exe -m unittest discover -s tests -v
& ./validation/run_live.ps1

The validation script copies INPUT.json into a new local key-value store, runs .venv/Scripts/python.exe -m src, and records counts and timings in validation/results.json. Calling python -m src directly reads Apify's INPUT record, not the root file automatically. Set APIFY_LOCAL_STORAGE_DIR and place your JSON at key_value_stores/default/INPUT.json beneath that directory, or use the supplied script. Docker uses the Apify Python 3.12 base image.

Inputs

companies is a required array of strings. Supported prefixes are greenhouse:stripe, lever:spotify, ashby:openai, workable:careers, and smartrecruiters:BoschGroup. Slugs identify ATS boards and do not always match the company's legal or trading name. Direct board URLs work too. The bundled input also includes https://www.workable.com/careers to demonstrate discovery.

autoDetect defaults to true. For a general careers URL, the actor retrieves HTML and finds the first supported board link or embedded script. It recognizes boards.greenhouse.io, job-boards.greenhouse.io, jobs.lever.co, jobs.ashbyhq.com, apply.workable.com, and jobs.smartrecruiters.com, including Greenhouse's embed script and escaped URLs in JavaScript. This is HTML inspection; it does not execute JavaScript or navigate an entire website. A page that builds its links only after browser execution may require an explicit slug. Disable discovery to accept only identifiers and direct board URLs.

includeDescription defaults to false. Enable it for both full HTML and derived plain text. Workable and SmartRecruiters require extra detail requests for matching jobs; the other providers supply descriptions in their listings. Lever's description includes its list sections and additional information. HTML is source content, not sanitized for embedding in a website.

locationFilter applies a case-insensitive substring to all normalized locations. titleFilter accepts a Python regular expression; prefix it with (?i) for case-insensitive matching. Invalid patterns fail before requests. postedAfter accepts an ISO calendar date, such as 2026-09-01, and keeps jobs posted strictly after that date. Missing or unparseable posting dates are excluded when this filter is enabled. Provider timestamps retain their reported timezone; the date comparison uses the reported calendar date.

maxJobsPerCompany defaults to 500 and caps matching unique rows per input, after filtering. timeoutSecs defaults to 30 and applies to individual HTTP operations, not the total actor run. Accepted ranges are 1–100000 jobs and 1–300 seconds. Workable uses POST requests with a continuation token; SmartRecruiters uses pages of 100 with increasing offsets.

Results and billing

Each job contains company, ats, jobId, title, department, team, location, locations, remote, employmentType, postedAt, updatedAt, url, applyUrl, descriptionHtml, descriptionText, salary, and source. Missing scalar fields are null. Locations are strings, and salary preserves the provider's structure when present. Company names fall back to the board slug. The source identifies the public listing API, including pagination. An application URL falls back to the public job page when none is exposed.

Remote status is a best-effort boolean based on provider flags, workplace type, or location text. False means no remote signal was detected; it is not a guarantee that a role is office-only. Employment types are preserved as provider labels. Departments, teams, timestamps, and structured salaries are not published consistently by every board.

One job-scraped event costs $0.0015 per saved job: 1,000 jobs cost $1.50 in event fees. Configure the event from .actor/pay_per_event.json in Console before publication, with synthetic start and dataset events disabled. Local runs do not bill. The SDK's charged dataset write respects the event budget. IDs are deduplicated within each input; entering the same board twice produces and charges separate results for each input.

A failed board produces one uncharged row with error. Previously collected jobs remain available if a later page fails. An empty board produces zero job rows. HTTP 429 receives two retries with exponential backoff; Retry-After is honored up to 60 seconds. Other HTTP errors are reported directly. Key-value records SUMMARY-N and SUMMARY record counts, durations, and failures. Review these records even when the process exits successfully. See VALIDATION.md for tested boards, measured results, and remaining deployment checks. No private recruiting records or candidate information are requested.

Example output

One real dataset row from storage/live-20260926-041557/datasets/default/000000001.json, trimmed by omitting fields without changing retained values:

{
"company": "Stripe",
"ats": "greenhouse",
"jobId": "8172510",
"title": "Abuse Investigator",
"location": "Seattle, San Francisco, New York City",
"remote": false,
"postedAt": "2026-09-09T10:50:29-04:00",
"url": "https://stripe.com/jobs/search?gh_jid=8172510",
"descriptionText": null
}

This Stripe job is from the saved 1,581-row validation run. Omitted fields remain available in the full dataset. Description collection was disabled, explaining the null descriptionText; the listing is historical and may no longer be open.

Use cases

  • Recruiting teams can track public openings at selected employers to guide account research.
  • Job-board operations teams can collect matching vacancies from known ATS boards for editorial review.
  • Workforce research teams can compare role titles and locations across a fixed employer panel.
  • Career-services teams can build location-filtered vacancy digests for students from selected company boards.

Pricing example: 2,000 saved matching jobs x $0.0015 per job-scraped event = $3.00 in event fees, using .actor/pay_per_event.json. Local runs do not bill.

Limitations

Coverage is limited to the five supported public ATS sources and discoverable board links. Caps, filters, unavailable dates, and partial failures can reduce results. Remote flags and job counts should be checked against the source before publication; the validation sample is not a completeness or accuracy benchmark.