# Job Board Scraper: Greenhouse, Lever, Ashby & More (`everyotherfriday/ats-jobs`) Actor

Extract normalized public jobs from Greenhouse, Lever, Ashby, Workable and SmartRecruiters using board slugs or supported careers pages. Filter by title, location and posted date. Export job details for recruiting, job boards and market research.

- **URL**: https://apify.com/everyotherfriday/ats-jobs.md
- **Developed by:** [Paul Vasquez](https://apify.com/everyotherfriday) (community)
- **Categories:** Jobs, Lead generation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.50 / 1,000 job scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ATS Job Aggregator

Collect public vacancies from Greenhouse, Lever, Ashby, Workable, and
SmartRecruiters in one consistent dataset. Supply board identifiers or company
careers pages, then export jobs as JSON, CSV, Excel, or another Apify dataset
format. The actor supports recruiting research, job alerts, and comparisons
across employers without requiring an ATS account or private API credentials.

### Quick start

Use Python 3.12. From this directory, create an isolated environment and install
the dependencies:

```powershell
python -m venv .venv
& .venv/Scripts/python.exe -m pip install -r requirements.txt
apify validate-schema .actor/input_schema.json
& .venv/Scripts/python.exe -m unittest discover -s tests -v
& ./validation/run_live.ps1
```

The validation script copies `INPUT.json` into a new local key-value store,
runs `.venv/Scripts/python.exe -m src`, and records counts and timings in
`validation/results.json`. Calling `python -m src` directly reads Apify's INPUT
record, not the root file automatically. Set `APIFY_LOCAL_STORAGE_DIR` and place
your JSON at `key_value_stores/default/INPUT.json` beneath that directory, or
use the supplied script. Docker uses the Apify Python 3.12 base image.

### Inputs

`companies` is a required array of strings. Supported prefixes are
`greenhouse:stripe`, `lever:spotify`, `ashby:openai`, `workable:careers`, and
`smartrecruiters:BoschGroup`. Slugs identify ATS boards and do not always match
the company's legal or trading name. Direct board URLs work too. The bundled
input also includes `https://www.workable.com/careers` to demonstrate discovery.

`autoDetect` defaults to true. For a general careers URL, the actor retrieves
HTML and finds the first supported board link or embedded script. It recognizes
boards.greenhouse.io, job-boards.greenhouse.io, jobs.lever.co, jobs.ashbyhq.com,
apply.workable.com, and jobs.smartrecruiters.com, including Greenhouse's embed
script and escaped URLs in JavaScript. This is HTML inspection; it does not
execute JavaScript or navigate an entire website. A page that builds its links
only after browser execution may require an explicit slug. Disable discovery
to accept only identifiers and direct board URLs.

`includeDescription` defaults to false. Enable it for both full HTML and derived
plain text. Workable and SmartRecruiters require extra detail requests for
matching jobs; the other providers supply descriptions in their listings.
Lever's description includes its list sections and additional information.
HTML is source content, not sanitized for embedding in a website.

`locationFilter` applies a case-insensitive substring to all normalized
locations. `titleFilter` accepts a Python regular expression; prefix it with
`(?i)` for case-insensitive matching. Invalid patterns fail before requests.
`postedAfter` accepts an ISO calendar date, such as `2026-09-01`, and keeps jobs
posted strictly after that date. Missing or unparseable posting dates are
excluded when this filter is enabled. Provider timestamps retain their
reported timezone; the date comparison uses the reported calendar date.

`maxJobsPerCompany` defaults to 500 and caps matching unique rows per input,
after filtering. `timeoutSecs` defaults to 30 and applies to individual HTTP
operations, not the total actor run. Accepted ranges are 1–100000 jobs and
1–300 seconds. Workable uses POST requests with a continuation token;
SmartRecruiters uses pages of 100 with increasing offsets.

### Results and billing

Each job contains `company`, `ats`, `jobId`, `title`, `department`, `team`,
`location`, `locations`, `remote`, `employmentType`, `postedAt`, `updatedAt`,
`url`, `applyUrl`, `descriptionHtml`, `descriptionText`, `salary`, and `source`.
Missing scalar fields are null. Locations are strings, and salary preserves
the provider's structure when present. Company names fall back to the board
slug. The source identifies the public listing API, including pagination.
An application URL falls back to the public job page when none is exposed.

Remote status is a best-effort boolean based on provider flags, workplace
type, or location text. False means no remote signal was detected; it is not
a guarantee that a role is office-only. Employment types are preserved as
provider labels. Departments, teams, timestamps, and structured salaries are
not published consistently by every board.

One `job-scraped` event costs **$0.0015 per saved job**: 1,000 jobs cost $1.50
in event fees. Configure the event from `.actor/pay_per_event.json` in Console
before publication, with synthetic start and dataset events disabled. Local
runs do not bill. The SDK's charged dataset write respects the event budget.
IDs are deduplicated within each input; entering the same board twice produces
and charges separate results for each input.

A failed board produces one uncharged row with `error`. Previously collected
jobs remain available if a later page fails. An empty board produces zero job
rows. HTTP 429 receives two retries with exponential backoff; Retry-After is
honored up to 60 seconds. Other HTTP errors are reported directly. Key-value
records `SUMMARY-N` and `SUMMARY` record counts, durations, and failures.
Review these records even when the process exits successfully. See
`VALIDATION.md` for tested boards, measured results, and remaining deployment
checks. No private recruiting records or candidate information are requested.

### Example output

One real dataset row from [storage/live-20260926-041557/datasets/default/000000001.json](storage/live-20260926-041557/datasets/default/000000001.json), trimmed by omitting fields without changing retained values:

```json
{
  "company": "Stripe",
  "ats": "greenhouse",
  "jobId": "8172510",
  "title": "Abuse Investigator",
  "location": "Seattle, San Francisco, New York City",
  "remote": false,
  "postedAt": "2026-09-09T10:50:29-04:00",
  "url": "https://stripe.com/jobs/search?gh_jid=8172510",
  "descriptionText": null
}
```

This Stripe job is from the saved 1,581-row validation run. Omitted fields remain available in the full dataset. Description collection was disabled, explaining the null descriptionText; the listing is historical and may no longer be open.

### Use cases

- Recruiting teams can track public openings at selected employers to guide account research.
- Job-board operations teams can collect matching vacancies from known ATS boards for editorial review.
- Workforce research teams can compare role titles and locations across a fixed employer panel.
- Career-services teams can build location-filtered vacancy digests for students from selected company boards.

**Pricing example:** 2,000 saved matching jobs x $0.0015 per `job-scraped` event = **$3.00 in event fees**, using `.actor/pay_per_event.json`. Local runs do not bill.

### Limitations

Coverage is limited to the five supported public ATS sources and discoverable board links. Caps, filters, unavailable dates, and partial failures can reduce results. Remote flags and job counts should be checked against the source before publication; the validation sample is not a completeness or accuracy benchmark.

# Actor input Schema

## `companies` (type: `array`):

Public careers URLs or ats:slug strings for greenhouse, lever, ashby, workable, smartrecruiters.

## `autoDetect` (type: `boolean`):

Fetch careers URLs and find supported board links or embedded scripts.

## `includeDescription` (type: `boolean`):

Return both HTML and plain text; fetch individual Workable and SmartRecruiters details.

## `locationFilter` (type: `string`):

Optional case-insensitive substring matched against every location.

## `titleFilter` (type: `string`):

Optional Python regex. Use (?i) for case-insensitive matching.

## `maxJobsPerCompany` (type: `integer`):

Stop after this many matching unique jobs for each input.

## `postedAfter` (type: `string`):

Optional ISO date YYYY-MM-DD, exclusive. Jobs without a usable posting date are excluded.

## `timeoutSecs` (type: `integer`):

Timeout for each network operation, including detail requests.

## Actor input object example

```json
{
  "companies": [
    "greenhouse:stripe",
    "lever:spotify"
  ],
  "autoDetect": true,
  "includeDescription": false,
  "maxJobsPerCompany": 10,
  "timeoutSecs": 30
}
```

# Actor output Schema

## `jobs` (type: `string`):

All normalized jobs and error rows.

## `jobsCsv` (type: `string`):

Same rows as CSV.

## `summaries` (type: `string`):

SUMMARY and SUMMARY-N keys with counts, timings and failures.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "greenhouse:stripe",
        "lever:spotify"
    ],
    "maxJobsPerCompany": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("everyotherfriday/ats-jobs").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "greenhouse:stripe",
        "lever:spotify",
    ],
    "maxJobsPerCompany": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("everyotherfriday/ats-jobs").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "greenhouse:stripe",
    "lever:spotify"
  ],
  "maxJobsPerCompany": 10
}' |
apify call everyotherfriday/ats-jobs --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,everyotherfriday/ats-jobs"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pDORu4gLvBeOphqt2/builds/t62mSPyCJbxdjrbxN/openapi.json
