# Careers Page Jobs Scraper - Greenhouse, Lever, Workday & More (`herbcoder/careers-page-jobs-scraper`) Actor

Turn any company careers page into structured job data. Auto-detects Greenhouse, Lever, Ashby, Workday, SmartRecruiters, Personio, Recruitee, BambooHR and Workable boards and returns title, location, remote flag, salary, description and apply link in one schema. Pay per job.

- **URL**: https://apify.com/herbcoder/careers-page-jobs-scraper.md
- **Developed by:** [HerbCode LLC](https://apify.com/herbcoder) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Careers Page Jobs Scraper (Greenhouse, Lever, Ashby, Workday, SmartRecruiters, Personio, Recruitee, BambooHR, Workable)

Turn any company careers page into clean, structured job data. Paste the careers URL (or the ATS job-board URL) and get every open position with title, location, remote flag, department, employment type, posting date, salary (when published) and the full description - all in **one unified schema, whichever ATS the company uses**.

Works with the nine most common applicant tracking systems used on tech and enterprise careers pages:

| ATS | Example source you can paste |
|---|---|
| Greenhouse | `https://boards.greenhouse.io/stripe` or `greenhouse:stripe` |
| Lever | `https://jobs.lever.co/spotify` |
| Ashby | `https://jobs.ashbyhq.com/ramp` |
| Workday | `https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite` |
| SmartRecruiters | `https://jobs.smartrecruiters.com/SmartRecruiters` |
| Personio | `https://demo.jobs.personio.de` |
| Recruitee | `https://company.recruitee.com` |
| BambooHR | `https://company.bamboohr.com/careers` |
| Workable | `https://apply.workable.com/company` |
| **Any company careers page** | `https://www.example.com/careers` - the page is scanned for an embedded ATS board and the right API is used automatically |

No proxies, no browser, no login: every provider is read through the public job-board endpoint the ATS vendor publishes for careers-site embedding, so runs are fast (hundreds of jobs in seconds) and reliable.

### What you can do with it

- **Job alerts** - schedule a daily run over the companies you care about, filter by keyword/location/remote, push new rows to Slack, email or a Google Sheet via Apify integrations.
- **Recruiting and sales intelligence** - track which companies are hiring for which roles (headcount signals, tech-stack signals from descriptions, expansion into new locations).
- **Job boards and aggregators** - keep a niche board fresh with direct links to the original posting and apply page.
- **Market research** - salary ranges, remote policies and role mix across hundreds of employers in one dataset.
- **AI agents** - the plain-text description field is ready for LLM summarization or matching against a CV.

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `sources` | array of strings | required | ATS board URLs, `provider:slug` shorthands, or company careers pages. One per company. |
| `keywords` | array | `[]` | Keep jobs whose title contains any of these (case-insensitive). |
| `excludeKeywords` | array | `[]` | Drop jobs whose title contains any of these. |
| `locations` | array | `[]` | Keep jobs whose location contains any of these strings. |
| `remoteOnly` | boolean | `false` | Keep only remote jobs. |
| `postedAfter` | date | - | Drop jobs published before this date (jobs without a date are kept). |
| `includeDescription` | boolean | `true` | Include the job description. Turn off for faster Workday / SmartRecruiters / BambooHR runs. |
| `descriptionFormat` | `text` / `html` / `both` | `text` | Plain text is best for spreadsheets and LLMs. |
| `parseSalaryFromText` | boolean | `true` | Parse `$120,000 - $150,000`-style ranges out of descriptions when the ATS gives no structured pay. |
| `maxJobsPerSource` | integer | `0` (all) | Cap per company. |
| `maxJobs` | integer | `0` (all) | Cap for the whole run - hard cost ceiling. |
| `searchText` | string | - | Server-side search for Workday boards only. |
| `sourceConcurrency` | integer | `3` | Companies fetched in parallel. |

Example input:

```json
{
    "sources": [
        "https://boards.greenhouse.io/stripe",
        "https://jobs.lever.co/spotify",
        "https://jobs.ashbyhq.com/ramp",
        "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite"
    ],
    "keywords": ["engineer", "developer"],
    "remoteOnly": true,
    "postedAfter": "2026-09-01",
    "maxJobsPerSource": 200
}
```

### Output

One dataset item per job. Every provider is mapped to the same fields; a field is `null` when the ATS does not expose it.

```json
{
    "id": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
    "provider": "ashby",
    "company": "ramp",
    "companySlug": "ramp",
    "title": "Security Engineer, Cloud",
    "department": "Engineering",
    "team": "Backend",
    "location": "New York, NY (HQ)",
    "locations": ["New York, NY (HQ)", "Remote (Canada)", "Remote (US)", "Miami, FL"],
    "remote": true,
    "workplaceType": "hybrid",
    "employmentType": "FullTime",
    "postedAt": "2026-04-07T17:12:35.753Z",
    "updatedAt": null,
    "salary": { "min": 211400, "max": 290600, "currency": "USD", "interval": "year", "raw": "$211.4K - $290.6K" },
    "descriptionText": "ABOUT RAMP\n\nRamp is building the smart infrastructure for finance teams...",
    "descriptionHtml": null,
    "url": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
    "applyUrl": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245/application",
    "sourceUrl": "https://jobs.ashbyhq.com/ramp",
    "scrapedAt": "2026-09-21T16:43:08.071Z"
}
```

A `RUN_SUMMARY` record in the key-value store lists how many sources succeeded, how many jobs were fetched/filtered/emitted and any per-source errors, so a failing board never silently disappears.

Export as JSON, CSV, Excel, XML or RSS from the Dataset tab, or read it through the API / MCP.

### Pricing

Pay per event, no subscription:

| Event | Price | What it means |
|---|---|---|
| Actor start | $0.005 | Once per run. |
| `job-result` | $0.0015 (= $1.50 per 1,000 jobs) | One per job saved to the dataset. Jobs removed by your filters are free. |

Example: watching 25 companies with ~1,500 open jobs costs about $2.26 per full refresh. Use `keywords`, `locations`, `remoteOnly`, `postedAfter` and `maxJobs` to pay only for the rows you want. Free-plan users can run it within their free platform credit.

### Limits and notes

- Only **published, public** postings are returned - exactly what the company shows on its careers site. Nothing behind a login is accessed.
- Posting dates: Greenhouse, Lever, Ashby, SmartRecruiters, Personio, Recruitee, BambooHR and Workable expose exact timestamps; Workday exposes the posted date on the detail page (fetched when `includeDescription` is on) and otherwise a relative "Posted 3 Days Ago" string that is converted to a date.
- Structured salary is available from Lever, Ashby and Recruitee when the employer publishes it; for other providers the description is parsed with `parseSalaryFromText`.
- Workday boards can list thousands of jobs; each job needs one detail request for the description. Use `searchText`, `maxJobsPerSource` or `includeDescription: false` to keep runs short.
- Auto-detection of a company careers page works when the page links or embeds one of the supported ATS boards. If it does not, pass the board URL directly.
- Teamtailor, iCIMS, Taleo, SuccessFactors and Jobvite are not supported yet - open an issue to request one.

### FAQ

**Is this legal?** The Actor reads job-board endpoints that ATS vendors publish specifically for displaying jobs on third-party sites (e.g. Greenhouse Job Board API, Lever Postings API, Ashby Job Posting API). Only public job postings are collected; no personal data and no login-protected content.

**How fresh is the data?** Each run reads live from the ATS, so the dataset reflects the board at the moment of the run. Schedule the Actor to keep a feed current.

**Can I deduplicate across runs?** Use the `id` + `provider` + `companySlug` fields as a stable key, or the dataset deduplication options in Apify integrations.

# Changelog

This Actor's version history is a separate document: https://apify.com/herbcoder/careers-page-jobs-scraper/changelog.md

# Actor input Schema

## `sources` (type: `array`):

One entry per company. Accepts ATS board URLs (https://boards.greenhouse.io/stripe, https://jobs.lever.co/spotify, https://jobs.ashbyhq.com/ramp, https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite, https://jobs.smartrecruiters.com/SmartRecruiters, https://demo.jobs.personio.de, https://company.recruitee.com, https://company.bamboohr.com/careers, https://apply.workable.com/gorgias), shorthand `provider:slug` (e.g. `greenhouse:stripe`, `workday:nvidia.wd5/NVIDIAExternalCareerSite`), or a company's own careers page URL, which is scanned for an embedded ATS.

## `keywords` (type: `array`):

Keep only jobs whose title contains at least one of these words (case-insensitive). Leave empty for all jobs.

## `excludeKeywords` (type: `array`):

Drop jobs whose title contains any of these words (case-insensitive).

## `locations` (type: `array`):

Keep only jobs whose location contains one of these strings, e.g. `New York`, `Berlin`, `United States`, `Remote`.

## `remoteOnly` (type: `boolean`):

Keep only jobs the ATS marks as remote, or whose title/location says remote.

## `postedAfter` (type: `string`):

ISO date (e.g. 2026-09-01). Jobs with a known posting date before this are dropped. Jobs without a date are kept.

## `includeDescription` (type: `boolean`):

Fetch and include the full job description. Turning this off makes Workday, SmartRecruiters and BambooHR runs much faster (no per-job detail request).

## `descriptionFormat` (type: `string`):

`text` = clean plain text (best for LLMs and spreadsheets), `html` = original HTML, `both` = both fields.

## `parseSalaryFromText` (type: `boolean`):

When the ATS does not expose structured pay, try to parse a salary range (e.g. `$120,000 - $150,000`) from the description text.

## `maxJobsPerSource` (type: `integer`):

Stop after this many jobs for each company (0 = no limit). Useful for very large Workday boards.

## `maxJobs` (type: `integer`):

Stop the whole run after this many jobs have been saved (0 = no limit). Caps your cost.

## `searchText` (type: `string`):

Server-side search string applied on Workday boards only (Workday boards can have thousands of jobs). Other providers use the title keyword filter.

## `sourceConcurrency` (type: `integer`):

How many companies to fetch at the same time.

## Actor input object example

```json
{
  "sources": [
    "https://boards.greenhouse.io/stripe",
    "https://jobs.lever.co/spotify",
    "https://jobs.ashbyhq.com/ramp"
  ],
  "keywords": [],
  "excludeKeywords": [],
  "locations": [],
  "remoteOnly": false,
  "includeDescription": true,
  "descriptionFormat": "text",
  "parseSalaryFromText": true,
  "maxJobsPerSource": 0,
  "maxJobs": 0,
  "sourceConcurrency": 3
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sources": [
        "https://boards.greenhouse.io/stripe",
        "https://jobs.lever.co/spotify",
        "https://jobs.ashbyhq.com/ramp"
    ],
    "keywords": [],
    "excludeKeywords": [],
    "locations": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("herbcoder/careers-page-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sources": [
        "https://boards.greenhouse.io/stripe",
        "https://jobs.lever.co/spotify",
        "https://jobs.ashbyhq.com/ramp",
    ],
    "keywords": [],
    "excludeKeywords": [],
    "locations": [],
}

# Run the Actor and wait for it to finish
run = client.actor("herbcoder/careers-page-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sources": [
    "https://boards.greenhouse.io/stripe",
    "https://jobs.lever.co/spotify",
    "https://jobs.ashbyhq.com/ramp"
  ],
  "keywords": [],
  "excludeKeywords": [],
  "locations": []
}' |
apify call herbcoder/careers-page-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,herbcoder/careers-page-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xfPhBIBYnlWSdfBhg/builds/Uh9ZcpnVf8qg3wJE8/openapi.json
