# Career Site Jobs Scraper: Greenhouse, Lever, Workday & more (`sourcedirect/career-site-jobs`) Actor

Every open job from any company's own career site, in one clean format: title, location, remote, salary, posting date, description. 12 job boards incl. Greenhouse, Lever, Ashby, Workday. Just give a website or careers URL.

- **URL**: https://apify.com/sourcedirect/career-site-jobs.md
- **Developed by:** [Kurt Landman](https://apify.com/sourcedirect) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Career Site Jobs Scraper: Greenhouse, Lever, Workday & more

Get **every open job from a company's own career site**, in one clean, consistent format. You don't need to know which recruiting system the company uses: give a website (`stripe.com`), a careers page URL, or just a name (`Airbnb`), and the Actor finds the job board and returns all its jobs.

Each job includes:

- title
- department and team
- locations
- remote / hybrid / on-site
- employment type
- salary range, when published
- posting date
- full description
- job and apply links

### Supported job boards (ATS)

| Job board | Example URL | Description | Salary |
|---|---|---|---|
| Greenhouse | `boards.greenhouse.io/airbnb` | ✅ | — |
| Lever | `jobs.lever.co/zoox` | ✅ | ✅ |
| Ashby | `jobs.ashbyhq.com/ramp` | ✅ | ✅ |
| Workday | `nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite` | ✅ | — |
| SmartRecruiters | `careers.smartrecruiters.com/Visa` | ✅ | — |
| Workable | `apply.workable.com/huggingface` | ✅ | — |
| Recruitee | `company.recruitee.com` | ✅ | ✅ |
| Personio | `company.jobs.personio.de` | ✅ | — |
| Teamtailor | `company.teamtailor.com` | ✅ | — |
| BambooHR | `company.bamboohr.com/careers` | ✅ | ✅ (as text) |
| Breezy HR | `company.breezy.hr` | ✅ | ✅ (as text) |
| Pinpoint | `company.pinpointhq.com` | ✅ | ✅ |

Salary appears whenever the company publishes it on its job board.

### Why use it

- **Straight from the source.** The data comes from each company's own public job board, not a stale aggregator. Postings are live, and closed jobs disappear.
- **One schema for 12 systems.** Workplace type, employment type, dates and salary are normalized, so jobs from Workday and Lever line up in the same columns.
- **Works from a website or a name.** It reads the careers page to find the job board. If the board isn't linked there, it checks the supported boards for the company's name.
- **Fast and cheap.** It uses public job-board APIs, with no browser and no proxy needed. 15 companies across all 12 systems took under a minute in our tests.

### Use cases

- **Recruiting and job boards:** keep a niche job board or newsletter filled with fresh jobs from your list of employers.
- **Sales intelligence:** spot companies hiring for roles that signal they need your product (e.g. "hiring data engineers").
- **Market and salary research:** compare pay ranges, remote policies and hiring volume across companies.
- **Job-search agents:** let an AI assistant check target companies for new matching roles every day.

### Input

| Field | What it does |
|---|---|
| `companies` | Websites, careers or job-board URLs, company names, or shorthands like `greenhouse:airbnb`, `lever:zoox`, `ashby:ramp`, `workday:nvidia.wd5/NVIDIAExternalCareerSite` |
| `keywords` / `excludeKeywords` | Keep or drop jobs by words in the title (optionally in the description too) |
| `locations` | e.g. `London`, `Germany`, `US`, `remote` |
| `remoteOnly` | Fully remote jobs only |
| `departments` | e.g. `engineering`, `sales` |
| `postedWithinDays` | Only recent jobs |
| `descriptionFormat` | `text`, `html` or `none` |
| `maxJobsPerCompany` / `maxJobs` | Limits |

#### Examples

All jobs from three companies:

```json
{ "companies": ["stripe.com", "https://jobs.lever.co/zoox", "https://jobs.ashbyhq.com/ramp"] }
```

Remote engineering jobs posted in the last 7 days, without descriptions:

```json
{
  "companies": ["Airbnb", "ramp.com", "workday:nvidia.wd5/NVIDIAExternalCareerSite"],
  "keywords": ["engineer"],
  "remoteOnly": true,
  "postedWithinDays": 7,
  "descriptionFormat": "none"
}
```

#### Output example

This is a real result from a test run, with the description shortened:

```json
{
  "type": "job",
  "jobId": "34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "title": "Security Engineer, Cloud",
  "company": "Ramp",
  "ats": "ashby",
  "atsCompanyId": "ramp",
  "department": "Engineering",
  "team": "Backend",
  "employmentType": "full-time",
  "seniority": null,
  "workplaceType": "hybrid",
  "isRemote": false,
  "location": "New York, NY (HQ)",
  "locations": [
    "New York, NY (HQ)",
    "Remote (Canada)",
    "Remote (US)",
    "Miami, FL"
  ],
  "city": "New York City",
  "region": "NY",
  "country": "USA",
  "countryCode": null,
  "salary": {
    "min": 211400,
    "max": 290600,
    "currency": "USD",
    "period": "year",
    "text": "$211.4K - $290.6K"
  },
  "postedAt": "2026-04-07T17:12:35.753Z",
  "postedAtIsApproximate": false,
  "updatedAt": null,
  "url": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245",
  "applyUrl": "https://jobs.ashbyhq.com/ramp/34413f8d-26bf-4bbc-8ade-eb309a0e2245/application",
  "sourceUrl": "https://api.ashbyhq.com/posting-api/job-board/ramp?includeCompensation=true",
  "companyInput": "https://jobs.ashbyhq.com/ramp",
  "description": "About Ramp\n\nRamp is building the smart infrastructure for finance teams, embedded in the t ...",
  "scrapedAt": "2026-10-01T04:08:39.958Z"
}
```

The run also saves a **per-company summary** in the key-value store record `COMPANIES`. It shows which job board was found and how (`url`, `careers-page` or `name-match`), plus how many jobs were listed and returned. The dataset has two ready-made table views: **Jobs** and **Jobs with salary**.

### Pricing

**$0.002 per job returned.** There are no platform-usage charges on top and no fee per run.

For example, a company with 150 open jobs costs $0.30, and 1,000 jobs cost $2.

You are **never charged** for companies that can't be found or fetched. They appear as items with `"type": "error"` and a reason. If you set a maximum cost, the run stops cleanly when it is reached.

### Using it from AI agents (MCP)

Available through the [Apify MCP server](https://mcp.apify.com) for Claude, ChatGPT, Cursor and other MCP clients. Agents can also pay per call through Apify's agentic payments.

Example prompts:

- "Find remote data-engineering jobs at Stripe, Ramp and Airbnb posted this week."
- "Which of these 20 companies are hiring salespeople in Germany?"
- "List NVIDIA's open architect roles with their locations."

Agents: batch many companies in one call. Set `descriptionFormat: "none"` when you only need titles and links.

### Good to know

- **Detection.** When you give a website, the Actor reads the homepage and common careers pages (`/careers`, `/jobs`, `careers.` and `jobs.` subdomains) for links to a supported board. If none is linked, it checks whether a board account with the company's name exists, and the summary marks this as `name-match`. For full certainty, pass the job-board URL.
- **Workday.** Workday's list shows relative dates ("Posted 3 Days Ago"). The exact date comes with the job details, which are fetched when you request descriptions or filter by location or date. Otherwise `postedAtIsApproximate` is `true`.
- **Not supported yet:** companies on systems outside the 12 above (for example iCIMS, Taleo, SuccessFactors) and fully custom career sites. They return an error item, at no cost.
- **Legal.** Employers publish these job postings publicly so they can be found and applied to. The Actor reads public job data only and no personal data.

### FAQ

**How fresh is the data?** It is fetched live at run time from each company's board.

**How many companies per run?** Up to 5,000. Use Apify Schedules to watch your list daily, and connect the results to Google Sheets, Slack, Zapier, Make or n8n.

**Something missing or wrong?** Open an issue on the **Issues** tab with the company URL. We add new job boards based on demand.

# Changelog

This Actor's version history is a separate document: https://apify.com/sourcedirect/career-site-jobs/changelog.md

# Actor input Schema

## `companies` (type: `array`):

One per line. Accepted: job-board URLs (boards.greenhouse.io/airbnb, jobs.lever.co/zoox, jobs.ashbyhq.com/ramp, nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite, ...), company websites (stripe.com), company names (Airbnb), or shorthands (greenhouse:airbnb, lever:zoox, ashby:ramp, workday:nvidia.wd5/NVIDIAExternalCareerSite).

## `keywords` (type: `array`):

Only return jobs whose title contains any of these words (not case-sensitive), e.g. engineer, data, sales. Leave empty for all jobs.

## `keywordsInDescription` (type: `boolean`):

Also return jobs whose description (not only title) mentions a keyword. Slower for boards that need one request per job.

## `excludeKeywords` (type: `array`):

Skip jobs whose title contains any of these words, e.g. intern, senior.

## `locations` (type: `array`):

Only jobs whose location mentions any of these, e.g. London, Germany, US, CA, remote.

## `remoteOnly` (type: `boolean`):

Only return jobs marked fully remote.

## `departments` (type: `array`):

Only jobs whose department or team contains any of these, e.g. engineering, marketing.

## `postedWithinDays` (type: `integer`):

Only jobs posted in the last N days. 0 = any date. Jobs whose board shows no posting date are excluded when this is set.

## `descriptionFormat` (type: `string`):

Plain text is best for reading and AI; HTML keeps formatting; None is fastest and smallest.

## `maxJobsPerCompany` (type: `integer`):

Upper limit of job listings read from each company's board.

## `maxJobs` (type: `integer`):

Stop after returning this many jobs across all companies. 0 = no limit.

## `detectFromWebsite` (type: `boolean`):

When you give a website or company name, look up the careers page and the supported job boards to find where the company posts its jobs. Turn off to accept only job-board URLs and shorthands.

## `maxConcurrency` (type: `integer`):

How many companies to fetch at the same time.

## `proxyConfiguration` (type: `object`):

Not needed for the job-board APIs. Only enable it if a company website blocks the careers-page lookup.

## Actor input object example

```json
{
  "companies": [
    "stripe.com",
    "https://jobs.lever.co/zoox",
    "workday:nvidia.wd5/NVIDIAExternalCareerSite"
  ],
  "keywords": [],
  "keywordsInDescription": false,
  "excludeKeywords": [],
  "locations": [],
  "remoteOnly": false,
  "departments": [],
  "postedWithinDays": 0,
  "descriptionFormat": "text",
  "maxJobsPerCompany": 1000,
  "maxJobs": 0,
  "detectFromWebsite": true,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `jobs` (type: `string`):

No description

## `overview` (type: `string`):

No description

## `salaries` (type: `string`):

No description

## `companies` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "https://jobs.ashbyhq.com/ramp",
        "https://boards.greenhouse.io/airbnb"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("sourcedirect/career-site-jobs").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "companies": [
        "https://jobs.ashbyhq.com/ramp",
        "https://boards.greenhouse.io/airbnb",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("sourcedirect/career-site-jobs").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "https://jobs.ashbyhq.com/ramp",
    "https://boards.greenhouse.io/airbnb"
  ]
}' |
apify call sourcedirect/career-site-jobs --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sourcedirect/career-site-jobs"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wwvAeEJFhtm8sM68f/builds/4dmYaP2tpBpF3VqI9/openapi.json
