# Career Site Jobs Scraper (`dektionstudio/career-site-jobs-scraper`) Actor

Open jobs from company career sites on Greenhouse, Lever, Ashby, Workday, SmartRecruiters, Recruitee and Workable in one run: title, location, remote flag, salary where posted, date and full description. Give a board URL or just a domain. $1 per 1,000 jobs.

- **URL**: https://apify.com/dektionstudio/career-site-jobs-scraper.md
- **Developed by:** [Dektion Studio](https://apify.com/dektionstudio) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Career Site Jobs Scraper

Give it a list of companies and get back every open role from their own career sites: title, location, remote flag, department, salary where the company posts one, posting date and the full description. Companies run those sites on a handful of job systems, and this Actor reads seven of them: Greenhouse, Lever, Ashby, Workday, SmartRecruiters, Recruitee and Workable. You get the company's own listing, not a copy from a job aggregator.

You don't need to know which system a company uses. Paste its job board URL or just its domain: given `figma.com`, the Actor found Figma's Greenhouse board and its 163 open jobs. The job systems publish these listings openly so companies can embed them on their own sites, which makes the data quick and steady to collect. A run on 27 September covered eight companies across all seven systems, including NVIDIA's 2,000 open roles on Workday and Bosch's 4,802 on SmartRecruiters.

### What you get for each job

- Company, job title, department and team
- Location, all listed locations, and whether the job is remote
- Employment type, such as full time or contract
- Salary: structured min, max, currency and period when the job board provides it, which is common on Ashby, otherwise a pay range found in the description, like "$120,000 - $150,000"
- Posting date, job page URL, apply URL
- The full description as plain text

### Input

```json
{
    "companies": [
        "https://jobs.ashbyhq.com/openai",
        "https://boards.greenhouse.io/stripe",
        "https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite",
        "figma.com"
    ],
    "keywords": ["engineer", "data"],
    "locations": ["London", "Remote"],
    "onlyNewJobs": false
}
```

- `companies`: a job board URL, or just the company's domain. With a domain, the Actor looks for a job board link on the company's careers pages and tries the domain name on each system. Up to 200 companies per run.
- `keywords`, `locations`, `remoteOnly` and `postedWithinDays` filter the jobs before you pay for them.
- `onlyNewJobs` is for scheduled runs. It returns only the jobs that weren't open the last time it ran, so a daily schedule gives you a feed of new openings.
- `includeDescription` can be turned off to make big Workday or SmartRecruiters runs faster, since those need one extra request per job for the description.

### Output

A real item from OpenAI's Ashby board, description shortened:

```json
{
    "company": "openai",
    "ats": "ashby",
    "title": "Research Engineer",
    "department": "Research",
    "location": "San Francisco",
    "remote": false,
    "employmentType": "FullTime",
    "salary": { "min": 250000, "max": 445000, "currency": "USD", "interval": "1 YEAR", "text": "$250K – $445K • Offers Equity" },
    "salaryText": "$250K – $445K • Offers Equity",
    "postedAt": "2025-04-05T00:03:20.653Z",
    "url": "https://jobs.ashbyhq.com/openai/240d459b-696d-43eb-8497-fab3e56ecd9b",
    "description": "By applying to this role, ..."
}
```

The Jobs view shows one row per job, ready to export to Excel or Google Sheets. A company whose board can't be found or read shows up as an error row, with the reason, and costs nothing.

### Pricing

$0.001 per job, so 1,000 jobs cost $1. Filters are applied before charging, and error rows are free. Apify adds $0.00005 per run for starting the Actor. Set a maximum charge per run in the run options and the Actor stops cleanly when it gets there.

### Limits

- Seven systems are supported. Companies on other systems, or with a home-made careers page, need to be added some other way. For a domain the Actor can't match, paste the job board URL instead.
- On Lever, Ashby, Workday and Workable the company name is the board's ID, like `openai` or `nvidia`, because those systems don't send a display name.
- Salary shows up only when the company publishes one. Greenhouse, Lever, Workday and Workable rarely have structured pay, so for them it comes from the description text.
- Workday gives posting dates like "Posted 3 Days Ago" in its list. With descriptions on, the Actor reads the exact date from the job page.
- Workable sometimes rate-limits Apify's servers. When a board answers "too many requests", the Actor retries it through a residential IP.

### Changelog

- 2026-09-27: first version.

# Actor input Schema

## `companies` (type: `array`):

One per line: a job board URL (boards.greenhouse.io/stripe, jobs.lever.co/palantir, jobs.ashbyhq.com/openai, nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite...) or just the company domain (stripe.com), and the Actor looks for its job board. Up to 200 per run.

## `keywords` (type: `array`):

Keep jobs whose title contains any of these words, for example engineer or sales. Empty keeps all jobs.

## `locations` (type: `array`):

Keep jobs whose location contains any of these, for example London, Germany or Remote. Empty keeps all.

## `remoteOnly` (type: `boolean`):

Keep only jobs the company marks as remote.

## `postedWithinDays` (type: `integer`):

Keep jobs posted in the last N days. 0 keeps all.

## `onlyNewJobs` (type: `boolean`):

For scheduled runs: return only jobs that were not open the last time this ran. The list of seen jobs is kept in a key-value store called career-site-jobs-seen in your account.

## `includeDescription` (type: `boolean`):

Full description text. Workday, SmartRecruiters and Workable need one extra request per job for it, so turning it off makes big runs faster.

## `maxJobsPerCompany` (type: `integer`):

Stop after this many matching jobs per company.

## Actor input object example

```json
{
  "companies": [
    "https://jobs.ashbyhq.com/openai",
    "https://jobs.lever.co/palantir"
  ],
  "keywords": [
    "engineer"
  ],
  "remoteOnly": false,
  "postedWithinDays": 0,
  "onlyNewJobs": false,
  "includeDescription": true,
  "maxJobsPerCompany": 20
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "https://jobs.ashbyhq.com/openai",
        "https://jobs.lever.co/palantir"
    ],
    "keywords": [
        "engineer"
    ],
    "maxJobsPerCompany": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("dektionstudio/career-site-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "https://jobs.ashbyhq.com/openai",
        "https://jobs.lever.co/palantir",
    ],
    "keywords": ["engineer"],
    "maxJobsPerCompany": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("dektionstudio/career-site-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "https://jobs.ashbyhq.com/openai",
    "https://jobs.lever.co/palantir"
  ],
  "keywords": [
    "engineer"
  ],
  "maxJobsPerCompany": 20
}' |
apify call dektionstudio/career-site-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dektionstudio/career-site-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BgscxjJ0en86lCCUg/builds/YVkYbD7zObSxkJD09/openapi.json
