# Career Page Jobs Scraper: Auto-Detects the Job Board (ATS) (`jtpalms/career-page-jobs`) Actor

Give company websites or careers pages and get every open job. It finds which job board each company uses (Greenhouse, Lever, Ashby, Workable, Recruitee, Personio, Breezy, Teamtailor, Gem, Pinpoint) and returns one row per job, with an only-new-jobs mode. USD 2 per 1,000 jobs.

- **URL**: https://apify.com/jtpalms/career-page-jobs.md
- **Developed by:** [JT Palms](https://apify.com/jtpalms) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.60 / 1,000 job listings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Career Page Jobs Scraper: open jobs from any company website

Give it company websites, careers pages or plain domains (`stripe.com`, `https://www.notion.com/careers`) and get back every open job. It finds which job board (ATS) each company uses, then reads that board's public feed and returns one clean row per job: title, department, locations, remote flag, employment type, pay range when published, dates, job and apply links and the description as plain text. Each row also says which job board was found (`ats`) and how (`detectedFrom`).

Supported job boards: **Greenhouse, Lever, Ashby, Workable, Recruitee, Personio, Breezy HR, Teamtailor, Gem and Pinpoint.**

**USD 2 per 1,000 jobs.** Companies whose job board cannot be found are reported free, with a list of what was checked.

### What people use it for

- **Hiring signals for your account list.** Your CRM has company domains, not job board names. Paste the domains, get every open role, and filter by department or keyword to spot companies that are growing a team you sell to.
- **Lead enrichment.** The `SUMMARY` record tells you, per company, which job board it uses and how many jobs are open. That is useful for segmenting accounts or for selling to HR teams.
- **Job boards and newsletters.** Turn a list of companies (climate startups, remote-first companies, your portfolio) into a job feed without looking up each company's job board by hand.
- **Job alerts.** Schedule it with **Only new jobs** turned on and receive only postings that appeared since the last run.

### Sample output

A real job row (description trimmed):

```json
{
  "source": "ashby",
  "ats": "Ashby",
  "detectedFrom": "https://notion.com/careers: links to jobs.ashbyhq.com/notion",
  "company": "notion",
  "companyName": null,
  "jobId": "06e1d3a8-2706-4ed9-9fe6-8b69607de18e",
  "title": "Security Engineer, Detection and Response",
  "department": "Security",
  "team": "Security",
  "location": "San Francisco, California",
  "locations": ["San Francisco, California", "New York, New York"],
  "workplaceType": "hybrid",
  "isRemote": false,
  "employmentType": "full_time",
  "compensation": null,
  "postedAt": "2026-09-28T16:20:26.330Z",
  "updatedAt": null,
  "jobUrl": "https://jobs.ashbyhq.com/notion/06e1d3a8-2706-4ed9-9fe6-8b69607de18e",
  "applyUrl": "https://jobs.ashbyhq.com/notion/06e1d3a8-2706-4ed9-9fe6-8b69607de18e/application",
  "descriptionText": "WHO WE ARE\n\nNotion is the collaborative AI workspace where teams and agents think together ...",
  "scrapedAt": "2026-09-29T03:20:19.921Z"
}
```

A real row for a company that could not be read (free; `error` and `checked` trimmed):

```json
{
  "source": "career-page",
  "company": "avidbots.com",
  "input": "avidbots.com",
  "error": "No supported job board found on avidbots.com. The company uses BambooHR, which is not supported: BambooHR does not allow it (its terms of service forbid robots and other automated means that monitor or copy content). Checked 9 place(s); see \"checked\". ...",
  "checked": [
    "https://avidbots.com/careers links to avidbots.bamboohr.com (BambooHR, not supported)",
    "https://avidbots.com/ (HTTP 200, no job board found)",
    "https://avidbots.com/jobs (HTTP 404)",
    "https://careers.avidbots.com/ (ENOTFOUND)"
  ],
  "scrapedAt": "2026-09-29T03:20:19.921Z"
}
```

Real runs: `notion.com` was found on Ashby (130 jobs), `stripe.com` on Greenhouse (704 jobs), `https://theblueground.com/careers` on Workable, `blenderbox.com` on Breezy HR, `ottonova.de` on Personio and `bunq.com` on Recruitee, in 1 to 11 seconds per company.

### How it finds the job board

For each company, fastest first. Every board it finds is confirmed by loading it before it is used.

1. **The input is already a job board URL** (`https://jobs.lever.co/spotify`, `https://blenderbox.breezy.hr`): used as is.
2. **The company's pages.** It loads the page you gave (or the home page, `/careers` and `/jobs` for a bare domain) and looks for links, iframes, embed scripts and API calls of a supported job board, and for redirects to one.
3. **Careers links.** It follows careers-looking links on the same site ("Careers", "Join us", "Jobs") and tries `careers.` and `jobs.` subdomains. A site that redirects to a new domain is followed there.
4. **The company name as a board name** (optional, on by default). It tries the name from the domain on every supported job board, and keeps a board only when it clearly belongs to the company: the same company name, job links on the company's domain, or job descriptions that name it.

`detectedFrom` tells you which step found the board and the evidence, for example `https://notion.com/careers: links to jobs.ashbyhq.com/notion` or `board name guess "stripe" on Greenhouse (company name "Stripe" matches)`.

### How to use it

1. Add companies to **Companies**, one per line. Any of these work:
   - a domain: `stripe.com`
   - a careers page: `https://www.notion.com/careers`
   - a job board URL: `https://jobs.lever.co/spotify`
   - a company name: `Coursera` (found by the name guess only)
2. Optionally narrow the results:
   - **Title keywords**: `engineer`, `AI`, `/product (manager|owner)/`
   - **Locations**: `London`, `New York`, `Remote`
   - **Remote jobs only**
   - **Departments**: `Engineering`, `Sales`
   - **Posted within (days)**: `7`
   - **Max jobs per company**: newest first
3. Click **Start**, then export as CSV, JSON or Excel, or read the dataset through the Apify API.

#### Job alerts on a schedule

1. Turn on **Only new jobs** and set your filters.
2. Save the input as a task and add a schedule (daily or weekly).
3. Connect the results to email, Slack, Google Sheets or a webhook in the task's **Integrations** tab.

### Input

| Field | What it does |
|---|---|
| `companies` | Domains, careers page URLs, job board URLs or company names. Required. |
| `keywords` | Keep jobs whose title matches any keyword. Words match at their start (`engineer` finds "Engineering"); keywords of 3 letters or fewer must be a whole word. `/regex/` is supported. |
| `locations` | Keep jobs whose location contains any of these texts. |
| `remoteOnly` | Keep only jobs that can be done remotely. |
| `departments` | Keep jobs whose department (or team, function or parent department, where the job board has them) contains any of these texts. |
| `postedWithinDays` | Keep jobs first published in the last N days. `0` = any date. |
| `maxJobsPerCompany` | Newest jobs first, at most this many per company. `0` = no limit. |
| `onlyNewJobs` | Output only jobs not delivered by an earlier run (see FAQ). |
| `includeDescription` | Add the description as plain text, up to 5,000 characters. Default on. |
| `guessBoardNames` | Step 4 above. Turn it off to use only boards the company's own pages point to. Default on. |
| `maxConcurrency` | Companies processed in parallel. Default 5. |

### Output fields

Job rows have the same fields as the single job board scrapers (Greenhouse, Lever, Ashby, Workable, Recruitee, Personio and Breezy HR); Teamtailor, Gem and Pinpoint rows use the same format, plus:

| Field | Notes |
|---|---|
| `source`, `ats` | The job board, as a key (`greenhouse`) and as a name (`Greenhouse`). |
| `detectedFrom` | How the board was found (see above). |
| `company`, `companyName` | The board name on that job board, and the company's display name when the board has one. |
| `jobId`, `title`, `department`, `team` | As the job board has them. |
| `location`, `locations`, `workplaceType`, `isRemote` | Location as shown, every location, and `remote`, `hybrid`, `onsite` or null. |
| `employmentType` | `full_time`, `part_time`, `contract`, `temporary`, `internship` or null. |
| `compensation` | `min`, `max`, `currency`, `interval`, `summary` when the company publishes pay. |
| `postedAt`, `updatedAt` | ISO 8601 in UTC. |
| `jobUrl`, `applyUrl`, `descriptionText` | Job page, application link, plain-text description (up to 5,000 characters). |
| `error`, `checked` | Only on rows for companies that could not be read: why, and every page and board that was checked. These rows are free. |

A `SUMMARY` record in the run's key-value store lists every company with its job board, how it was found, its number of open jobs, matches, new jobs and errors.

### Pricing

| What | Price |
|---|---|
| Job listing in the results | USD 0.002 (USD 2 per 1,000) |
| Company whose job board cannot be found, or an invalid input | Free |

Example: 200 company domains with 15 open jobs each is 3,000 jobs, USD 6. Detection itself is not charged. If you already know a company's job board, the single job board scrapers cost USD 1.50 per 1,000 jobs.

Set a maximum cost per run in the run options and the actor stops cleanly when it gets there; everything saved so far stays in the dataset.

### Limits

- **Supported job boards:** Greenhouse, Lever, Ashby, Workable, Recruitee, Personio, Breezy HR, Teamtailor, Gem and Pinpoint.
- **Not supported on purpose:** SmartRecruiters and BambooHR. Their terms forbid automated access, so the actor never contacts them. A company that uses one gets a free error row that says so.
- Other systems (Workday, iCIMS, Taleo, SuccessFactors, Teamtailor and in-house job pages) are not read. Neither are sites that load their jobs with JavaScript from their own servers without linking to a supported job board.
- At most 9 pages are loaded per company, each up to 3 MB.
- A board found by the name guess can, rarely, belong to a different company with the same name. `detectedFrom` shows the evidence; turn off `guessBoardNames` if you need only boards the company's own pages point to.

### Related actors

- [Career Site Jobs Index: Greenhouse, Ashby, Lever & More](https://apify.com/JTPalms/ats-jobs-index): Fresh jobs from thousands of company career sites, indexed twice a day from the official public job-board APIs of Greenhouse, Ashby,...
- [Hiring Signals: Companies Hiring by Role, Team & Location](https://apify.com/JTPalms/hiring-signals): Find companies hiring right now, one row per company: open roles by team, seniority and location, roles posted in the last 7 and 30...
- [Company Profile From Domain: Tech Stack, Socials & Hiring](https://apify.com/JTPalms/company-profile): Turn a list of website domains into company profiles: name, description, official social links, tech stack (CMS, ecommerce, analytics),...
- [Greenhouse Jobs Scraper: Any Company's Job Board API](https://apify.com/JTPalms/greenhouse-jobs): Get every open job from any company's Greenhouse job board through the official public API: title, department, location, remote flag,...

### FAQ

**What does it read?** The public pages you give it (a few per company) and then only the public job feeds of the supported job boards: the Greenhouse, Lever and Ashby job board APIs, Workable's jobs widget feed, Recruitee's Careers Site API, Personio's XML feed and Breezy HR's careers site feed. It collects no personal data and never logs in or submits anything. If you republish jobs, link to `jobUrl` or `applyUrl` so candidates apply with the company.

**Can it be pointed at internal addresses?** No. Before any request, and again at every redirect, it refuses loopback, private, link-local and other internal addresses (for example `localhost`, `10.x.x.x`, `192.168.x.x` or `169.254.169.254`), including names that resolve to them.

**Why was my company not found?** Open the error row's `checked` list: it shows every page and board that was tried. The usual reasons are an unsupported job board, or a careers page that loads jobs with JavaScript. If you know the company's job board URL, put that in **Companies** instead.

**How does Only new jobs work?** The actor keeps, per company and job board, the IDs of the jobs it has already delivered, in a key-value store named `career-page-jobs-monitor` in your own Apify account. Each run outputs only matching jobs whose ID is not in that list, then adds them. A company that moves to another job board starts a fresh history. Jobs left out only by **Max jobs per company** are marked as seen; jobs cut off by your cost limit stay new for the next run. To start over, delete the store in **Storage**.

**Something missing or wrong?** Open an issue with the company and what you expected.

# Actor input Schema

## `companies` (type: `array`):

Company websites or domains (stripe.com), careers page URLs (https://www.notion.com/careers), job board URLs (https://jobs.lever.co/spotify) or company names (Coursera). The actor finds the job board by itself. One per line.

## `keywords` (type: `array`):

Keep jobs whose title matches any of these. A keyword matches at the start of a word, so engineer also finds Engineering. Keywords of 3 letters or fewer (AI, QA, SRE) must be a whole word. Wrap in slashes for a regular expression, for example /product (manager|owner)/.

## `locations` (type: `array`):

Keep jobs whose location contains any of these texts, for example London, New York or Remote.

## `remoteOnly` (type: `boolean`):

Keep only jobs that can be done remotely (remote workplace type, or a location or title that says remote).

## `departments` (type: `array`):

Keep jobs whose department (or team, function or parent department, where the job board has them) contains any of these texts, for example Engineering or Sales.

## `postedWithinDays` (type: `integer`):

Keep jobs first published in the last N days. 0 means any date.

## `maxJobsPerCompany` (type: `integer`):

Newest jobs first. 0 means no limit. Caps the cost on very large boards.

## `onlyNewJobs` (type: `boolean`):

Remember which jobs were already delivered for each company (in the key-value store career-page-jobs-monitor in your account) and output only jobs that are new since the last run. The first run outputs all matching jobs. Turn on for scheduled job alerts.

## `includeDescription` (type: `boolean`):

Add the job description as plain text (HTML removed, up to 5,000 characters).

## `guessBoardNames` (type: `boolean`):

When the company pages do not reveal the job board, try the company name (from the domain) as a board name on every supported platform. Such a board is used only when it clearly belongs to the company: same company name, job links on its domain, or job descriptions that name it.

## `maxConcurrency` (type: `integer`):

How many companies to process at the same time.

## Actor input object example

```json
{
  "companies": [
    "notion.com",
    "https://theblueground.com/careers"
  ],
  "remoteOnly": false,
  "postedWithinDays": 0,
  "maxJobsPerCompany": 50,
  "onlyNewJobs": false,
  "includeDescription": true,
  "guessBoardNames": true,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

All output rows in the default dataset (JSON, CSV, Excel via the format parameter).

## `summary` (type: `string`):

Counts and per-input status for this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "notion.com",
        "https://theblueground.com/careers"
    ],
    "maxJobsPerCompany": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("jtpalms/career-page-jobs").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companies": [
        "notion.com",
        "https://theblueground.com/careers",
    ],
    "maxJobsPerCompany": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("jtpalms/career-page-jobs").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "notion.com",
    "https://theblueground.com/careers"
  ],
  "maxJobsPerCompany": 50
}' |
apify call jtpalms/career-page-jobs --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jtpalms/career-page-jobs"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QgSbTl1ZM5Iyhh7X1/builds/ekwUUKjbUVUOSNtEE/openapi.json
