# Career Site Job Scraper – Live Jobs from Company Career Pages (`datataskflow/career-site-job-scraper`) Actor

Get live job postings straight from company career sites on Greenhouse, Lever, Ashby, SmartRecruiters and Recruitee. Filter by keyword, location, remote and date. Optional salary ranges. Clean JSON, fair pay per job.

- **URL**: https://apify.com/datataskflow/career-site-job-scraper.md
- **Developed by:** [DataTaskFlow](https://apify.com/datataskflow) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.90 / 1,000 job postings

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Career Site Job Scraper – Live Jobs from Company Career Pages

Get **live job postings straight from company career sites** as clean JSON, CSV or Excel. Point the Actor at a company and it reads the official public job board the company uses: **Greenhouse, Lever, Ashby, SmartRecruiters and Recruitee**. Every job is fetched at run time, so you never get stale or already-closed roles.

You only pay for jobs that were returned with data. Failed or empty companies are never charged.

### What you get

| Feature | Details |
|---|---|
| 🏢 Live from the source | Jobs come directly from the employer's own board, with a direct apply link |
| 🔎 Auto-detection | Paste a board URL, `greenhouse:stripe`, a plain company slug, or even the company's own career page – the Actor works out which system it uses |
| 🎯 Filters | Title keywords, excluded keywords, locations, remote only, posted within N days, search inside descriptions |
| 💰 Salary (add-on) | Structured pay ranges when published, plus ranges detected in the job text (min, max, currency, interval) |
| 🧼 Clean output | One flat item per job, same field names for every source, ISO dates, plain-text description (HTML optional) |
| 🛡️ Reliability | Automatic retries, clear error rows instead of crashes, a `RUN_STATS` summary of every run |

### Use cases

- **Job boards and aggregators:** feed fresh, direct-apply jobs from the employers you choose.
- **Recruiters and sales teams:** see which companies are hiring for which roles, and how fast they grow.
- **Market research:** track hiring by department, location and pay across competitors.
- **Job alerts and AI agents:** schedule a daily run with `postedWithinDays` and push new roles to Slack, Sheets, n8n or Make.

### How to use it

1. Add companies, one per line (see formats below).
2. Optionally set keywords, locations, remote only or a date limit.
3. Click **Start** and download results in JSON, CSV, Excel or through the API.

Accepted company formats:

- Board URL: `https://boards.greenhouse.io/stripe`, `https://jobs.lever.co/palantir`, `https://jobs.ashbyhq.com/ramp`, `https://jobs.smartrecruiters.com/Visa`, `https://acme.recruitee.com`
- Explicit: `greenhouse:stripe`, `lever:palantir`, `ashby:ramp`, `smartrecruiters:Visa`, `recruitee:acme`
- Company slug: `stripe` (all supported systems are tried)
- A company career page URL (the Actor looks for the embedded job board)

#### Input example

```json
{
  "companies": ["greenhouse:stripe", "https://jobs.lever.co/palantir", "ashby:ramp"],
  "keywords": ["engineer", "data"],
  "excludeKeywords": ["intern"],
  "locations": ["Remote", "United States"],
  "postedWithinDays": 14,
  "includeSalary": true,
  "maxJobsPerCompany": 50,
  "maxItems": 200
}
```

#### Output example (one item per job)

```json
{
  "ats": "greenhouse",
  "companySlug": "stripe",
  "company": "Stripe",
  "jobId": "8172510",
  "title": "Abuse Investigator",
  "url": "https://stripe.com/jobs/search?gh_jid=8172510",
  "applyUrl": "https://stripe.com/jobs/search?gh_jid=8172510",
  "location": "Seattle, San Francisco, New York City",
  "locations": ["US"],
  "country": null,
  "remote": null,
  "workplaceType": null,
  "employmentType": null,
  "department": "Security Analytics",
  "team": null,
  "postedAt": "2026-09-09T14:50:29Z",
  "updatedAt": "2026-09-25T20:45:00Z",
  "salaryMin": 120000,
  "salaryMax": 160000,
  "salaryCurrency": "USD",
  "salaryInterval": "year",
  "salarySource": "description",
  "descriptionText": "About the team ...",
  "scrapedAt": "2026-10-03T08:00:00Z"
}
```

Fields a source system does not publish are `null`. Salary fields appear only when **Include salary** is on. If a company cannot be read, you get a row with the `input` and an `error` instead.

### Run summary

Every run saves a `RUN_STATS` record (Storage → Key-value store): companies with jobs, without jobs and failed (with reasons), jobs found, filtered out, returned, and jobs per source system.

### Pricing

Pay per result, no start fee and no platform usage fees:

| Event | When it is charged |
|---|---|
| Job posting | Once per job returned with data |
| Salary data (add-on) | Per job where a salary range was found, only if you turned **Include salary** on |

Jobs that are filtered out, companies without a board and failed items are **never charged**. Use *Max jobs in total* to cap the cost of a run. Exact prices are on the Pricing tab and get lower on higher Apify plans.

### FAQ

**Which systems are supported?** Greenhouse, Lever, Ashby, SmartRecruiters and Recruitee. Companies on other systems (for example Workday) are not covered; you get a clear error row for them.

**Can I search all companies by keyword?** No. This Actor reads the companies you list, so results are always live and complete for those employers. It does not search a pre-built database.

**Is the pay shown exact?** Pay is only returned when the employer publishes it, either as structured data or as a range in the job text. Ranges detected from text are marked with `salarySource: "description"`.

**Can I schedule it?** Yes. Use Apify Schedules, set `postedWithinDays` to 1 and connect the output to Slack, Sheets, n8n, Make or your own API.

**Is this legal?** The Actor uses only the public job-board endpoints that employers publish for job seekers and job boards, with no login and no personal data. You are responsible for using the data in line with applicable laws and the terms of the career sites.

### Support

Found a bug or need another job board supported? Open an issue on the **Issues** tab – fixes usually ship within a few days.

# Actor input Schema

## `companies` (type: `array`):

One entry per line. Accepts a career-board URL (boards.greenhouse.io/stripe, jobs.lever.co/palantir, jobs.ashbyhq.com/ramp, jobs.smartrecruiters.com/Visa, acme.recruitee.com), a company's own career page URL (the Actor looks for the job board it uses), 'ats:company' such as greenhouse:stripe, or just a company name slug (all supported systems are tried). Up to 500 per run.

## `keywords` (type: `array`):

Only return jobs whose title contains at least one of these words (case-insensitive). Leave empty for all jobs.

## `excludeKeywords` (type: `array`):

Skip jobs whose title contains any of these words, for example intern or director.

## `searchInDescription` (type: `boolean`):

Match keywords and exclude keywords in the full job description as well as the title.

## `locations` (type: `array`):

Only return jobs whose location contains one of these texts, for example Berlin, Germany, United States. Leave empty for all locations.

## `remoteOnly` (type: `boolean`):

Only return jobs flagged as remote (or with 'remote' in the location).

## `postedWithinDays` (type: `integer`):

Only return jobs published in the last N days. 0 = no limit. Jobs without a publish date are skipped when this is set.

## `includeSalary` (type: `boolean`):

Adds salaryMin, salaryMax, salaryCurrency and salaryInterval. Uses the structured pay range when the company publishes one, otherwise detects ranges written in the job text (for example $120,000 - $150,000). Charged as an add-on only for jobs where a salary was found.

## `includeDescription` (type: `boolean`):

Include the job description as clean plain text.

## `includeDescriptionHtml` (type: `boolean`):

Also include the original HTML description.

## `maxJobsPerCompany` (type: `integer`):

Newest jobs first. Limits how many jobs are returned for each company.

## `maxItems` (type: `integer`):

Hard limit of jobs returned in one run (and therefore of cost).

## `maxConcurrency` (type: `integer`):

How many companies are fetched at the same time.

## Actor input object example

```json
{
  "companies": [
    "greenhouse:stripe",
    "jobs.lever.co/palantir",
    "ramp"
  ],
  "keywords": [],
  "excludeKeywords": [],
  "searchInDescription": false,
  "locations": [],
  "remoteOnly": false,
  "postedWithinDays": 0,
  "includeSalary": false,
  "includeDescription": true,
  "includeDescriptionHtml": false,
  "maxJobsPerCompany": 50,
  "maxItems": 100,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `runStats` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "greenhouse:stripe",
        "jobs.lever.co/palantir",
        "ramp"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("datataskflow/career-site-job-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "companies": [
        "greenhouse:stripe",
        "jobs.lever.co/palantir",
        "ramp",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("datataskflow/career-site-job-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "greenhouse:stripe",
    "jobs.lever.co/palantir",
    "ramp"
  ]
}' |
apify call datataskflow/career-site-job-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datataskflow/career-site-job-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8YgZTUxygfPAQLZ53/builds/4geNQrnAD8ldeoPyw/openapi.json
