# Greenhouse Jobs Scraper (`scrapyx/greenhouse-jobs-scraper`) Actor

Jobs from any company's Greenhouse job board via the public Job Board API: title, location, departments, offices, publish and update dates, requisition id, apply URL and the full description as text and HTML. Give board tokens or boards.greenhouse.io URLs; filter by keyword, department, location.

- **URL**: https://apify.com/scrapyx/greenhouse-jobs-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Greenhouse Jobs Scraper

Every open job on any company's **Greenhouse** job board — Stripe, Airbnb,
GitLab and thousands of others — through Greenhouse's public Job Board API:
title, location, departments, offices, first-published and updated dates,
requisition id, the apply URL, custom metadata, and the full description as
text and HTML.

HTTP only, no login, no key, no browser. **One request per board** returns
the whole board.

### What it is for

- **Company watchlists** — a list of boards on a schedule, diff by `jobId`.
- **Multi-company search** — many boards, one keyword, one dataset.
- **Department / location slices** — filtered locally from the full board.

### Input

| field | what it does |
| --- | --- |
| `boards` | Board tokens (`stripe`) or URLs (`boards.greenhouse.io/stripe`, `job-boards.greenhouse.io/<token>`, embed `?for=<token>`). |
| `searchTerms`, `departments`, `locations` | Local filters on the full list (the API has no search). |
| `includeDescription` | On by default; same response, bigger rows. |
| `maxItems`, `maxConcurrency`, `minRequestInterval`, `proxyConfiguration` | Limits. |

Each board gets a `BOARD_SUMMARY` with its total, what the filters dropped,
and the board's own department and office vocabularies.

### Three things about this API worth knowing before you trust a run

#### 1. There is no pagination, and the page parameters pretend otherwise

`?page=2&per_page=10` answers with all 654 Stripe jobs, same as no
parameters. A client that walks pages gets the whole board again on every
"page". This Actor makes one request per board and applies `maxItems`
locally.

#### 2. The description is HTML that was escaped before it was put in the JSON

`"content": "&lt;h2&gt;&lt;strong&gt;Who we are..."` — decoding the JSON
leaves the entities in place. Stored raw, the "description" is a wall of
`&lt;`. The Actor unescapes once (`descriptionHtml`) and strips to text
(`description`).

#### 3. Without `content=true`, departments and offices are null

The lighter call does not just omit the description; it drops the two
fields a jobs dataset is usually filtered on. `content=true` is always sent;
`includeDescription: false` only drops the description fields from the rows.

### Other things measured

- An unknown board answers `404 {"error":"Job not found"}` — the message
  names a job, not a board. Reported as `board_not_found`.
- Timestamps arrive with offsets in two shapes (`-04:00` and `-0400`);
  `firstPublished` / `updatedAt` are normalised to UTC, raw copies kept.
- `requisitionId` can be a placeholder string ("See Opening ID").
- No throttling or anti-bot layer; responses up to ~5 MB per board.

### Output

- **`JOB`** — `jobId`, `internalJobId`, `requisitionId`, `title`, `url`,
  `companyName`, `location`, `departments`, `offices`, `officeLocations`,
  `firstPublished`, `updatedAt`, `applicationDeadline`, `metadata`,
  `description`, `descriptionHtml`, `board`, `resultPosition`.
- **`BOARD_SUMMARY`** — `board`, `boardUrl`, `totalJobsOnBoard`,
  `jobsReturned`, `stoppedReason`, `filteredOut`, `departmentsOnBoard`,
  `officesOnBoard`.
- **`ERROR`** — `invalid_input`, `board_not_found`, `payload_shape_changed`,
  `fetch_failed`, with detail.

### Known limits

- Salary is not a Job Board API field; it appears only inside descriptions
  or custom `metadata` when the company adds it.
- Filters are substring matches on the board's own words (`departments`
  vocabulary is in the summary).
- Boards that require a login (internal boards) are out of scope.

# Actor input Schema

## `boards` (type: `array`):

Greenhouse board tokens or URLs: 'stripe', 'airbnb', 'gitlab', https://boards.greenhouse.io/stripe, https://job-boards.greenhouse.io/<token>, or an embed URL with ?for=<token>. One request per board returns the whole board. An unknown token is reported as board\_not\_found.

## `searchTerms` (type: `array`):

Keep jobs whose title, department or description contains any of these (case-insensitive). Applied locally - the API has no search. Empty = all jobs.

## `departments` (type: `array`):

Keep jobs whose department name contains any of these. departmentsOnBoard in the summary row lists the vocabulary.

## `locations` (type: `array`):

Keep jobs whose location, office or office location contains any of these ('remote', 'dublin', 'united states').

## `includeDescription` (type: `boolean`):

On by default: the job's full description as text (description) and HTML (descriptionHtml). Costs nothing extra - it comes in the same response - but makes rows large.

## `maxItems` (type: `integer`):

Overall cap on JOB rows across every board. A board arrives whole in one response (Stripe: 654 jobs); the cap is applied locally.

## `maxConcurrency` (type: `integer`):

Boards fetched in parallel. Responses can be 5 MB each.

## `minRequestInterval` (type: `integer`):

Politeness delay between request starts. One request per board; rarely needed.

## `proxyConfiguration` (type: `object`):

Optional. A public API published for embedding; no anti-bot layer. Enable Apify's free datacenter proxy only if a cloud run reports fetch\_failed.

## Actor input object example

```json
{
  "boards": [
    "stripe"
  ],
  "includeDescription": true,
  "maxItems": 1000,
  "maxConcurrency": 3,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "boards": [
        "stripe"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/greenhouse-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "boards": ["stripe"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/greenhouse-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "boards": [
    "stripe"
  ]
}' |
apify call scrapyx/greenhouse-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/greenhouse-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/tEnfYk68wz9jjlldm/builds/oRpX9PDJUQjQtqLT5/openapi.json
