# Greenhouse Jobs Scraper: Open Jobs from Any Board (`tinlark/greenhouse-jobs-scraper`) Actor

Open jobs from any company's Greenhouse job board (boards.greenhouse.io): title, department, location, pay range when published, posting date, link, optional full description. Filters, hiring signals and new-only monitoring. No login.

- **URL**: https://apify.com/tinlark/greenhouse-jobs-scraper.md
- **Developed by:** [Tinlark](https://apify.com/tinlark) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Greenhouse Jobs Scraper: open jobs from any company's Greenhouse board

Give it the name or address of a company's Greenhouse job board and get one row per open job: title, department, location, pay range when the employer publishes one, posting date and a link. Filter by keyword, department, location and age, add full descriptions, get one hiring summary per company, or schedule it to return only jobs that are new since the last run.

Many companies host their job list on Greenhouse (`boards.greenhouse.io/<name>`). Greenhouse documents a public Job Board API for exactly this list. The Actor reads it, one request per board, and returns uniform records. It needs no login, no browser and no proxy.

For Workday, Lever, Ashby and other vendors in the same run, see the multi-vendor [ATS Jobs Scraper by Tinlark](https://apify.com/tinlark/ats-jobs-hiring-signals) in the section "More than Greenhouse?" below.

### Who it is for

Recruiters and sourcers who track competitors' openings, sales and RevOps teams who read hiring as a buying signal, job-board and newsletter builders, and analysts of hiring trends.

### How to find a company's Greenhouse board name

The board name (Greenhouse calls it the board token) is the last part of the board address. It is usually the company's name in lowercase.

1. Open the company's careers page and click an open job.
2. If the address looks like `https://boards.greenhouse.io/stripe/jobs/1234567` or `https://job-boards.greenhouse.io/figma/jobs/1234567`, the board name is `stripe` or `figma`.
3. Some companies show jobs on their own domain and hide Greenhouse: the address may be `careers.example.com/positions/1234567?gh_jid=1234567`. The `gh_jid` in the address is the sign that Greenhouse sits behind the page. View the page source and search for `greenhouse.io`: an embed such as `boards.greenhouse.io/embed/job_board?for=example` gives the board name `example`. Paste that address or the name.
4. Not sure? Try the name. A wrong name gives a free error row that says so.

#### What you can paste

| You paste | Result |
|---|---|
| `stripe` | board `stripe` (a bare name is looked up on Greenhouse) |
| `greenhouse:stripe` | the same, explicit |
| `https://boards.greenhouse.io/stripe` | board `stripe` |
| `https://job-boards.greenhouse.io/figma` | board `figma` |
| `https://boards.greenhouse.io/embed/job_board?for=cloudflare` | board `cloudflare` |
| `https://boards.greenhouse.io/stripe/jobs/1234567` | the whole board `stripe`, not just that job |
| `figma.com` | the Actor looks for a Greenhouse link on the company's home and careers pages, where robots.txt allows (best effort) |

Domain lookup is hit and miss: among seven well-known company domains in Tinlark's tests, it found the board for two (Figma and Anthropic). It cannot see boards that a page loads with JavaScript. When it fails, use the board name.

### What Greenhouse gives, and what it does not

| Field | What it holds |
|---|---|
| `company` | The name set on the board by the employer |
| `jobId` | Greenhouse's job id, the number in the job address |
| `title` | The job title as posted |
| `department` | The first department of the job. Greenhouse allows several; only the first is returned |
| `locations` | The location text of the job, or its offices when no location text is set |
| `isRemote` | True when the location or title says remote. Greenhouse has no remote flag |
| `postedAt` | When the job was first published (UTC) |
| `updatedAt` | When the posting was last changed |
| `salaryMin`, `salaryMax`, `salaryCurrency`, `salaryRanges` | The pay range, only when the employer publishes one on the posting. `salaryRanges` keeps all ranges, with their labels, when there are several |
| `jobUrl`, `applyUrl` | The job's address. For companies that show Greenhouse jobs on their own site, this is that page (with `gh_jid`) |
| `descriptionText` | Only with `includeDescription`: the posting text as plain text |
| `seniority` | A guess from the title: `senior`, `manager`, `intern` and so on |

Not available from Greenhouse's public board: team (stays empty), employment type, a workplace-type field, applicant counts, recruiter names and anything behind an applicant login. `salaryInterval` is empty when the pay label does not name one.

Pay ranges are common on some boards and rare on others. In a cloud run over ten large technology boards (3,572 jobs), 1,804 jobs carried a pay range. Do not expect that share on other boards.

### Input

| Field | What it does |
|---|---|
| `companies` (required) | Board names, `greenhouse:name`, board URLs or company domains |
| `mode` | `jobs` (default), `signals` (one summary row per company) or `new-only` |
| `keywords`, `excludeKeywords` | Matched against title, department and team, not case sensitive |
| `departments` | Text the department must contain, for example `Sales` or `Engineering` |
| `locations` | Text that a location must contain, for example `Berlin` or `United States` |
| `remoteOnly` | Keep jobs whose location or title says remote |
| `postedWithinDays` | Keep jobs first published within N days |
| `includeDescription` | Add the plain-text description to each row (much larger output) |
| `maxJobsPerCompany` | Newest first, default 25 |
| `maxItems` | Stop after this many billable rows, default 1,000 |
| `stateName` | A name such as `sales-watch`. Required for `new-only`. Remembers which jobs were seen |
| `emitExistingOnFirstRun`, `includeClosed` | Options for runs with a `stateName` |

#### Example: sales jobs from three boards, last 14 days

```json
{
  "companies": ["stripe", "airbnb", "https://boards.greenhouse.io/cloudflare"],
  "departments": ["Sales"],
  "postedWithinDays": 14,
  "maxJobsPerCompany": 100
}
```

#### Example: pay ranges only

Pay ranges are part of the output, not a filter. Run with `maxJobsPerCompany` set high and keep the rows where `salaryMin` is not empty, for example with a spreadsheet filter on the exported CSV. You do not need `includeDescription` for this.

#### Example: a daily watch

```json
{
  "companies": ["figma", "databricks", "anthropic"],
  "mode": "new-only",
  "stateName": "ai-watch",
  "keywords": ["engineer", "research"]
}
```

The first run records the open jobs and returns none. Later runs return only jobs that are new.

### Output

Results go to the default dataset (JSON, CSV, Excel, XML or the API). The `SUMMARY` record of the run's key-value store holds counts, errors and the number of requests.

#### Job row (real output from a cloud run)

```json
{
  "recordType": "job",
  "atsVendor": "greenhouse",
  "company": "Airbnb",
  "boardToken": "airbnb",
  "jobId": "8247303",
  "title": "Software Engineer, Passport & Commerce, Android",
  "department": "Software Engineering",
  "locations": ["Remote, USA"],
  "isRemote": true,
  "postedAt": "2026-10-01T21:35:44Z",
  "updatedAt": "2026-10-01T21:35:44Z",
  "salaryMin": 162000,
  "salaryMax": 186000,
  "salaryCurrency": "USD",
  "jobUrl": "https://careers.airbnb.com/positions/8247303?gh_jid=8247303",
  "sourceUrl": "https://boards-api.greenhouse.io/v1/boards/airbnb/jobs",
  "status": "open"
}
```

#### Hiring summary (signals mode, trimmed)

```json
{
  "recordType": "signal",
  "company": "Cloudflare",
  "boardToken": "cloudflare",
  "openRoles": 405,
  "rolesByDepartment": { "Solution Engineering": 69, "Engineering": 56, "Field Sales": 53, "Mid Market": 31 },
  "seniorRoles": 293,
  "remoteRoles": 3,
  "newLast7d": 36,
  "newLast30d": 161,
  "hiringSpike": false
}
```

### Limits and honest notes

- **Public boards only.** Boards whose employer switched the public API off answer 404 and give an error row.
- **No cap per board.** Greenhouse returns the whole board, with posting texts, in one request; the largest board in Tinlark's tests had 885 jobs, and the peak memory of a ten-board run was 137 MB.
- **The department is the first one** if a job has several, and team is empty.
- **Remote is a text match.** A job in "Remote, USA" is remote; a job whose location says only a city is not, even if the employer allows remote work.
- **`jobUrl` may not be on greenhouse.io.** Some employers rewrite job addresses to their own career site.
- **Closed jobs** are reported only against your own earlier run with the same `stateName`. A board that suddenly returns no jobs after returning five or more is treated as an outage.
- **Pace.** At most 4 requests a second to Greenhouse's API, one request per board, retries with backoff on HTTP 429, 5xx and network errors, a declared `TinlarkBot` User-Agent.

### Pricing

**Free during launch (until 31 October 2026).** You pay only Apify's own platform usage for your runs.

Measured platform usage: a cloud run over ten large boards returned 3,572 jobs in 6 seconds with a peak of 137 MB of memory and cost $0.018 in platform usage, about $0.005 per 1,000 jobs. The three-board prefill returns 75 jobs in about 5 seconds. The default memory is 512 MB.

From 1 November 2026: pay per event, $1.50 per 1,000 job rows ($0.0015 each on the Apify Free and Bronze plans; Silver $1.30, Gold $1.10 per 1,000) in `jobs` and `new-only` modes, $5.00 per 1,000 company summary rows in `signals` mode and $3.00 per 1,000 boards found from a company domain (not charged when none is found, and not for names or URLs). Error rows, closed-job rows and the baseline run of a `new-only` watch are free. For example, 50 boards with 40 jobs each is 2,000 jobs, or $3.00. Set *Maximum cost per run* to cap spending once prices are on.

### More than Greenhouse?

This Actor is the Greenhouse-only version of [ATS Jobs Scraper by Tinlark](https://apify.com/tinlark/ats-jobs-hiring-signals), which reads Greenhouse, Workday, Lever, Ashby, Workable, Recruitee and Personio in one run and tries a bare name on all of them. Both run the same code and return the same fields. A Workday-only version is [Workday Jobs Scraper](https://apify.com/tinlark/workday-jobs-scraper).

### Use it from code or an AI agent

Start the Actor through the Apify API, any Apify client library or Apify's MCP server, and read the dataset items. The input and output schemas are defined, so tools can describe the fields themselves.

```bash
curl -X POST "https://api.apify.com/v2/acts/tinlark~greenhouse-jobs-scraper/run-sync-get-dataset-items" \
  -H "Authorization: Bearer <YOUR_APIFY_TOKEN>" -H "Content-Type: application/json" \
  -d '{"companies": ["stripe"], "departments": ["Sales"], "maxJobsPerCompany": 50}'
```

### Data source, terms and your responsibility

The data comes from the Greenhouse Job Board API (`boards-api.greenhouse.io/v1/boards/<name>/jobs`). Greenhouse documents it as public: job board data can be read with GET requests and no authentication, so that companies and third parties can list postings. The Actor identifies itself as `TinlarkBot`, keeps to a gentle request rate and does not log in or read applicant data. The postings are the employers' own advertisements and carry no personal data about applicants. They belong to the employers: check their terms and the laws that apply to you before you reuse or republish them. Each row carries `jobUrl` and `sourceUrl` so you can link back.

This Actor is not affiliated with Greenhouse Software, Inc. or with any employer named above.

### FAQ

**Where do I find a board name?** See the first section. In `boards.greenhouse.io/<name>` or `job-boards.greenhouse.io/<name>`, the name is the first part of the path.

**Can I get salaries?** Only where the employer publishes a pay range on the posting. Those rows have `salaryMin` and `salaryMax`. Greenhouse has no salary field for the other jobs.

**Why did a company return an error row?** Typical reasons: the board name is spelled differently from the company, the company uses another vendor, or the employer turned its public board off. The row says which request failed.

**Can I run it for hundreds of boards?** Yes. Each board is one request and the Actor sends at most four per second, so 300 boards take a little over a minute. The limit is 5,000 entries per run.

**Is it legal to use this data? Disclaimers and legality.** The Actor reads only what employers publish on their own public job boards, through the API that Greenhouse provides for that purpose, at a slow rate. It does not log in or get around blocks. The postings belong to the employers. You are responsible for how you use them, including the laws on unsolicited contact and on republishing content that apply to you. Tinlark is not affiliated with Greenhouse or any employer, and this is not legal advice.

### Related Tinlark Actors

- [ATS Jobs Scraper](https://apify.com/tinlark/ats-jobs-hiring-signals): all seven vendors in one run.
- [Workday Jobs Scraper](https://apify.com/tinlark/workday-jobs-scraper): the same fields for Workday career sites.
- [Remote Jobs API](https://apify.com/tinlark/remote-jobs-feed): remote jobs from four job boards.

### Support

Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the board name or URL) and what you expected.

# Actor input Schema

## `companies` (type: `array`):

One entry per company. Accepted: the board token (stripe, the last part of boards.greenhouse.io/stripe), greenhouse:stripe, a board URL (https://boards.greenhouse.io/stripe or https://job-boards.greenhouse.io/stripe) or a company domain (figma.com). For a domain the Actor reads the company's home and careers pages, where robots.txt allows, to find a linked Greenhouse board.

## `mode` (type: `string`):

jobs: one row per open job. signals: one summary row per company (open roles by department, new and closed roles, hiring spike flag). new-only: only jobs not seen in earlier runs (needs a State name).

## `keywords` (type: `array`):

Keep jobs whose title, department or team contains any of these words (not case sensitive).

## `excludeKeywords` (type: `array`):

Drop jobs whose title, department or team contains any of these words.

## `departments` (type: `array`):

Keep jobs whose department or team contains any of these texts, for example Sales or Engineering.

## `locations` (type: `array`):

Keep jobs with a location containing any of these texts, for example Berlin, Germany or United States.

## `remoteOnly` (type: `boolean`):

Keep only jobs whose location or title says remote.

## `postedWithinDays` (type: `integer`):

Keep jobs first published within this many days. Leave empty for no limit. Workday's list only says "Posted N days ago" up to 30 days, so jobs older than that have no date there.

## `includeDescription` (type: `boolean`):

Add the full job description as plain text. Makes the output much larger.

## `maxJobsPerCompany` (type: `integer`):

Newest jobs first. Applies to jobs and new-only modes.

## `maxItems` (type: `integer`):

Stops the run once this many billable rows (jobs or company signals) were produced.

## `stateName` (type: `string`):

The name under which the Actor remembers what it has already seen, kept in a named storage of your account. Lowercase letters, digits and hyphens. Use the same name on every scheduled run and a different name for each watch list. Required for the new-only mode.

## `emitExistingOnFirstRun` (type: `boolean`):

By default the first new-only run for a company only records the current jobs as a baseline and returns nothing. Turn this on to receive all current jobs the first time.

## `includeClosed` (type: `boolean`):

With a State name, add a free row (status closed) for each job that disappeared since the last run.

## Actor input object example

```json
{
  "companies": [
    "stripe",
    "greenhouse:airbnb",
    "https://boards.greenhouse.io/cloudflare",
    "https://job-boards.greenhouse.io/figma",
    "figma.com"
  ],
  "mode": "jobs",
  "remoteOnly": false,
  "includeDescription": false,
  "maxJobsPerCompany": 25,
  "maxItems": 1000,
  "emitExistingOnFirstRun": false,
  "includeClosed": false
}
```

# Actor output Schema

## `jobs` (type: `string`):

Job rows, newest first within each company.

## `signals` (type: `string`):

One hiring summary per company (signals mode).

## `errors` (type: `string`):

Inputs that could not be resolved to a job board.

## `summary` (type: `string`):

Counts, errors, request totals.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companies": [
        "stripe",
        "airbnb",
        "cloudflare"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tinlark/greenhouse-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "companies": [
        "stripe",
        "airbnb",
        "cloudflare",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tinlark/greenhouse-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companies": [
    "stripe",
    "airbnb",
    "cloudflare"
  ]
}' |
apify call tinlark/greenhouse-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tinlark/greenhouse-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/AfjCBeZQvGg1WJW64/builds/EQVYbrwTEUh5o5w5k/openapi.json
