Greenhouse Jobs Scraper — Company Job Boards avatar

Greenhouse Jobs Scraper — Company Job Boards

Pricing

from $1.45 / 1,000 posting saveds

Go to Apify Store
Greenhouse Jobs Scraper — Company Job Boards

Greenhouse Jobs Scraper — Company Job Boards

Every open role on any company's Greenhouse job board as structured data: title, department, location with country, pay range where stated, posted and updated dates, full text on request. One request per board, one stable schema. Incremental mode charges only for what changed. No personal data.

Pricing

from $1.45 / 1,000 posting saveds

Rating

0.0

(0)

Developer

Adderley Data

Adderley Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 hours ago

Last modified

Share

What does Greenhouse Jobs Scraper do?

Greenhouse Jobs Scraper reads every open role on any company's public Greenhouse job board and returns it as structured data you can load straight into a spreadsheet, a database or a model. Thousands of companies run their careers page on Greenhouse; if the "Apply" button leads to boards.greenhouse.io or job-boards.greenhouse.io, this Actor reads it.

Give it a board token (acme) or any Greenhouse link for the company, and for each posting you get the title, department, location parsed to city, region and country where the posting states them, the pay range where the posting states one, the dates it was published and last updated, the application deadline if there is one, and a link to apply — and, if you ask for it, the full text of the posting.

What makes it different:

  • One request per board. The whole board, and every posting's full text with it, comes back in a single call to Greenhouse's public Job Board API. No HTML parsing, no page-by-page crawling, nothing to break when a careers page is redesigned. A thousand postings across fifty companies is fifty requests.
  • Pay ranges as numbers. Where a posting states a range — Greenhouse's pay-transparency block or a sentence in the text — it comes back as min, max, currency and period, with the original wording kept beside them. Where no range is stated, the fields are null; nothing is estimated.
  • Incremental mode. Put the Actor on a schedule and each run returns only postings that are new, changed, back again or gone. Greenhouse's own updated_at drives it, so an edit anywhere in a posting is caught. Unchanged postings are skipped and not charged.
  • One stable schema. Every row has every field, every time. Unknown is null, never a missing key, so nothing downstream breaks on a sparse posting. The schema is versioned (job.v1) and every Adderley Data jobs Actor uses it, so a Greenhouse board and a job board sit in the same table.
  • No personal data. There is no recruiter name, email or phone field, Greenhouse's custom metadata fields are never read, and contact details inside posting text are redacted by default.
  • Bounded cost. You set a maximum number of results; the run stops there. It also stops at the spending limit you set on the run in Apify.

What Greenhouse data can you extract?

FieldWhat it holds
idStable across runs: source:market:sourceJobId. Use it as your primary key.
titleJob title as listed.
company.nameThe hiring company, where the listing names one.
advertiser.nameThe business that placed the listing — often a recruitment agency. Never a person.
location.rawLocation text as listed. Several locations are joined with |.
location.suburbSuburb, when the listing states one.
location.cityCity or area, when the listing states one.
location.regionState or region, e.g. VIC.
location.postcodePostcode, when the source provides it.
location.countryISO 3166-1 alpha-2 country code.
workArrangementon_site, hybrid, remote or unknown.
employmentTypesNormalised: full_time, part_time, contract, casual, temporary, internship, volunteer.
salary.rawThe salary text exactly as shown, or null when the listing shows none.
salary.minLower bound as a number, when the text contains one.
salary.maxUpper bound as a number. Equal to min for a single figure.
salary.currencyISO 4217. Taken from the text, otherwise the market default.
salary.periodhour, day, week, month or year; null when the text does not say.
salary.includesSupertrue / false when the text says so ("plus super", "inc. super"); otherwise null.
classificationsThe source's category and subcategory pairs.
teaserThe short summary shown on the results page.
bulletPointsSelling points shown on the results page.
postedAtWhen the listing was posted, ISO 8601 UTC.
updatedAtThe source's own last-modified time, ISO 8601 UTC. Published by ATS and API sources; null where the site does not show one.
expiresAtExpiry, ISO 8601 UTC, where the source states one.
isPromotedtrue for paid placements. A listing shown both promoted and organic is returned once.
urlLink to the listing.
descriptionNull unless requested. text, optional sanitised html, and contactsRedacted.
changeTypeIncremental runs: NEW, UPDATED, REAPPEARED, EXPIRED (or UNCHANGED if you ask for those). Otherwise null.
firstSeenAtIncremental runs: when this monitor first saw the listing.
contentHashSHA-256 over the fields that define a change. Compare it to detect edits yourself.
scrapedAtWhen this row was produced, ISO 8601 UTC.
sourceSource key, e.g. seek.
marketMarket key, e.g. au, nz.
sourceJobIdThe source's own identifier for the listing.
company.sourceCompanyIdThe source's identifier for the company, when exposed.
company.urlThe company's page on the source site, when exposed.
advertiser.sourceAdvertiserIdThe source's identifier for the advertiser.
schemaVersionAlways job.v1. Breaking changes ship as job.v2 in a new Actor version, never silently.

market is always global for this source: a Greenhouse board is the company's, not a country's. The country comes from what the posting or its office states.

How much does it cost to scrape Greenhouse job boards?

You pay per posting saved to your dataset — $1.75 per 1,000 postings on Apify's Starter plan — plus $0.005 each time a run starts. There is no monthly rental.

Apify planPricePer listing
Free$1.75 per 1,000 listings$0.00175
Starter (Bronze)$1.75 per 1,000 listings$0.00175
Scale (Silver)$1.60 per 1,000 listings$0.00160
Business (Gold)$1.45 per 1,000 listings$0.00145

Plus $0.005 per run start. Compute and proxy are included in these prices.

What you runCost (USD, Starter plan)
100 listings, one run$0.18
1,000 listings, one run$1.75
10,000 listings, one run$17.50
50,000 listings, one run$87.50
A daily incremental monitor finding about 150 new or changed listings a day, for a month$8.03

Descriptions cost nothing extra here: they arrive in the same request as the listing. Use incremental mode for anything you run more than once — after the first run you pay only for what changed.

How to scrape a Greenhouse job board

  1. Find the company's board. On its careers page, the "Apply" link or the embedded job list points at boards.greenhouse.io/<token> or job-boards.greenhouse.io/<token>. The token is what you need; the whole link works too.
  2. Open the Actor in Apify Console and go to the Input tab. Paste one token or link per line into Job boards. Up to 500 boards per run.
  3. Optionally filter: Title keywords, Locations, Departments, Published within (days).
  4. Set Maximum results. This is also your cost cap.
  5. Press Start. When the run finishes, open the Output tab and export as JSON, CSV, Excel, XML or HTML, or read the dataset through the Apify API.

A posting that matches more than one board or filter is returned once.

Input

FieldTypeDefaultWhat it does
boardsarray—One entry per company: the board token (the first part of the path in https://boards.greenhouse.io/acme is "acme"), or any boards.greenhouse.io, job-boards.greenhouse.io or embedded job-board link. Up to 500 boards per run; each is one request.
keywordsarray—Keep postings whose title contains every word of any keyword, in any order — "engineer data" matches "Senior Data Engineer". Leave empty for all titles.
locationsarray—Keep postings whose location or office contains this text, e.g. "Melbourne", "Canada", "Remote". Leave empty for all locations.
departmentsarray—Keep postings in a department whose name contains this text, e.g. "Engineering". Leave empty for all departments.
postedWithinDaysinteger—Keep postings first published in the last N days. Leave empty for any time.
maxResultsinteger100The run stops once this many postings are saved. You are charged per posting saved, so this is also your cost cap.
includeDescriptionbooleanfalseOn: the full posting text comes back with each row, read from the same request as the listing, so it costs no extra requests. Off: listing fields only.
descriptionFormattext, text_and_html"text"Plain text, or plain text plus sanitised HTML.
redactContactsbooleantrueOn by default: email addresses and phone numbers inside description text are replaced with [redacted]. This Actor never outputs recruiter names or contact fields.
incrementalbooleanfalseRemember what earlier runs saw and save only postings that are new, changed or gone. Unchanged postings are skipped and not charged. Put the Actor on a schedule with this on.
stateKeystring—Optional name for this monitor, e.g. "competitor-engineering". Runs with the same key share memory. Left empty, a key is derived from the boards and filters themselves.
emitExpiredbooleantrueIncremental mode only. When a complete run no longer finds a posting it saw before, save one row with changeType EXPIRED.
emitUnchangedbooleanfalseIncremental mode only. Saves (and charges for) every posting, labelled UNCHANGED where nothing moved.
proxyConfigurationobject{"useApifyProxy":true}Apify Proxy, automatic group, is the default and is what this Actor is tested with.
maxConcurrencyinteger4Parallel requests. The default is deliberately modest.
maxRequestsPerMinuteinteger90An upper bound on request rate across the whole run.

A typical input:

{
"boards": [
"greenhouse"
],
"keywords": [
"engineer"
],
"maxResults": 100
}

Output

One row per posting. This is a synthetic example in the exact shape the Actor returns:

{
"schemaVersion": "job.v1",
"id": "greenhouse:global:7000001",
"source": "greenhouse",
"market": "global",
"sourceJobId": "7000001",
"url": "https://job-boards.greenhouse.io/example-freight/jobs/7000001",
"title": "Data Analyst",
"company": {
"name": "Example Freight Co",
"sourceCompanyId": "example-freight",
"url": "https://job-boards.greenhouse.io/example-freight"
},
"advertiser": {
"name": "Example Freight Co",
"sourceAdvertiserId": "example-freight"
},
"location": {
"raw": "Melbourne, Victoria, Australia",
"suburb": null,
"city": "Melbourne",
"region": "Victoria",
"postcode": null,
"country": "AU"
},
"workArrangement": "hybrid",
"employmentTypes": [],
"salary": {
"raw": "The salary range for this role is A$95,000 – A$110,000 per year plus super.",
"min": 95000,
"max": 110000,
"currency": "AUD",
"period": "year",
"includesSuper": false
},
"classifications": [
{
"category": "Data",
"subcategory": null
},
{
"category": "Analytics",
"subcategory": null
}
],
"teaser": null,
"bulletPoints": [],
"postedAt": "2026-09-20T22:14:05.000Z",
"updatedAt": "2026-09-21T03:02:11.000Z",
"expiresAt": null,
"isPromoted": false,
"description": null,
"changeType": "NEW",
"firstSeenAt": "2026-09-21T19:30:12.000Z",
"contentHash": "9a41c07e2b6d5f3e8c1a0b7d6e5f4a3b2c1d0e9f8a7b6c5d4e3f2a1b0c9d8e7f",
"scrapedAt": "2026-09-21T19:30:12.000Z"
}

The Output tab has two table views: Overview (the fields most people want, flattened) and Changes (for incremental runs).

Incremental mode: monitor new roles across companies

Turn on Incremental mode and run the same input on a schedule — hourly, daily, weekly. The Actor keeps a small record of what it has seen and every row tells you what happened:

changeTypeMeaning
NEWFirst time this monitor has seen the posting
UPDATEDSeen before, and the title, company, location, department, work arrangement, pay range, dates or Greenhouse's own updated_at has changed
REAPPEAREDWas reported as expired and is back
EXPIREDSeen before and no longer on the board. One row, once
UNCHANGEDOnly if you turn on Also save unchanged postings

How it behaves, so there are no surprises:

  • The first run returns everything as NEW. From the second run you pay only for the difference.
  • EXPIRED is only ever reported by a complete run. If a run hits your result cap or your spending limit, or a board cannot be read, nothing is declared expired — a posting on a board the run never read is not gone.
  • Runs share memory when they share a State key. Leave it empty and the key is derived from the boards and filters themselves, so the same input always continues the same monitor. Name it (competitor-engineering) if you want to change filters later without starting again.
  • A posting not seen for 45 days is forgotten.

Descriptions and contact details

Full descriptions are off by default. Turn on Include full descriptions and each row carries description.text (and sanitised description.html if you choose that format). Because Greenhouse returns the posting text in the same response as the listing, this costs no extra requests and no extra time.

Postings sometimes contain a recruiter's email address or phone number. With Redact contact details on — the default — those are replaced with [redacted] and description.contactsRedacted is true; so is a personal profile address (linkedin.com/in/…). Neither description.text nor the optional HTML carries the address behind a link: the HTML keeps each link's words and drops its address, and drops images. The Actor never returns recruiter names or contact details as fields, under any setting, and never reads Greenhouse's custom metadata fields, which can hold anything a company put there. If your use case is contacting individuals, this is the wrong tool.

What people use it for

  • Competitor and market hiring signals. Which companies are opening which roles, in which departments and locations, and how often — a daily monitor across a list of boards is one scheduled run.
  • Pay-transparency datasets. Stated ranges across companies, roles and locations, parsed to numbers with the original wording kept.
  • Job aggregators and alert products. A clean feed of new postings from a curated list of employers, deduplicated and labelled by change.
  • Sales and partnership research at company level. Growth signals from hiring, without collecting anything about the individuals involved.
  • Research and teaching. A clean, repeatable dataset with a documented schema.

Using the API

Run it from code with the Apify client, using your own API token:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('adderleydata/greenhouse-jobs-scraper').call({"boards":["greenhouse"],"keywords":["engineer"],"maxResults":100});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.length, items[0]?.classifications);
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("adderleydata/greenhouse-jobs-scraper").call(run_input={"boards":["greenhouse"],"keywords":["engineer"],"maxResults":100})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["title"], item["location"]["raw"], item["updatedAt"])

Schedules, webhooks and the Make, Zapier, n8n and Google Sheets integrations all work the way they do for any Apify Actor. The Actor runs with limited permissions and is priced per event, so AI agents can call it through Apify's MCP server as well.

The Actor reads Greenhouse's public Job Board API — the same feed the company's own careers page reads, published by Greenhouse for exactly this purpose, with no login, no key and no rate tricks — and returns facts about job postings. It does not log in, does not solve CAPTCHAs, and does not collect personal information.

What you do with the data is your responsibility. Each company's postings are its own copyright — analyse them, do not republish them. If your project touches personal information, privacy law applies to you wherever you are. This is general information, not legal advice.

Questions

Where do I find a company's board token? On its careers page, follow any job's "Apply" link: the address is boards.greenhouse.io/<token>/jobs/<id> or job-boards.greenhouse.io/<token>/jobs/<id>. Paste either the whole link or just the token. Company careers pages that embed Greenhouse use boards.greenhouse.io/embed/job_board?for=<token>; that link works too.

A board I gave came back as "not found". Greenhouse answers with 404 when the token does not exist, when the company has switched off its public board, or when the board is hosted on Greenhouse's EU region and not served by the main API — that last case is reported, not silently emptied. Check the token on the company's careers page. A missing board is asked for once, not retried, and a run in which every board is missing finishes with an empty dataset and a message naming each one.

Does it need a Greenhouse login or API key? No. The Job Board API is public.

Why is employmentTypes often empty? Greenhouse's public feed has no employment-type field. The Actor fills it only when the title says so ("Intern", "Contract", "Part-time"); it never guesses.

Can I get recruiter emails or phone numbers? No, by design.

How current is the data? It is read from Greenhouse while your run is in progress. updatedAt is Greenhouse's own last-modified time for the posting; scrapedAt records when the row was produced.

The field I need is not there. Open an issue on the Issues tab. Fields are added to the schema without breaking existing ones.

Support

Use the Issues tab on this page. We read it every day. Include the run ID and what you expected to see.

Other Adderley Data Actors

Every Actor in a vertical returns the same fields, so adding a source needs no new code on your side.

About

Made by Adderley Data, Melbourne — https://adderleydata.com. Not affiliated with, endorsed by or sponsored by Greenhouse Software, Inc. Greenhouse is a trade mark of its owner and is used here only to describe what this Actor reads.