Greenhouse Jobs Scraper avatar

Greenhouse Jobs Scraper

Pricing

from $0.84 / 1,000 results

Go to Apify Store
Greenhouse Jobs Scraper

Greenhouse Jobs Scraper

Jobs from any company's Greenhouse job board via the public Job Board API: title, location, departments, offices, publish and update dates, requisition id, apply URL and the full description as text and HTML. Give board tokens or boards.greenhouse.io URLs; filter by keyword, department, location.

Pricing

from $0.84 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Every open job on any company's Greenhouse job board — Stripe, Airbnb, GitLab and thousands of others — through Greenhouse's public Job Board API: title, location, departments, offices, first-published and updated dates, requisition id, the apply URL, custom metadata, and the full description as text and HTML.

HTTP only, no login, no key, no browser. One request per board returns the whole board.

What it is for

  • Company watchlists — a list of boards on a schedule, diff by jobId.
  • Multi-company search — many boards, one keyword, one dataset.
  • Department / location slices — filtered locally from the full board.

Input

fieldwhat it does
boardsBoard tokens (stripe) or URLs (boards.greenhouse.io/stripe, job-boards.greenhouse.io/<token>, embed ?for=<token>).
searchTerms, departments, locationsLocal filters on the full list (the API has no search).
includeDescriptionOn by default; same response, bigger rows.
maxItems, maxConcurrency, minRequestInterval, proxyConfigurationLimits.

Each board gets a BOARD_SUMMARY with its total, what the filters dropped, and the board's own department and office vocabularies.

Three things about this API worth knowing before you trust a run

1. There is no pagination, and the page parameters pretend otherwise

?page=2&per_page=10 answers with all 654 Stripe jobs, same as no parameters. A client that walks pages gets the whole board again on every "page". This Actor makes one request per board and applies maxItems locally.

2. The description is HTML that was escaped before it was put in the JSON

"content": "&lt;h2&gt;&lt;strong&gt;Who we are..." — decoding the JSON leaves the entities in place. Stored raw, the "description" is a wall of &lt;. The Actor unescapes once (descriptionHtml) and strips to text (description).

3. Without content=true, departments and offices are null

The lighter call does not just omit the description; it drops the two fields a jobs dataset is usually filtered on. content=true is always sent; includeDescription: false only drops the description fields from the rows.

Other things measured

  • An unknown board answers 404 {"error":"Job not found"} — the message names a job, not a board. Reported as board_not_found.
  • Timestamps arrive with offsets in two shapes (-04:00 and -0400); firstPublished / updatedAt are normalised to UTC, raw copies kept.
  • requisitionId can be a placeholder string ("See Opening ID").
  • No throttling or anti-bot layer; responses up to ~5 MB per board.

Output

  • JOB — jobId, internalJobId, requisitionId, title, url, companyName, location, departments, offices, officeLocations, firstPublished, updatedAt, applicationDeadline, metadata, description, descriptionHtml, board, resultPosition.
  • BOARD_SUMMARY — board, boardUrl, totalJobsOnBoard, jobsReturned, stoppedReason, filteredOut, departmentsOnBoard, officesOnBoard.
  • ERROR — invalid_input, board_not_found, payload_shape_changed, fetch_failed, with detail.

Known limits

  • Salary is not a Job Board API field; it appears only inside descriptions or custom metadata when the company adds it.
  • Filters are substring matches on the board's own words (departments vocabulary is in the summary).
  • Boards that require a login (internal boards) are out of scope.