Greenhouse Jobs Scraper
Pricing
from $0.84 / 1,000 results
Greenhouse Jobs Scraper
Jobs from any company's Greenhouse job board via the public Job Board API: title, location, departments, offices, publish and update dates, requisition id, apply URL and the full description as text and HTML. Give board tokens or boards.greenhouse.io URLs; filter by keyword, department, location.
Pricing
from $0.84 / 1,000 results
Rating
0.0
(0)
Developer
Ibnu Adzim
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Every open job on any company's Greenhouse job board — Stripe, Airbnb, GitLab and thousands of others — through Greenhouse's public Job Board API: title, location, departments, offices, first-published and updated dates, requisition id, the apply URL, custom metadata, and the full description as text and HTML.
HTTP only, no login, no key, no browser. One request per board returns the whole board.
What it is for
- Company watchlists — a list of boards on a schedule, diff by
jobId. - Multi-company search — many boards, one keyword, one dataset.
- Department / location slices — filtered locally from the full board.
Input
| field | what it does |
|---|---|
boards | Board tokens (stripe) or URLs (boards.greenhouse.io/stripe, job-boards.greenhouse.io/<token>, embed ?for=<token>). |
searchTerms, departments, locations | Local filters on the full list (the API has no search). |
includeDescription | On by default; same response, bigger rows. |
maxItems, maxConcurrency, minRequestInterval, proxyConfiguration | Limits. |
Each board gets a BOARD_SUMMARY with its total, what the filters dropped,
and the board's own department and office vocabularies.
Three things about this API worth knowing before you trust a run
1. There is no pagination, and the page parameters pretend otherwise
?page=2&per_page=10 answers with all 654 Stripe jobs, same as no
parameters. A client that walks pages gets the whole board again on every
"page". This Actor makes one request per board and applies maxItems
locally.
2. The description is HTML that was escaped before it was put in the JSON
"content": "<h2><strong>Who we are..." — decoding the JSON
leaves the entities in place. Stored raw, the "description" is a wall of
<. The Actor unescapes once (descriptionHtml) and strips to text
(description).
3. Without content=true, departments and offices are null
The lighter call does not just omit the description; it drops the two
fields a jobs dataset is usually filtered on. content=true is always sent;
includeDescription: false only drops the description fields from the rows.
Other things measured
- An unknown board answers
404 {"error":"Job not found"}— the message names a job, not a board. Reported asboard_not_found. - Timestamps arrive with offsets in two shapes (
-04:00and-0400);firstPublished/updatedAtare normalised to UTC, raw copies kept. requisitionIdcan be a placeholder string ("See Opening ID").- No throttling or anti-bot layer; responses up to ~5 MB per board.
Output
JOB—jobId,internalJobId,requisitionId,title,url,companyName,location,departments,offices,officeLocations,firstPublished,updatedAt,applicationDeadline,metadata,description,descriptionHtml,board,resultPosition.BOARD_SUMMARY—board,boardUrl,totalJobsOnBoard,jobsReturned,stoppedReason,filteredOut,departmentsOnBoard,officesOnBoard.ERROR—invalid_input,board_not_found,payload_shape_changed,fetch_failed, with detail.
Known limits
- Salary is not a Job Board API field; it appears only inside descriptions
or custom
metadatawhen the company adds it. - Filters are substring matches on the board's own words (
departmentsvocabulary is in the summary). - Boards that require a login (internal boards) are out of scope.