Turn a list of companies into sales and recruiting signals: new job postings matching your skills or tools (Rust, Kafka, Snowflake), hiring surges, and technologies added to or dropped from their website. Reads Greenhouse, Lever, Ashby, Recruitee, Personio and Teamtailor. Pay only for signals.
Versions follow MAJOR.MINOR.PATCH (src/version.py); Apify shows MAJOR.MINOR from .actor/actor.json.
Every run logs its version and records it in the RUN_STATS key-value record.
About future failures: this actor reads each company's public job-board feed and home page. Job-board vendors
change their feeds without notice. If a company that used to work starts failing, suspect a format change first:
the run log names the company and says what went wrong, and one broken company never affects the others.
1.0.1 (2026-09-26)
Re-check each home page every (days) (stackCheckEveryDays, default 7): with monitoring on, a home page
checked less than that many days ago isn't fetched again; its last stack is reported with techStackCheckedAt, and
a change shows up at the next check. The home-page check is most of a company's cost, and stacks change slowly, so
this keeps daily monitors cheap. Set 1 to check every run.
New field techStackCheckedAt.
1.0.0 (2026-09-26)
First release (deployed for its confirming run, not published).
Input: a list of companies (a name, a careers page or a job-board URL, as in Company Career Page Jobs Scraper),
each optionally followed by | <website domain>; signal keywords (skills, tools or role words, matched as whole
words in the title and description, or the title only); optional departments and locations; "only signals since my
last run" (monitoring); thresholds for hiring surges; an optional list of technologies that count as stack signals.
One result per company with a signal, with signalTypes (new_roles_matching, hiring_surge, stack_added,
stack_removed), a one-line summary, the new matching roles (title, link, location, department, posting date,
which keywords matched and where), open and matching role counts now and at the last run, the home page's tech
stack and what it added or dropped. Companies with nothing new are listed in the run's OUTPUT record and are
not results, so they're never charged.
Monitoring: one memory per search (keywords, where they're matched, departments, locations) and per saved task,
in the hiring-signals-memory store in the user's account. A company's entry is only updated once its result is in
the dataset (or it was quiet), so a signal cut by the run's limits comes back next run. Guards against false
signals: a failed board isn't evaluated; a role that disappears is remembered for 30 days; a board that suddenly
lists nothing or loses half its roles is held until the next run confirms it; a home page that fails, shows a bot
check or suddenly shows nothing never reports technologies removed.
Up to 500 companies per run. Charged per company with a signal through Apify's standard
apify-default-dataset-item event; Max results per run and the maximum cost per run stop the run cleanly (the
companies not checked are listed in OUTPUT.notChecked).
Shares its code with Company Career Page Jobs Scraper (mms_ats) and Website Technology Detector (mms_stack),
so a fix to a job board or a fingerprint reaches all of them.
Sources and terms (checked before building; re-read 2026-09-26)
Same use as Company Career Page Jobs Scraper, whose terms audit (2026-09-24) cleared these six platforms: each
company's public job feed, fetched on the user's behalf for the companies the user lists; no index is built or
resold. Re-read for this actor:
Greenhouse — docs (docs.greenhouse.io/job-board.html): "Job Board data is publicly available, so authentication
is not required for any GET endpoints." greenhouse.com/legal has no terms for website visitors or feed readers;
the only automated-access clause found (my.greenhouse.io user agreement) binds job-seeker accounts, not the feed.
Lever — docs (github.com/lever/postings-api): "all job postings in the published state are publicly viewable.
These jobs may be scraped by third parties." Terms (lever.co/legal/terms-of-service): no scraping or
automated-access clause; access limits name only Lever's direct competitors.
Ashby — docs (developers.ashbyhq.com/docs/public-job-posting-api): "This API allows you to get data for all
currently published Job Postings for your organization", no key. Terms (ashbyhq.com/resources/terms) bind
customers only; nothing forbids reading the public feed.
Recruitee — docs (docs.recruitee.com/reference/intro-to-careers-site-api): "The Recruitee Careers Site API
allows to view company's jobs ... from candidate perspective." Terms (recruitee.com/terms) bind the Subscriber;
"An End-User that is not the Subscriber does not derive any rights from these Terms"; no scraping clause.
Personio — docs (developer.personio.de/docs/retrieving-open-job-positions): "Current open job postings can be
retrieved in XML format under myaccount.jobs.personio.de/xml". Terms: cleared in the 2026-09-24 audit; the re-read
on 2026-09-26 got HTTP 429 from personio.com (rate limit), so that audit stands. Re-check when the site answers.
Teamtailor — docs (support.teamtailor.com RSS guide): the .rss feed serves "metadata that can be shared
elsewhere". Terms (teamtailor.com/en/terms-and-conditions) bind the customer and its users, not feed readers.
Company home pages — one page per company per run, as Website Technology Detector reads them (cleared for it):
a normal logged-out visit, robots.txt followed, no protection bypassed, cookie values never output. The domain is
one the user typed or the company itself publishes (careers page, job links, career-site host); it is never
guessed from a name.
All requests go through mms_common (honest User-Agent HumbleEchidnaApify, robots.txt per RFC 9309, private-network
guard, ports 80/443 only). The platforms refused by the ats-jobs audits (Workday, Workable, SmartRecruiters, Breezy,
Rippling and 18 more) are refused here with the same reasons.