Get only the jobs a company posted since your last run. Reads the public job feeds of Greenhouse, Lever, Ashby, Workable, Recruitee, Workday, Phenom and Eightfold; SmartRecruiters is closed by its host's robots.txt (measured 2026-09-07). Pay per job row; status rows and quiet days are free.
0.1 — 2026-09-08 (an equal-length robots.txt tie goes to Allow — STANDARDS R13, the owner's ruling)
A robots.txt tie is read the way the standard writes it.src/robots.js combines every group naming this client (or *) and lets the longest matching rule win, as before; the change is the tie-break: when an Allow and a Disallow of the same length both match a path, the Allow wins — RFC 9309 §2.2.2's own rule, which is how Google's reference parser reads the same file. The stricter reading shipped the day before (R12, an equal-length tie closing the path) is superseded by R13, decided by the owner on 2026-09-08. Nothing else moves: a host whose only rule is Disallow: / has nothing to tie against and is still never asked — api.smartrecruiters.com stays closed, and this actor still reads no SmartRecruiters board.
Harness: 125 cases, 0 failures. robots-equal-length-allow-and-disallow-close-the-path became robots-equal-length-tie-goes-to-allow-R13 — the same measured viGlobal file (two User-agent: * groups, Allow: / under a Content-Signal line in the first, Disallow: / in the second) now expects the feed to be read, with one robots.txt read; the grammar case's two tie keys flip to open in both rule orders. Mutation-checked: give the tie back to Disallow and those four checks fail; robots-disallow-stops-the-probe-before-any-request and robots-allow-overrides-a-shorter-disallow stay green either way.
0.1 — 2026-09-07 (an equal-length robots.txt tie closes the path — STANDARDS R12)
A robots.txt that contradicts itself at equal specificity is no longer read as permission.src/robots.js still combines every group naming this client (or *) and lets the longest matching rule win, as RFC 9309 says; the one change is the tie: when an Allow and a Disallow of the same length both match the feed path, the path is now closed. The RFC's tie-break (the least restrictive rule wins, §2.2.2) is a SHOULD and the same section permits a crawler to be stricter; the house rule (R12) takes that option, because the measured shape of such a file — a CDN's managed User-agent: * / Allow: / header placed above the host's own User-agent: * / Disallow: / (12 of the 17 viglobalcloud.com tenants and portal.velaw.com, measured 2026-09-07 by the legal-jobs-monitor build; the other five tenants end their file with Disallow: /Admin/ and stay readable) — is the host opting out, and a header its CDN wrote does not overrule the line the host wrote. Nothing else in the reader changed: a longer Allow still beats a shorter Disallow, so every Eightfold site (Disallow: / + Allow: /api/pcsx) reads exactly as before, and no host in selftest.json or the Store prefill carries an equal-length tie, so their rows are unchanged. The README's robots paragraph, the FAQ and the code comments in src/robots.js and src/fetchers.js now state the tie rule as shipped.
README figures measured, not local: the intro's "150 rows in 18 seconds in a local run" is now the platform figure (run vBhYaKgqE2Mwgcx0z, 2026-09-07, 150 rows in 14.8 s); the FAQ no longer lists SmartRecruiters among the feeds read (its host closes it — the same paragraph already said so). The previous entry now records its two platform run ids.
Harness: 125 cases (was 124). New: robots-equal-length-allow-and-disallow-close-the-path models the measured viGlobal file (two User-agent: * groups, Allow: / in the first under a Content-Signal line, Disallow: / in the second) — 0 feed requests and the refusal text naming the path and host; the grammar case's tie expectation flips to closed, in both rule orders. Mutation-checked: give the tie back to Allow and the new case fails on both checks (a feed request is made, no refusal) and the grammar case fails on both tie keys; robots-allow-overrides-a-shorter-disallow stays green either way.
Platform runs on this build (0.1.41, 2026-09-07): selftest.json run pG4O3htdMlrJt6RKZ — SUCCEEDED, 30 rows (5 each from Greenhouse, Lever, Recruitee, Workday, Phenom, Eightfold), 0 status rows, 18.2 s; the Store prefill run dZ7Vwx0Qy2igalvrk — SUCCEEDED, 150 rows (25 each from Greenhouse, Lever, Ashby, Workday, Phenom, Eightfold), 0 status rows, 14.6 s. Both unchanged from 0.1.40, as expected: no host in either input carries an equal-length tie. The build that records these ids is the re-push that followed them (0.1.42), same source otherwise.
The Store listing no longer offers SmartRecruiters as readable. The actor description, the seoDescription, package.json, the input schema's intro and the README's intro, first step and paging note now list the eight platforms whose feeds are read (Greenhouse, Lever, Ashby, Workable, Recruitee, Workday, Phenom, Eightfold) and say that SmartRecruiters is closed by api.smartrecruiters.com's own robots.txt (User-agent: * / Disallow: /, only LinkedInBot allowed /v1/companies/; measured 2026-09-07 and unchanged when re-read for this build). The code was already refusing it and saying so in the free row; the listing had not caught up. Nothing about how a SmartRecruiters company is handled changed.
Wording, not behaviour:api.lever.co's Crawl-delay: 1 is met by the fixed one-second per-host spacing every host gets; the Crawl-delay directive itself is not parsed, and the README, the previous entry below and the code comment now say exactly that instead of "honoured".
Harness: 124 cases (was 122). Two model a robots.txt that redirects to another host, which is the shape of the self-test's own Recruitee board (bunq.recruitee.com answers 302 to careers.bunq.com/robots.txt, Allow: /, measured 2026-09-07): the target's rules are what apply, the feed request still goes to the origin host, and a Disallow on the target closes the origin with a refusal that names the origin host and the feed path. Mutation-checked: a client that stops following redirects fails both.
Platform runs on this build (0.1.40, 2026-09-07): selftest.json run 37EaK1vGsr3Q9MnG6 — SUCCEEDED, 30 rows in 13.6 s; the Store prefill run vBhYaKgqE2Mwgcx0z — SUCCEEDED, 150 rows in 14.8 s.
robots.txt is now read and obeyed on every host this actor reads, not only on employers' own domains: the seven vendor API hosts (Greenhouse, Lever, Ashby, SmartRecruiters, Workable, Recruitee, Workday) get the same one-read-per-host-per-run check before their first feed request, with the same RFC 9309 matching, the same "4xx = no rules, unreadable = not asked this run" and the same free not-found row naming the rule. Measured 2026-09-07 with this actor's identity against the exact feed paths: six hosts permit their feeds (api.lever.co also asks Crawl-delay: 1, met by the fixed one-second spacing — Crawl-delay itself is not parsed); api.smartrecruiters.com does not (User-agent: * / Disallow: /, only LinkedInBot allowed /v1/companies/), so SmartRecruiters boards are refused with robots.txt disallows /v1/companies/<Company>/postings on api.smartrecruiters.com and no request is made until that host permits it. Bare-name detection skips the platform without a request and a name found nowhere says 1 of 6 feeds not permitted by the host's robots.txt instead of claiming an answer. No alternative path is tried.
About one request a second to every host, vendor API hosts included; Workday's per-posting description pages, which were fetched six at a time, are now one a second inside the same per-board time budget as Phenom and Eightfold, with descriptions past the budget counted in the free partial row. Hosts do not wait on each other. Measured locally: a 3-page Lever read at a 60 ms test spacing leaves ≥ 60 ms between requests with one in flight; live, a 220-request Workday board is about 220 s.
A request that got no answer is asked again (up to 3 attempts, 700 ms / 2.5 s back-off, 60 s budget per request), on every feed — it was one attempt everywhere except the robots.txt read and Eightfold's rate-limit path, so a first-request DNS/TLS blip on Greenhouse, Lever, Workday and the rest became a not-found row. An answered "not now" (408/425/429/500/502/503/504) is asked again on requests safe to repeat; an answered "no" (404, 400, 403) is never asked twice — the paging loop no longer re-sends a page the source answered. A host that never answers is reported as <host> did not answer after 3 attempt(s) (<code>: <error>), never as a bare "fetch failed"; a name that does not exist (ENOTFOUND) is an answer and is not retried.
A bot-verification challenge is named, never mistaken for a broken feed: every non-2xx or non-JSON answer is tested for Cloudflare's cf-mitigated: challenge, AWS WAF's action header and the shared interstitial marker list (the same list the daily audit and the factory's probe use; the harness fails if the two copies in this actor drift). The free row and RUN_SUMMARY (reason: walled, walled: true, sources_walled) say a challenge answered this client's network, and that no workaround was attempted.
The run's exit code now means something: it fails only when every company produced no board AND at least one was a genuine outage (a host that never answered, or an error answer). A run whose companies were all challenged, all refused by robots.txt or all on no platform exits 0 with every row explained. Before, no run could fail for source reasons — six dead boards were a green run with six not-found rows.
RUN_SUMMARY: each not-found entry carries reason (absent, walled, disallowed, unreachable, error) beside detail; top level adds sources_ok, sources_failed, sources_walled, sources_robots_disallowed, sources_unreachable, all_sources_failed. The free not-found note is worded per reason.
Quality audit: a feed the host's robots.txt refuses is reported DISALLOWED in its own list (never a FAIL, never a wall). Harness: 122 cases (was 73) — per vendor fetcher: disallow → no request + reason, 404 → read, unreachable → not asked, first-request blip → retried, challenge → named; plus the seven hosts' robots.txt as measured, host-never-answers wording, ENOTFOUND, "not now" vs answered "no", Workday's stateless POST, WAF 202 / cf-mitigated, pacing, Workday descriptions one at a time and budgeted, marker-list identity; end to end: SmartRecruiters refusal, walled company (both halves assert the exit code), all-outage run fails / mixed run does not, first-request blip during detection. Every fix mutation-checked.
robots.txt is read and obeyed on every employer-owned host before its feed is asked — Phenom and Eightfold sites live on the employer's domain, and until now a careers URL on an unrecognised host was probed (/widgets, /api/pcsx/search) without reading that host's rules. Now: one robots.txt request per host per run, RFC 9309 matching (the group naming DockhandCareerJobsMonitor outranks *; longest rule wins, Allow beats Disallow on a tie; * wildcard and $ anchor), a disallowed path is never requested and the company gets the free not-found row with robots.txt disallows <path> on <host> in RUN_SUMMARY; a robots.txt that cannot be read (network error, 5xx) is not permission and says so; a 404 is no rules. The same check runs before the first listing request on every Phenom and Eightfold board, including the description endpoints. Measured live: every Phenom tenant leaves /widgets open, every Eightfold tenant is Disallow: / + Allow: /api/pcsx — the honest answer on both is unchanged.
An Eightfold careers URL ending in a page name (…/jobs.html, …/Default.aspx) is no longer read as an explicit tenant domain; only a domain-shaped segment is, so the probe asks for domain=example.com instead of domain=jobs.html.
The prefilled company list and the self-test now include a Phenom board (careers.lilly.com, 620 postings) and a firm-owned Eightfold board (careers.micron.com, 2,772 postings), so Apify's daily health check exercises both new platforms: six boards, 150 rows in 12 s measured locally (cap 25).
README: Phenom's CVS read is "about 46 requests" (45–46 measured); Eightfold's postedTs is documented as date precision (epoch seconds at midnight UTC), not a time of day.
The robots.txt read gets one retry after a 2 s back-off when it got no answer (network error, timeout, 5xx), so a transient blip no longer costs a host its whole run; a 4xx is an answer and is not retried. A robots.txt group with an empty User-agent: is no longer mistaken for one naming this client. FAQ and code comments state the CVS read as "about 46 requests (45–46 across reads)" and the prefill as 10–12 s, as measured.
0.1 — 2026-09-07
Two more platforms: Phenom (employers' own career sites such as careers.lilly.com, jobs.cvshealth.com) and Eightfold (<company>.eightfold.ai or an employer's site fronting it, such as careers.micron.com). Both are reached by pasting any page of the careers site — a URL on an unrecognised host is now asked whether it is a Phenom or an Eightfold site — or by the explicit phenom: / eightfold: prefix. Same columns as every other platform, nulls where a feed publishes nothing (see the README's column table).
Phenom's hard 9,999-row window is read through by category (then US state) facet partitions, so jobs.cvshealth.com's 19,036 postings come back whole (19,013–19,028 measured, the rest seam churn from the feed's unstable ordering, de-duplicated by job id); a bucket no facet can split is reported as a source ceiling with both numbers.
Eightfold's server-fixed 10-row pages are walked newest first at one request a second, matching the feed's own count (jobs.ericsson.com 464 = 464); its rate limiter (429 / 401 Please try again later) is backed off and retried, and a limit that never clears is reported as a lost page, never absorbed.
A per-board time budget (boardTimeBudgetMinutes, default 20) bounds a read too slow to finish inside a run — Eightfold's 22,985-posting Starbucks board is ~38 minutes at that pace — and reports the stop as this actor's own, with the board's size, instead of a timed-out run. Descriptions on Phenom/Eightfold that do not fit the budget are counted in a free partial row, never left blank silently.
Requests to an employer's own domain are spaced at least one second apart.
0.1 — 2026-09-06
The audit also recognises a challenge that arrives as a 202 with an interstitial (AWS WAF's Challenge action) or part-way through a paged walk, and reports it BLOCKED rather than as a short row count.
The quality audit now tells a human-verification wall apart from a broken feed: a refusal that carries a challenge page is reported as BLOCKED for as long as it lasts, never as a failure, and the audit publishes the URL so the factory can ask Apify's own network the same question before anyone is troubled.
The actor now introduces itself on the wire as DockhandCareerJobsMonitor with a link to this listing (it was already an honest non-browser string; this is the house shape every Dockhand bot uses). No behaviour change; every feed re-measured.
Output schema added (.actor/output_schema.json): the run's Output tab and the API's output field now link straight to the rows as the overview table, JSON and CSV. Apify Store requires it before an Actor can be published. Nothing about the rows, the sources or the price changed.
0.1 — 2026-09-03
Reads the public job feeds of Greenhouse, Lever, Ashby, SmartRecruiters, Workable, Recruitee and Workday: bare company names for the first six, a careers-page URL for Workday.
Monitor mode: the first run delivers a company's current openings, every later run delivers only the postings that appeared since. Memory lives in a key-value store on your own account.
Title filters and a per-company cap, applied before anything is charged; pay per job row delivered, nothing per run.
A free status row for every company with nothing to show: not-found, no-open-jobs, filtered-out, no-new-jobs, charge-limit-reached, results-truncated, partial.
Every feed is read to its end. A board stopped short is reported with the numbers behind it, and RUN_SUMMARY states each board's own size, never the number a cap returned.