Attorney, lawyer, paralegal and solicitor openings from the law-firm portals and ATS platforms firms post on — viRecruit/viGlobal, cvMail, AllHires, WizeHire, iCIMS. 45 portals: 43 readable, 2 closed by their own robots.txt (2026-09-08). Paste any portal URL. Monitor mode returns only NEW roles.
0.1 — 2026-09-23 (a cvMail board already read is no longer thrown away to ask for it twice — STANDARDS F5)
A read that had already succeeded was being decided by a second answer to the same question. cvMail's landing page is page one. Measured 2026-09-23 across all eight cvMail tenants in the directory, the landing carries exactly what the page=jobBoard&rcd= hop carries: the same rows (Mishcon 46, Squire Patton Boggs 30, Ashurst 30, RPC 23, Bird & Bird 23, and the platform's own empty-board page on Travers Smith, Norton Rose Fulbright and K&L Gates), the same "Jump To List" total select, and the same next-page form with its own fresh rcd. src/sources.js asked for that hop unconditionally and then judged the whole read on the hop's answer alone, so one unrecognised page destroyed a board that had already been served. That is what incident 21 was: on 2026-09-23 at 11:39 the daily quality audit filed cvmail/rpc: cvMail: no vacancy board found for this slug as a data-quality failure, and a fresh run of that same case read the board's 23 rows with every field at its measured fill. The hop is now asked for when, and only when, the landing is not a board this reader can read — which is the platform's other landing shape, a splash page holding nothing but the "Vacancies" link.
A page the reader did not recognise is no longer reported as a fact about the buyer's slug.cvMail: no vacancy board found for this slug was a catch-all: it fired on any 200 page that was neither a vacancy table nor the empty-board page, named the slug for an event that says nothing about the slug, and threw the page itself away — so the incident could not be told apart from a genuinely dead board, and no one can now say what arrived. The failure names the URL, the status, the byte count and the page's own <title>, says it is neither of the two things this reader can read, and states plainly that nothing about the firm's vacancies was measured (F9/F10/F17). The three wordings that are diagnoses — a firm that has moved off the platform, the portal's own service-error page, and a board that is genuinely empty — are unchanged.
Rows are unchanged and the request count falls. Live re-read of all eight tenants: 46 / 71 / 141 / 23 / 23 / 0 / 0 / 0 rows, identical to before the change, with one request fewer each (Ashurst's 141 of its stated 141 over 5 requests, not 6). That is one less read per board against a host eight of the directory's boards share, at this actor's one-request-a-second pacing.
Live audit on this source: node src/qa.js — 0 failures, 0 blocked, 0 transient; cvmail/rpc 23 rows in 1 request, paging/cvmail-ashurst 141 of its stated 141 over 5 requests, both PASS.
Harness: 218 checks (was 215), 0 failures, mutation-checked (T4) — restore the unconditional hop and the catch-all wording, and exactly the three named F5 cases fail: the board read from a landing returns 0 rows instead of 128, the hop is asked for again, and the failure reproduces incident 21's exact string, cvMail: no vacancy board found for this slug. test/stub-http.mjs gains both live landing shapes and the unrecognised page, so the splash-landing cases keep proving the hop is still made where it is the only way through.
0.1 — 2026-09-15 (a gap the audit cannot see is stated, never worked around — the owner's ruling)
Documentation only; no code, row, or price change. The daily quality audit's host is answered HTTP 403 + a Cloudflare managed challenge by Gibson Dunn's portal (selfapplygibsondunn.viglobalcloud.com, on viGlobal's shared viglobalcloud.com zone), so the audit learns nothing about that portal's job list and reports the source as BLOCKED (F10/T8) — while the same URL answers runs from the Apify platform normally. The owner ruled 2026-09-15 (incident 17): keep the source and state the gap plainly. The README gains a Known gaps section naming the blocked source, what the audit can and cannot vouch for, and what a buyer sees instead: a platform run that is answered returns the openings as usual, and a challenged run returns a free error row naming the challenge and stating that no workaround was attempted. src/qa.js keeps reporting it as blocked every day the challenge persists; nothing about fetching changed.
0.1 — 2026-09-14 (a lost page is OUR unconfirmed read, never the source's ceiling — STANDARDS F14)
A read that lost a page or a job category is no longer recorded as a board that came back whole. The paging driver (src/paging.js, the one verdict function every fetcher in this actor shares) answered F8 alone: a lost unit set partial: true and returned truncated: false, truncation_reason: null, so the free partial row fired and nothing said the read had stopped short of the board — the same hole the sibling actor measured on 2026-09-12 (incident 16: one unanswered request of a 101-request walk, 1,580 of the 2,657 postings the feed states held, results_truncated: false). partial and truncated are two facts about one read (F14) and both are reported now: the read is truncated: true, truncation_reason: "unknown", worded as ours — the source refused nothing and the rest was never asked for (F4) — and never source-ceiling, however wide the gap the source states. Where the source stated a total and the walk reached it anyway, a lost unit that cost nothing stays partial only.
The free row says which of the two things that are ours it was.unknown used to have one sentence ("the feed repeated a page"); it is now reached two ways, so the fetch contract carries unconfirmed — "lost-request" or "wrap", the fetcher's own word, never guessed from a request or error count — and the results-truncated note is worded from it. A lost request says that a request part-way through the read failed, that the openings behind it were never read, that the source refused nothing, and that what went unread was not marked as seen so the next run returns it; the partial row beside it names the request and why. The note claims no retry, because a stateful POST the board answered is never re-sent (F9), and it holds for both shapes this actor has (a cursor walk stops at the page it lost; the viGlobal category sweep skips the category and carries on). The README status table, the input schema and the dataset schema describe both causes (P3/P4).
The daily quality audit's contract check asks the same question (P1).src/qa.js now fails a result that is partial, short of a stated total and not truncated; one that is "unknown" without naming what was unconfirmed; and one that names it on any other reason.
This is the tail of the Medic's incident-16 fix, found uncommitted in the tree: the code and the checks were in place, but the note's wording and the checks had not been aligned (the note said "nothing unread was marked as seen" where the check looked for "not marked as seen", and the cvMail lost-page check still asserted the pre-F14 !truncated). Aligned here, and no check was weakened: the cvMail check now asserts the F14 truth — truncated, unknown, the note worded as ours and never as a ceiling — where it used to assert the hole.
Harness: 210 checks (was 209), 0 failures, mutation-checked (T4) both ways — restore the pre-F14 "partial only" verdict, or leave unconfirmed unnamed so the note falls back to "repeated a page", and exactly the two named cases fail (P1 a lost category is ALSO a read that stopped short, and says so as ours, P2 and a lost page is never re-labelled as the source's own ceiling); src/paging.js restored byte-identical after each. src/robots.js stays byte-identical with the sibling actor's copy. The Medic's own daily audit ran this tree's src/qa.js at 10:57 today: PASS, 4 blocked — the standing viGlobal challenge of the audit host, unchanged (F10/T8). Build 0.1.44 is this source; its selftest.json run is recorded in the commit that ships it.
0.1 — 2026-09-11 (a non-answer is not a measurement, and our own timeout is called by its name — STANDARDS T9)
The daily quality audit no longer files one lost packet as broken data. On 2026-09-11 src/qa.js filed cvmail/rpc and wizehire/divorce-with-a-plan as data-quality failures because a robots.txt read timed out three times — on two hosts that were answering normally the whole run (the audit's own wall check re-asked fsr.cvmailuk.com a second later and was answered, which is the only way paging/cvmail-ashurst could then read its 134 rows from that same host in the same run, and it did; a re-run that afternoon measured 23 and 10 rows on the two "failed" cases, and 16 of 16 paced reads of the two robots.txt files answered in 128-646 ms and 250-1108 ms). A case whose failure carried no answer at all now asks once more, a moment later, before it is filed — and never for anything the source actually said: a Disallow is an answer, a 4xx is an answer ("no rules"), a 5xx is the host saying "not now", and a wall stays BLOCKED. A source still silent on the second ask fails with the message it failed with the first time, and every second ask is reported (answered_on_retry on the case, transient in the report and its count in summary) so a host that needs re-asking daily is a measurement of its own, never absorbed.
Our own timeout is named TimeoutError, never the number 23.AbortSignal.timeout rejects with a DOMException that sets no cause and whose legacy code is the DOM spec's NUMBER 23, so the e.cause?.code || e.code || e.name picker every client used printed robots.txt … could not be read (23: The operation was aborted due to timeout) — fetch failed one spelling later (F9), which is why the incident could not say whether the blip was DNS, TCP, TLS or our own clock. errorCode() in src/robots.js now takes a STRING code from the cause, else a string code, else the name — never a number — and every message in the actor goes through it. src/robots.js stays byte-identical with the sibling actor's copy, which carries the same fix.
Harness: 209 checks (was 200), 0 failures, mutation-checked (T4): restore the numeric picker and R28 reproduces the incident's exact string; remove the second ask, or let it re-ask something the source answered, and the named R29 cases fail. Live audit on this source: node src/qa.js — 16 of 16 cases PASS, 0 failures, 1,421 rows audited, transient empty (today's blip was over, exactly the class's shape); the 4 blocked entries are the standing viGlobal challenge, unchanged (F10/T8).
0.1 — 2026-09-10 (a robots.txt read that got no answer is no longer remembered as one — STANDARDS F13)
One lost packet closed eight law firms.fsr.cvmailuk.com hosts eight of the directory's boards (Mishcon de Reya, Squire Patton Boggs, Ashurst, RPC, Bird & Bird, Travers Smith, Norton Rose Fulbright, K&L Gates UK). On 2026-09-10 two GETs of that host's robots.txt failed at the transport layer, and src/robots.js cached the failure in its per-host map exactly as if it were a set of rules — so all eight boards were refused for the rest of the run without another request being made. The host was never down: re-measured the same day with this actor's identity, 40 paced reads of that file answered 40 of 40 in 126-646 ms, and all 38 hosts the directory reads answered a cold robots.txt GET in 360-1471 ms, 38 of 38.
Only an answer is cached now. Rules, or a 4xx meaning "no rules" (RFC 9309 §2.3.1.3), are a fact about the host and are kept for the run. A read that got no answer is returned to the path that asked and thrown away, so the next board on that host asks again. A genuinely dead host still cannot eat the run one path at a time: reads are bounded per host per run by count and by elapsed time, whichever trips first, after which the verdict does stick.
The robots.txt read gets the same retry as every other request (F9). It was the one request in the actor that did not: 2 attempts, no time budget, and a message reading fetch failed, twice with no underlying code — exactly the string F9 calls "not a diagnosis". It is now 3 attempts with a growing back-off inside a 60 s budget, a 10 s per-attempt timeout (≈7x the slowest cold read measured, holding one read to 38 s where the door's 45 s timeout allowed 92 s), and an error naming the host, the code and the attempt count.
Live audit on this build's source: node src/qa.js — 16 of 16 cases PASS, 0 failures, 1,435 rows audited, 38 robots hosts read, 0 unreadable. Both cvMail cases (cvmail/rpc, paging/cvmail-ashurst) are green. The 4 blocked entries are the standing viGlobal Cloudflare challenge of this audit host, unrelated and unchanged (F10/T8) — the same URLs answer 200 from the platform.
Harness: 200 checks, 0 failures, and the new behaviour is mutation-checked (T4) — cache the non-answer like an answer and the named R14 cases fail: the second board on the host returns 0 rows instead of 60, and the host's read count stops at one read's attempts. src/robots.js stays byte-identical with the sibling actor's copy, which carries the same fix.
0.1 — 2026-09-08 (an equal-length robots.txt tie goes to Allow — STANDARDS R13, the owner's ruling)
A robots.txt that contradicts itself at the same level is read again.src/robots.js (byte-identical with the sibling actor's, as it must be) combines every group naming this client or * and lets the longest matching rule win, as RFC 9309 says; the change is the tie-break. When an Allow and a Disallow of equal length both match, the Allow wins — §2.2.2's own rule, the least restrictive one, and how Google's reference parser reads the same bytes. Yesterday's build took the stricter reading the RFC also permits (R12) and closed those paths; the owner ruled on 2026-09-08 that the standard's own tie-break applies (R13), so a CDN's managed Allow: / above a tenant's Disallow: / leaves the path open. A host whose only rule is a bare Disallow: /, with no Allow anywhere, plainly says no and stays closed. A longer Allow still beats a shorter Disallow.
Re-measured live 2026-09-08 with this actor's own identity, one robots.txt read on each of the 38 hosts behind the 45 directory portals: 33 hosts permit the portal path, 3 publish no robots.txt (a 404 is "no rules"), 2 close it. The two closures are the bare Disallow: / hosts refused since 0.1.35 — recruiting.dorsey.com (Dorsey & Whitney) and recruiting.lowenstein.com (Lowenstein Sandler). Thirteen hosts carry the tie and are read again: twelve viglobalcloud.com tenants (Gibson Dunn, Dechert, Troutman Pepper, Fisher Phillips, Akin Gump, Simpson Thacher, ArentFox Schiff, Nutter, Bracewell, Stoel Rives, Jones Day, Davies Ward) and Vinson & Elkins' portal.velaw.com. The five shared-zone tenants that end their file with Disallow: /Admin/ (Steptoe, Patterson Belknap, Baker Botts, Paul Hastings, Blakes) were never affected. So 43 of the 45 directory portals (39 of 41 firms) are readable today and 2 are refused with a free robots-txt row — up from 30 and 26 yesterday. firms.jsROBOTS_CLOSED, the README (directory list, Honest limits, FAQ), the input schema and the Store description all say which, and the daily quality audit fails the day the live verdicts drift from that record.
The quality audit (src/qa.js) puts its primaries back on the widest boards now that they are readable, with every threshold re-measured live on 2026-09-08 rather than inherited: viGlobal template A returns to Gibson Dunn (127 rows over 3 requests; title/location/url/job_id/detail_url/category all 100%) with Willkie kept on as its firm-owned alternate; template B returns to Jones Day (96 rows over 5 requests; detail_url 0% by the template's own design) with Mintz and Faegre Drinker as alternates; the category-sweep paging case returns to Jones Day (4 categories, 96 rows over 5 requests) with Blakes and Faegre Drinker behind it. Every case now has at least one alternate, which is a wider bench than before the closure. The Store prefill and selftest.json stay on Faegre Drinker.
Platform run on this source (build 0.1.38, 2026-09-08 07:34 UTC): the full directory sweep n1DVFrXkkesLGbkQG — SUCCEEDED, 223.4 s, 1,121 job rows from the 43 readable portals, 7 free status rows: exactly the 2 closed firms as robots-txt rows (each naming its host and path, each with portal_url in RUN_SUMMARY) plus 5 no-openings (Travers Smith, Norton Rose Fulbright, K&L Gates UK, Farrer & Co, Kingsley Napley), sources_ok 43, sources_failed 0, sources_walled 0, nothing partial or truncated. All 45 directory portals are accounted for in RUN_SUMMARY.sources; none vanished. The thirteen firms the tie brought back each served openings: Gibson Dunn 127, Jones Day 96, Dechert 46, ArentFox Schiff 40, Troutman Pepper 37, Fisher Phillips 34, Vinson & Elkins 19, Akin Gump 15, Stoel Rives 7, Nutter 4, Simpson Thacher 2, Davies Ward 2, Bracewell 1 — 430 of the run's rows. That figure replaces the 690-in-130 s sweep measured earlier the same day under the stricter tie reading. Build 0.1.39 is this source plus these numbers; the selftest.json run on 0.1.39 is recorded in the commit that ships it.
Harness: 196 checks (was 195), 0 failures. The equal-length tie is modelled in both rule orders and as the measured Cloudflare file (two * groups, Allow: / under a Content-Signal line, Disallow: / at the end) and must be READ — six job rows, one robots.txt read, no status row, nothing counted as refused. The control sits beside it: the same file with the managed Allow: / struck out is a bare Disallow: /, and it must still cost the portal zero requests, still emit the free row naming host and path, and still carry portal_url in its RUN_SUMMARY entry. Mutation-checked in both directions — a parser that gives the tie to Disallow fails the tie cases, one that ignores a bare Disallow: / fails the control.
0.1 — 2026-09-07 (an equal-length robots.txt tie closes the path — STANDARDS R12)
A robots.txt that contradicts itself at equal specificity is no longer read as permission.src/robots.js (byte-identical with the sibling actor's, as it must be) still combines every group naming this client (or *) and lets the longest matching rule win, as RFC 9309 says; the one change is the tie: when an Allow and a Disallow of the same length both match the portal path, the path is now closed. RFC 9309 §2.2.2 says a crawler should take the least restrictive rule on such a tie and permits a crawler to be stricter; the house rule (R12) takes that option, because the measured shape of such a file — Cloudflare's managed User-agent: * / Allow: / placed above the tenant's own User-agent: * / Disallow: / — is the tenant opting out, and a header its CDN wrote does not overrule the line the tenant wrote. A longer Allow still beats a shorter Disallow.
Re-measured 2026-09-07 with this actor's own identity, one robots.txt read on each of the 38 hosts behind the 45 directory portals: 20 hosts permit the portal path, 3 publish no robots.txt (a 404 is "no rules"), 15 close it. Two are the bare Disallow: / hosts already refused since 0.1.35 (Dorsey & Whitney, Lowenstein Sandler); thirteen are the Cloudflare-managed tie — twelve viglobalcloud.com tenants (Gibson Dunn, Dechert, Troutman Pepper, Fisher Phillips, Akin Gump, Simpson Thacher, ArentFox Schiff, Nutter, Bracewell, Stoel Rives, Jones Day, Davies Ward) and Vinson & Elkins' portal.velaw.com. The other five shared-zone tenants (Steptoe, Patterson Belknap, Baker Botts, Paul Hastings, Blakes) end their file with Disallow: /Admin/, not Disallow: /, and stay readable — the earlier assumption that all 17 shared-zone hosts carry the tie was wrong by five, which is why this was measured rather than copied. So 30 of the 45 directory portals (26 of 41 firms) are readable today and 15 are refused with a free robots-txt row; firms.jsROBOTS_CLOSED, the README (directory list, Honest limits, FAQ), the input schema and the Store description all say which, and the daily quality audit fails the day the live verdicts drift from that record.
Every sources entry in RUN_SUMMARY now carries portal_url — the robots-txt and error entries too, at the front-door URL the portal would have been read at — so a refusal can be checked against the host's own file without looking the portal up. The WizeHire probe URL is now built from the board slug the fetcher itself extracts (a pasted career-site URL used to double up).
The quality audit (src/qa.js) reads no host whose robots.txt closes it: the viGlobal template-A case moved from Gibson Dunn to Willkie (careers.willkie.com, 25 rows; this tenant publishes no job id and no per-job link, so those two carry no floor there), the template-B case from Jones Day to Mintz (careers.mintz.com, 18 rows) with Faegre Drinker as its alternate, and the category-sweep paging case from Jones Day to Blakes (5 categories, 8 rows over 6 requests) with Faegre Drinker (2 categories, 3 requests) as its firm-owned alternate — all thresholds measured on those hosts on 2026-09-07. The template-A case has no alternate: every other template-A tenant in the directory is closed, and a closed host is never read, so a wall on Willkie leaves that case honestly BLOCKED. The Store prefill and selftest.json stay on Faegre Drinker, which is open.
Harness: 195 checks (was 189). The equal-length tie is modelled in both rule orders and as the measured Cloudflare file (two * groups, Allow: / under a Content-Signal line, Disallow: / at the end): 0 portal requests, one robots.txt read, the free row naming host and path, and a RUN_SUMMARY refusal with portal_url; the same file ending in Disallow: /Admin/ must stay readable; robots-txt and error entries are held to carrying portal_url. npm test now runs the harness (npm run verify still does). The "41 hosts" wording in 0.1.35's entry, the README and src/http.js was the portal count minus a miscount; the measured host count is 38.
Platform run on this source (build 0.1.36, 2026-09-08 04:55 UTC): the full directory sweep 2onnLHvJoelJGIuNT — SUCCEEDED, 130.3 s, 690 job rows from the 30 readable portals, 20 free status rows: exactly the 15 closed firms as robots-txt rows (each naming its host and path, each with portal_url in RUN_SUMMARY) plus 5 no-openings, sources_ok 30, sources_failed 0, sources_walled 0, nothing partial or truncated. That figure replaces the 2026-09-06 one (1,146 openings, 134 s) in the README, which predated pacing and the closures. Build 0.1.37 is this source plus these numbers; the selftest.json run on 0.1.37 is recorded in the commit that ships it.
0.1 — 2026-09-07
robots.txt is now read and obeyed on every host this actor touches. Until this build it never read one. Measured 2026-09-07 across the 38 hosts the directory reads (this entry originally said 41; 38 is the count of distinct hosts behind the 45 portals), two publish User-agent: * / Disallow: / — recruiting.dorsey.com (Dorsey & Whitney) and recruiting.lowenstein.com (Lowenstein Sandler) — and both were being read on every run, Dorsey on the default input. Now, before a host is asked anything, its robots.txt is read once per run (this actor's own identity, paced like any other request, one retry if it did not answer), and the exact path is checked: a disallowed path is never requested and the firm gets a free robots-txt row saying robots.txt disallows <path> on <host>; a host whose rules could not be read is not requested that run, with that reason; a 4xx robots.txt is no rules. This applies to directory firms, to pasted portal URLs (both probe candidates), to WizeHire's API hosts, and to any host a portal redirects to — redirects are now followed by hand so each hop is checked and paced. RUN_SUMMARY gains sources_robots_refused; a refusal is never a failure, never fails the run and does not trip failOnSourceError. The reader (src/robots.js) is byte-identical with the sibling actor's; groups are matched by exact product token (DockhandLegalJobsMonitor, or the family token Dockhand), never by prefix — cvMail's host lists User-agent: DOC under Disallow: /, which a prefix match would have read as closing all eight cvMail boards. The 18 viGlobal hosts on Cloudflare-managed robots.txt (two * groups, Allow: / above Disallow: /) were read under RFC 9309's tie-to-Allow rule in this build; the next entry reverses that under R12.
The default input and the daily self-test moved off Dorsey & Whitney to Faegre Drinker (apply.faegredrinker.com, firm-owned viGlobal domain, robots.txt closes only /Admin/; measured 2026-09-07: 18 openings over 2 categories, 3 requests). The quality audit's viGlobal alternates moved with it, with thresholds measured on that portal.
About one request a second per host. Requests were sequential but back to back; a category sweep or a five-page cvMail board now waits out a one-second spacing per host, and the harness measures the gap at the shipped setting.
Builds 0.1.31–0.1.33 (2026-09-07, previously unlisted here): the entity-decoder and Latin-1 fixes listed under 2026-09-06 shipped in them, and the bot-challenge marker list is now one list shared with the sibling actor and the factory's network probe (the harness fails if they drift), with a Cloudflare cf-mitigated: challenge header recognised on any status.
0.1 — 2026-09-06
A run whose every portal was bot-challenged no longer fails. Watching a single firm whose portal served a challenge crashed the run with "Every portal failed" — reporting the actor as broken for something that was neither the actor's nor the portal's doing, and that the same run from another network would not have seen at all. A challenge is now separated from a portal failure everywhere it is counted: RUN_SUMMARY carries sources_walled and all_sources_failed, each walled entry in sources carries walled: true, the wall is logged in full as an error, and the run succeeds with its free error rows. A portal that genuinely would not answer still fails the run when it is the only one, and failOnSourceError still fails on anything, challenge included.
A portal that answers with a bot-verification challenge now says so. viGlobal enabled Cloudflare bot management on its shared viglobalcloud.com domain on 2026-09-06, and the 17 directory firms hosted there began serving some clients an HTTP 403 challenge page instead of their listings — while the same URLs answered normally from the Apify platform (Jones Day: 96 openings, measured the same hour). The free error row for such a portal now names the challenge and states that no workaround was attempted, instead of reading "HTTP 403" as though the portal had broken. The 7 viGlobal firms on firm-owned domains are unaffected. See Honest limits in the README.
The default input runs on a firm-owned viGlobal domain (dorsey; moved to faegre-drinker on 2026-09-07 when Dorsey's robots.txt was found to close the portal) rather than the vendor's shared one, so a bot-management change on that shared domain cannot decide whether the sample run returns rows.
Latham & Watkins is live. The actor now introduces itself as DockhandLegalJobsMonitor instead of a Chrome browser string. That string was the whole "iCIMS bot wall": the firewall in front of Latham's five career sites challenges a request that claims to be a browser but is not one, and had answered every build so far with HTTP 405 from any network. Named honestly, all five tenants answer normally: 129 US openings measured over 3 pages (matching the 129 job URLs in the tenant's own sitemap), 27 UK, 9 Europe/Middle East, 8 Europe/Asia, 1 France. The quality audit now calibrates iCIMS like every other platform and walks its paging live against the sitemap count. The directory stands at 45 verified portals; a full platform sweep on 2026-09-06 returned 1,146 openings in 134 seconds with no errors.
The entity decoder answers only names it actually knows; an entity spelled like a JavaScript object property (&toString;) stays as written instead of splicing function text into a row.
Latin-1 HTML entities now decode with their case kept: Latham's French tenant posts À temps plein, which reached the dataset raw because the decoder knew a short hand-picked list; it now knows the whole –ÿ set and tells À from à.
Output schema added (.actor/output_schema.json): the run's Output tab and the API's output field now link straight to the rows as the overview table, JSON and CSV. Apify Store requires it before an Actor can be published. Nothing about the rows, the portals or the price changed.
0.1 — 2026-09-04
Law-firm openings read straight from four live hiring-portal platforms — viRecruit/viGlobal, cvMail, AllHires and WizeHire — for 40 built-in firm portals, or any firm on those platforms by pasting its portal URL. iCIMS (Latham & Watkins) is included but blocked from most networks; see Honest limits in the README.
One 17-column row layout across every firm: firm, platform, job_id, title, practice_area, department, location, remote, pqe, experience_level, employment_type, salary, category, posted_at, url, portal_url, scraped_at. A field the source does not publish is null, never invented.
Monitor mode: the first run returns everything currently open; every later run returns only openings that appeared since. Openings you were not delivered are never marked as seen, so they come back.
Title and location filters run before you are charged; maxJobsPerFirm caps what each portal delivers, never how far its feed is read.
Every "nothing to show" case gets a free status + note row — nine values, all listed in the README — and only delivered job rows are ever charged.
A request that never completed — a dropped connection, a timeout, or a 429/5xx — is asked again up to three times over a few seconds before a portal is called broken. A host that really is down still gets its free error row, and that row now names the host, the attempts and the underlying reason instead of a bare "fetch failed". A portal that answers "no" is never asked twice.