Law Firm Jobs Monitor — Legal Job Alerts from BigLaw Portals
Pricing
$2.00 / 1,000 job rows
Law Firm Jobs Monitor — Legal Job Alerts from BigLaw Portals
Law firm job openings straight from the portals firms hire through — viRecruit/viGlobal, cvMail, AllHires, WizeHire, iCIMS. 45 portals, robots.txt obeyed: 43 readable, 2 closed by theirs (2026-09-08). Or paste a portal URL. Monitor mode returns only NEW openings; status rows free.
Pricing
$2.00 / 1,000 job rows
Rating
0.0
(0)
Developer
Dockhand
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Law firms publish their openings on their own hiring portals. This actor reads those portals directly and, in monitor mode, tells you only what is new.
Give it a list of law firms and get their current openings as clean rows — title, practice area, office, apply link. Turn on monitor mode and every run after the first returns only the openings that appeared since your last check, so a scheduled daily watch costs you a fraction of a cent per new role and nothing on a quiet day. The fastest way to try it: press Start with the three prefilled firms (Faegre Drinker, Farrer & Co, RPC) and read the rows.
What you get
- Openings from five live law-firm portal platforms — viRecruit/viGlobal, cvMail, AllHires, WizeHire and iCIMS — with 45 firm portals (41 firms) built in and verified live, plus any other firm on those platforms by pasting its portal URL. Every host's
robots.txtis read and obeyed before it is asked anything: 43 of the 45 portals are readable today and 2 are closed by their ownrobots.txtas of 2026-09-08 (see Honest limits — the closed ones return a free row saying so, never a workaround). - One row layout across every firm.
firm,platform,job_id,title,practice_area,department,location,remote,pqe,experience_level,employment_type,salary,category,posted_at,url,portal_url,scraped_at. CSV and Excel exports line up whatever the source. A field the source does not publish isnull— never a guess. - Only-new monitor mode. The first run returns everything currently open; every later run returns only openings that appeared since. What you were shown and charged for is never shown or charged again.
- Filters that run before you pay. Keep titles containing your words, drop titles containing others, keep only the offices you care about. Rows your filters remove are never charged.
- A free status row for every "nothing to show" case. A portal with no openings, no new openings, nothing matching your filters, a portal that did not answer, a run that hit your charge limit, a read that stopped early — each says so in the dataset, in plain words, for free. You never stare at an unexplained empty table.
- Public, logged-out pages only, and every host's
robots.txtobeyed. No logins, no proxies, no headless browser, no bot-wall evasion, about one request a second per host, and a path a host'srobots.txtdisallows is never requested. When a portal blocks — or itsrobots.txtsays no — the run says so in a free row instead of working around it.
Which law firms and job boards it reads
| Platform | Who hires through it | Coverage |
|---|---|---|
| viRecruit / viGlobal | US and Canadian firms — Gibson Dunn, Jones Day, Dechert, Akin Gump, Baker Botts, Troutman Pepper, Vinson & Elkins, Willkie, Simpson Thacher, Blakes and more | 24 firm portals verified live; 22 readable today, 2 closed by their own robots.txt as of 2026-09-08 (see Honest limits) |
| cvMail (Thomson Reuters) | UK and international firms — Ashurst, Squire Patton Boggs, Mishcon de Reya, Bird & Bird, RPC, Norton Rose Fulbright and more | 8 firm portals verified live |
| AllHires | UK firms and the London offices of US firms — Skadden, Debevoise, Shoosmiths, Birketts, Farrer & Co and more | 8 firm portals verified live |
| WizeHire | Small and boutique US law firms, searchable network-wide by job title (attorney, paralegal, legal assistant) | live |
| iCIMS (Latham & Watkins) | Latham's five regional career sites — US, UK, Europe/Middle East, Europe/Asia, France | 5 tenants verified live (174 openings on 2026-09-06) |
A full sweep of the built-in directory returned 1,121 open listings from the 43 readable portals in 223 seconds (measured on the platform 2026-09-08 07:34 UTC, run n1DVFrXkkesLGbkQG, with one request a second per host and every host's robots.txt read first; the 2 closed portals came back as free robots-txt rows, no portal errored, none was challenged, nothing was truncated or partial). That is a snapshot, not a promise: firms post and close roles daily, so your totals will differ. Five of the 43 readable portals listed nothing that day (Travers Smith, Norton Rose Fulbright, K&L Gates UK, Farrer & Co, Kingsley Napley) — that is the firm, not the fetcher, and the run said so in a free status row rather than leaving a gap. The earlier figure of 690 listings in 130 seconds (2026-09-08 04:55 UTC) was measured while thirteen of these portals were closed under the stricter tie reading, and is no longer what this build does.
The built-in directory
Type any of these short handles into firms (spaces or hyphens both work, and the firm's full name as written here works too):
- viRecruit / viGlobal (US & Canada):
gibson-dunn(Gibson Dunn & Crutcher),willkie(Willkie Farr & Gallagher),vinson-elkins,dorsey† (Dorsey & Whitney),dechert,troutman-pepper(Troutman Pepper Locke),fisher-phillips,mintz,akin-gump,faegre-drinker,davis-wright-tremaine,steptoe,patterson-belknap,simpson-thacher,baker-botts,paul-hastings,arentfox-schiff,nutter(Nutter McClennen & Fish),bracewell,lowenstein-sandler†,stoel-rives,jones-day,blakes(Blake, Cassels & Graydon),davies-ward(Davies Ward Phillips & Vineberg)- † closed by its own
robots.txtas of 2026-09-08 —dorseyandlowenstein-sandler, the two hosts whose file saysUser-agent: * / Disallow: /and nothing else — each returns a freerobots-txtrow instead of openings until its host's rules change (see Honest limits for the rule and the reading). The other 22 viRecruit portals are read normally.
- † closed by its own
- cvMail (UK & international):
mishcon-de-reya,squire-patton-boggs,ashurst,rpc,bird-and-bird,travers-smith,norton-rose-fulbright,kl-gates(K&L Gates UK) - AllHires (UK):
birketts,shoosmiths,skadden(London),tees-law,farrer(Farrer & Co),debevoise(London),broadfield,kingsley-napley - iCIMS (Latham & Watkins, one tenant per region):
latham-watkins(US),latham-watkins-uk,latham-watkins-eume,latham-watkins-euas,latham-watkins-fr(France)
A firm that is not in the directory
Paste its portal URL into firms and the platform is detected from the address: viRecruit/viGlobal (…/viRecruitSelfApply/ReDefault.aspx or RecDefault.aspx, or any viglobalcloud.com host), cvMail (fsr.cvmailuk.com/<slug>/…), AllHires (<firm>.allhires.com), WizeHire (wizehire.com/cmp/<slug>) and iCIMS (<tenant>.icims.com). These portals have no guessable addresses — Squire Patton Boggs is spb, Bird & Bird is twobirds, Jones Day is jonesdaylegalrecruitselfapply — which is why the directory exists and why the actor asks for a URL instead of guessing.
Cannot find the portal, or the firm hires through a platform not listed here? Open an issue on this actor naming the firm. Portal requests are answered within a day, and portal coverage is how this actor grows.
How much does it cost?
This actor is priced per event, and the only billable event is job-result: one job row delivered to your dataset, at $0.002. There is no fee for starting a run, no fee for a run that finds nothing, and status rows are never charged.
| What happens | Job rows charged | Cost |
|---|---|---|
| A daily monitor run finds 50 new openings across your watchlist | 50 | $0.10 |
| A first run on a large watchlist returns 500 openings | 500 | $1.00 |
| A first sweep of every readable portal in the directory (690 openings on 2026-09-08) | 690 | $1.38 |
| A scheduled run on a day when nothing new was posted | 0 | $0.00 — you still get free no-new-openings rows saying so |
| A portal is down, or none of its roles match your filters | 0 | $0.00 — a free error or no-match row explains it |
Three rules make the price predictable:
- You pay for rows you asked for, after filtering. Title and location filters run before billing; each portal's feed is de-duplicated before billing; and in monitor mode an opening you were charged for is never charged again on later runs. The overlap to know about: WizeHire searches are separate from each other and from any WizeHire firm board you also watch, so a posting matched by two keywords — or by a keyword and a board — is delivered under each.
- Your run's maximum charge is respected before anything is pushed. Set a maximum charge on the run (
maxTotalChargeUsd) and the actor delivers only what fits, says so in a freecharge-limit-reachedrow, counts the held-back openings in RUN_SUMMARY (rows_held_back_by_charge_limit), and leaves those openings unmarked so your next run returns them. It never delivers rows it cannot bill and never keeps rows it did not deliver. - Status and error rows are always free. They never touch billing at all.
How to use it
- Try it. Open the actor and press Start. The prefilled input watches three firms (Faegre Drinker, Farrer & Co, RPC) and returns their current openings, with a free status row for any portal that has nothing to show. The prefilled three-firm run returned 40 openings in 16 seconds on the platform (measured 2026-09-08 UTC, run
RQVkRoYh1424zAznM); runs that add a WizeHire keyword search took 18–25 seconds (2026-09-03); the full directory sweep — 43 readable portals plus 2 freerobots-txtrows — took 223 seconds (2026-09-08, platform runn1DVFrXkkesLGbkQG). - Build your watchlist. Put firm handles from the directory, or portal URLs, in
firms. For small US firms addwizehireKeywordssuch asattorneyorparalegal, which search every WizeHire board at once. SetallFirmsto sweep the whole directory. - Narrow it. Use
titleIncludes(associate,counsel,paralegal),titleExcludes(summer,intern,trainee) andlocationIncludes(new york,london). Filtering happens before you are charged. - Turn on
monitorModeand schedule it. Create a schedule for it in Apify Console — daily, or hourly if you want to be first to a posting. The first scheduled run returns everything currently open; every run after that returns only openings posted since. The seen-list lives in a key-value store namedlegal-jobs-monitor-stateon your own account, so it survives between runs. - Send new openings somewhere useful. In the Integrations tab, connect Slack, email, a webhook, Zapier or Make so each finished run hands its dataset onward. Because a monitor run's dataset holds only new openings (plus free status rows), what arrives is your alert.
- Export any run. Download the dataset as CSV, Excel or JSON from the run's Storage tab, or read it through the Apify API.
Input
{"firms": ["dechert", "farrer", "rpc"],"monitorMode": true,"titleIncludes": ["associate", "counsel"],"locationIncludes": ["new york", "london"]}
firms— directory handles (dechert,gibson dunn,mishcon de reya,skadden,birketts) or a portal URL for any firm not in the directory. URLs are auto-detected for all five platforms.allFirms— sweep every portal in the directory instead of naming firms.wizehireKeywords— search every WizeHire board at once by job title (attorney,paralegal,legal assistant). This is how the small-firm market shows up.monitorMode— after the first run, deliver only openings that appeared since the last one.titleIncludes/titleExcludes/locationIncludes— case-insensitive "contains" filters that run before you are charged.maxJobsPerFirm— caps how many openings each portal delivers per run, and therefore the cost. It never shortens a fetch, and in monitor mode the openings over the cap are not marked as seen, so they arrive next run instead of disappearing.maxPagesPerSource— a runaway guard, in pages. Leave it at 0: each source ships a guard set far above the widest board ever measured on it, and it exists only so a broken feed cannot loop forever. Set it lower and you get a deliberately short read — which tells you so, in a freeresults-truncatedrow that names this input.includeStatusRows(on by default) — a portal with nothing to report explains itself in the results instead of leaving a silent gap. These rows are free.failOnSourceError— fail the whole run if any portal errors, so a scheduled watch (and Apify's failure notification) alerts you the moment one breaks. Off by default: the broken portal gets a freeerrorrow and the rest still return. With it off, the run itself fails only when every portal you asked for failed and at least one of them refused to answer at all — a run in which every portal served a bot challenge instead is reported loudly and still succeeds, because a challenge is a fact about the network the run was made from, not about the portals.
What monitor mode remembers (and what it deliberately doesn't)
Monitor mode keeps a per-watchlist "already seen" list in the legal-jobs-monitor-state store on your account. What lands on that list is what you are done with, not everything the portal showed:
| Rows | Remembered? | Why |
|---|---|---|
| Openings delivered to you and charged | Yes | You have them. You will not be shown — or charged for — them again. |
Openings your titleIncludes / titleExcludes / locationIncludes filters excluded | Yes | Your filter is a decision about that opening, not a failure to deliver it. So widening a filter later does not resurface openings you previously filtered out — delete the legal-jobs-monitor-state store to start a fresh watch. |
Openings cut by maxJobsPerFirm | No | The cap limits how many roles are delivered per run, not which roles exist. A busy portal drains over several runs instead of vanishing. |
Openings held back because the run hit your maxTotalChargeUsd | No | You never received them and were never charged for them. Re-run, or raise the limit, and they arrive. |
| Openings pushed but that the platform declined to charge for | No | Only rows actually paid for are remembered, so a partly-charged batch comes back in full next run instead of half-vanishing. |
The list is saved after every portal, not at the end of the run. If a run is cut short — a timeout, a memory limit, a platform restart — everything already delivered is already remembered, so the next run picks up where it stopped instead of re-delivering and re-charging you for the same openings.
The list is keyed to your exact watchlist. Add or remove a firm, a URL or a WizeHire keyword and the next run starts a fresh watch for the new list: it returns (and charges for) everything currently open on that list once, then goes back to only-new. Filters do not affect the key — change titleIncludes freely.
Output
Every row has the same columns regardless of which platform it came from:
{"firm": "Dechert","platform": "viglobal","job_id": "446","title": "Bankruptcy Associate","practice_area": "Financial Restructuring","department": null,"location": "New York","remote": null,"pqe": null,"experience_level": null,"employment_type": null,"salary": null,"category": "Associate","posted_at": "Mar 17, 2026","url": "https://dechertselfapply.viglobalcloud.com/viRecruitSelfApply/RecDefault.aspx","portal_url": "https://dechertselfapply.viglobalcloud.com/viRecruitSelfApply/RecDefault.aspx","scraped_at": "2026-09-03T03:50:55.713Z"}
A field the source does not publish is null. Nothing is ever invented to fill a column.
Which columns each platform fills
Every measured platform provides firm, platform, job_id, title, location, url, portal_url and scraped_at (90–100% filled on every source measured). The rest depend on what the portal publishes:
| Platform | practice_area | department | category | posted_at | remote | salary | pqe | experience_level | employment_type |
|---|---|---|---|---|---|---|---|---|---|
| viRecruit / viGlobal, older template | — | — | yes (the portal's job category) | — | — | — | — | — | — |
| viRecruit / viGlobal, newer template | yes | — | yes | yes (display text, e.g. Mar 17, 2026) | — | — | — | — | — |
| cvMail | yes (the board's own Department / Practice Area / Expertise column) | — | — | — | — | — | — | — | — |
| AllHires | yes | yes | yes | — | — | — | where stated (7–29% of rows measured) | where stated (60–65% measured) | — |
| WizeHire, single firm | — | — | — | yes (ISO 8601) | where stated | where stated (71–100% measured) | — | — | — |
| WizeHire, keyword search | — | — | yes (your keyword) | yes (ISO 8601) | yes | where stated (100% on 180 measured rows) | — | — | — |
| iCIMS | yes (Department) | — | — | — | where stated | where stated | — | — | yes (Position Type) |
On the newer viRecruit template the cards carry no per-job link (measured on Jones Day and Dechert), so url is the portal address; the older template gives a per-job link, which expires with the portal session (see Honest limits). iCIMS was measured 2026-09-06 on Latham's US tenant, 129 rows: practice_area 99% (the card's Department), employment_type 99% (Position Type), remote 100% (Work Arrangement), salary 100% there because Latham posts US pay ranges — the UK tenant posts none, so salary is null on its 27 rows.
You never get an unexplained empty table
Any target that produces no job row says why, in the dataset. These rows carry status and note, and they are never charged. This is the complete list — ten values, and the actor emits no others:
status | What it means |
|---|---|
no-openings | The portal answered normally and currently lists no open roles. |
no-new-openings | Monitor mode: the portal has roles, but none new since your last run. A watch that finds nothing still tells you so. |
no-match | The portal has roles, but none passed your title/location filters. The note says how many it held. |
error | The portal did not answer, after up to three tries — the note carries the host, the attempts and the exact reason. |
unknown-firm | The name is not in the directory and is not a recognised portal URL. Paste the portal URL instead. |
charge-limit-reached | The run hit the maximum charge you set for it, so it stopped delivering. Fires whenever openings were held back — including when the portal delivered some of them. Held-back openings are not marked as seen: re-run, or raise the limit, and they come back. |
charge-failed | The platform could not record a charge, so the run stopped delivering there rather than sweeping the rest of your watchlist for free. Nothing was marked as seen. |
results-truncated | The fetch stopped before the feed ended, and the row's truncation_reason says which of three things happened. our-guard — this actor's own runaway guard stopped it; the note names maxPagesPerSource, the input that raises it, and nothing was cut by the source. source-ceiling — the source itself stated it holds N openings and then served fewer; both numbers are in the note. unknown — the feed repeated a page instead of advancing, so paging stopped and some openings may be missing. |
partial | A page or job category of that portal could not be read (a 500 on page 3, a dead category in a viRecruit sweep). The rows you got are real but incomplete, and the note says how many parts were lost and why. |
robots-txt | The portal's host publishes a robots.txt that does not permit this actor to read that path — or the rules could not be read at all, and rules that cannot be read are not permission — so nothing was requested from it. The note names the path and the host (robots.txt disallows /viRecruitSelfApply/RecDefault.aspx on recruiting.dorsey.com). It is not a failure and no workaround is attempted; the rules are re-read on every run, so the portal comes back the day they allow it. |
Where the page size is ours to choose (both WizeHire feeds), a source-ceiling note also carries its receipt — the page sizes we asked the source for — and says which of three things it measured: we asked for the whole feed in one request and were still refused the rest; or the source states a maximum page size below its own total, so that request cannot be made at all and it refuses any page above the maximum it stated; or the source stated no maximum and our own largest request still fell below its total, in which case the note says plainly that this is where the actor stopped asking, not a limit the source set. A number the source never stated is never worded as the source refusing anything.
Set includeStatusRows: false if you want job rows only.
RUN_SUMMARY
Written to the run's key-value store. Every field below is present on every run, so a 0 is a measurement rather than a missing key.
| Field | Meaning |
|---|---|
sources | One entry per target. A target that answered carries: status ok, jobs_found, jobs_output, jobs_held_back_by_charge_limit, charge_failed, source_total, source_page_max, page_sizes_tried, requests, results_truncated, truncation_reason, partial, source_errors, portal_url. A target that did not carries status (error, robots-txt, unknown or charge-limit-reached) and, for errors and refusals, detail with the reason. jobs_found is what was read from the source, never what a delivery cap returned. source_total is what the source itself says it holds (null when it states none). requests is how many HTTP requests that fetch made. |
total_jobs_output | Job rows delivered — the billable ones. |
status_rows | Free explanation rows delivered. |
sources_ok / sources_failed | How many targets answered, and how many did not. |
sources_walled / all_sources_failed | How many of the failures were a bot challenge rather than a portal failure (those entries also carry walled: true in sources), and whether no target answered at all. A run where every portal was challenged has all_sources_failed: true with sources_walled equal to sources_failed, and does not fail: nothing there says a job list changed. |
sources_robots_refused | How many targets this actor declined to read because the host's robots.txt does not permit it (or could not be read). Not a failure and never counted as one: those entries carry status: "robots-txt" in sources, and the same run's sources_failed does not include them. |
monitorMode | Whether this run returned only new openings. |
pay_per_event | Whether this run had pay-per-event pricing at all. |
job_rows_charged | Job rows actually charged. 0 on any run without pay-per-event pricing. |
charge_limit_reached / rows_held_back_by_charge_limit | Whether your maxTotalChargeUsd stopped the run, and how many openings it held back. |
charge_errors | How many portals hit a charge failure. |
results_truncated / truncated_sources | Whether any source stopped short of the end of its feed, and how many did. |
partial_sources | How many portals could only be read in part. |
page_guards | The runaway guards in force, in pages: per_platform (the measured defaults) or override_all when you set maxPagesPerSource, plus the name of that input. |
Who uses a law firm jobs monitor
- Legal recruiters and search firms — know a firm is hiring the day the role goes up on its own portal, not when it reaches LinkedIn.
- Law-firm business development and competitive-intelligence teams — track competitors' lateral hiring and practice-group build-outs by office.
- Associates, counsel and paralegals with a target list — your ten firms, filtered to the titles and offices you would actually apply to, checked every morning for you.
- Law-school career offices — a live feed of what your students' target firms are actually hiring for.
- Vendors selling to law firms and legal-market analysts — hiring velocity by firm, office and practice area as a growth signal.
The data itself is not new — Leopard Solutions and Firm Prospects have sold law-firm hiring intelligence for years, behind a demo request and an annual contract. What is new is buying it by the row, on a public marketplace, with no sales call. The job scrapers on this store read aggregators (LinkedIn, Indeed, Lawjobs, TotallyLegal); none reads the viRecruit, cvMail or AllHires portals (checked 2026-09-03); generic iCIMS career-site scrapers exist and will read a Latham tenant if you hand them its address, but none of them carries a law-firm directory (checked 2026-09-06).
Honest limits
Read this before you buy — it is cheaper than a refund.
-
Not every field exists on every platform. cvMail, AllHires and iCIMS publish no posted date; viRecruit publishes one only on its newer template; WizeHire has no employment-type field; salary appears only on WizeHire and iCIMS. Missing fields are
null, never invented. -
posted_atfollows the source and is never normalised across platforms:Platform posted_atFormat WizeHire (keyword search and single-firm boards) Yes, ~100% filled (measured on 180 rows for paralegal, 10 rows for a single firm)ISO 8601 timestamp straight from the API — sortable as-is viRecruit / viGlobal, newer "gridviewList" template Yes, 100% filled (measured on 96 rows) The portal's own display string, e.g. "Mar 17, 2026"— not ISOviRecruit / viGlobal, older template No null— the template has no date columncvMail No null— the board has no date columnAllHires No null— not in the public API payloadiCIMS No null— not on the job cardsA WizeHire-only run gives you a genuinely sortable date; a mixed run does not. viRecruit's
"Mar 17, 2026"is deliberately not converted to an ISO timestamp — that would invent a precision the source never gave. -
No full job descriptions. Descriptions need a separate request per listing, which would multiply run time and cost. Titles, practice area, location and the apply link are what you get.
-
viRecruit apply links expire. Those portals mint a per-session token, so a detail link works now but not next week;
portal_urlis the durable one. On the newer viRecruit template there is no per-job link at all, andurlis the portal address. -
Latham & Watkins (iCIMS) answers a client that says what it is. Until 2026-09-06 every Latham tenant returned HTTP 405 to this actor — not because iCIMS blocks datacenter traffic, as the earlier build believed, but because that build introduced itself as a Chrome browser, and iCIMS's firewall challenges a request that claims to be a browser and does not behave like one. The actor now identifies itself plainly (
DockhandLegalJobsMonitor), and all five tenants answer normally from every network measured, an office connection and the Apify platform alike. Should a challenge ever return, it is still reported as a freeerrorrow naming the challenge, and never worked around. -
A bot challenge is reported as a challenge, and the network you run from can decide whether you see one. On 2026-09-06 viGlobal switched on Cloudflare bot management for its shared
viglobalcloud.comdomain. The 17 directory firms hosted there — Gibson Dunn, Jones Day, Dechert, Akin Gump, Baker Botts, Troutman Pepper, Fisher Phillips, Steptoe, Patterson Belknap, Simpson Thacher, Paul Hastings, ArentFox Schiff, Nutter, Bracewell, Stoel Rives, Blakes, Davies Ward — now answer some clients with HTTP 403 and a "Just a moment…" challenge instead of their listings. The 7 viGlobal firms on firm-owned domains (Willkie, Vinson & Elkins, Dorsey, Mintz, Faegre Drinker, Davis Wright Tremaine, Lowenstein Sandler) are unaffected. Measured within the same hour on 2026-09-06: all 17 answered normally from the Apify platform — Jones Day returned its full 96 openings — while the same URLs were challenged from an ordinary office connection. So this is about which network a request comes from, not about portals going away, and the actor does not attempt to defeat a challenge. When one is served, that firm gets a freeerrorrow that names the challenge and says no workaround was attempted; the run's other portals are unaffected. (All 17 are readable as of 2026-09-08. Twelve of them carry the equal-length tie the next bullet explains, and a tie is read; the other five — Steptoe, Patterson Belknap, Baker Botts, Paul Hastings, Blakes — end theirrobots.txtatDisallow: /Admin/, which never covered the portal path. Measured that day with this actor's own identity.) -
Two directory portals are closed by their own
robots.txt, and this actor obeysrobots.txt. Before any host is asked anything, itsrobots.txtis read once (with this actor's own identity, paced like any other request, retried once if it did not answer) and the exact path is checked; a disallowed path is never requested, and a host whose rules could not be read is not requested that run. Measured 2026-09-08 with this actor's own identity, onerobots.txtread on each of the 38 hosts the directory reads: 33 hosts permit their portal path, 3 publish norobots.txtat all (a 404 is "no rules") and 2 close it — 2 portals, 2 firms: Dorsey & Whitney and Lowenstein Sandler (handlesdorseyandlowenstein-sandler). Both sayUser-agent: * / Disallow: /and nothing else: the site plainly says no, so the portal is not asked. Thirteen other hosts — twelve tenants on viGlobal's sharedviglobalcloud.comzone and Vinson & Elkins'portal.velaw.com— publish a file that contradicts itself: twoUser-agent: *groups, Cloudflare's managedAllow: /at the top and the tenant's ownDisallow: /at the bottom, both about the whole site. A site file that contradicts itself at the same level is read. That is what the standard says to do — RFC 9309 §2.2.2 combines every group that names the same agent, lets the longest matching rule win, and breaks a tie between rules of equal length in favour of theAllow— and it is how Google's own reference parser reads the same file, so those thirteen portals are read, exactly as a search engine reads them. A site that plainly says no is not read: the two bareDisallow: /hosts above have noAllowanywhere in the file, so there is no tie to break and nothing is asked of them. (This is house rule R13, the owner's ruling of 2026-09-08. An earlier build, 2026-09-07 only, took the stricter reading the RFC also permits and closed all fifteen; that reading was superseded.) The same file withDisallow: /Admin/as the tenant's line, which is what the other five shared-zone tenants (Steptoe, Patterson Belknap, Baker Botts, Paul Hastings, Blakes) publish, closes only the admin area and leaves the portal open. Every closed firm returns a freerobots-txtrow naming the host and the path instead of openings, and comes back the day its rules allow it; the rules are re-read on every run, so a host that closes is dropped the same day. The default input sat on Dorsey until 2026-09-07 and now sits on Faegre Drinker (apply.faegredrinker.com, whoserobots.txtcloses only/Admin/). Which group is "ours" is decided by exact product token —DockhandLegalJobsMonitor, or the family tokenDockhand— never by prefix, so a blocklist line such asUser-agent: DOC(cvMail's host lists it among 316 agents) does not close a board to us. -
Lawcruit is not supported, deliberately. Its firm portals are application forms — the job lists live on each firm's own website, which is a per-firm scraper, not a platform integration. Claiming it would be a lie.
-
An incomplete read is labelled, never smoothed over. If a page or a job category dies mid-fetch, you get the rows that did come back plus a free
partialrow saying how many parts were lost and why. The count inRUN_SUMMARYis then what was actually read, and it is flagged. -
Some portals are stale. A handful of firms leave old postings up (a 2022 summer-associate slot, even a "Test Position"). We report what the portal says and surface
posted_atwhere it exists so you can judge, rather than silently dropping rows. -
Coverage is portal-by-portal, not "all of BigLaw." 45 verified portals (41 firms): 43 portals (39 firms) readable, 2 closed by their own
robots.txtas of 2026-09-08. Several marquee New York firms (Cravath, Wachtell, Davis Polk, Sullivan & Cromwell) do not use any of these platforms publicly, and no scraper can invent a portal that is not there. -
Nothing is sorted by date. Portals list roles in their own order and WizeHire groups results by company, so the newest postings sit scattered through a result set. Sort on
posted_atwhere it exists, or use monitor mode and let the actor do the "what is new" for you.
How each source is read
Before a portal is asked anything, its host's robots.txt is read once and obeyed for every path (see Honest limits), including any host the portal redirects to, and requests to one host are spaced about a second apart. Then every source is paged to the end of its feed, and every fetcher reports the same facts: what the source said it holds, what it served, and what (if anything) stopped the read early. Where the page size is a request parameter it is treated as ours: we ask for a size that clears the feed, and a walk that still ends below the source's own stated total is re-requested at a bigger size and merged before any shortfall is reported. Guards exist only so a broken feed cannot loop forever; each sits far above the widest board that source has ever shown us, and one that fires produces a free results-truncated row saying our guard did it.
| Source | How it pages | How we know the feed ended | Widest board measured | Runaway guard |
|---|---|---|---|---|
| viRecruit / viGlobal | Not paged. What multiplies requests is the job-category sweep: the dropdown defaults to one category and the rest are only reachable via ?FilterJobCategoryID=N. | The category list is exhausted. Measured 2026-09-03 across five tenants: the positions list carries no pager markup of any kind — no "Page N of M", no __doPostBack('…Page$…'), no page-size selector. | 4 categories, 96 rows over 5 requests (Jones Day) | 100 categories |
| cvMail | 30 rows a page, advanced by a stateful Next >> POST whose action URL carries a fresh token each time. The board states its own total on page 1 in the "Jump To List" pager (1 - 30, 31 - 60, … 121 - 128), reported as source_total. | Every row of the stated total in hand, or the served page drops the Next control. Measured 2026-09-03: Ashurst, 128 rows over 5 pages, last page 8 rows, matching its stated 128. A board that drops Next before its own stated total is flagged source-ceiling rather than presented as complete; one that re-serves a page it already served stops and is reported as unknown. A board that fits on one page prints no pager, so it states no total (RPC: 23 rows). | 128 rows / 5 pages (Ashurst) | 200 pages (~6,000 rows) |
| AllHires | Not paged. One /webapi/candidate/Positions call returns the whole board. | The payload has no paging keys and states no total (measured 2026-09-03: Shoosmiths 31, Birketts 47). If the API ever claims more than it served, that is reported instead of dropped. | 47 rows in one payload (Birketts) | none needed — no loop |
| WizeHire, single firm | Offset paging. The response carries paging.count (the board's own total), has_more, and a server-side clamp: the page size is capped at 100 — asking for 500 comes back as limit: 100. | has_more: false, an empty page, or every row of paging.count in hand. Verified by walking a 10-row board 3 rows at a time and matching the unpaged result exactly. A walk that ends below paging.count is re-requested at a size that can hold the board, and the results merged, before anything is reported. | 10 rows (law-firm boards are small) | 100 pages (10,000 rows) |
| WizeHire, network keyword search | Keyset cursor (after / endCursor), with pagination.total stating the source's own total. The page size is ours — first= is honoured — and we ask for 500, which holds every legal keyword measured in a single request. | hasNextPage: false, or the stated total reached. Pages overlap and, at small page sizes, the cursor also skips rows, so rows are deduped and any shortfall is re-requested at a size that can hold the whole feed before anything is reported. A cursor that stops advancing ends the walk and is reported as unknown. | 352 of a stated 352 in one request (attorney, 2026-09-03) | 50 pages (25,000 rows) |
| iCIMS | pr= page parameter against the tenant's own "Page X of Y" marker. | The declared page count is read out, or a page adds no new job ids. Verified live 2026-09-06 on Latham's US tenant: 3 pages of 50, 129 rows, matching the 129 job URLs in the tenant's own sitemap.xml — an independent count the source publishes. | 50 rows | 100 pages (5,000 rows) |
Two details of the WizeHire keyword search are worth knowing, because they are where a lesser reader would quietly lose rows.
It delivers the full number it advertises, because the page size is ours. Measured 2026-09-03: at 100 rows a page, attorney ended a four-page walk on 344 distinct rows against its own stated pagination.total: 352 — the keyset cursor drops rows at page boundaries. Asking for 500 in one request returned attorney 352/352, paralegal 181/181, legal 214/214 and law 176/176, whole. So the actor asks for 500, and if a walk still ends below the stated total it re-requests at a page size big enough to hold the feed and merges the results before reporting anything: manager (879 stated) came back 858 at 500 a page and all 879 after a re-request at first=879. A shortfall at a page size we chose is never reported as the source's ceiling.
What is and is not claimed about its page-size limits. The only size this API refuses outright is one above the maximum it states for itself: first=10000 answers 400 … "Number must be less than or equal to 9999", and that stated 9,999 is the only figure ever reported as the source's own maximum. Everything below it is accepted (measured first=301, 500, 880, 2000 and 9999 on manager, all HTTP 200), but a response carrying thousands of rows fails: on the unfiltered feed (stated total 5,461) requests for 1,000 and 1,500 rows were served whole every time, 2,000 was a flaky mix of 200 and 502, and 4,000 and above answered 503 every time. So the actor never asks for more than 1,500 rows in one response — the largest size measured reliably served — and pages by cursor above that. A gateway failure is never read as the source stating a limit, because it does not repeat.
FAQ
Is this legal? The actor reads public, logged-out job listings that firms publish specifically so people will find them — the same pages any visitor sees. It collects no personal data (no names, emails or candidate details; the rows are job postings), uses no logins or credentials, no proxies, no headless browser and no bot-wall evasion. It reads and obeys every host's robots.txt — a path the host disallows is never requested, and the run says so in a free robots-txt row — identifies itself honestly on every request (DockhandLegalJobsMonitor), and asks each host about one request a second. When a portal challenges the run instead of answering it, the run reports that in a free error row rather than routing around it. Whether and how you may use the data downstream is your responsibility under your own jurisdiction's rules and each portal's terms.
How fresh is the data? Each run reads the portals live at run time — nothing is cached between runs. Freshness is your schedule: an hourly schedule sees a posting within the hour it appears. posted_at carries the portal's own posting date where the portal publishes one (WizeHire and the newer viRecruit template); the others publish none, so monitor mode is the only reliable "what is new" signal for them.
What happens when a portal blocks or breaks? First, a request that never completed — a dropped connection, a timeout, or a 429/5xx — is asked again up to three times over a few seconds, because a momentary blip is not a broken portal. Only a portal that still will not answer gets a free error row, and the row's note names the host, how many attempts it took and the underlying reason (for example HTTP 405, or did not answer after 3 attempt(s) (ECONNRESET)). The other portals in your watchlist return normally and sources_failed in RUN_SUMMARY counts the failure.
A bot challenge is separated from a portal failure everywhere, because it is not the same event: it says the portal is refusing this client on this network, and says nothing about the openings it holds. Those rows name the challenge, are counted in sources_walled, and — this is the part that decides whether your schedule goes red — a run in which every portal was challenged still succeeds, with the wall written into the log and into RUN_SUMMARY. The run fails when every portal failed and at least one of them genuinely would not answer, or whenever you turn on failOnSourceError. Nothing is retried through proxies, and a portal that answers "no" is never asked twice.
What if a firm's robots.txt says no? Then that portal is not read, full stop — not the directory entry, not a pasted URL, not a host it redirects to. You get a free robots-txt row naming the path and the host, sources_robots_refused counts it in RUN_SUMMARY, the rest of your watchlist is unaffected and the run does not fail (failOnSourceError does not fire on it either: the portal did not fail, this actor declined). The rules are re-read on every run. As of 2026-09-08 that is two of the 45 directory portals, both on viRecruit/viGlobal: Dorsey & Whitney and Lowenstein Sandler. Both files say User-agent: * / Disallow: / and nothing else — the site plainly says no, so it is not asked. Thirteen other portals publish a file that contradicts itself, allowing the whole site in one group and forbidding it in another; a file like that is read, because that is what the standard says to do on a tie (RFC 9309 §2.2.2 takes the Allow) and it is how Google's own parser reads the same file (see Honest limits).
Does it work for firms outside the directory? Yes, on any of the five platforms: paste the portal URL. A firm on a platform this actor does not read yet (or whose portal you cannot find) is a portal request — open an issue on this actor and it will be answered within a day.
Do I get the full job description? No. Rows carry title, practice area, department, office, level and salary where published, and the apply link. Fetching descriptions would need one extra request per listing, multiplying run time and cost for a field most buyers do not need to watch.
I changed my firm list and everything came back. Why? The seen-list is keyed to the exact watchlist. A new list is a new watch, so its first run returns everything currently open once. Filters (titleIncludes, titleExcludes, locationIncludes) are not part of the key and can be changed without a reset.
How do I reset monitor mode? Delete the legal-jobs-monitor-state key-value store from your account's Storage. The next run behaves like a first run.
Why does a run cost nothing some days? Because there was nothing new to deliver. The only billable event is a job row delivered to you; status rows, empty portals and quiet days are free by design.
Need another source? Open an issue on this actor naming the firm or the platform. Requests are answered within a day; portals on the five supported platforms are usually added after a live probe.
Related
Watching companies outside the legal market too? Company Jobs Monitor: Greenhouse, Lever, Workday, Ashby & More is the sibling actor: the same only-new monitor design across Greenhouse, Lever, Ashby, SmartRecruiters, Workable, Recruitee, Workday, Phenom and Eightfold. Both are on our publisher page.
About Dockhand
Dockhand builds data tools that count every row against the source and say exactly what is missing. Questions or a platform request: open an issue on this actor.