California Labor Commissioner enforcement data with the dollar amounts: every judgment with its total owed, status, court and NAICS, the LC 2810.4 port drayage unsatisfied-judgment list with its history, the wage-claim index, and free debarment feeds from WA, NJ, CO, IL and OR.
README "Personal data" section now carries the FCRA framing: the dataset is
not a consumer report, the publisher is not a consumer reporting agency, and
no eligibility use is permitted; the actor republishes state records without
enrichment or edits, and corrections flow through from the source. Citation
checked against the live text of 15 U.S.C. 1681a.
docs/STORE-LISTING.md: the Apify Store listing package as published on
2026-09-30 (publication fields with character counts, PPE pricing with the
per-result price scaled to the corpus, Console defaults, sample dataset,
personal-data posture).
docs/assets/icon-golden-seal.png and .svg, the family icon shared with the
other CA DIR actors.
snapshots/ is git-ignored; it holds local full-sweep archives such as the
675,015-record acceptance baseline.
Changed
Project status files record the 0.3.0 acceptance run (675,015 records,
46 minutes, $3.54 measured cost, peak memory 566MB of 1024MB) and the Store
publication (PPE result price $0.00003, Console timeout 7200s).
[0.3.0] - 2026-09-29
Fixed
Full runs no longer hold the whole corpus in memory. Records mode streams
each slice and keeps only per-slice labelled counts and dedupe keys, so the
~670k-record corpus runs in 1024 MB instead of being OOM-killed at 512 MB.
A resumed run (resurrection or platform migration) no longer re-pushes
records. The 0.1.1 build re-enumerated the source on resume and wrote the
corpus roughly twice; verification now reconciles checkpointed slices by
their stored counts and never merges inside a checkpointed window.
The checkpoint is saved when each primary sweep finishes, on the platform's
MIGRATING and PERSIST_STATE events, and every 50 slices, under a lock with
detached snapshots so saves land in order. Once a migration begins the run
stops pushing, pushes only the verification merges its checkpoint covers
(recording their counts first), and reboots to resume instead of exiting.
The port drayage fetch honors the checkpoint, so a resumed run no longer
duplicates its rows.
An empty judgment status picklist now fails verification loudly instead of
vacuously reporting a match.
Colorado workbook tabs are validated per sheet (claim, employer, amount and
address columns) and skipped with a warning naming what is missing; a
judgment case with disagreeing amount rows now deterministically keeps the
largest amount.
Changed
Default actor memory is 1024 MB. Entities mode still holds the corpus and
needs 4096 MB, now documented in the input schema and README, along with
run-timeout guidance for full runs.
Checkpoints use schema version 2 (labelled per-slice counts, fully
validated); an old or malformed checkpoint logs a warning and starts fresh.
The CLI resume file is written every 50 slices and on exit, matching the
actor's checkpoint cadence.
[0.2.0] - 2026-09-29
Changed
Project status files record the 0.1.0 release, the Apify push (build 0.1.1,
actor mCxAWNmBS0krvFSgv), and the first full platform run: 670,559 unique
records measured (wage claims 629,776 back to 1997, judgments 31,027, all
six spine feeds match=True, zero slice errors).
TODO.md carries the 0.3.0 platform-fix plan the full run exposed: stream
records instead of retaining the corpus in memory (512MB OOM), resume-aware
verification (a resumed run re-pushed ~660k duplicate rows through the
union-merge path), a larger default memory plus run-timeout guidance, and
done_slices for fetch_port_drayage.
[0.1.0] - 2026-09-29
Added
Aura client for the CA DIR Lightning sites: form builder for the
unauthenticated aura.ApexAction.execute endpoint, framework-config harvester
that decodes the URL-encoded fwuid and loaded hash out of a site landing
page, self-healing from the framework id every response reports, one
re-harvest and retry on an unusable body, and a row-cap detector.
California judgment sweep with per-judgment dollar amounts: entry-date year
slices from 2000, a catch-all slice for everything before the corpus starts,
and a far-future slice that catches DLSE's data-entry typos in 2029 and 2205.
A row is one liable party, so a judgment naming three defendants keeps all
three, keyed on the judgment id plus the normalized employer name. The state's
own judgmentPartyId is not stable across queries or runs, so it is published
as an informational field only.
Port drayage unsatisfied-judgment list (Labor Code 2810.4) in one request,
plus an asOfDate back-fill that replays one snapshot per month and
reconstructs the whole add-and-remove history.
Wage-claim sweep, one request per docket day from 2010 plus one fixed
catch-all slice for the earlier corpus, which reaches back to 1997 and holds
about 31,000 records the catch-all collects in roughly 37 requests. Defaults to
defendant-side rows only, so the output does not carry the workers who filed
the claims. Narrowing the first docket date narrows the run: the catch-all is a
constant, never derived from the input, so it can never widen into the corpus.
Free multi-state spine: WA debarred contractors (form POST download; a plain
GET returns HTTP 500), NJ WALL spreadsheet, CO wage theft transparency sheet
read from the xlsx export across all nine tabs (the CSV export serves only
the active tab), IL debarred contractors parsed from free text, OR labor
contractor licensees found by the spreadsheet link the page publishes each
month, and Cal/OSHA penalties over $100,000.
Self-verification per dataset: judgments re-enumerated by judgment status over
exactly the same date windows as the primary sweep,
wage claims recounted by calendar month with an explicit zero-date-gap
assertion, the port drayage history checked for missing months, and every
feed checked for a parsed row count with its Content-Length and
Last-Modified. Records only the second pass found are merged into the
output. Records a status pass structurally cannot reach (no status at all, or
a status DLSE omits from its own picklist, such as Pending/Open) are counted
apart rather than reported as a coverage miss. The whole comparison is written
to the VERIFICATION key-value record, including when the picklist request
itself fails.
Cap handling: a response of exactly 5,000 rows is treated as the cap, the
slice is halved and re-fetched down to a single day, and a single day still at
the cap is recorded as an error instead of being silently truncated.
Two output modes: records (one flat row per source record, identical fields
across all ten sources) and entities (one row per employer keyed on a hash
of the normalized name and address, with total judgment amount, open judgment
count, and port drayage, WA debarment and NJ WALL flags).
Checkpoint and resume: completed slice keys are written to the key-value store
every 50 slices and on Apify's migration event, and a run resumes from them
when the input fingerprint still matches. A slice that failed is never
checkpointed as complete, and rows that no slice callback carried (feeds,
partially covered slices, verification merges) are emitted once at the end
from an explicit list rather than by position.
Apify actor wrapper (src/main.py), local CLI runner (src/cli.py), input,
dataset and output schemas, and an offline pytest suite that runs entirely
against saved fixtures through httpx.MockTransport.