Returns used industrial machinery for sale on Exapro.com, a European B2B marketplace, with title, category, year, asking price, condition, country, and the seller's business type and account ID, and no seller names, emails or phone numbers.
First build of exapro-machinery, from the candidate write-up in
candidates/exapro-industrial-machinery-2026-09-27.md.
CheerioCrawler over category pages (discovered via the site's own
sitemap_index.xml where useful, and by following ?page=N pagination
found on each category page — not blocked by robots.txt) and per-listing
detail pages.
Real compliance issue found and fixed before any data shipped: some
category-page cards are not Exapro listings at all — an "auction" badge
marks a listing syndicated from a partner site (observed: netbid.com),
linked with rel="nofollow" target="_blank". Those cards are skipped
outright (extractProductLinksFromCategoryPage only follows same-origin
exapro.com URLs); a fixture-based test
(test/extract.test.js) locks this in so a future refactor can't
silently start following them.
Real rate-limit found and fixed before any data shipped: exapro.com's
robots.txt states no Crawl-delay, but sustained or concurrent
detail-page requests are answered with HTTP 429. Measured during this
build: a single request always succeeds; maxConcurrency: 5 with no
domain delay produces 429s within the first page of detail requests.
Settings were tightened twice during development (5→2→1 concurrent,
0→1→2s same-domain delay, 120→30→20 requests/minute) as heavier testing
in one session made the safe rate look progressively lower — this may
partly reflect this session's own unusually concentrated test traffic
rather than a hard permanent ceiling, and is worth re-measuring against
real daily-run history rather than assumed fixed. Default maxItems was
set to 12 (comfortably above the Rule 2 minimum of 10) so a polite crawl
at this rate still finishes within the house's 5-minute default-run
budget.
Detail pages that return HTTP 200 with no product markup (a removed
listing redirected to a category/search page, or a transient response)
now make the request handler throw rather than silently skip, so
Crawlee's normal retry path gets a chance to recover a transient case
before giving up — per house Rule 1 ("failing the request" on
unexpected content rather than a quiet skip).
One real person's name (product_agent_fullname, seen on more than one
sampled listing page under different values — most likely an Exapro sales
agent assigned to that listing, not the machine's actual seller) is never
extracted, under any field name. Confirmed on both a "Factory" and a
"Private person / Without company" seller-type listing.
25 tests: 17 fixture-based (test/extract.test.js, test/scrub.test.js,
no network) plus 6 live tests against the real site
(test/live.test.js) asserting item count, dataset_schema.json
validation, no empty strings, no email/phone-shaped strings anywhere,
and no Rule 5 forbidden key anywhere in the output. All passed against
real, live data during this build (see this file's rate-limit note for
why the live test may run slower than its configured rate suggests).
Not yet published to the Apify Store — this environment has no Apify
platform credentials to run apify push (only a read/consumption-scoped
Apify MCP connector was available: Store search and running existing
Actors, not creating or deploying a new one). Code, tests, schemas, and
documentation are complete and committed; publishing is the one
remaining step for a future session with real platform credentials.
Rate-limit finding, continued: later in this same build session, after
roughly 150+ cumulative test requests while developing and calibrating
the settings above, exapro.com stopped answering detail-page requests
within the crawler's request-handler timeout at all — even at the most
conservative settings tested (maxConcurrency: 1, 2s same-domain delay,
20 requests/minute), a 3-item test run delivered nothing in over two
minutes. This looks like an escalating cooldown keyed to sustained
request volume from one identity over a session, not a fixed per-minute
ceiling — a single request in isolation kept succeeding throughout (spot
checked with curl between crawler runs). This was not verified further
within this session, in favor of not continuing to push against a site
that was visibly asking to be left alone; a real daily-run cadence
(once every 24h, not dozens of test runs in one afternoon) should not
reproduce this, but a future cycle publishing this Actor should watch
its first few real runs' timing rather than assume the 5-minute default
budget is met, and lower maxConcurrency/maxRequestsPerMinute further
in src/main.js if it isn't.
Provenance note: this session, like prior ones (see RESEARCH.md's
own provenance notes), had no configured access to the house's GitHub
repository. src/identity.js, src/scrub.js, and src/charge.js were
written fresh against STANDARD.md's specification rather than copied
byte-identical from the house's actual scaffold/ directory, because
this session never saw it. If a future cycle has real repo access and
finds the canonical scaffold differs from these, the canonical scaffold
wins per Rule 3 ("a change to a shared module is made in the scaffold
and copied to every Actor") — replace these three files with the real
copies and re-run this Actor's tests to confirm nothing here depended on
a difference. Likewise, npm run lint here is node --check on each
source file (a syntax check only) rather than the house's actual
ESLint/Prettier template config, which this session also never saw.
This Actor's code itself was written by hand to the conventions
STANDARD.md describes, but hasn't been run through the real linter.