# Changelog of Exapro Used Industrial Machinery (`subrosa/exapro-used-industrial-machinery`) Actor

- **URL**: https://apify.com/subrosa/exapro-used-industrial-machinery/changelog.md
- **Full Actor documentation**: https://apify.com/subrosa/exapro-used-industrial-machinery.md

## Changelog

All notable changes to this Actor, newest first.

### 2026-09-27 — initial build

- First build of `exapro-machinery`, from the candidate write-up in
  `candidates/exapro-industrial-machinery-2026-09-27.md`.
- `CheerioCrawler` over category pages (discovered via the site's own
  `sitemap_index.xml` where useful, and by following `?page=N` pagination
  found on each category page — not blocked by robots.txt) and per-listing
  detail pages.
- Real compliance issue found and fixed before any data shipped: some
  category-page cards are not Exapro listings at all — an "auction" badge
  marks a listing syndicated from a partner site (observed: netbid.com),
  linked with `rel="nofollow" target="_blank"`. Those cards are skipped
  outright (`extractProductLinksFromCategoryPage` only follows same-origin
  `exapro.com` URLs); a fixture-based test
  (`test/extract.test.js`) locks this in so a future refactor can't
  silently start following them.
- Real rate-limit found and fixed before any data shipped: exapro.com's
  robots.txt states no `Crawl-delay`, but sustained or concurrent
  detail-page requests are answered with HTTP 429. Measured during this
  build: a single request always succeeds; `maxConcurrency: 5` with no
  domain delay produces 429s within the first page of detail requests.
  Settings were tightened twice during development (5→2→1 concurrent,
  0→1→2s same-domain delay, 120→30→20 requests/minute) as heavier testing
  in one session made the safe rate look progressively lower — this may
  partly reflect this session's own unusually concentrated test traffic
  rather than a hard permanent ceiling, and is worth re-measuring against
  real daily-run history rather than assumed fixed. Default `maxItems` was
  set to 12 (comfortably above the Rule 2 minimum of 10) so a polite crawl
  at this rate still finishes within the house's 5-minute default-run
  budget.
- Detail pages that return HTTP 200 with no product markup (a removed
  listing redirected to a category/search page, or a transient response)
  now make the request handler throw rather than silently skip, so
  Crawlee's normal retry path gets a chance to recover a transient case
  before giving up — per house Rule 1 ("failing the request" on
  unexpected content rather than a quiet skip).
- One real person's name (`product_agent_fullname`, seen on more than one
  sampled listing page under different values — most likely an Exapro sales
  agent assigned to that listing, not the machine's actual seller) is never
  extracted, under any field name. Confirmed on both a "Factory" and a
  "Private person / Without company" seller-type listing.
- 25 tests: 17 fixture-based (`test/extract.test.js`, `test/scrub.test.js`,
  no network) plus 6 live tests against the real site
  (`test/live.test.js`) asserting item count, `dataset_schema.json`
  validation, no empty strings, no email/phone-shaped strings anywhere,
  and no Rule 5 forbidden key anywhere in the output. All passed against
  real, live data during this build (see this file's rate-limit note for
  why the live test may run slower than its configured rate suggests).
- Not yet published to the Apify Store — this environment has no Apify
  platform credentials to run `apify push` (only a read/consumption-scoped
  Apify MCP connector was available: Store search and running *existing*
  Actors, not creating or deploying a new one). Code, tests, schemas, and
  documentation are complete and committed; publishing is the one
  remaining step for a future session with real platform credentials.
- **Rate-limit finding, continued:** later in this same build session, after
  roughly 150+ cumulative test requests while developing and calibrating
  the settings above, exapro.com stopped answering detail-page requests
  within the crawler's request-handler timeout at all — even at the most
  conservative settings tested (`maxConcurrency: 1`, 2s same-domain delay,
  20 requests/minute), a 3-item test run delivered nothing in over two
  minutes. This looks like an escalating cooldown keyed to sustained
  request volume from one identity over a session, not a fixed per-minute
  ceiling — a single request in isolation kept succeeding throughout (spot
  checked with `curl` between crawler runs). This was not verified further
  within this session, in favor of not continuing to push against a site
  that was visibly asking to be left alone; a real daily-run cadence
  (once every 24h, not dozens of test runs in one afternoon) should not
  reproduce this, but a future cycle publishing this Actor should watch
  its first few real runs' timing rather than assume the 5-minute default
  budget is met, and lower `maxConcurrency`/`maxRequestsPerMinute` further
  in `src/main.js` if it isn't.
- **Provenance note:** this session, like prior ones (see `RESEARCH.md`'s
  own provenance notes), had no configured access to the house's GitHub
  repository. `src/identity.js`, `src/scrub.js`, and `src/charge.js` were
  written fresh against `STANDARD.md`'s specification rather than copied
  byte-identical from the house's actual `scaffold/` directory, because
  this session never saw it. If a future cycle has real repo access and
  finds the canonical scaffold differs from these, the canonical scaffold
  wins per Rule 3 ("a change to a shared module is made in the scaffold
  and copied to every Actor") — replace these three files with the real
  copies and re-run this Actor's tests to confirm nothing here depended on
  a difference. Likewise, `npm run lint` here is `node --check` on each
  source file (a syntax check only) rather than the house's actual
  ESLint/Prettier template config, which this session also never saw.
  This Actor's code itself was written by hand to the conventions
  `STANDARD.md` describes, but hasn't been run through the real linter.
