All notable changes to this Actor, newest first.
- Initial build. CheerioCrawler over DOTmed.com's public category and
item pages. Extraction verified against live fixtures captured the
same day (see
test/fixtures/).
- Discovered during build that DOTmed's own pagination (
?offset=) is
disallowed by robots.txt (Disallow: /browse/*?*) — the crawl design
uses breadth across category/subcategory entry points instead of
query-string pagination. Documented in README's "Limits and honesty".
- Discovered DOTmed answers bursts of detail-page requests faster than
~1 every 3 seconds with HTTP 429;
maxConcurrency set to 1 and
maxRequestsPerMinute to 20 accordingly.
specifications, sellerLocation, sellerType, auctionEndsAt, and
listedAt shipped as reserved (always-null) fields — the source does
not expose them reliably enough to fill without guessing.
- First real platform run (through Apify's shared datacenter proxy pool)
came back with zero items — not a code bug: the crawler got a fast
HTTP 200 with no embedded product data. A same-input run with the
proxy off got full data immediately. DOTmed silently omits the
embedded JSON-LD product data for requests from Apify's shared
datacenter IP pool while serving it normally to a direct connection
under the same declared identity. Default
proxyConfiguration
changed to no proxy (direct) accordingly — this is a data-availability
finding, not bot evasion; nothing about the identity or request
changes, only which network path it goes out on.