# Changelog of Clutch Scraper: Agency Directory (`danthedataman/clutch-agency-directory`) Actor

- **URL**: https://apify.com/danthedataman/clutch-agency-directory/changelog.md
- **Full Actor documentation**: https://apify.com/danthedataman/clutch-agency-directory.md

## Changelog

### 0.6 — 2026-09-10

- Local fallbacks for `maxRetriesPerRequest` and `rotateAfterRequests` now
  match the schema and README defaults (20 and 1), so an API caller who omits
  the fields gets the documented retry ladder instead of the 0.3-era 8/24
  values.
- Pricing sentence scoped: the run's own compute and its result writes are
  covered by the event prices; storage retention and reads of the dataset
  after the run are billed according to the customer's plan.

### 0.5 — 2026-09-10

- **README and Actor schemas aligned with the shipped contract.** Input-table
  defaults now match `input_schema.json` for every documented field, including
  `maxRetriesPerRequest` 20, `rotateAfterRequests` 1, empty `sortBy`, and the
  RESIDENTIAL proxy group. Pricing names the automatic, memory-scaled Actor
  Start event as a separate charge from result rows; platform usage is still
  not billed on top, and failed pages are still never billed as rows. A
  key-value store schema and an output schema declare the run's `ERRORS`
  record.
- Partial-failure wording corrected: an unreadable page is written to `ERRORS`
  and that directory's walk stops; remaining directories still land; a failed
  location lookup is logged and is not written to `ERRORS`.

### 0.4 — 2026-07-26

- **Retry ladder widened to 20 with per-request IP rotation.** Clutch refuses a
  large share of residential proxy exits, so the previous 8-retry ladder failed
  half of all platform runs. Measured: 8 retries succeeded 2 of 4 runs; 20
  retries with per-request rotation succeeded 3 of 3, typically in under 30s.

### 0.3 — 2026-07-26

- **Default to the RESIDENTIAL proxy group.** Clutch answers 403 to Apify's
  datacenter proxy IPs, so the previous default made every out-of-the-box run
  fail — including the daily platform health check. Measured on the platform:
  datacenter 403s all eight attempts; residential returns 50 companies in 10s.

### 0.2 — 2026-07-26

- **Fixed silent data loss.** Clutch's pagination is 1-based with page 1 at the
  bare URL; the walk treated the loop index as the page number, which requested
  page 2 twice and stopped one page early. A 216-match query returned 200 rows
  and exited green. It now returns all 216.
- `listingPage` was one higher than the page the row actually came from for
  every row past the first hundred, contradicting `sourceUrl` in the same row.
- `sortBy` now defaults to sending no sort parameter at all. Its previous
  default was Clutch's own ordering, so it changed no results while putting
  every default run on a path Clutch's robots.txt disallows.
- De-duplication is scoped to a single directory walk. It was global, so a
  company listed in two directories you asked for was returned only once.
- A non-numeric rating can no longer abort a run whose earlier pages are
  already billed; invalid `maxItems` ends with a status message, not a
  traceback.

### 0.1 — 2026-07-26

Initial release.

- Scrapes any Clutch.co directory: company name, Clutch profile, website,
  rating, review count, minimum project size, hourly rate, employee count,
  location, postal address, phone, service mix and review summary.
- Filters that map to Clutch's own: location, sort order, minimum reviews,
  project budget, hourly rate band, company size, service lines, industry focus
  and verified-only. Any filter not exposed as a field can be passed by pasting a
  filtered directory URL into `startUrls`.
- Rows are de-duplicated on the Clutch company id, because the site serves one
  block of companies under more than one page number — a repeated page never
  becomes a repeated or repeatedly billed row.
- Sponsored placements are labelled (`isSponsored`) rather than silently mixed in
  with ranked listings.
- The displayed listing location and the listing's structured postal address are
  kept as separate fields, because they disagree for multi-office companies on a
  location-filtered run.
- Resilient transport: browser-grade sessions, exponential backoff with jitter,
  and proactive proxy/session rotation before the site throttles an identity.
- Failed pages go to the run's `ERRORS` record, never into the billable dataset;
  an empty result finishes cleanly with guidance instead of a stack trace.
