# Changelog of Substack Newsletter Sponsorship Prospect Finder (`datagrit/substack-sponsor-prospector`) Actor

- **URL**: https://apify.com/datagrit/substack-sponsor-prospector/changelog.md
- **Full Actor documentation**: https://apify.com/datagrit/substack-sponsor-prospector.md

## Changelog

### 0.8

- A post archive or recommendation list that arrives with HTTP 200 but a body that is not JSON (an HTML error page, for example) is now judged like any other failed read of that one publication: the publication is skipped without billing (archive) or gets `dataGaps: ["recommendations"]` (recommendations), it is counted in the same "most cannot be read" threshold (at least 5 answers, more than half failed) and never ends the run by itself. Before, one such response ended the whole run with exit 91: on the 9-publication fixture an unreadable archive at positions 1-6 gave 0 rows and at positions 7-9 six rows already billed in a failed run; an unreadable recommendation list at position 2 gave 0 rows. Now each of those runs exits 0 with the 8 or 9 healthy rows, and an unreadable body for every publication still fails the run with the message that names how many answered with an error or an unreadable body.
- `maxSubscribers` description and README now say, as `minSubscribers` did, that publications that hide their subscriber count are dropped when a limit is set.

### 0.7

- Silent addresses no longer decide whether the source is broken: the "most post archives or recommendation lists cannot be read" check now judges only requests that were answered (an error status or an unreadable body) and needs at least 5 of them. Before, the first 3 silent newsletters of a ranking (the first batch) ended the run with exit 91 and no rows, although the 22 others were healthy and had not been requested; with this version they cost their own 30 seconds each and the run delivers the ordered rows. Measured on a copy of the Actor with the silent addresses redirected to a local server that accepts connections and never answers (everything else live; default test input, technology, 25 per category, 20 results; healthy run 12.2 s, 20 rows): positions 1+2+3 silent exit 0 in 39.8 s with 20 rows (was exit 91, 0 rows); positions 1, 5, 8, 12, 14 silent exit 0 in 158.6 s with 20 rows (was 10); positions 1-6 silent exit 0 in 68.7 s with 19 rows (all healthy ones); the example input with positions 1+2+3 silent exit 0 in 45.1 s with 17 rows (all healthy ones). In `publications` mode, 6 silent addresses followed by 2 healthy ones now exit 0 in 58.9 s with the 2 rows and 6 status rows (was exit 91 and no rows).
- Total failure still ends the run honestly: 12 or more archive requests with no answer and none that worked, or no readable archive at the end of the run, fail with a message that names missing answers (timeouts or refused connections), not a changed source layout. Rows are not released before 5 answered archive requests.
- Run budget with a reserve: once the 150 s of failed-request time of the run is used up, an address that has not failed yet still gets one attempt of up to 8 seconds per request, and an address that has failed gets none. Before, six silent addresses used the whole pool and two healthy addresses behind them received no request at all.
- README and run summary: the 150 s are described as failed-request time summed over parallel requests (not clock time); the summary no longer claims that affected publications carry `dataGaps` (a publication without a readable archive has no row and is not billed) and reports how many archive addresses gave no answer.
- `differentiator`: the second ranking is `paid`, there is no `trending` mode.

### 0.6

- Failure budget per address: the time a run loses to failed requests is now limited per address (30 s), per Substack's shared ranking and recommendation host (75 s) and for the run as a whole (150 s). Before, one 75-second budget was shared by every address, so one newsletter whose site never answered used it up and the healthy publications queued behind it were skipped; the default test input (technology, 25 per category, 20 results) failed with exit 91 and no rows. A silent newsletter now costs about 30 seconds and only its own row.
- Honest reasons: a request that was never sent because a budget was already used up is no longer reported as a source failure. Status rows say "was not checked in this run ... a limit of this run, not a verdict on the address", the run summary counts such publications separately, and the fatal "could not read the post archive" message counts only archives that were actually requested and names how many more were never requested.
- Request tracking: an address whose first request already failed and was then abandoned stays "could not be reached or checked"; a publication whose archive request was never sent is skipped without billing with its own note.
- README: the rate-limit FAQ describes the three budgets.

### 0.5

- Every request has a hard 25-second deadline (the underlying HTTP client's own timeout does not fire when a connection is accepted and never answered), and the failure budget counts the time spent in failed attempts as well as the pauses between them.
- A domain that does not resolve (`ENOTFOUND`) is reported as a missing address after one attempt, without retries; an address that could not be reached or checked at all is reported as `could not be reached or checked in this run, so it is unknown whether it is a Substack publication` and the run summary lists it under "Could not be checked", instead of claiming that a publication exists.
- `dataGaps` description matches the code: a publication without a readable post archive is skipped and not billed, so `archive` never appears on a billed row.

### 0.4

- Rate limits: every pause of a run (retries after 429 and 5xx, `Retry-After` included) is drawn from one shared 75-second waiting budget held by the HTTP layer. A pause longer than what is left is not taken: the request is abandoned with an explanatory error, the publication gets `dataGaps` and the run summary counts the abandoned requests. Before, one 429 with `Retry-After: 120` cost up to 363 s on a single request and a mild rate limit on every request added up to over five minutes in the default test run. The run still fails (exit 91) when most post archives or recommendation lists cannot be read.
- Schema: the `paidSubscribersLabel` description now says what the code, README and 0.3 changelog say: the publication's own band comes first and the author's bestseller tier fills in only when there is none.
- README: the example shows `recommendedBy30dCapped`, the output section explains it, and the FAQ describes the waiting budget.

### 0.3

- Paid subscribers: the band is read from the numeric order of magnitude in the publication's own record (`rankingDetailOrderOfMagnitude`), which equals the ranking label whenever one is shown. The author's bestseller tier now only fills in when the publication has no band of its own and never overrides it (a publication with order 0 and a 10,000 badge is no longer reported as "Tens of thousands"). `paidSubscribersSource` has a third value, `ranking-order-of-magnitude`; the band 1 (single digits) is reported too.
- Removed `activeSponsorSlots` and `sponsorSlotDomains`: Substack's own sponsorship-campaign tool is used by almost no publication, so the fields were empty for nearly every row and did not show who sponsors a newsletter. The differentiator and README no longer promise it.
- Billing: a publication whose post archive could not be read is skipped and not billed, whether or not the recommendations were read (the row would otherwise carry null in seven of the nine measured fields). Nothing is emitted until at least five post archives have been tried, so a source that stops serving fails the run (exit 91) before the first billed row.
- Duplicate publications across categories are dropped before any archive request; a missing publication address no longer counts as a failed archive read.

### 0.2

- Profile: when the homepage redirects to a custom domain without profile data (for example Pirate Wires), the profile is read from the `/archive` page. A publication that exists but cannot be read is reported as such, separately from an address that is not a Substack publication.
- Paid subscribers: `paidSubscribersLabel` and `paidSubscribersAtLeast` no longer depend on the ranking context; the author's bestseller tier is the fallback and `paidSubscribersSource` says which one was used. New fields `bestsellerTier`, `activeSponsorSlots` and `sponsorSlotDomains`.
- Billing: a publication whose post archive and recommendations both failed is skipped and not billed. A publication with posts that returns an empty archive gets `dataGaps: ["archive"]` instead of being treated as dead.
- Failures: the run fails when the archive or the recommendations cannot be read for most publications. A zero result caused by the source is reported as a source problem, not as "raise maxPerCategory".
- Price lowered from $0.006 to $0.0045 per publication (2.26x the median of the top-3 competitors). The differentiator now acknowledges that the competitor also returns the recommendation network and the bestseller tier.
- Verified with three inputs (see the smoke report):
  - Typical: the technology leaderboard, nine publications, with paid band and source, cadence, medians and recommendations.
  - Edge: Pirate Wires profile from `/archive`, archive unreadable for one publication, empty archive with `has_posts`, bestseller-tier fallback and an unreadable profile next to a readable one.
  - No results: filters that match nothing, an activity filter with an unreadable archive and a non-existent address each return one unbilled status row with the reason.

### 0.1

- Initial release: Substack publications from category leaderboards, a list you provide or recommendation links, with subscribers, plan prices, posting cadence, engagement per 1,000 subscribers and new recommendations in the last 30 days.
- Verified with three inputs (see the smoke report):
  - Typical: the technology leaderboard, nine publications; every row has subscribers, prices, cadence from the 90-day archive, median reactions and recommendation counts.
  - Edge: subscriber range, paid plan with a price limit, activity filters, lookalikes of one publication and a publication without posts.
  - No results: filters that match nothing return one unbilled status row with `found: false`; when no requested publication exists the run fails with the reason.
