# Substack Newsletter Sponsorship Prospect Finder (`datagrit/substack-sponsor-prospector`) Actor

Substack publications ranked for sponsorship: subscriber count, paid-tier size, plan prices, posting cadence and engagement per 1,000 subscribers.

- **URL**: https://apify.com/datagrit/substack-sponsor-prospector.md
- **Developed by:** [datagrit](https://apify.com/datagrit) (community)
- **Categories:** Marketing, Lead generation, Social media
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Substack Newsletter Sponsorship Prospect Finder do?

Substack Newsletter Sponsorship Prospect Finder turns Substack category rankings, a list of newsletters you name, or the recommendation links of a newsletter into one analyzed row per publication. Each row combines the public profile (subscriber count, paid-subscriber label, monthly, annual and founding prices) with numbers measured from the publication's own post archive: posts in the last 30 and 90 days, posts per week, the share of paid-only posts, median reactions and comments per post, reactions per 1,000 subscribers, and how many newsletters started recommending it in the last 30 days.
Export the result as JSON, CSV or Excel, call it through the Apify API, or plug it into n8n, Make and AI agents through MCP.

### Why use Substack Newsletter Sponsorship Prospect Finder?

- **Build a sponsorship shortlist.** Rank newsletters in a niche by audience size, engagement per 1,000 subscribers and posting consistency instead of guessing from subscriber counts alone.
- **Find growing newsletters early.** The recommendation count of the last 30 days shows which publications are gaining readers through other writers, before they reach the top of a ranking.
- **Qualify outreach lists.** Filter by subscriber range, paid plan, monthly price, language, days since the last post and posts per month, so you contact only active publications that fit your budget.
- **Expand from newsletters you already like.** Give one or more publications and get the newsletters they recommend, each with the same metrics.
- **Track the market over time.** Schedule the same input weekly and compare ranks, subscriber counts and cadence between runs.

### Example output

| name | subscribers | monthlyPriceUsd | rank | posts30d | postsPerWeek | reactionsPer1kSubscribers | recommendedBy30d |
|---|---|---|---|---|---|---|---|
| The Pragmatic Engineer | 1100000 | 15 | 2 | 12 | 2.41 | 0.11 | 41 |

```json
{
  "publicationId": 458709,
  "name": "The Pragmatic Engineer",
  "url": "https://newsletter.pragmaticengineer.com",
  "language": "en",
  "authorName": "Gergely Orosz",
  "subscribers": 1100000,
  "paidSubscribersLabel": "Tens of thousands of paid subscribers",
  "paidSubscribersAtLeast": 10000,
  "paidSubscribersSource": "ranking-label",
  "bestsellerTier": 10000,
  "paymentsState": "enabled",
  "monthlyPriceUsd": 15,
  "annualPriceUsd": 150,
  "annualDiscountPercent": 17,
  "foundVia": "leaderboard",
  "category": "technology",
  "rank": 2,
  "lastPostAt": "2026-09-30T16:30:05.153Z",
  "daysSinceLastPost": 0,
  "posts30d": 12,
  "posts90d": 31,
  "postsPerWeek": 2.41,
  "paidPostShare90d": 0.61,
  "medianReactionsPerPost": 122,
  "medianCommentsPerPost": 7,
  "reactionsPer1kSubscribers": 0.11,
  "recommendedBy30d": 41,
  "recommendedBy30dCapped": false,
  "dataGaps": [],
  "scrapedAt": "2026-10-01T04:07:14.402Z"
}
```

### How much does it cost?

You pay a small fee per run and a price per analyzed publication that is lower on paid Apify plans. The Apify free plan includes monthly credit to try it. Set a maximum spend on the run and the Actor stops when it is reached.
Only publications that pass your filters are billed. The Actor uses plain HTTP requests, so runs are fast and cheap on platform resources.

### Input

- **Source** – read a category leaderboard, analyze publications you list, or find lookalikes of publications you list.
- **Categories** and **Leaderboard type** – which of the 32 Substack categories to read, and whether to use the ranking of all publications or of publications with paid plans.
- **Publications to scan per category** – how many ranking positions to read before filters apply.
- **Publications** and **Seed publications** – Substack URLs, custom domains or subdomains such as `lenny`.
- **Filters** – minimum and maximum subscribers (publications that hide their subscriber count are dropped when either limit is set), only publications with paid plans, maximum monthly price, language, posted within days, minimum posts in 30 days.
- **Maximum results** – total limit of publications in the run.
- **Proxy configuration** – optional; only needed if Substack blocks your region.

### Output fields

Each record contains the profile fields (`publicationId`, `name`, `url`, `language`, `authorName`, `subscribers`, `paidSubscribersLabel`, `paidSubscribersAtLeast`, `paidSubscribersSource`, `bestsellerTier`, `paymentsState`, plan prices), the origin of the row (`foundVia`, `category`, `rank`, `lookalikeOf`), the measured activity fields (`lastPostAt`, `daysSinceLastPost`, `posts30d`, `posts90d`, `postsPerWeek`, `paidPostShare90d`, `medianReactionsPerPost`, `medianCommentsPerPost`, `reactionsPer1kSubscribers`), the recommendation fields (`recommendedBy30d`, `recommendedBy30dCapped`, `recommendedByLatestAt`; `recommendedBy30dCapped` is `true` when the 30-day window may extend beyond the recommendations Substack returned, so read the count as a lower bound), `dataGaps`, `sourceUrl` and `scrapedAt`.
Substack exposes the paid-subscriber count only as a band such as "Thousands of paid subscribers", so `paidSubscribersAtLeast` is a lower bound. The band comes from the publication's own record: its label when Substack shows one (`paidSubscribersSource` = `ranking-label`), or the numeric order of magnitude in the same record when the label is not shown (`ranking-order-of-magnitude`). Only when the publication has no band of its own does the author's bestseller badge fill in (`author-bestseller-tier`); the badge is set per author and can sit one step above or below the publication, so it never overrides the publication's band and is always returned separately as `bestsellerTier`. Check `paidSubscribersSource` before relying on a figure. The Actor does not report who sponsors a newsletter.
Publications that hide their subscriber count have `subscribers` empty. When the recommendations of one publication cannot be read, its recommendation fields stay empty and `dataGaps` names what is missing; a publication whose post archive cannot be read is skipped without billing. If the archive or the recommendations cannot be read for most publications, the run fails instead of returning empty columns.
Rows with a status instead of data (`found: false`) are never billed: a requested address that does not exist or is not a Substack publication, an address that could not be reached or checked in this run (the note says it is unknown whether it is a Substack publication), a publication confirmed to exist whose profile could not be read, and a publication whose post archive could not be read, because the activity metrics are what a row is billed for. The `note` of each row and the run summary give the reason, and a run that returns nothing because the source could not be read says so instead of blaming your filters.

### Is it legal to scrape this data?

The Actor reads only publicly available information and does not log in or bypass access controls. You are responsible for using the data in line with applicable laws
(including data protection rules) and the source's terms. If you find an issue, open it in the Issues tab; problems are answered within one business day.

### FAQ

**How fresh is the data?** Every run reads Substack live. Ranks change between runs, so the same leaderboard input can return slightly different publications.

**How is the posting cadence calculated?** From the post archive, over the last 90 days; for publications younger than 90 days the window is their age. Reactions are medians over posts older than three days, so a fresh post does not distort them.

**Why is a number empty?** Either it could not be read in this run, which `dataGaps` and the run summary name, or Substack shows no figure for that publication. An empty `paidSubscribersLabel` means the publication's record carries no paid-subscriber band (its order of magnitude is 0 or absent) and the author has no bestseller badge to fall back on; it does not mean the publication has no paid subscribers.

**What happens when Substack rate-limits the run, or one newsletter's site stops answering?** The Actor retries failed requests with a growing pause and follows the `Retry-After` header, but the time a run loses to failures is limited in three layers: 30 seconds per newsletter address (the pauses between attempts plus the time spent in attempts that failed or timed out; each attempt waits at most 25 seconds for an answer), 75 seconds for Substack's own ranking and recommendation service, and 150 seconds of failed-request time for the whole run, counted as a sum over requests that run in parallel, so it is used up in less than 150 seconds of clock time. Healthy requests never count. One newsletter whose site does not answer therefore costs at most about 30 seconds of that budget and affects only that newsletter. Newsletters that have not failed yet are protected from the others: once the run budget is used up, every address that has not failed in this run still gets one attempt of up to 8 seconds per request, and an address that has failed gets no further request and no pause that no longer fits. When a post archive stays unread, the publication is skipped without billing and, in `publications` mode, gets a status row. A post archive or recommendation list that arrives as something other than JSON (for example an HTML error page) is treated like any other failed read of that one newsletter: a publication without a readable archive is skipped without billing, one without a readable recommendation list gets `dataGaps: ["recommendations"]`, and neither ends the run. A rate-limited run, or a source that accepts connections and never answers, therefore ends in minutes, not hours. The run summary and the status rows say which case it was: a request that failed is reported as a source or network problem, a publication for which no request was sent because the budget was already used up is reported as a limit of the run, not as a problem of the address. The run fails with a clear message in three cases only: most post archives or recommendation lists that were answered with an error or an unreadable body (judged from at least 5 answers, so a few silent sites at the start of a ranking do not end the run), 12 or more archive requests that got no answer at all and not one that worked, or, at the end of the run, no post archive readable at all. Addresses that do not answer are never counted as a sign that the source changed. An address whose domain does not exist is reported at once, without retries.

**Can I schedule runs?** Yes, use Apify schedules or call the Actor from your own workflow.

**Something looks wrong.** Open an issue with the input you used; layout changes at the source are fixed quickly.

### Related Actors

See other data Actors from the same publisher on the Store profile.

# Changelog

This Actor's version history is a separate document: https://apify.com/datagrit/substack-sponsor-prospector/changelog.md

# Actor input Schema

## `source` (type: `string`):

Where the publications come from. "leaderboard" reads the Substack category rankings, "publications" analyzes the publications you list, "lookalikes" finds publications that the ones you list recommend.

## `categories` (type: `array`):

Leaderboard categories to read when the source is "leaderboard". Available: culture, technology, business, us-politics, finance, food, sports, art, world-politics, health-politics, news, fashionandbeauty, music, faith, climate, science, literature, fiction, health, design, travel, parenting, philosophy, comics, international, crypto, history, humor, education, film-and-tv, home-garden, games.

## `leaderboardType` (type: `string`):

Which ranking to read: "all" is the ranking of all publications of the category, "paid" is the ranking of publications with paid subscriptions.

## `maxPerCategory` (type: `integer`):

How many leaderboard positions to scan per category before filters apply. One page of the ranking holds 25 positions.

## `publications` (type: `array`):

Used when the source is "publications": Substack URLs, custom domains or subdomains such as "lenny".

## `lookalikesOf` (type: `array`):

Used when the source is "lookalikes": publications whose recommendations are followed to find similar newsletters. Accepts URLs, custom domains or subdomains.

## `minSubscribers` (type: `integer`):

Keep publications with at least this many displayed subscribers. Zero disables the filter. Publications that hide their count are dropped when a limit is set.

## `maxSubscribers` (type: `integer`):

Keep publications with at most this many displayed subscribers. Zero disables the filter. Publications that hide their count are dropped when a limit is set.

## `requirePaidPlan` (type: `boolean`):

Keep only publications that have paid subscriptions switched on and a price.

## `maxMonthlyPriceUsd` (type: `number`):

Keep publications whose regular monthly plan costs at most this many US dollars. Zero disables the filter. Publications without a monthly plan are dropped when a limit is set.

## `language` (type: `string`):

Keep publications whose language code starts with this value, for example "en" or "de". Empty keeps all languages.

## `activeWithinDays` (type: `integer`):

Keep publications whose newest post is at most this many days old. Zero disables the filter.

## `minPosts30d` (type: `integer`):

Keep publications that published at least this many posts in the last 30 days. Zero disables the filter.

## `maxItems` (type: `integer`):

Stop after this many publications in total.

## `proxyConfiguration` (type: `object`):

Optional proxy. Leave disabled unless Substack blocks datacenter traffic; residential proxy raises the platform cost of the run.

## Actor input object example

```json
{
  "source": "leaderboard",
  "categories": [
    "technology"
  ],
  "leaderboardType": "all",
  "maxPerCategory": 25,
  "publications": [],
  "lookalikesOf": [],
  "minSubscribers": 0,
  "maxSubscribers": 0,
  "requirePaidPlan": false,
  "maxMonthlyPriceUsd": 0,
  "language": "",
  "activeWithinDays": 0,
  "minPosts30d": 0,
  "maxItems": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All extracted records as a dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "source": "leaderboard",
    "categories": [
        "technology"
    ],
    "leaderboardType": "all",
    "maxPerCategory": 25,
    "publications": [],
    "lookalikesOf": [],
    "minSubscribers": 0,
    "maxSubscribers": 0,
    "requirePaidPlan": false,
    "maxMonthlyPriceUsd": 0,
    "language": "",
    "activeWithinDays": 0,
    "minPosts30d": 0,
    "maxItems": 20,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagrit/substack-sponsor-prospector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "source": "leaderboard",
    "categories": ["technology"],
    "leaderboardType": "all",
    "maxPerCategory": 25,
    "publications": [],
    "lookalikesOf": [],
    "minSubscribers": 0,
    "maxSubscribers": 0,
    "requirePaidPlan": False,
    "maxMonthlyPriceUsd": 0,
    "language": "",
    "activeWithinDays": 0,
    "minPosts30d": 0,
    "maxItems": 20,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("datagrit/substack-sponsor-prospector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "source": "leaderboard",
  "categories": [
    "technology"
  ],
  "leaderboardType": "all",
  "maxPerCategory": 25,
  "publications": [],
  "lookalikesOf": [],
  "minSubscribers": 0,
  "maxSubscribers": 0,
  "requirePaidPlan": false,
  "maxMonthlyPriceUsd": 0,
  "language": "",
  "activeWithinDays": 0,
  "minPosts30d": 0,
  "maxItems": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call datagrit/substack-sponsor-prospector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagrit/substack-sponsor-prospector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/O8EMl9s39i47ngShd/builds/b4ZHxGLWwbR1izaX3/openapi.json
