# Substack Newsletter Directory & Sponsor Discovery (`automation_studio/substack-sponsor-scraper`) Actor

Discover high-engagement Substack newsletters, average likes/comments, verified sponsor history, promotional ad links, and creator contact info across 30+ niches.

- **URL**: https://apify.com/automation\_studio/substack-sponsor-scraper.md
- **Developed by:** [Asad Naeem](https://apify.com/automation_studio) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 verified newsletter & sponsor leads

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 📬 Substack Newsletter Directory & Sponsor Discovery

> **Discover high-growth Substack newsletters, analyze audience engagement, and instantly identify existing sponsors and creator contact details.**

***

### 🚀 The Opportunity: Newsletter Sponsorship Intelligence

Newsletter advertising is one of the highest-converting marketing channels in the world today. Modern decision-makers, software engineers, founders, and investors read curated newsletters like *The Pragmatic Engineer*, *Lenny's Newsletter*, *Not Boring*, and *Platformer* rather than scrolling social feeds.

However, **Substack has no public sponsor marketplace or advertiser directory**. Brands, B2B SaaS growth teams, and PR agencies typically spend hundreds of hours manually opening newsletters, skimming posts, checking links, and trying to find author contacts.

**Substack Newsletter Directory & Sponsor Discovery** automates this entire pipeline:

1. **Discovers Top Newsletters** across 35+ Substack categories (Technology, Business, Finance, Crypto, Health, AI, etc.) or custom leaderboards.
2. **Audits Audience Scale & Tier**: Extracts subscriber badges (*"Tens of thousands of paid subscribers"*, *"Over 100,000 subscribers"*).
3. **Calculates Real Engagement**: Analyzes recent posts for average likes, comments, and publishing cadence.
4. **Detects Active Sponsors & Brands**: Scans post bodies and outbound links for UTM sponsor tracking, affiliate links, and brand mentions (e.g. Sentry, Google Cloud, Notion, HubSpot).
5. **Extracts Author & Contact Details**: Captures writer names, handles, Twitter/X profiles, and direct publication URLs.
6. **Computes Sponsorship Feasibility Score**: Automatically rates each publication as `High`, `Medium`, or `Niche/Emerging` based on engagement velocity and sponsorship history.

***

### ✨ Key Features

| Feature | Description |
| :--- | :--- |
| 🏷️ **35+ Curated Categories** | Target Tech, Business, Finance, AI, Marketing, Health, Crypto, Food, Politics, and more. |
| 👥 **Audience & Paid Tier Badges** | Filter by subscriber tiers (e.g., minimum 1,000+ or 10,000+ paid subscribers). |
| 📈 **Live Engagement Benchmarks** | Computes exact average likes and comments per post across recent issues. |
| 🔍 **Automated Sponsor Detection** | Identifies companies already sponsoring the newsletter with brand domains, promo UTMs, and anchor text. |
| 🎯 **Sponsorship Readiness Scoring** | AI-informed heuristic scoring based on sponsor frequency and community activity. |
| 🐦 **Social & Author Contacts** | Extracts author names, handles, Twitter/X profiles, and custom domain endpoints. |
| 💸 **Cost-Effective Pay-Per-Event (PPE)** | You only pay for what you extract—$0.002 to initiate + $0.004 per enriched publication lead. |

***

### 📥 Input Parameters

| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `category` | `string` | `"technology"` | Category to explore (e.g., `technology`, `business`, `finance`, `crypto`, `health-and-wellness`, `all`). |
| `rankingType` | `string` | `"all"` | Leaderboard ranking filter: `"all"` (All publications), `"paid"` (Top Grossing Paid), or `"free"` (Top Free). |
| `maxPublications` | `integer` | `25` | Maximum number of publications to analyze and extract. |
| `samplePostsCount` | `integer` | `5` | Number of recent posts to inspect for engagement calculations and sponsor link discovery (1–20). |
| `detectSponsors` | `boolean` | `true` | Deep-crawl recent posts to detect outbound sponsor links, UTM campaigns, and brand partnerships. |
| `minSubscribersTier` | `string` | `"all"` | Minimum subscriber tier filter: `"all"`, `"hundreds"`, `"thousands"`, or `"tens_of_thousands"`. |
| `proxyConfiguration` | `object` | `{ "useApifyProxy": true }` | Proxy settings. Default Apify datacenter proxies ensure 100% reliable uptime. |

***

### 📤 Sample Output

Each output dataset item contains comprehensive publication intelligence, author metadata, engagement metrics, and identified sponsors:

```json
{
  "publicationId": 25779,
  "subdomain": "blog.pragmaticengineer.com",
  "name": "The Pragmatic Engineer",
  "url": "https://blog.pragmaticengineer.com",
  "customDomain": "blog.pragmaticengineer.com",
  "logoUrl": "https://substackcdn.com/image/fetch/w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F09199342-d619-487c-86e5-2ae34969fa4a_1024x1024.png",
  "heroText": "Big tech and startups: from the inside.",
  "about": "The #1 technology newsletter on Substack. Pragmatic, actionable advice for software engineers, engineering managers and tech leads. Highly rated advice on career growth, engineering culture, and tech trends.",
  "primaryCategory": "technology",
  "subscriberCountBadge": "Over 700,000 subscribers",
  "rankingPosition": 1,
  "rankingType": "all",
  "author": {
    "name": "Gergely Orosz",
    "handle": "gergelyorosz",
    "twitter": "GergelyOrosz",
    "profileUrl": "https://substack.com/@gergelyorosz"
  },
  "metrics": {
    "samplePostsAnalyzed": 5,
    "avgLikesPerPost": 342.6,
    "avgCommentsPerPost": 58.2,
    "avgEngagementScore": 400.8,
    "latestPostDate": "2026-09-08T14:15:00.000Z"
  },
  "sponsorshipFeasibility": {
    "score": "High",
    "reason": "Active sponsor detected across multiple issues with high engagement (>300 interactions/post)",
    "hasActiveSponsors": true,
    "detectedSponsorsCount": 3
  },
  "detectedSponsors": [
    {
      "brand": "sentry",
      "sampleUrl": "https://sentry.io/welcome/?utm_source=pragmaticengineer&utm_medium=newsletter&utm_campaign=q3",
      "anchorText": "Sentry: Error Tracking and Performance Monitoring",
      "postTitle": "Engineering Hiring Trends for 2026"
    },
    {
      "brand": "antithesis",
      "sampleUrl": "https://antithesis.com/?ref=pragmaticengineer",
      "anchorText": "Antithesis: Autonomous Software Testing",
      "postTitle": "How Big Tech Tests Distributed Systems"
    }
  ],
  "recentPosts": [
    {
      "title": "Engineering Hiring Trends for 2026",
      "url": "https://blog.pragmaticengineer.com/p/engineering-hiring-trends-2026",
      "publishedAt": "2026-09-08T14:15:00.000Z",
      "likes": 412,
      "comments": 74
    }
  ],
  "scrapedAt": "2026-09-09T23:30:00.000Z"
}
```

***

### 💼 High-ROI Use Cases

#### 1. B2B SaaS Growth & Performance Marketing

Stop guessing which newsletters your target buyers read. Search your industry category (e.g. `technology` or `finance`), extract who is currently sponsoring them, and reach out to the authors with relevant proposals.

#### 2. PR & Influencer Marketing Agencies

Build proprietary sponsor databases for your clients. Pitch your clients' products to verified newsletter creators with audience tier proof and historical engagement numbers.

#### 3. Competitor Intelligence

Want to know where your competitors are placing their ad spend? Run this scraper to discover which newsletters feature competitor UTM referral links.

#### 4. Creator Partnerships & Podcast Guests

Identify top voices in any niche with verified Twitter and Substack profiles for podcast invitations, co-marketing webinars, or syndicated guest posts.

***

### 💰 Pay-Per-Event (PPE) Pricing

This Actor utilizes transparent, pay-per-event pricing. You never pay for idle container run time or empty queries:

- **Actor Start Event**: `$0.002` per run execution.
- **Verified Publication Dataset Item**: `$0.004` per fully enriched publication lead (including sponsor detection and metrics).

*Example*: Extracting 50 verified publications with sponsor intelligence costs only **~$0.20**!

***

### 🛡️ Best Practices & Proxy Configuration

- Substack employs standard rate limiting. For best performance, keep the default **Apify Proxy** enabled.
- Setting `samplePostsCount` between **3 and 5** provides optimal speed and high sponsor discovery accuracy.

# Actor input Schema

## `category` (type: `string`):

Select the target content vertical. Scrapes publications directly from Substack's official leaderboards.

## `rankingType` (type: `string`):

Choose which segment of publications to target based on Substack's revenue rankings.

## `maxPublications` (type: `integer`):

Total number of newsletters to extract and analyze in this run.

## `postsPerPublication` (type: `integer`):

Number of recent issues analyzed per publication to calculate engagement and detect sponsor blocks.

## `onlyWithSponsors` (type: `boolean`):

When enabled, outputs only newsletters that have proven active sponsors or promotional partner links.

## `minAvgLikes` (type: `number`):

Filter for high-engagement publications with at least this many average reader likes / reactions per post.

## `proxyConfiguration` (type: `object`):

Use Apify Proxy to ensure uninterrupted scraping without rate limiting.

## Actor input object example

```json
{
  "category": "technology",
  "rankingType": "paid",
  "maxPublications": 50,
  "postsPerPublication": 10,
  "onlyWithSponsors": false,
  "minAvgLikes": 0,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset containing publications, author profiles, average likes, comment volume, and detected sponsors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("automation_studio/substack-sponsor-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("automation_studio/substack-sponsor-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call automation_studio/substack-sponsor-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation_studio/substack-sponsor-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BcFI1JGbqibyEIGVV/builds/mweszADZe0qbTNC6M/openapi.json
