# Substack Scraper - Newsletters, Posts & Sponsorship Leads (`ahmed-data-tools/substack-scraper`) Actor

Scrape Substack newsletters and posts without login: subscriber tier, paid plans, posting frequency, engagement, reactions and comments. Build newsletter sponsorship lead lists and research creators.

- **URL**: https://apify.com/ahmed-data-tools/substack-scraper.md
- **Developed by:** [ahmed abdelmenem](https://apify.com/ahmed-data-tools) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Substack Scraper: Newsletters, Posts & Sponsorship Leads

Scrape **Substack newsletters and posts** without logging in. This Substack scraper turns any Substack publication (on `*.substack.com` or a custom domain) into clean, structured data: post titles, dates, reactions, comments, restacks and paid/free status, plus **publication profiles** with subscriber tier, paid plan prices, posting frequency and engagement averages. Use it to build **newsletter sponsorship lead lists**, research creators, analyze content, or monitor competitors.

It works like a lightweight **Substack API**: you give it newsletter URLs or keywords and get JSON, CSV, Excel or HTML back.

### What data you get

#### Publications (sponsorship and outreach leads)

| Field | Example |
|---|---|
| `name`, `url`, `subdomain`, `customDomain` | Lenny's Newsletter, https://www.lennysnewsletter.com |
| `description`, `language`, `logo` | Newsletter tagline and logo |
| `authorName`, `authorHandle`, `authorProfileUrl`, `authorBio`, `twitterHandle` | Public author info |
| `subscriberTier`, `subscriberCountText`, `paidSubscriberTier` | "Millions of subscribers", "Over 1,200,000 subscribers" (as shown publicly by Substack) |
| `hasPaidPlan`, `paidPrices` | Monthly, yearly and founding-member prices |
| `postsLast30Days`, `postsLast90Days` | Publishing cadence |
| `avgReactionsRecent`, `avgCommentsRecent`, `paidPostShareRecent` | Engagement over the 10 most recent posts |
| `lastPostAt`, `firstPostAt`, `createdAt` | Activity and age |
| `hasPodcast`, `communityEnabled`, `sampleSize`, `scrapedAt` | Extras |

No personal email addresses are collected.

#### Posts

`publication`, `publicationUrl`, `postId`, `title`, `subtitle`, `slug`, `url`, `postType`, `author`, `authors` (name, handle, profile URL), `publishedAt`, `audience` (`everyone`, `only_paid`, `founding`, `only_free`), `isPaid`, `wordcount`, `reactionCount`, `commentCount`, `restacks`, `section`, `tags`, `coverImage`, `podcastDurationSec`, `bodyText` (optional, free public posts only), `scrapedAt`.

### Use cases

- **Newsletter sponsorships**: find newsletters in your niche, check audience size tiers, posting cadence and engagement before you pitch an ad placement.
- **Influencer and creator research**: build lists of Substack writers by topic, with their profile links and social handles.
- **Content analysis**: see which topics, formats and lengths get the most reactions, comments and restacks.
- **Competitor monitoring**: track what competing newsletters publish, how often, and what goes behind the paywall.
- **Pricing research**: compare paid subscription prices across a niche.

### How to use

1. Add **Substack URLs**: publication homepages (`https://example.substack.com`, `https://www.lennysnewsletter.com`), single post URLs (`.../p/post-slug`) or author profiles (`https://substack.com/@handle`).
2. Optionally add **search keywords** (e.g. `product management`, `sourdough`) to discover publications automatically.
3. Choose the **output**: posts, publications, or both.
4. Run and download your dataset. The **Posts** and **Publications** views in the dataset show each type as a table.

#### Example input

```json
{
  "startUrls": [
    { "url": "https://www.lennysnewsletter.com" },
    { "url": "https://newsletter.pragmaticengineer.com" }
  ],
  "searchQueries": ["product management"],
  "maxPublicationsPerQuery": 20,
  "outputMode": "both",
  "maxPostsPerPublication": 20,
  "postedWithinDays": 90,
  "includeBody": false,
  "maxItems": 1000
}
```

#### Example output (publication)

```json
{
  "type": "publication",
  "name": "Lenny's Newsletter",
  "url": "https://www.lennysnewsletter.com",
  "authorName": "Lenny Rachitsky",
  "authorProfileUrl": "https://substack.com/@lenny",
  "subscriberTier": "Millions of subscribers",
  "subscriberCountText": "Over 1,200,000 subscribers",
  "hasPaidPlan": true,
  "paidPrices": [{ "name": "$20 a month", "amount": 20.0, "currency": "USD", "interval": "month", "isFounding": false }],
  "postsLast30Days": 3,
  "postsLast90Days": 10,
  "avgReactionsRecent": 481.2,
  "avgCommentsRecent": 9.7,
  "lastPostAt": "2026-09-29T13:15:57+00:00"
}
```

### Input options

| Option | Description |
|---|---|
| `startUrls` | Publication, post or author profile URLs. |
| `searchQueries` | Keywords to discover publications through Substack's public search. |
| `maxPublicationsPerQuery` | Publications per keyword (default 20). |
| `outputMode` | `posts`, `publications` or `both` (default). |
| `maxPostsPerPublication` | Most recent posts saved per publication (default 20). |
| `postedWithinDays` | Only save posts from the last N days. |
| `includeBody` | Add plain-text body (max 8,000 characters) for **free, public posts only**. |
| `maxItems` | Cap on total results. |

### Pricing: pay per result

This Actor uses **pay-per-result** pricing: you are charged per item saved to the dataset (each post or publication counts as one result). Use `maxItems` and `maxPostsPerPublication` to control cost. If you set a maximum charge for the run, the scraper stops as soon as it is reached.

### Notes and limitations

- Only publicly available data is collected. No login is used and **paid-only content is never extracted**; for paid posts you get metadata (title, date, engagement) but no body text.
- Subscriber numbers are the public tier text that Substack shows (e.g. "Over 1,000 subscribers"); exact counts are not public. Some publications hide them.
- `postsLast30Days` / `postsLast90Days` are computed from a recent sample of up to 120 posts and are `null` if the sample doesn't cover the period.
- Keyword discovery uses Substack's public profile search, which ranks writers by relevance; results depend on Substack.
- Requests are rate-limited (max 3 concurrent) with automatic retries.

This Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Substack Inc.

# Actor input Schema

## `startUrls` (type: `array`):

Publication URLs (e.g. https://example.substack.com or a custom domain like https://www.lennysnewsletter.com), post URLs (…/p/post-slug) or author profile URLs (https://substack.com/@handle).

## `searchQueries` (type: `array`):

Keywords to discover Substack publications via Substack's public search, e.g. "product management", "crypto", "fitness". Each matching publication is then scraped.

## `maxPublicationsPerQuery` (type: `integer`):

Maximum number of publications discovered for each search keyword.

## `outputMode` (type: `string`):

What to save: posts, publication profiles (sponsorship leads with stats), or both.

## `maxPostsPerPublication` (type: `integer`):

Maximum number of most recent posts saved per publication.

## `postedWithinDays` (type: `integer`):

Only save posts published within this many days. Leave empty for no date limit. Publication stats are unaffected.

## `includeBody` (type: `boolean`):

Add plain-text body (HTML stripped, max 8,000 characters) for free, public posts. Paid-only post content is never included.

## `maxItems` (type: `integer`):

Maximum total number of dataset items (posts + publications). Leave empty for no limit.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.lennysnewsletter.com"
    },
    {
      "url": "https://newsletter.pragmaticengineer.com"
    }
  ],
  "maxPublicationsPerQuery": 20,
  "outputMode": "both",
  "maxPostsPerPublication": 20,
  "includeBody": false
}
```

# Actor output Schema

## `posts` (type: `string`):

No description

## `publications` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.lennysnewsletter.com"
        },
        {
            "url": "https://newsletter.pragmaticengineer.com"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ahmed-data-tools/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        { "url": "https://www.lennysnewsletter.com" },
        { "url": "https://newsletter.pragmaticengineer.com" },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("ahmed-data-tools/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.lennysnewsletter.com"
    },
    {
      "url": "https://newsletter.pragmaticengineer.com"
    }
  ]
}' |
apify call ahmed-data-tools/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ahmed-data-tools/substack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/whTi1oLezS23oEJvM/builds/ZOLEt7eejR0lUrfao/openapi.json
