# Substack Scraper (`labrat011/substack-scraper`) Actor

Scrape Substack newsletters: every post with title, date, engagement, paywall status and full text for free posts, plus comments, publication details and subscriber counts. Or search all of Substack by keyword. No login.

- **URL**: https://apify.com/labrat011/substack-scraper.md
- **Developed by:** [mick\_](https://apify.com/labrat011) (community)
- **Categories:** News, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.51 / 1,000 post (metadata)s

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<img src="https://apify-image-uploads-prod.s3.us-east-1.amazonaws.com/wCP1WauwRX2Gr3Gir-actor-JBzOEMsNPLVCGewTq-TuobEgtpfn-substack-scraper.png" alt="Substack Scraper logo" width="120">

## Substack Scraper

Scrape any Substack newsletter: every post with title, date, reactions, comments count, restacks, paywall status and the **full text of free posts**, plus comments and the newsletter's subscriber count. Or search all of Substack by keyword. No login, no API key.

| At a glance | |
|---|---|
| **You give it** | Newsletter URLs or search keywords, plus optional date, free-only and comment settings |
| **You get** | One row per post (full text for free posts, reactions, comments count, paywall status) and optionally per comment |
| **Price** | $0.005 per run + $0.00098 per post, $0.002 per post with full text, $0.00049 per comment (Free plan) |
| **Speed** | 10 posts and 35 comments in about 11 seconds |
| **Needs** | Nothing: no login, no API key, no proxy |

### What you get

- **Whole archives,** newest first, with a date range that stops paging as soon as it reaches older posts.
- **Full text for free posts** as HTML and as clean plain text, ready for AI. Paid posts come with their public preview.
- **Fair billing on paid posts.** A paywalled post is always charged the lower metadata price, because there is no full text to deliver.
- **Comments and replies,** one row each with the parent comment id and nesting level, so threads can be rebuilt.
- **Publication details on every post:** name, URL, author, language, description, and subscriber count (for example "1.2M+" and 1,200,000).
- **Keyword search across all of Substack:** posts on a topic from many newsletters, with each newsletter's details.
- **Filters:** free posts only, post type (newsletter, podcast, thread), date range. Filters apply to newsletters, single posts and keyword search alike.
- **Cheap to run:** 10 posts plus 35 comments used $0.0004 of platform usage.

### Use cases

- **Content research.** What the top newsletters in your niche publish, and which posts get the most reactions and comments.
- **Newsletter sponsorship and partnership prospecting.** Find newsletters on a topic with their subscriber counts and authors.
- **AI and RAG.** Clean full text from the writers you trust, updated on a schedule.
- **Audience research.** Read what commenters ask and argue about.
- **Competitive monitoring.** Get every new post from competitor newsletters.

### Example inputs

#### Latest 20 posts from two newsletters, with comments

```json
{ "urls": ["https://www.lennysnewsletter.com", "importai"], "maxPostsPerNewsletter": 20, "includeComments": true }
```

#### A whole archive since January, metadata only

```json
{ "urls": ["lenny"], "maxPostsPerNewsletter": 0, "dateFrom": "2026-01-01", "includeContent": false }
```

#### Search Substack for a topic

```json
{ "searchKeywords": ["ai agents", "vertical saas"], "maxSearchResultsPerKeyword": 50, "onlyFree": true }
```

#### Specific posts

```json
{ "urls": ["https://www.lennysnewsletter.com/p/announcing-lennys-jobs-the-best-place"], "includeComments": true }
```

### Input

| Field | What it does |
|---|---|
| `urls` | Newsletter URLs (`https://www.lennysnewsletter.com`, `lenny.substack.com`, or `lenny`) and post URLs (anything with `/p/`). |
| `searchKeywords` | Search all of Substack. |
| `maxPostsPerNewsletter` | Newest first. `0` means the whole archive. Default 50. |
| `maxSearchResultsPerKeyword` | Default 20. |
| `includeContent` | Full text of free posts. Default on. |
| `includeComments`, `maxCommentsPerPost` | Comments and replies, best first. |
| `onlyFree`, `postType`, `dateFrom`, `dateTo` | Filters. |

### Output

Each row has a `type`: `post` or `comment`. The **Posts** and **Comments** tabs split them. Real rows from a test run on 2026-09-26 (long text shortened here):

#### Post

```json
{
  "type": "post",
  "postId": 208730073,
  "title": "Announcing Lenny’s Jobs: The best place in the world to find, vet, and land your dream job",
  "subtitle": "Where product managers, engineers, designers, and growth/marketing professionals discover high-quality open roles at tech companies",
  "url": "https://www.lennysnewsletter.com/p/announcing-lennys-jobs-the-best-place",
  "slug": "announcing-lennys-jobs-the-best-place",
  "publishedAt": "2026-08-18T15:40:06.921Z",
  "postType": "newsletter",
  "audience": "everyone",
  "isPaid": false,
  "wordCount": 933,
  "reactionCount": 378,
  "commentCount": 18,
  "restacks": 9,
  "tags": [
    "Career"
  ],
  "section": null,
  "description": "Where product managers, engineers, designers, and growth/marketing professionals discover high-quality open roles at tech companies",
  "previewText": "👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs ...",
  "bodyHtml": "<p><em><span>👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: ...",
  "bodyText": "👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs | Lenny’s Podcast | How I AI | Lennybot | Become an AI-Native Builder and some o...",
  "coverImage": "https://substackcdn.com/image/fetch/$s_!PFZn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f696879-6374-42d6-a171-47933f6eb126_1456x970.png",
  "authors": [
    "Lenny Rachitsky"
  ],
  "authorName": "Lenny Rachitsky",
  "authorHandle": "lenny",
  "podcastUrl": null,
  "podcastDurationSecs": null,
  "hasVoiceover": false,
  "publicationId": 10845,
  "publicationName": "Lenny's Newsletter",
  "publicationUrl": "https://www.lennysnewsletter.com",
  "publicationAuthor": "Lenny Rachitsky",
  "publicationAuthorHandle": "lenny",
  "publicationLanguage": "en",
  "publicationDescription": "Deeply researched product, growth, and career advice for product leaders, founders, and ambitious bu",
  "publicationLogo": "https://substackcdn.com/image/fetch/$s_!8MSN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441213db-4824-4e48-9d28-a3a18952cbfc_592x592.png",
  "publicationPaidEnabled": true,
  "freeSubscriberCount": 1200000,
  "subscriberCountText": "1.2M+",
  "searchKeyword": null
}
```

#### Comment

```json
{
  "type": "comment",
  "commentId": 317389445,
  "parentCommentId": null,
  "level": 0,
  "postId": 208730073,
  "postTitle": "Announcing Lenny’s Jobs: The best place in the world to find, vet, and land your dream job",
  "postUrl": "https://www.lennysnewsletter.com/p/announcing-lennys-jobs-the-best-place",
  "body": "The Lenny 100 list is GENIUS 👏",
  "date": "2026-08-18T15:48:15.832Z",
  "editedAt": null,
  "authorName": "Tom Orbach",
  "authorHandle": "marketingideas",
  "reactionCount": 7,
  "restacks": 0,
  "replyCount": 0,
  "publicationName": "Lenny's Newsletter"
}
```

### Pricing

Pay per event. Each post is charged one of the first two, never both:

- **$0.005** per run start
- **$0.00098** per post with metadata (paid posts, or when full text is off)
- **$0.002** per post with full text (free posts, full text on)
- **$0.00049** per comment

Lower on higher Apify plans. Apify platform usage is billed separately to your account and is tiny: 10 posts plus 35 comments used $0.0004.

| What you scrape | Actor cost |
|---|---|
| 50 posts, metadata only | about $0.054 |
| 50 free posts with full text | about $0.105 |
| 50 posts + 500 comments | about $0.30 |

### Automate it with n8n

Each workflow uses n8n's official **Apify** node, operation **Run actor and get dataset**, actor `labrat011/substack-scraper`. Paste the input into **Input JSON**; switch it to **Expression** when it contains `{{ }}`.

#### 1. New posts from your favorite newsletters, summarized by AI

```
Schedule Trigger (daily 07:00)
  > Apify: Run actor and get dataset
      { "urls": ["lenny", "importai"], "dateFrom": "{{ $now.minus({ days: 1 }).toISODate() }}", "onlyFree": true }
  > OpenAI / Anthropic: "Summarize this post in 3 bullets" on bodyText
  > Slack or Gmail: one message per post with the summary and url
```

#### 2. Sponsorship prospect list

```
Manual Trigger
  > Apify: Run actor and get dataset   { "searchKeywords": ["b2b marketing", "product management"], "maxSearchResultsPerKeyword": 100 }
  > Remove Duplicates: on publicationUrl
  > Filter: freeSubscriberCount > 5000
  > Google Sheets: publicationName, publicationUrl, publicationAuthor, freeSubscriberCount
```

#### 3. Newsletter RAG knowledge base

```
Schedule Trigger (weekly)
  > Apify: Run actor and get dataset   (your newsletters, last 7 days, full text on)
  > Filter: type is post and bodyText is not empty
  > Embeddings + Vector Store insert (Pinecone, Qdrant, Supabase), metadata: title, url, publishedAt
```

#### 4. Top-performing posts report

```
Schedule Trigger (monthly)
  > Apify: Run actor and get dataset   (competitor newsletters, dateFrom = first of last month, includeContent false)
  > Sort: reactionCount descending
  > Google Sheets: top 20 with title, reactions, comments, restacks, url
```

#### 5. Comment insights

```
Manual Trigger
  > Apify: Run actor and get dataset   (one newsletter, 20 posts, includeComments true, 100 per post)
  > Filter: type is comment
  > Aggregate: all comment bodies
  > OpenAI / Anthropic: "What questions and objections come up most? Quote examples."
```

### For AI agents

- **Actor:** `labrat011/substack-scraper`
- **Smallest input:** `{ "urls": ["https://www.lennysnewsletter.com"] }`
- **One row = one post (`type: "post"`) or one comment (`type: "comment"`).** Key fields: posts: `title`, `url`, `publishedAt`, `isPaid`, `reactionCount`, `commentCount`, `bodyText`; comments: `commentId`, `postId`, `body`, `date`.
- **Billing:** `apify-actor-start` once per run, `post-with-content` or `post-metadata` per post (never both), `comment` per comment row. Cap spend with `maxPostsPerNewsletter`, `maxCommentsPerPost` or a maximum cost per run.
- **Run it:** `POST https://api.apify.com/v2/acts/labrat011~substack-scraper/run-sync-get-dataset-items` with the input as the JSON body, or call it from the Apify MCP server.
- **Done signal:** the run's status message reads `Saved N posts and M comments in R requests.`

### FAQ

**Can it read paid posts?** No. Paid posts return their metadata and the public preview, and are charged the lower metadata price.

**Where does the subscriber count come from?** From the number Substack itself shows on the newsletter's page. Substack rounds it, for example "1.2M+".

**Do I need a proxy?** No. The actor spaces its requests and backs off if Substack slows it down.

### Support

Open an issue on the actor's Issues tab with the run ID and it will be looked at.

# Changelog

This Actor's version history is a separate document: https://apify.com/labrat011/substack-scraper/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Newsletters: https://www.lennysnewsletter.com, lenny.substack.com, or just lenny. Single posts: any URL with /p/ in it.

## `searchKeywords` (type: `array`):

Search all of Substack for posts on a topic. Posts from many different newsletters.

## `maxPostsPerNewsletter` (type: `integer`):

Newest first. 0 means the whole archive.

## `maxSearchResultsPerKeyword` (type: `integer`):

Search results to save per keyword.

## `includeContent` (type: `boolean`):

Adds bodyHtml and clean bodyText. Paid posts always get their public preview only.

## `includeComments` (type: `boolean`):

Comments and replies, one row each, with the parent comment id.

## `maxCommentsPerPost` (type: `integer`):

Best comments first.

## `onlyFree` (type: `boolean`):

Skip paywalled posts.

## `postType` (type: `string`):

Only one kind of post.

## `dateFrom` (type: `string`):

Only posts published on or after this date, e.g. 2026-01-01. Paging stops at the first older post.

## `dateTo` (type: `string`):

Only posts published before this date.

## `proxyConfiguration` (type: `object`):

Not needed normally.

## Actor input object example

```json
{
  "urls": [
    "https://www.lennysnewsletter.com"
  ],
  "maxPostsPerNewsletter": 10,
  "maxSearchResultsPerKeyword": 20,
  "includeContent": true,
  "includeComments": false,
  "maxCommentsPerPost": 50,
  "onlyFree": false,
  "postType": "",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.lennysnewsletter.com"
    ],
    "maxPostsPerNewsletter": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("labrat011/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.lennysnewsletter.com"],
    "maxPostsPerNewsletter": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("labrat011/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.lennysnewsletter.com"
  ],
  "maxPostsPerNewsletter": 10
}' |
apify call labrat011/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,labrat011/substack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JBzOEMsNPLVCGewTq/builds/mGF3tQVi04jgyR0dv/openapi.json
