# Substack Scraper: Posts, Full Text, Comments, Newsletter Data (`deriverge/substack-scraper`) Actor

Scrape Substack newsletters: posts with likes, comments, restacks, word count and audience, full post text on request, comments as rows, and newsletter details with author and subscriber tier. Custom domains, new posts only; JSON, CSV, API.

- **URL**: https://apify.com/deriverge/substack-scraper.md
- **Developed by:** [deriverge s.r.o.](https://apify.com/deriverge) (community)
- **Categories:** News, Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 post returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Substack Scraper

### What does Substack Scraper do?

**Substack Scraper** reads Substack newsletters through the public interface every Substack site serves to its own readers, and returns three kinds of clean rows:

- **Newsletter header**: name, author, description, launch date, first post date, subscriber tier as Substack shows it (for example "Tens of thousands of paid subscribers"), free subscriber count where published, language, whether it sells paid plans and whether it has a podcast.
- **Posts** from the archive, newest first or most liked first: title, subtitle, date, audience (free or paid), type, section, authors, word count, likes, comments, restacks, cover image, podcast audio, tags and the preview text. With `includeBody`, the full text as plain text and as HTML.
- **Comments** under each post, with author name and handle, date, text, likes, replies and thread depth.

Name newsletters any way you like: a subdomain such as `lenny`, a `substack.com` address, a custom domain such as `www.noahpinion.blog`, or a link to any post. No browser, no proxies, no login, no API key.

You pay only for posts returned, from $0.40 per 1,000, with no start fee. The $5 of monthly credit in Apify's Free plan covers about 6,200 posts, so you can try it for free.

### Fields

| Post row | What you get |
|---|---|
| `title`, `subtitle`, `description`, `previewText` | What Substack shows in the archive. |
| `publishedAt`, `audience`, `postType`, `section` | Date, `everyone` or `only_paid`, `newsletter` / `podcast` / `thread` / `video`. |
| `likes`, `comments`, `restacks`, `wordCount` | The metrics on the post. |
| `authors` | Bylines as name and handle. |
| `bodyText`, `bodyHtml`, `bodyIsPreview` | With `includeBody`. Paid posts return the free preview and `bodyIsPreview` is `true`. |
| `podcastUrl`, `podcastDurationSeconds`, `coverImage`, `tags` | When the post has them. |

| Comment row | What you get |
|---|---|
| `author`, `authorHandle`, `publishedAt`, `body` | The comment as published. |
| `likes`, `replies`, `depth`, `parentId` | Thread structure; replies are their own rows. |

Nothing is guessed. A field Substack did not publish is `null`.

### A new-post alert, not an export

Turn on `newOnly`, give the run a watch name or save it as a task, and schedule it daily. Each run compares against the previous snapshot and returns only the posts and comments that appeared since. You pay for the new rows and nothing else.

### Filters that stop the noise before it is charged

- **`publishedAfter`** cuts the archive off at a date.
- **`audience`** keeps free posts only or paid posts only.
- **`keywords`** keeps only posts whose title, subtitle, description or preview mention one of your words.
- **`maxPostsPerNewsletter`** takes the newest N; archives go back years.

Posts removed by a filter are never charged.

### How much does it cost to scrape Substack?

You pay per result. There is no start fee and no charge for compute time or proxies.

| | Free plan | Starter | Scale | Business |
|---|---|---|---|---|
| 1,000 posts | $0.80 | $0.64 | $0.52 | $0.40 |
| Full post text, per 1,000 posts (optional) | $0.50 | $0.40 | $0.33 | $0.25 |
| 1,000 comments (optional) | $0.30 | $0.24 | $0.20 | $0.15 |
| Newsletter header, per 1,000 newsletters | $1.00 | $0.80 | $0.65 | $0.50 |

For example, a batch of 1,000 posts with their metrics costs $0.80 on the Free plan and $0.40 on the Business plan. The $5 of monthly credit in Apify's Free plan covers about 6,200 posts.

Posts removed by your filters and rows already returned in new-only mode are never charged.

### How to scrape Substack

1. Click **Try for free** (or **Start** if you are signed in) to open the actor in Apify Console.
2. In **Newsletters**, add Substack subdomains, custom domains or post links, one per line.
3. Click **Start**. Rows appear in the **Output** tab within seconds.
4. Download the results as JSON, CSV, Excel or HTML, or read them through the API.
5. To repeat it, click **Save as a task** and add a schedule. A scheduled task keeps its own snapshot, so change and new-only modes work without any setup.

### Input

```json
{
  "newsletters": ["lenny", "https://www.noahpinion.blog", "astralcodexten"],
  "sort": "new",
  "maxPostsPerNewsletter": 50,
  "audience": "all",
  "includeBody": false,
  "includeComments": false,
  "includeNewsletterInfo": true,
  "newOnly": false
}
```

### Output

A post row:

```json
{
  "type": "post",
  "key": "post:216826341",
  "id": 216826341,
  "slug": "the-problems-with-utilitarianism",
  "url": "https://www.noahpinion.blog/p/the-problems-with-utilitarianism",
  "newsletter": "Noahpinion",
  "newsletterSubdomain": "noahpinion",
  "newsletterUrl": "https://www.noahpinion.blog",
  "title": "The problem(s) with utilitarianism",
  "subtitle": "Why I am not a utilitarian, and why you probably should not be either",
  "description": "Why I am not a utilitarian, and why you probably should not be either",
  "publishedAt": "2026-09-22T09:31:07.000Z",
  "audience": "everyone",
  "postType": "newsletter",
  "section": null,
  "authors": [{ "name": "Noah Smith", "handle": "noahpinion" }],
  "wordCount": 2646,
  "likes": 349,
  "comments": 51,
  "restacks": 38,
  "coverImage": "https://substackcdn.com/image/fetch/...",
  "podcastUrl": null,
  "podcastDurationSeconds": null,
  "tags": [],
  "previewText": "Utilitarianism is the idea that ...",
  "bodyText": null,
  "bodyHtml": null,
  "bodyIsPreview": null
}
```

A newsletter row:

```json
{
  "type": "newsletter",
  "key": "newsletter:noahpinion",
  "id": 35345,
  "name": "Noahpinion",
  "subdomain": "noahpinion",
  "customDomain": "www.noahpinion.blog",
  "url": "https://www.noahpinion.blog",
  "author": "Noah Smith",
  "authorHandle": "noahpinion",
  "description": "Economics and other interesting stuff",
  "logo": "https://substackcdn.com/image/fetch/...",
  "language": "en",
  "launchedAt": "2020-03-28T03:32:51.086Z",
  "firstPostAt": "2020-11-24T18:26:23.401Z",
  "paidEnabled": true,
  "subscriberTier": "Tens of thousands of paid subscribers",
  "freeSubscribers": 458000,
  "hasPodcast": true,
  "explicit": false,
  "checkedAt": "2026-09-23T14:20:00.000Z"
}
```

### What it does not collect

Nothing about readers. Authors are the public bylines; commenters appear with the display name and handle they chose to publish under. Subscriber lists, emails and paywalled text are not collected: a paid post returns the free preview that Substack itself shows to everyone.

### Integrations and API

Connect the actor to **Make**, **Zapier**, **n8n**, **Google Sheets**, **Slack** or any webhook, or call it from your own code. With the Python client:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("deriverge/substack-scraper").call(run_input={"newsletters": ["lenny", "https://www.noahpinion.blog"]})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(row["type"], row.get("title"))
```

The **API** tab on this page has the same call for Node.js and cURL, and AI agents can run the actor through the Apify MCP server.

### Is it legal to scrape Substack?

The actor reads only what Substack shows publicly to every visitor: public posts, public comments and newsletter pages. Paywalled text is not collected; paid posts return the free preview. Respect authors' copyright if you republish.

### FAQ

#### Does it work with custom domains?

Yes. `lenny`, `lenny.substack.com` and `www.lennysnewsletter.com` are the same newsletter; the actor follows Substack's redirect and reads from the domain the newsletter actually uses.

#### What about paid posts?

The archive lists them with all metrics. With `includeBody`, Substack returns the free preview that it shows to non-subscribers, and the row carries `bodyIsPreview: true`. Comments under paid posts are visible only to subscribers, so they come back empty.

#### How many posts can I get?

The whole archive. `maxPostsPerNewsletter` and `publishedAfter` keep a run affordable; a monitor uses `newOnly`.

#### Can I get subscriber numbers?

Substack publishes a tier, not a number, for paid subscribers, and a free subscriber count for some newsletters. Both come through as published.

#### Where do I find the leaderboards?

In **Substack Newsletter Directory**, the sibling actor that lists every category leaderboard. **Substack Comments Scraper** does comments only.

### Support and feedback

Missing a field or found something that does not work? Open an issue in the **Issues** tab and it will be answered, usually within a day. If the actor saves you time, a short review helps other people find it.

# Actor input Schema

## `newsletters` (type: `array`):

Substack newsletters in any form: a subdomain such as lenny, a substack.com address, a custom domain or a link to any of their posts.

## `sort` (type: `string`):

Newest first is what a monitor wants; most liked first finds the classics.

## `maxPostsPerNewsletter` (type: `integer`):

Newest first. Archives go back years, so this is what keeps a broad run affordable.

## `publishedAfter` (type: `string`):

Drop posts older than this date (YYYY-MM-DD). Dropped posts are never charged.

## `audience` (type: `string`):

Substack marks each post as free for everyone or for paying subscribers.

## `keywords` (type: `array`):

Keep only posts whose title, subtitle, description or preview contain at least one of these words, case-insensitive. Dropped posts are never charged.

## `includeBody` (type: `boolean`):

Fetch each post and return its text as plain text and as HTML. Paid posts return the free preview and the row says so. One extra request per post, charged as a post detail.

## `includeComments` (type: `boolean`):

Return the comments under each post as separate rows with author, date, likes and replies. Charged per comment.

## `maxCommentsPerPost` (type: `integer`):

Replies count towards the cap.

## `commentSort` (type: `string`):

The order Substack returns them in; the cap applies after ordering.

## `includeNewsletterInfo` (type: `boolean`):

One row per newsletter with name, author, description, launch date, subscriber tier, language and whether it has paid plans or a podcast. Charged once per newsletter.

## `newOnly` (type: `boolean`):

Keeps a snapshot per watch name (or per saved task) and returns only posts, comments and newsletters that were not there before. Schedule it daily and you have a new-post alert that bills only for what is new.

## `watchName` (type: `string`):

Name of the snapshot used by the new-only mode, for example "my-newsletters". Runs from a saved task get a snapshot automatically even without a name.

## `maxItems` (type: `integer`):

Hard cap on returned rows in the run.

## Actor input object example

```json
{
  "newsletters": [
    "lenny",
    "https://www.noahpinion.blog"
  ],
  "sort": "new",
  "maxPostsPerNewsletter": 50,
  "audience": "all",
  "includeBody": false,
  "includeComments": false,
  "maxCommentsPerPost": 100,
  "commentSort": "best_first",
  "includeNewsletterInfo": true,
  "newOnly": false,
  "maxItems": 10000
}
```

# Actor output Schema

## `rows` (type: `string`):

One row per newsletter header, one per post with metrics and text on request, one per comment.

## `changes` (type: `string`):

Posts, comments and newsletters that appeared since the previous snapshot of the same watch name or task.

## `summary` (type: `string`):

Counts per row type, what each filter dropped, what was charged and notes.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "newsletters": [
        "lenny",
        "https://www.noahpinion.blog"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("deriverge/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "newsletters": [
        "lenny",
        "https://www.noahpinion.blog",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("deriverge/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "newsletters": [
    "lenny",
    "https://www.noahpinion.blog"
  ]
}' |
apify call deriverge/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,deriverge/substack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JjI0ZiTULJjkvGTmc/builds/uHK26ysTeCpbbKWOY/openapi.json
