# Substack Scraper (`lergassy/substack-scraper`) Actor

\[$0.60/1K posts] Posts of any Substack publication: title, subtitle, date, free or paid, likes, comments, restacks, word count, podcast length, cover, URL — plus full text of free posts and a publication row with author, subscriber tier and plans. Custom domains work. No run fee, no login.

- **URL**: https://apify.com/lergassy/substack-scraper.md
- **Developed by:** [Matvey](https://apify.com/lergassy) (community)
- **Categories:** News, Social media, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $0.45 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Substack Scraper** turns any Substack publication into a table of its posts: **title, subtitle, date, free or paid, likes, comments, restacks, word count, post type, podcast length, cover image and URL** — newest first, as deep into the archive as you want. Add the **full text of free posts** with one switch, and get a **publication row** with the author, description, subscriber tier, launch date, paid plans and benefits.

**$0.0006 per post** — a thousand posts for sixty cents. No start fee, no login, no browser. Custom domains work; publications that do not exist are free.

### What is Substack Scraper?

A research tool for anyone who studies newsletters: content teams tracking competitors, analysts sizing a niche, writers picking topics by what gets likes, and AI agents that need a publication's archive as data. Substack serves its archive as JSON to its own pages; this Actor reads that directly, so a 500-post archive takes seconds and costs 30 cents.

### What is in every row

| Field | Meaning |
|---|---|
| `title`, `subtitle`, `url`, `publishedAt`, `authors`, `section` | The post |
| `audience`, `isPaid` | `everyone`, `only_paid` or `founding` |
| `likes`, `comments`, `restacks` | Engagement as Substack counts it |
| `wordCount`, `postType`, `podcastDurationSec`, `hasVideo`, `hasPodcast` | Format and length |
| `description`, `truncatedBody`, `coverImageUrl` | Preview text and image |
| `contentHtml`, `contentText`, `contentIsPreview` | Full text when **Add the full text** is on; a paid post gives its free preview, flagged and not charged |

The publication row adds `authorName`, `authorHandle`, `authorBio`, `description`, `subscriberTier` ("Hundreds of thousands of subscribers"), `paidSubscriberTier`, `subscribersOrderOfMagnitude`, `createdAt`, `language`, `twitter`, `hasPaidPlan`, `plans` (price, currency, interval) and the free and paid benefits.

### How much does it cost?

| Event | Price |
|---|---|
| Post row | **$0.0006** |
| Full text of a free post (optional) | **$0.0006** |
| Publication row (optional) | **$0.002** |

- Posts skipped by the date filter or the type filter: **$0**.
- Paid posts when full text is on: the preview is delivered **free**.
- Publications that do not exist or are not on Substack: **$0**.

#### Bulk export: what 50,000 rows actually cost

This Actor is built for bulk jobs — hundreds of publications in one run, whole archives, or a schedule
that reads every new post daily. There is **no fee per run, no fee per page and no proxy charge**:
you pay for the rows you keep.

| Job | This Actor | Most-used Actor in this category |
|---|---|---|
| 50,000 post rows across 100 runs | **$30** | $57.50 + start fees |

Checked on the Apify Store on 23 September 2026 against the Actor with the most monthly users in
this category ($0.00115 per post plus a fee per run). Some Actors here ask less per row — this table
compares against the one buyers actually use most.

### How to use it in three steps

1. Paste publication URLs or subdomains into **📰 Publications** — `https://www.lennysnewsletter.com`, `platformer.substack.com`, or just `lenny`.
2. Set **Posts per publication** (newest first) and, if you like, **Only posts published after** and **Post type**. Turn on **Add the full text of free posts** for the body.
3. Press **Start**, then download JSON, CSV or Excel — or call the run from the API, n8n, Make, Zapier or an AI agent.

### ⬇️ Input

```json
{
  "publications": ["https://www.lennysnewsletter.com", "platformer.substack.com"],
  "maxPostsPerPublication": 100,
  "publishedAfter": "2026-01-01",
  "includeContent": true
}
```

### ⬆️ Output

```json
{
  "type": "post",
  "publication": "Lenny's Newsletter",
  "title": "Advanced evals: How to find (and fix) hidden AI failures in your product",
  "subtitle": "Why you should never skip error discovery",
  "url": "https://www.lennysnewsletter.com/p/advanced-evals-how-to-find-and-fix",
  "postType": "newsletter",
  "audience": "only_paid",
  "isPaid": true,
  "publishedAt": "2026-09-22T12:45:14.998Z",
  "likes": 240,
  "comments": 2,
  "restacks": 10,
  "wordCount": 3807,
  "authors": ["Hamel Husain", "Shreya Shankar"],
  "coverImageUrl": "https://substackcdn.com/image/fetch/…",
  "scrapedAt": "2026-09-23T12:04:11.004Z"
}
```

### Use cases

#### Competitive content research

Fifty publications in your niche, the last 100 posts each — sort by likes per word and see what actually lands.

#### Newsletter monitoring

Schedule daily with **Only posts published after** yesterday; new posts flow into Sheets or Slack.

#### Sizing a market

The publication row gives subscriber tiers, launch dates and paid plans for hundreds of newsletters in one run.

#### AI agents and RAG

Full text of free posts, clean `contentText`, through the Apify MCP server — *"summarize what Platformer wrote about AI this month"* returns sourced rows.

### Integrations

Every run is available through the Apify API, and the Actor works out of the box with n8n, Make, Zapier, Google Sheets, LangChain and the Apify MCP server. Schedule it, webhook it, or export straight to CSV.

### FAQ

**Does it read paid posts?** No. Paid posts come back with all their metadata and engagement, and with the free preview when full text is on — never the paywalled body.

**Are subscriber counts exact?** Substack publishes tiers, not numbers ("Tens of thousands of paid subscribers"); the row carries the tier text and its order of magnitude.

**Custom domains?** Yes — paste the domain; the Actor follows Substack's own redirects and reads the same JSON.

**Comments and notes?** Not in this Actor; it is posts and publications.

**Is this legal?** The Actor reads public pages and the JSON behind them, the same thing a browser loads without logging in. Use the content in line with each publication's terms and copyright.

# Actor input Schema

## `publications` (type: `array`):

Substack URLs or subdomains, one per line — <b>https://www.lennysnewsletter.com</b>, <b>lenny.substack.com</b>, <b>lenny</b>. Custom domains work.

## `startUrls` (type: `array`):

The same thing from a file or a Google Sheet, for long lists.

## `maxPostsPerPublication` (type: `integer`):

Newest first. Set high to read a whole archive.

## `publishedAfter` (type: `string`):

ISO date, e.g. 2026-01-01. Older posts are skipped and not charged.

## `postType` (type: `string`):

Keep only one kind of post.

## `includeContent` (type: `boolean`):

Fetches each post and adds contentHtml and contentText. Charged per free post; paid posts return their free preview, flagged and not charged.

## `includePublicationRow` (type: `boolean`):

Author, description, subscriber tier, launch date, paid plans and benefits — one row per publication.

## `proxyConfiguration` (type: `object`):

Datacenter proxy by default — \*.substack.com refuses some cloud addresses, custom domains do not. A challenged request is retried on another address automatically.

## Actor input object example

```json
{
  "publications": [
    "https://www.lennysnewsletter.com",
    "platformer.substack.com"
  ],
  "startUrls": [],
  "maxPostsPerPublication": 50,
  "postType": "all",
  "includeContent": false,
  "includePublicationRow": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `posts` (type: `string`):

One row per post (and one per publication): title, date, audience, likes, comments, restacks, word count, URL, optional full text.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publications": [
        "https://www.lennysnewsletter.com",
        "platformer.substack.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("lergassy/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "publications": [
        "https://www.lennysnewsletter.com",
        "platformer.substack.com",
    ],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("lergassy/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publications": [
    "https://www.lennysnewsletter.com",
    "platformer.substack.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call lergassy/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lergassy/substack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cd4Rw1lMbIu5SPrYW/builds/zM4cBFbMXLAQymCi9/openapi.json
