# Substack Scraper, Posts, Full Text and Comments, No Login (`george.the.developer/substack-scraper`) Actor

Collect public Substack newsletter posts, full text and comments through public JSON endpoints. Support custom domains and post URLs without login, cookies, a proxy or a browser.

- **URL**: https://apify.com/george.the.developer/substack-scraper.md
- **Developed by:** [George Kioko](https://apify.com/george.the.developer) (community)
- **Categories:** Social media, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 post delivereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Substack Scraper, Posts, Full Text and Comments, No Login

Substack Scraper is an Apify Actor that collects structured posts from public Substack newsletters through their public JSON endpoints.

**No login. No cookies. No proxy. No browser.** Collect publication metadata, titles, dates, authors, reaction counts, public body HTML, plain text and optional comments.

Supply newsletter hosts, custom domains or individual post URLs. Each dataset row represents one post. Comments and replies stay inside that row. Post IDs are deduplicated across the run.

### Pricing

You pay for each distinct post row successfully delivered and each public comment actually included in that row. Empty results, failed fetches, duplicates and skipped paid posts have no result event charge.

| Event | Price | You pay when |
| --- | --- | --- |
| `apify-actor-start` | $0.00005 | The platform records its synthetic start event |
| `post-result` | $0.002 per post | A distinct post row is delivered |
| `comment-result` | $0.0003 per comment | A comment or reply is included in a delivered row |

**The price is $2 per 1,000 delivered posts and $0.30 per 1,000 included comments.** Ten posts without comments have $0.02 in result charges plus the platform start event.

The SDK checks the post budget, persists the row and charges the post event. The Actor charges included comments after the post has been persisted. Before fetching comments it reserves the next post price and trims the comment allowance to the remaining budget. It stops when another post would exceed the maximum charge limit. Local and free runs write rows without paid event charges. Actor code never charges the synthetic start event.

### Collect newsletter posts without login

Provide one or more publication URLs and a result limit for each publication.

```json
{
  "publications": ["https://www.lennysnewsletter.com"],
  "maxPostsPerPublication": 10,
  "includeFullText": true,
  "includeComments": false
}
```

### How it works

1. Normalize newsletter hosts, homepage URLs and individual post URLs.
2. Fetch archive pages sequentially with a page limit of 50 and a desktop Chrome User-Agent.
3. Optionally request public post detail and comments, then normalize the result.
4. Push each distinct post and charge result events on paid runs.
5. Stop at the per publication result limit, an empty or repeated page, the date cutoff, the charge limit or one minute before the run timeout.

```text
Input -> Publication archive or single post
                         |
                         v
                Public detail and comments
                         |
                         v
                Normalize and deduplicate
                         |
                         v
                 Dataset -> Event charges
```

Request starts are spaced at least one second apart. Each request has a 20 second timeout, shortened near the timeout reserve. Transient failures receive up to three retries with increasing delays of 1, 2 and 4 seconds. Retry-After seconds and HTTP dates take precedence when supplied.

A custom domain returning HTTP 404 triggers publication metadata and feed self link discovery. If those do not identify the Substack host, the Actor tries the custom domain's base label as a subdomain once. Arbitrary Substack links inside articles are ignored. This fallback worked for Platformer during the spike but does not guarantee a mapping for every custom domain.

### What data does it extract?

Each row contains `type`, `publication`, `publicationHost`, `postId`, `title`, `subtitle`, `slug`, `url`, `publishedAt`, `audience`, `isPaid`, `authors`, `authorHandles`, `wordCount`, `reactionCount`, `commentCount`, `coverImage`, `description`, `postType`, `bodyHtml`, `bodyText`, `isTruncated` and `scrapedAt`.

`restacks` and `podcastUrl` appear when provided. Missing scalar metadata is null and missing list metadata is an empty array. Publication names come from publication metadata when available, otherwise the host is used. `publicationHost` is the API host used, which can change after fallback. `url` retains the canonical post URL supplied by Substack.

With `includeComments`, the `comments` array contains `id`, `parentId`, `author`, `body`, `date`, `reactionCount` and `depth`. Replies follow their parent in depth first order. Deleted and suppressed comments are omitted. The included array can be shorter than the reported `commentCount` because of visibility, the comment limit and the charge budget.

`bodyText` is HTML stripped of tags, scripts and styles with entity decoding and block breaks. It is plain text rather than a faithful rendering. `bodyMarkdown` is not produced.

### What is NOT returned

- Private content, authenticated subscriber content or material beyond the public preview of paid posts.
- A global search across Substack publications.
- A guaranteed complete archive or every comment on a post.
- Markdown, media downloads or private author profile information.

### Start a run

1. Enter up to 50 newsletter URLs, hosts or post URLs.
2. Set the maximum posts per publication and optionally add publication search or a date cutoff.
3. Choose whether to include public full text, comments or only free posts.
4. Run the Actor and export the dataset as JSON, CSV or Excel.

The prefilled input requests 10 posts from Lenny's Newsletter with public full text and no comments. A live local SDK run on October 6, 2026 delivered exactly 10 rows in 11.621 seconds with body HTML and plain text verified. The normal result default is 50 per publication. Default platform resources are 256 MB and a 900 second timeout.

### Input

| Field | Type | Default | Meaning |
| --- | --- | --- | --- |
| `publications` | string array | Required, prefill Lenny's Newsletter | 1 to 50 newsletter URLs, hosts or individual post URLs |
| `maxPostsPerPublication` | integer | 50, prefill 10 | Delivered posts per supplied publication, from 1 to 5,000 |
| `sortBy` | enum | newest | `newest` or `top` |
| `searchQuery` | string | None | Search phrase within each publication archive |
| `publishedAfter` | ISO date string | None | Stop at older posts, only valid with newest order |
| `includeFullText` | boolean | true | Add public body HTML and plain text |
| `includeComments` | boolean | false | Include public comments inside each post row |
| `maxCommentsPerPost` | integer | 100 | Limit included comments and replies, zero disables fetching |
| `freeOnly` | boolean | false | Skip paid and founding audience posts |

```json
{
  "publications": ["platformer.substack.com", "https://www.lennysnewsletter.com"],
  "maxPostsPerPublication": 50,
  "sortBy": "newest",
  "searchQuery": "AI",
  "publishedAfter": "2026-01-01",
  "includeFullText": true,
  "includeComments": true,
  "maxCommentsPerPost": 100,
  "freeOnly": false
}
```

A `/p/post-slug` URL returns just that post and does not fetch its archive. Search applies to archives, so it does not filter a directly supplied post URL. The date and free audience filters still apply. Overlapping publication and post inputs share a deduplication set.

Invalid publication entries are logged and skipped when at least one valid input remains. Input with zero valid publications fails before fetching. A valid hostname whose endpoint fails is skipped so later publications can still deliver rows.

### Output

This row comes from the live local prefill dataset on October 6, 2026. The body fields are shortened excerpts.

```json
{
  "type": "post",
  "publication": "Lenny's Newsletter",
  "publicationHost": "www.lennysnewsletter.com",
  "postId": 215694124,
  "title": "All of the Lenny & Friends Summit talks are now online!",
  "subtitle": "Plus, some reflections and takeaways from the day",
  "slug": "all-of-the-lenny-and-friends-summit",
  "url": "https://www.lennysnewsletter.com/p/all-of-the-lenny-and-friends-summit",
  "publishedAt": "2026-09-29T13:15:57.512Z",
  "audience": "everyone",
  "isPaid": false,
  "authors": [
    "Lenny Rachitsky"
  ],
  "authorHandles": [
    "lenny"
  ],
  "wordCount": 1276,
  "reactionCount": 279,
  "commentCount": 6,
  "coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/9fed347c-0e02-4190-a590-1236719b99de_9213x6142.jpeg",
  "description": "Plus, some reflections and takeaways from the day",
  "postType": "newsletter",
  "bodyHtml": "<p><em>...",
  "bodyText": "Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For m...",
  "isTruncated": false,
  "scrapedAt": "2026-10-06T20:33:54.942Z",
  "restacks": 4
}
```

### Use from MCP agents and API clients

After deployment and publication, add this Actor through the Apify MCP server to let an agent collect newsletter posts.

```text
https://mcp.apify.com/?tools=george.the.developer/substack-scraper
```

From Node.js with the Apify client and an Apify token.

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('george.the.developer/substack-scraper').call({
  publications: ['https://www.lennysnewsletter.com'],
  maxPostsPerPublication: 10,
  includeFullText: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

With curl, wait for the run and return dataset rows.

```bash
curl -X POST "https://api.apify.com/v2/acts/george.the.developer~substack-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"publications":["https://www.lennysnewsletter.com"],"maxPostsPerPublication":10}'
```

Apify authentication runs the Actor. No Substack authentication is required.

### Integrations

**Clay.** Read dataset items as an HTTP data source and map publication, title, author and post URL to columns.

**n8n and Make.** Select the published Actor, supply newsletter input and process dataset rows in the next step.

**Google Sheets.** Export CSV or use the Apify Google Sheets integration. JSON preserves the full nested comment arrays.

### Use cases

- **Researchers** collect public newsletter coverage and publication dates.
- **Readers** organize links and publicly available text from their chosen newsletters.
- **Editorial teams** monitor topics, reactions and public discussion.
- **Analysts** compare posts across known publications without relying on global search.

### Limits

Paid posts return only the public response available without authentication and are marked `isTruncated`. A long public body does not prove that the full paid article is available. The Actor does not log in, send cookies or bypass paywalls.

There is no global Substack publication search. `searchQuery` searches within the supplied publication archives only. Public endpoints and search behavior can change.

The maximum measured archive limit was 50. A short initial page still had further pages, so the Actor advances by the number of returned items and continues until an empty or repeated page. Paging was verified at offset 1000 on Noahpinion, but no unlimited archive guarantee was established. A result cap of 5,000 does not make more posts available.

Newest date cutoff stopping assumes the endpoint returns newest order. Feed changes, pinned posts and changing archives can affect completeness. Publication names may fall back to the hostname when public bylines do not identify the publication.

One request per second succeeded for the tested detail calls without HTTP 429. Rate limits can vary by publication and network. No real 429 was observed during the spike; Retry-After handling was verified with offline responses.

Custom domain discovery can fail when the publication name differs from its Substack subdomain or when the custom domain has moved to another publishing system. Failure after the fallback is logged and skipped.

Public comments can be incomplete, unavailable or restricted. A failed detail call still delivers archive metadata with null body fields and `fullTextError`. Failed comment calls deliver an empty comment array and `commentsError`. Storage and billing failures stop the run as visible errors.

### FAQ

#### Do I need a Substack account or proxy?

No. The Actor uses unauthenticated public endpoints with a desktop User-Agent.

#### Can I retrieve complete paid articles?

Only the publicly available preview is returned. Paid audience posts are marked as truncated, and no paywall bypass is attempted.

#### Are duplicate posts charged twice?

No. Each post ID is delivered once per run, even across overlapping inputs.

#### What happens when a publication is empty or fails?

An empty publication has no result charges. Exhausted publication fetch failures are logged and skipped. The run exits successfully with whatever it delivered unless input, storage or billing fails.

#### Can I run this locally?

Yes. Install Node 22 or newer, run `npm install --omit=dev --omit=optional`, place input in `storage/key_value_stores/default/INPUT.json` and run `npm start`. Read rows in `storage/datasets/default`. Run offline checks with `npm test` and reproduce the live prefill check with `node test/live-prefill.mjs`.

# Actor input Schema

## `publications` (type: `array`):

Newsletter hosts, homepage URLs or single /p/post-slug URLs. Supply 1 to 50. A post URL returns only that post.

## `maxPostsPerPublication` (type: `integer`):

Maximum delivered posts for each supplied publication, from 1 to 5000. Actual archive availability may be lower.

## `sortBy` (type: `string`):

Newest posts or top posts according to the publication feed.

## `searchQuery` (type: `string`):

Optional phrase sent to the archive search within each publication. No global Substack search.

## `publishedAfter` (type: `string`):

Optional ISO date such as 2026-01-01. Only valid with newest sort. Paging stops when an older post is encountered.

## `includeFullText` (type: `boolean`):

Fetch public body HTML and plain text with one extra request per archive post. Paid posts provide only their public preview.

## `includeComments` (type: `boolean`):

Fetch public comments and flatten replies inside each post row. Comments are a separate billing event.

## `maxCommentsPerPost` (type: `integer`):

Maximum included comments and replies per post. Reduced when the remaining charge budget requires it. Zero disables comment fetching.

## `freeOnly` (type: `boolean`):

Skip posts whose audience is restricted to paid or founding subscribers.

## Actor input object example

```json
{
  "publications": [
    "https://www.lennysnewsletter.com"
  ],
  "maxPostsPerPublication": 10,
  "sortBy": "newest",
  "includeFullText": true,
  "includeComments": false,
  "maxCommentsPerPost": 100,
  "freeOnly": false
}
```

# Actor output Schema

## `posts` (type: `string`):

Public posts with publication metadata, optional public text and comments inside each post row.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publications": [
        "https://www.lennysnewsletter.com"
    ],
    "maxPostsPerPublication": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("george.the.developer/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "publications": ["https://www.lennysnewsletter.com"],
    "maxPostsPerPublication": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("george.the.developer/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publications": [
    "https://www.lennysnewsletter.com"
  ],
  "maxPostsPerPublication": 10
}' |
apify call george.the.developer/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,george.the.developer/substack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JffmaZP2OzQwZYJ6v/builds/fqRdTbc0G3bUqoGhF/openapi.json
