# Substack Scraper: Newsletter Posts & Stats (`axiorasolutions/substack-archive-scraper`) Actor

Export a Substack newsletter's full archive with engagement data. Returns title, subtitle, publish date, audience tier, reactions, comments, restacks, word count, tags, bylines and canonical URL per post, plus optional full post text. Accepts subdomains and custom domains.

- **URL**: https://apify.com/axiorasolutions/substack-archive-scraper.md
- **Developed by:** [Axiora Solutions](https://apify.com/axiorasolutions) (community)
- **Categories:** News, Marketing, For creators
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.83 / 1,000 newsletter posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Substack Scraper — export any newsletter archive with engagement data

**Substack scraper** that exports a publication's entire newsletter archive as structured rows: every post with its reactions, comments, restacks, word count, bylines and canonical URL. It works on both `*.substack.com` subdomains and **custom domains** (`www.noahpinion.blog`, `www.slowboring.com`), and free posts need **no login or API key**. The fastest way to try it: leave the prefilled publications in place and click **Start**.

### What you get

- `title`, `subtitle`, `slug`, `publishedAt` and canonical `postUrl` for every post.
- `reactionTotal` plus the raw `reactions` breakdown by emoji.
- `commentCount`, `childCommentCount` (replies) and `restackCount`.
- `audience` and `isFree` — the exact paywall tier each post sits behind.
- `wordCount`, `tags`, `sectionName`, `postType` and `isPodcast`.
- `bylines` (with `isGuest`) and `firstAuthorName`, plus optional clean `bodyText`.

### Quick start

1. Open the Actor on Apify and list your **Publications** — a subdomain or a custom domain, one per line.
2. Set **Max posts per publication** and **Max posts for the whole run** (defaults are fine), or enable **Audience filter**, tags, keywords and **Published after** to trim the run.
3. Click **Start**, then read the **Posts**, **Engagement** and **Content** dataset tabs.
4. Turn on **Fetch full post text** only if you need the `bodyText` field.

Minimal input:

```json
{
  "publications": ["astralcodexten.substack.com", "newsletter.pragmaticengineer.com"],
  "maxPostsPerPublication": 100,
  "maxPostsTotal": 500,
  "audienceFilter": "all",
  "includePostText": false
}
```

### Example output

One representative dataset row:

```json
{
  "ok": true,
  "errorCode": null,
  "postId": "218431605",
  "slug": "our-ai-midwife",
  "title": "Our AI Midwife",
  "subtitle": "A birth story",
  "publication": "astralcodexten.com",
  "publicationUrl": "https://astralcodexten.substack.com",
  "canonicalUrl": "https://www.astralcodexten.com/p/our-ai-midwife",
  "postUrl": "https://www.astralcodexten.com/p/our-ai-midwife",
  "publishedAt": "2026-10-02T01:08:13.190Z",
  "audience": "everyone",
  "isFree": true,
  "postType": "newsletter",
  "isPodcast": false,
  "coverImageUrl": "https://substackcdn.com/image/fetch/...",
  "wordCount": 1556,
  "commentCount": 96,
  "childCommentCount": 41,
  "restackCount": 18,
  "reactions": { "❤": 328, "🔥": 12 },
  "reactionTotal": 340,
  "bylines": [
    { "id": 12345, "name": "Scott Alexander", "handle": "astralcodexten", "isGuest": false }
  ],
  "firstAuthorName": "Scott Alexander",
  "tags": [],
  "truncatedBodyText": "The story of a birth with an AI doula...",
  "bodyText": null,
  "bodyChars": 0,
  "sectionName": null,
  "sectionSlug": null,
  "contentHash": "7cb41a09e35d8260",
  "scrapedAt": "2026-10-02T12:00:00.000Z"
}
```

### Why engagement data is the whole point

An RSS feed gives you the title and the date. A CMS export gives you the text. **Neither tells you what worked.** Reaction counts, comments, replies and restacks are the signal that turns an archive into a strategy: which topics land, which formats earn a conversation, and where a publication's paywall sits.

### What this Substack scraper returns

- 📊 **Engagement per post** — `reactionTotal` plus the raw `reactions` breakdown by emoji, `commentCount`, `childCommentCount` (replies) and `restackCount`. Sort on any of them.
- 🚪 **Paywall visibility** — `audience` is `everyone` for a free post or the paid tier the author chose. `isFree` gives you a boolean. Filter to free posts to study the free funnel, or to paid posts to see what a publication decides is worth charging for.
- ✍️ **Bylines** — every author with id, name, handle and an `isGuest` flag, so multi-author and guest-post publications remain analysable.
- 🎙️ **Podcast posts are handled** — `postType` and `isPodcast` distinguish audio issues from written ones, and `wordCount` plus `sectionName` round out the metadata.
- 📰 **Multiple publications in one run** — the main input is an array, so a competitive set of twenty newsletters is one run, one dataset, with `publication` on every row to keep them apart.
- 📄 **Optional full text** — `bodyText` gives clean plain text from the post page for research or embedding. Paywalled posts return the same truncated preview a non-subscriber sees: **this Actor does not attempt to bypass paywalls.**
- 🎯 **Filters that cut your bill** — audience tier, tags, keywords and `publishedAfter` all run before rows are written, so filtered posts are never charged.
- 🔁 **Built for schedules** — `contentHash` covers the post plus its engagement, so a weekly run tells you which posts gained traction *after* publication, not just which ones are new.
- 🛟 **Honest failures** — a private or empty publication gives an `ok: false` row explaining that, and every other publication still lands.

Running on Apify adds scheduling, webhooks, monitoring, API and SDK access, and one-click export to JSON, CSV, Excel, Google Sheets and 20+ integrations.

### How to use it

1. Add publications to **Publications** — subdomain or custom domain, either works.
2. Set **Max posts per publication** and **Max posts for the whole run**.
3. Use **Audience filter** to study the free tier, the paid tier, or podcasts only.
4. Turn on **Fetch full post text** if you need `bodyText`; leave it off if `wordCount`, `title` and `truncatedBodyText` are enough.
5. Click **Start**, then use the **Posts**, **Engagement** and **Content** dataset tabs.

#### How do I find a publication's best-performing posts ever?

Set **Max posts per publication** high, leave **Audience filter** on All, then sort the **Engagement** view by `reactionTotal`. That is the canonical "greatest hits" analysis, and it costs one run.

### How much does it cost to scrape a Substack archive?

Pricing is **pay per event** with one event:

| Event | What triggers it | Billed |
|---|---|---|
| Newsletter post | One post written to the dataset | per post |
| Actor start | Once per run, platform fee | per run |

**Fetching full post text is included in the per-post price** — it costs you run time, not money. Posts removed by filters or de-duplication are **not** billed, and a publication that cannot be read is **not** billed.

A 200-post archive is 200 billed events, and twenty publications at 200 posts each is 4,000 — that is the whole calculation. Compute, bandwidth and storage are included; there is no separate platform-usage charge on top.

Set **Max cost per run** in the run options for a hard ceiling. Higher Apify plans get progressively lower per-post pricing through Apify Store tier discounts.

Evaluating? Set **Max posts per publication** to `10` with the prefilled publications.

### Example input

A fuller run with filters:

```json
{
  "publications": [
    "astralcodexten.substack.com",
    "newsletter.pragmaticengineer.com",
    "https://newsletter.pragmaticengineer.com"
  ],
  "maxPostsPerPublication": 150,
  "maxPostsTotal": 400,
  "publishedAfter": "1 year",
  "audienceFilter": "all",
  "tags": ["AI"],
  "keywords": ["LLM"],
  "includePostText": false
}
```

### Error rows

A publication that cannot be read does not fail the run; it produces one row like this:

```json
{
  "ok": false,
  "errorCode": "NOT_FOUND",
  "requestedInput": "not-a-real-publication.substack.com",
  "error": {
    "code": "NOT_FOUND",
    "message": "Archive API failed for https://not-a-real-publication.substack.com: HTTP 404 from https://not-a-real-publication.substack.com/api/v1/archive. Check that the publication is public and that the URL is correct.",
    "httpStatus": 404
  }
}
```

### Use cases

- **Content strategy research** — find which topics, lengths and formats actually earn reactions and comments.
- **Competitive newsletter analysis** — publishing cadence, paywall placement and growth of a peer set in one dataset.
- **Creator benchmarking** — compare engagement per word across publications.
- **Sponsorship and media planning** — quantify a newsletter's engaged audience before buying.
- **Research corpora** — `bodyText` plus `canonicalUrl` and `postId` is a clean, citable archive for study or embedding.
- **Trend detection** — schedule weekly and diff on `contentHash` to catch posts that keep gaining traction.

### Related Actors by Axiora Solutions

| Actor | Use it for |
|---|---|
| **News & RSS Feed Scraper** | News coverage feeding the same research, including Google News |
| **App Store Review Scraper** | Voice-of-customer data to pair with creator and product research |
| **Domain Contact Enricher** | Contact and company profiles for the operators behind the publications |

### Frequently asked questions

#### Do custom domains work, or only substack.com?

Both. The Actor uses the publication's own origin, so `newsletter.pragmaticengineer.com` and `astralcodexten.substack.com` behave identically. This matters because many of the largest publications have moved to custom domains.

#### Can it read paywalled posts?

It reads what Substack serves publicly. For a paywalled post you get the metadata, the engagement counts, and the same truncated preview a non-subscriber sees. **This Actor does not bypass paywalls**, and the README says so rather than leaving you to discover it. Metadata and engagement numbers for paid posts are still fully available, which is usually the analysable part anyway.

#### Why is `bodyText` null?

Two reasons. Either **Fetch full post text** is off, in which case nothing is fetched by design, or the fetch failed for that post — in which case `bodyText` is null and the row still carries `truncatedBodyText` plus full metadata. The post is still counted and still useful.

#### How many posts can I export from one publication?

The whole archive. The Actor pages the publication's archive API until it reaches your **Max posts per publication** limit or the archive ends. Large publications with thousands of posts are bounded only by your limits and run-cost ceiling.

#### What is a restack?

Substack's equivalent of a share or repost: a reader rebroadcasting the post to their own subscribers. It is the strongest distribution signal a post can have, which is why it is a first-class field here.

#### Do I need a Substack account or API key?

No. The publication archive API is public, and the Actor reads it directly. Nothing in the input is a credential, so an autonomous agent can call this Actor without a human.

#### Can I run this on a schedule?

Yes. Weekly is the useful cadence: new posts appear, and existing posts accumulate reactions and comments. Because `contentHash` includes the reaction total, a diff between runs surfaces posts that are still gaining traction, which is a signal you cannot get from the publish date alone.

#### Is scraping Substack archives legal?

The Actor requests the same public JSON API that a browser uses when you visit a publication's archive page, and it identifies itself. Post titles, URLs and engagement counts are public data. Full post text is the author's copyrighted work — **check their licensing before republishing it or using it to train a model.** This is not legal advice.

#### Something looks wrong — how do I report it?

Open the **Issues** tab on this Actor page with the publication URL and what you expected to see.

***

Runnable examples and how-to guides for these Actors: [github.com/batow133/axiora-apify-actors](https://github.com/batow133/axiora-apify-actors)

# Actor input Schema

## `publications` (type: `array`):

One publication per entry. Substack subdomain or custom domain both work: astralcodexten.substack.com, www.noahpinion.blog, www.slowboring.com. Bare domains, full URLs and subpaths are all accepted.

## `maxPostsPerPublication` (type: `integer`):

Stop after this many posts from each publication, newest first.

## `maxPostsTotal` (type: `integer`):

Hard ceiling across all publications.

## `publishedAfter` (type: `string`):

Keep only posts published on or after this date. Accepts 2026-01-01 or a relative value such as 6 months.

## `includePostText` (type: `boolean`):

Fetch each post page and extract the main body as clean plain text. Adds one request per free post. Paywalled posts return truncated text from the API, the same as a non-subscriber sees.

## `postTextMaxChars` (type: `integer`):

Truncate extracted post text at this many characters.

## `audienceFilter` (type: `string`):

all keeps every post. free keeps only posts visible without a subscription. paid keeps only subscriber-only posts. podcast keeps only audio posts.

## `tags` (type: `array`):

Keep only posts carrying at least one of these tags, case-insensitive.

## `keywords` (type: `array`):

Keep only posts whose title, subtitle or truncated preview contains one of these words, case-insensitive.

## `proxyConfiguration` (type: `object`):

Optional. The Substack API normally answers direct requests.

## Actor input object example

```json
{
  "publications": [
    "astralcodexten.substack.com",
    "www.noahpinion.blog",
    "https://newsletter.pragmaticengineer.com"
  ],
  "maxPostsPerPublication": 100,
  "maxPostsTotal": 500,
  "publishedAfter": "6 months",
  "includePostText": false,
  "postTextMaxChars": 8000,
  "audienceFilter": "all",
  "tags": [
    "AI",
    "startups"
  ],
  "keywords": [
    "LLM",
    "funding"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `posts` (type: `string`):

Every post collected, newest first per publication.

## `runSummary` (type: `string`):

Per-publication posts found and written, filter effects, billing and network totals.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publications": [
        "astralcodexten.substack.com",
        "newsletter.pragmaticengineer.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiorasolutions/substack-archive-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "publications": [
        "astralcodexten.substack.com",
        "newsletter.pragmaticengineer.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("axiorasolutions/substack-archive-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publications": [
    "astralcodexten.substack.com",
    "newsletter.pragmaticengineer.com"
  ]
}' |
apify call axiorasolutions/substack-archive-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiorasolutions/substack-archive-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/re4nRNpK5C77DDcId/builds/g5oa7l1D7jHu5khSg/openapi.json
