# Substack Posts Monitor - Posts, Likes & New-Post Alerts (`neverempty/substack-posts-monitor`) Actor

Post list for Substack publications from their public archive: title, date, free or paid, likes, comments, restacks, word count, authors and tags. Turn monitoring on and later runs return only the posts published since the last run.

- **URL**: https://apify.com/neverempty/substack-posts-monitor.md
- **Developed by:** [NeverEmpty](https://apify.com/neverempty) (community)
- **Categories:** Social media, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.80 / 1,000 post returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Substack Posts Monitor

Post list for Substack publications, read from each publication's **public archive list**: title, subtitle, URL, publish date, free or paid (`audience`), like count, comment count, restack count, word count, authors and tags. With monitoring on, **later runs return only the posts that were not there before**.

Paste publications as `https://name.substack.com`, the publication's own domain (`https://www.lennysnewsletter.com`), a post link, or just the name before `.substack.com`.

No login, no cookies, no subscription.

### What you get

One row per post, newest first:

| Column | Meaning |
|---|---|
| `publication`, `publicationId` | The host that answered (after a redirect to a custom domain, the custom domain) and Substack's numeric ID of the publication |
| `title`, `subtitle`, `postUrl`, `slug`, `postId` | The post |
| `postDate` | Publish date and time (UTC, ISO 8601) |
| `audience`, `isPaywalled` | `everyone` for a free post, `only_paid` (or another value Substack gives) for a paid one; `isPaywalled` is `true` when `audience` is not `everyone` |
| `reactionCount`, `commentCount`, `restackCount` | Likes, comments and restacks at the moment of the read |
| `wordCount` | Word count of the full post as Substack reports it |
| `authors`, `authorHandles`, `tags`, `sectionName`, `type`, `language` | Bylines, tags, section, post type (`newsletter`, `podcast`, ...) and language |
| `previewText` | The short preview text in the public list (the opening lines) |
| `coverImage`, `podcastDurationSeconds` | Cover image URL; duration for audio posts |
| `input`, `scrapedAt`, `status` | What you typed, when it was read, `ok` for a post row |

Monitoring columns (filled only in monitoring mode): `change` (`first-check` or `new-post`), `isFirstCheck`, `previousCheckedAt`.

A value that is not in the list is `null`, never 0.

The rows are the posts the publication's own archive list returns. If a publication keeps some post types out of that list, they are not in the rows.

### Input

- **`publications`** - one per line. `https://name.substack.com`, `name.substack.com`, `name`, a custom domain, or a post link from the publication. A `name.substack.com` address that redirects to a custom domain is followed. Profile links (`substack.com/@name`) are not publications and come back as an `invalid-input` row.
- **`maxPostsPerPublication`** - newest first, stop after this many posts (default 50, up to 5,000).
- **`postedAfter`** - optional date `YYYY-MM-DD` (UTC). Older posts are not returned and reading stops when the archive reaches them.
- **`monitoringMode`**, **`resetMonitoringState`** - see below.
- **`useProxy`** - requests go directly to each publication; a residential proxy is tried once per refused request, at most three times in a run, only when a site answers HTTP 403 or 429. After a refusal the run stops sending requests and lists what it did not read.

### Monitoring mode

Turn **monitoringMode** on and schedule the Actor. It remembers, per publication, the posts it returned and, on later runs:

- a publication with no new post returns **no post row**; it still costs one check (see "What is charged"), and if no publication had a new post you get one free `no-new-posts` row saying how many were compared;
- a new post comes back once, as `change: "new-post"`;
- a publication it has never seen returns its newest posts once (`first-check`, up to `maxPostsPerPublication`). Posts older than the oldest post returned on that first run are treated as already there and are never returned as new - also when the first run was cut short by a charge limit or a failed page. To start over, run once with `resetMonitoringState`.

On later runs new posts are not cut by `maxPostsPerPublication`. A change in likes, comments or restacks alone does not return a post again; the counts are simply reported in every returned row. A post that is published with a date earlier than the oldest post of the first run is not reported as new.

Publications are remembered one by one (by Substack's publication ID), so adding a publication to the list does not return the others again. Do not put the same publication in two schedules that run at the same moment: the platform's key-value store has no atomic update, so two runs finishing together can overwrite each other's record and return the same post twice.

### What is charged

- **`post-returned`** - one per post row.
- **`publication-checked`** - monitoring mode only: one per publication whose archive was read and has at least one post, whether or not it had a new post.

There is no start fee and no monthly fee. Rows that explain why nothing was returned are free: `no-such-publication`, `not-a-substack-archive`, `robots-disallowed`, `invalid-input`, `duplicate-input`, `unreadable`, `blocked`, `no-posts`, `no-new-posts`, `incomplete`, `not-read`, `budget-reached`. A publication that could not be read is not charged a check.

If you set a maximum total charge for a run, the Actor stops reading before it would exceed it, returns the posts that fit, and adds a free `budget-reached` row and a `not-read` row naming the publications it did not read. In monitoring mode a publication is read only if the limit has room for its check plus one post row.

### How it reads a publication

1. `GET https://<publication>/robots.txt`. If the rules for all crawlers (`User-agent: *`) disallow `/api/v1/archive`, the archive is **not requested** and you get a free `robots-disallowed` row.
2. `GET https://<publication>/api/v1/archive?sort=new&offset=N&limit=M` (at most 50 per request), newest first, until the number you asked for, the `postedAfter` date, an empty page, or - in monitoring mode - the posts it already returned. At most 150 archive requests per publication.

Requests carry an honest User-Agent (`neverempty-substack-posts-monitor`) and are sent one at a time with a pause between them. A refused request (HTTP 403 or 429) is sent once more through a residential proxy when `useProxy` is on, and is otherwise not retried; after a refusal the run stops. Individual post pages, comments, private feeds and anything behind a login are not requested.

### Rows that are not posts

| `status` | When |
|---|---|
| `no-such-publication` | The address answered HTTP 404 for the archive: a misspelled name, a removed publication, or a site that is not on Substack |
| `not-a-substack-archive` | The address answered, but not with a Substack post list |
| `robots-disallowed` | The publication's robots.txt disallows the archive request |
| `blocked` | HTTP 403 or 429, or a verification page. The run stops sending requests |
| `unreadable` | No answer, a server error after one retry, or an answer in an unexpected format |
| `incomplete` | A later page of a publication's archive could not be read; the posts read before it are returned |
| `no-posts` | The archive is empty, or has no post on or after `postedAfter` |
| `no-new-posts` | Monitoring: nothing new in any publication |
| `invalid-input`, `duplicate-input`, `not-read`, `budget-reached` | Input lines that were not read, and why |

### Notes

- Post bodies are not returned, for paid or free posts: a row carries what the public archive list shows for that post, including the short preview text.
- Like, comment and restack counts are the publication's public figures at the moment of the read; they keep moving on live posts.
- This Actor is not affiliated with or endorsed by Substack. You are responsible for how you use the data; respect Substack's Terms of Use and the rights of the writers. The data returned here is the public listing data only, not the posts themselves.

# Actor input Schema

## `publications` (type: `array`):

Publication addresses: https://name.substack.com, the publication's own domain (https://www.lennysnewsletter.com), a post link from that publication, or just the name before .substack.com. One publication per line. For each one the Actor first reads its robots.txt and then its public archive list (/api/v1/archive). Profile links (substack.com/@name) are not publications and come back as an "invalid-input" row that is not charged; an address with no Substack publication comes back as a "no-such-publication" row that is not charged.

## `maxPostsPerPublication` (type: `integer`):

Newest posts first; reading a publication stops after this many post rows. In monitoring mode this limits only the first run for a publication (how many of its newest posts are returned and remembered); on later runs every new post is returned.

## `postedAfter` (type: `string`):

Optional date (YYYY-MM-DD, UTC). Posts published before it are not returned and reading stops when the archive reaches them.

## `monitoringMode` (type: `boolean`):

Off = return the newest posts of every publication you listed, charged per post row. On = the posts returned for each publication are remembered, and later runs return only posts that were not there before (published after the oldest post returned on the first run and not returned yet). In monitoring mode each publication whose archive was read and has at least one post is charged as one check (publication-checked), and each post that is returned is also charged as one post row (post-returned). The first monitoring run for a publication returns its newest posts once. Like, comment and restack counts are reported in every returned row but a change in those counts alone does not return a post again. Publications are remembered one by one, so adding a publication does not return the others again. Do not put the same publication in two schedules that run at the same moment.

## `resetMonitoringState` (type: `boolean`):

Clears everything remembered for this Actor on your account before this run reads anything, so this run (with monitoring on) returns the newest posts of each publication once again and charges for them. Use it for one run and turn it off again: left on in a schedule, every run returns and charges the same posts. This affects all your monitoring runs, because publications are remembered one by one rather than per list.

## `useProxy` (type: `boolean`):

Requests go directly to each publication. If one answers HTTP 403 or 429, the request is sent one more time through an Apify residential proxy, at most three times in a run; if it is refused again the run stops sending requests and lists what it did not read. Turn this off to never use a proxy.

## Actor input object example

```json
{
  "publications": [
    "https://www.lennysnewsletter.com",
    "https://platformer.substack.com"
  ],
  "maxPostsPerPublication": 50,
  "monitoringMode": false,
  "resetMonitoringState": false,
  "useProxy": true
}
```

# Actor output Schema

## `results` (type: `string`):

One row per Substack post from the publication's public archive list: publication, title, subtitle, URL, publish date, audience (free or paid), like, comment and restack counts, word count, authors, tags and the public preview text and, in monitoring mode, whether it is a first check or a new post. Post bodies are not included. Addresses with no Substack publication, lines that are not publication addresses, archives that could not be read or are disallowed by robots.txt, and runs with no new post come back as their own rows and are not charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publications": [
        "https://www.lennysnewsletter.com",
        "https://platformer.substack.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neverempty/substack-posts-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "publications": [
        "https://www.lennysnewsletter.com",
        "https://platformer.substack.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("neverempty/substack-posts-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publications": [
    "https://www.lennysnewsletter.com",
    "https://platformer.substack.com"
  ]
}' |
apify call neverempty/substack-posts-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neverempty/substack-posts-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JFFMcJoB033LxdYgD/builds/UET1uwwu3eEwJNFQr/openapi.json
