# Substack Scraper — Posts to JSON & AI-Ready Markdown (`harmonious_handball/substack-scraper-markdown`) Actor

Scrape every post from any Substack newsletter (including custom domains). Get titles, authors, dates, likes, comments, paywall status, tags and full post text as clean Markdown for AI/RAG. No login, no browser.

- **URL**: https://apify.com/harmonious\_handball/substack-scraper-markdown.md
- **Developed by:** [Clean Data Labs](https://apify.com/harmonious_handball) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 post scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Substack Scraper: Posts to JSON & AI-Ready Markdown

Scrape every post from any Substack newsletter, including publications on **custom domains**. You get clean, structured data plus the **full text of free posts as Markdown**, ready to drop into ChatGPT, Claude, a vector database, or a RAG pipeline.

No login. No browser. Fast and cheap.

### What you get for each post

| Field | Example |
|---|---|
| `title`, `subtitle` | "Shea Serrano did the math on podcasting..." |
| `url`, `slug` | `https://on.substack.com/p/shea-serrano-podcast` |
| `postDate` | `2026-09-10T17:58:26Z` |
| `authors` | `["Arielle Swedback", "Shea Serrano"]` |
| `likes`, `comments`, `restacks` | `1350`, `0`, `63` |
| `wordCount` | `3649` |
| `audience`, `isPaywalled` | `everyone`, `false` |
| `section`, `tags` | `Stories`, `[]` |
| `type` | `newsletter` / `podcast` / `thread` |
| `markdown` | Full post body as clean Markdown (free posts) |
| `coverImage`, `podcastUrl` | URLs |

### Use cases

- **AI & RAG**: turn a newsletter's whole archive into a Markdown knowledge base
- **Competitor & content research**: see which topics get the most likes and comments
- **Sponsorship & partnership research**: find active newsletters in your niche
- **Monitoring**: schedule daily runs with `since` to capture only new posts

### How to use

1. Add one or more publications: `lenny`, `on.substack.com`, or `https://www.lennysnewsletter.com`
2. (Optional) Set **Max posts**, **Only posts after** a date, or **Keyword filter**
3. Click **Start** and export as JSON, CSV, Excel, or via API

#### Example input

```json
{
  "publications": ["https://www.lennysnewsletter.com", "on.substack.com"],
  "maxPostsPerPublication": 100,
  "since": "2026-01-01",
  "includeMarkdown": true
}
```

### Pricing

You pay per post saved. There are no monthly fees, and filtered-out posts cost nothing.

### Notes

- Paywalled posts return metadata only. Full text is included for free posts.
- Only publicly available data is collected.

### Support

Found a bug or need a field added? Open an issue on the **Issues** tab. Replies usually come within 24 hours.

# Actor input Schema

## `publications` (type: `array`):

Publication names or URLs. Examples: 'lenny', 'on.substack.com', 'https://www.lennysnewsletter.com'. Custom domains work.

## `maxPostsPerPublication` (type: `integer`):

0 = no limit (whole archive).

## `since` (type: `string`):

Optional date, e.g. 2026-01-01.

## `keywords` (type: `array`):

Only keep posts whose title/subtitle/description contains any of these words.

## `freePostsOnly` (type: `boolean`):

Skip paywalled posts.

## `includeMarkdown` (type: `boolean`):

Full post body converted to clean Markdown (free posts only). Great for AI, RAG and summarization.

## `includeHtml` (type: `boolean`):

Also include the post body as raw HTML (free posts only). Leave off unless you need the original formatting.

## `delaySeconds` (type: `number`):

Pause between archive pages. Raise it if you scrape very large archives and hit rate limits.

## Actor input object example

```json
{
  "publications": [
    "on.substack.com"
  ],
  "maxPostsPerPublication": 50,
  "freePostsOnly": false,
  "includeMarkdown": true,
  "includeHtml": false,
  "delaySeconds": 0.5
}
```

# Actor output Schema

## `posts` (type: `string`):

All scraped posts from the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publications": [
        "on.substack.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("harmonious_handball/substack-scraper-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "publications": ["on.substack.com"] }

# Run the Actor and wait for it to finish
run = client.actor("harmonious_handball/substack-scraper-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publications": [
    "on.substack.com"
  ]
}' |
apify call harmonious_handball/substack-scraper-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,harmonious_handball/substack-scraper-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HLfRG1mFlI8alfDfd/builds/wFDDoZs843AqyraOO/openapi.json
