# Substack Scraper | Public Full Text, Archives & Comments (`peerless_columbine/substack-scraper`) Actor

Extract public Substack posts, readable text, author and publication metadata, nested comments and archive date filters. Track content access and source coverage explicitly. Search articles by keyword through the public Posts search interface.

- **URL**: https://apify.com/peerless\_columbine/substack-scraper.md
- **Developed by:** [tingyou333 zhuang](https://apify.com/peerless_columbine) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.40 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Substack Scraper | Public Full Text, Archives & Comments

Collect public Substack posts as structured data for publication monitoring, research datasets, editorial workflows, and content-change tracking. Get original public HTML, clean text, author and publication metadata, engagement, podcasts, and nested comments.

This independent tool is not affiliated with Substack. No Substack login, personal cookies, paid subscription, or proxy purchase is required. Paid and restricted article bodies are not unlocked.

### What you can collect

- A single article, a publication archive, or keyword search results, including custom publication domains.
- Public newsletter, podcast, and discussion-thread posts, with date and access filters applied before the result cap.
- Original public HTML plus readable text and a body hash for downstream change detection.
- Public comments and available replies, preserving the source's default order and nested structure.
- Explicit content availability, comment visibility gaps, request limits, and pagination termination reasons.

### Quick start

One public article:

```json
{
  "urls": ["https://www.oneusefulthing.org/p/the-overhang"],
  "maxPostsPerNewsletter": 1,
  "includeContent": true,
  "includePublicationInfo": true,
  "includeComments": true,
  "maxCommentsPerPost": 20,
  "maxRequests": 8
}
```

Find articles about a topic:

```json
{
  "keywords": ["artificial intelligence"],
  "maxSearchResultsPerKeyword": 5,
  "includeContent": true,
  "onlyFree": true,
  "maxRequests": 20
}
```

Run the Actor and open its default dataset. JSON preserves nested comments and publication objects; CSV and Excel are convenient for flat article fields. Read the `SUMMARY` key in the run's key-value store for source errors, warnings, and coverage. A partial result keeps useful rows and explains what was unavailable; a source failure never becomes a fake article row.

### Inputs

All 12 input names from `automation-lab/substack-scraper` are supported. Supply at least one URL or keyword. Each of the `urls` and `keywords` lists accepts at most 500 entries per run.

| Input | Default | Meaning |
|---|---:|---|
| `urls` | `[]` | HTTPS publication homepages, `/archive`, or `/p/slug`; custom domains accepted. |
| `keywords` | `[]` | Search public articles, preserving source relevance order. |
| `maxSearchResultsPerKeyword` | `20` | Maximum emitted matching posts per keyword; 1-100. |
| `maxPostsPerNewsletter` | `100` | Maximum emitted matching posts per archive. `0` removes this output cap. |
| `includeContent` | `true` | Return public HTML and readable text. Restricted requested bodies remain empty. |
| `includeComments` | `false` | Include public comment trees and available replies. |
| `maxCommentsPerPost` | `20` | Count every emitted parent and reply toward this cap; `0` removes the output cap. |
| `includePublicationInfo` | `true` | Include public publication details and subscriber labels when exposed. |
| `contentType` | `"all"` | `all`, `newsletter`, `podcast`, or `thread`. |
| `startDate` | unset | Inclusive UTC calendar date in `YYYY-MM-DD` format. |
| `endDate` | unset | Inclusive UTC calendar date, including the entire day. |
| `onlyFree` | `false` | Keep public `audience=everyone` posts without hidden/unlock requirements. |

Additional controls help keep jobs bounded:

| Input | Default | Meaning |
|---|---:|---|
| `pageSize` | `20` | 1-50 requested archive items per page. |
| `maxRequests` | `1000` | Maximum source attempts including retries and redirects; 1-100,000. DNS resolution may make separate requests. |
| `requestDelaySeconds` | `0.4` | Minimum delay between source requests; 0.1-60 seconds. |
| `requestTimeoutSeconds` | `30` | Per-request timeout; 1-120 seconds. |
| `maxRunSeconds` | `600` | Collection time budget; 1-86,400 seconds. Set the platform run timeout higher to allow the summary to finish. |

Filters apply before result caps and are rechecked against current post details. Search relevance is not chronological, so date-filtered searches may scan multiple pages. Archive offsets advance by the actual response length, including short nonterminal pages. `0` removes an output cap but does not remove request, time, source, or billing limits.

Unknown input fields are rejected rather than silently ignored. Reader-profile URLs, proxy settings, and concurrency controls are not part of the reference's public input contract. Requests are sequential. Tracking query parameters on article URLs are removed.

### Output and migration

Each dataset row is one post. These reference fields are preserved:

- Identity and dates: `postId`, `title`, `subtitle`, `slug`, `url`, `publishedAt`, `updatedAt`, `postType`.
- Access and content: `audience`, `isPaid`, `wordcount`, `coverImage`, `description`, `tags`, `bodyHtml`, `truncatedBodyText`.
- Engagement and media: `reactionCount`, `commentCount`, `childCommentCount`, `restacks`, `hasVoiceover`, `podcastUrl`, `podcastDuration`.
- People and publication: `authorName`, `authorHandle`, `authorPhotoUrl`, `authorBio`, `authorId`, `publicationId`, `publicationName`, `publicationUrl`, `subscriberCount`.
- `comments` and the actual observation timestamp `scrapedAt`.

Migration details matter: `bodyHtml` is **null** when content is not requested; `publicationUrl` is **null** when publication information is not requested; `comments` is **null** when comments are not requested. Requested but empty comments use an empty array. Requested restricted bodies use empty strings with an explicit access status.

Comments contain `id`, `body`, `date`, `editedAt`, `name`, `handle`, `photoUrl`, `reactionCount`, `restacks`, `isAuthor`, `isPinned`, and recursive `replies`. The source's default ordering is retained before applying the nested cap; no guessed ranking formula is applied.

Additional fields provide source context without changing the reference fields:

| Fields | Purpose |
|---|---|
| `bodyText`, `bodyTextLength`, `bodySha256` | Readable content and a hash of the exact returned public HTML. |
| `contentStatus` | `public_full`, `restricted`, `access_unknown`, `unavailable`, or `not_requested`. |
| `sourceUrl`, `discovery` | Public source URL and keyword discovery context. |
| `publication`, `language` | Additional public publication details and source-reported language. |
| `commentsStatus` | `not_requested`, `limit_reached`, `public_snapshot`, `source_partial`, or `unavailable`. |
| `commentsReturned`, `commentsAvailable`, `commentsTruncated` | Emitted count, visible count, and whether your cap removed visible comments. |
| `commentsReportedCount`, `commentsReportedTopLevelCount`, `commentsCountGap` | Source totals and gaps against the visible public snapshot. |
| `commentsSourceHasMissingReplies` | Whether returned reply counts indicate unavailable child comments. |

Unknown metrics remain null instead of becoming zero. Subscriber labels retain the source's strings or rounded values; they are not estimates of paying subscribers. Source `commentCount` already includes replies in the observed API and must not be added to `childCommentCount`.

### Coverage and limits

Public sources have been checked with single posts, keyword pagination, a 110-post archive ending in a real empty page, public podcasts and threads, and nested comments. Bounded private cloud checks include 100 archive rows across three pages, 20 public article bodies, and a 90-comment snapshot. These samples describe tested coverage, not a guarantee for every publication or future source response. Detailed version-specific evidence is retained in the maintainer's validation records.

Deleted, suppressed, and moderated-hidden comments are excluded. One observed post reported 91 comments but exposed 90 public comments; the output preserved the gap and a partial summary. A reported total is not a promise that hidden material can be retrieved. Very large threads may encounter source-side limits.

No paid body, subscriber-only content, login challenge, or paywall is bypassed. Source changes, custom-domain failures, restricted comments, and unavailable publication metadata are visible in `SUMMARY`. Requests retain TLS verification and public-address checks; local/private destinations, credentials in URLs, and nonstandard ports are rejected.

### Pricing and cost controls

One saved post is one billable result. Requested public HTML, readable text, publication information and available comments are included in the post price. There is no separate per-comment or full-text event. The Pricing tab is the source of truth for the effective price.

| Apify plan / discount | Price per 1,000 saved posts |
|---|---:|
| Free / no discount | $0.80 |
| Starter / Bronze | $0.65 |
| Scale / Silver | $0.50 |
| Business / Gold | $0.40 |

The startup event costs **$0.0002** at the default 256 MB memory. Apify charges one startup event up to 1 GB, then one per additional GB. Platform usage is included in the event prices. For example, 100 posts at the no-discount price and default memory cost $0.0802, including public content and requested comments. Even a run with no saved posts may incur the startup event.

Set **Max cost per run** to control event spending; its minimum is $0.001. The Actor stops collecting when the next post no longer fits the remaining budget. The default platform timeout is 660 seconds, leaving time for the default 600-second collection budget and final summary. If you raise `maxRunSeconds`, also raise the platform timeout. An API client's own request timeout is separate.

Content jobs require additional article requests, and comments can increase response size. Start with a small cap and inspect the run's actual usage. For repeat monitoring, use a date window and deduplicate results in your downstream store. `bodySha256` helps identify article changes.

### API and scheduling

With an Apify API token in your environment:

```python
import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("tGHdx3RdqGDv2BbaR").call(run_input={
    "urls": ["https://www.oneusefulthing.org/p/the-overhang"],
    "includeContent": True,
    "maxRequests": 8,
})
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())
summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("SUMMARY")
```

Save an input as an Apify task to schedule repeat runs, or use a webhook/integration to consume the dataset. Deduplication within a run uses `postId`. Cross-run deduplication belongs in your downstream store; summary offsets are diagnostics, not an automatic resume-token API.

### FAQ

**Can I collect an entire archive?** Set `maxPostsPerNewsletter` to `0` and allow sufficient request/time budgets. Collection stops on a real empty page or a stated limit/error. Only the public listing is accessible.

**Why is my run marked partial?** Inspect `SUMMARY`: visible comments may differ from reported totals, a requested source may be unavailable, or a safety budget may have stopped collection.

**Can I retrieve paid articles?** Public metadata and previews may be returned. Restricted article HTML and text remain empty; no subscription access is supplied.

**Do I need a Substack account or paid proxy?** No. The implementation uses anonymous public sources. Source availability can still change.

**Why are some fields null?** They were not requested or not exposed by the source. Nulls preserve that distinction instead of inventing values.

Substack branding identifies the supported source; all brand rights remain with their owners. This Actor is an independent integration.

# Actor input Schema

## `urls` (type: `array`):

Add publication homepages, archive URLs, or individual article URLs. Supports Substack subdomains and custom domains. Use keywords below to discover articles across publications. Provide at least one URL or keyword.

## `maxPostsPerNewsletter` (type: `integer`):

Maximum matching posts returned from each publication URL after filters. Use 0 to remove this result cap; request and time limits still apply. This limit does not restrict keyword search results.

## `keywords` (type: `array`):

Find public articles across Substack by topic, even without publication URLs. Results follow the public article search ranking, not newsletter-name lookup. Each keyword has its own result limit; complete coverage of every publication is not guaranteed.

## `maxSearchResultsPerKeyword` (type: `integer`):

Maximum matching posts emitted for each keyword after filters and run-wide deduplication (1–100). Separate from the limit for publication URLs.

## `includeContent` (type: `boolean`):

Return original publicly available HTML and readable text. When disabled, bodyHtml is null and readable text is empty. When enabled, paid, hidden or locked bodies remain empty; public metadata and previews may still be available.

## `includePublicationInfo` (type: `boolean`):

Include available publication name, owner details, language, public subscriber labels and subscription plans. When disabled, publicationName, publicationUrl, subscriberCount and publication are null. Unpublished values remain null.

## `includeComments` (type: `boolean`):

Collect accessible public comments and nested replies in the source default order. The comment limit counts parents and replies together. Deleted, moderated or otherwise hidden comments are excluded; any visible count gap is reported.

## `maxCommentsPerPost` (type: `integer`):

Maximum comments including all nested replies. 0 removes the output limit; source visibility still applies.

## `contentType` (type: `string`):

Filter by the exact source post type.

## `onlyFree` (type: `boolean`):

Keep only posts whose current audience is everyone, with no unlock or hidden flag.

## `startDate` (type: `string`):

Keep posts published on or after this date, including the whole UTC day. Format: YYYY-MM-DD. Applies to both URL collection and keyword search.

## `endDate` (type: `string`):

Keep posts published on or before this date, including the whole UTC day. Format: YYYY-MM-DD. Leave blank for no upper date bound.

## `pageSize` (type: `integer`):

Requested archive page size; actual response lengths determine offsets.

## `maxRequests` (type: `integer`):

Global request-attempt budget including retries and redirects; stopping is reported in SUMMARY.

## `requestDelaySeconds` (type: `number`):

Minimum delay between HTTP requests. Concurrency is one.

## `requestTimeoutSeconds` (type: `number`):

Per-request wall-clock timeout.

## `maxRunSeconds` (type: `number`):

Whole-run time limit; a partial result is reported when reached.

## Actor input object example

```json
{
  "urls": [
    "https://www.oneusefulthing.org/p/the-overhang"
  ],
  "maxPostsPerNewsletter": 1,
  "keywords": [],
  "maxSearchResultsPerKeyword": 2,
  "includeContent": true,
  "includePublicationInfo": true,
  "includeComments": false,
  "maxCommentsPerPost": 20,
  "contentType": "all",
  "onlyFree": false,
  "pageSize": 20,
  "maxRequests": 1000,
  "requestDelaySeconds": 0.4,
  "requestTimeoutSeconds": 30,
  "maxRunSeconds": 600
}
```

# Actor output Schema

## `posts` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.oneusefulthing.org/p/the-overhang"
    ],
    "maxPostsPerNewsletter": 1,
    "maxSearchResultsPerKeyword": 2
};

// Run the Actor and wait for it to finish
const run = await client.actor("peerless_columbine/substack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.oneusefulthing.org/p/the-overhang"],
    "maxPostsPerNewsletter": 1,
    "maxSearchResultsPerKeyword": 2,
}

# Run the Actor and wait for it to finish
run = client.actor("peerless_columbine/substack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.oneusefulthing.org/p/the-overhang"
  ],
  "maxPostsPerNewsletter": 1,
  "maxSearchResultsPerKeyword": 2
}' |
apify call peerless_columbine/substack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,peerless_columbine/substack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/tGHdx3RdqGDv2BbaR/builds/KJnp3J1ohaJ6S2TGk/openapi.json
