# Substack & Newsletter Archive Scraper (Ghost, Beehiiv) (`everyotherfriday/newsletter-archive`) Actor

Export publicly accessible newsletter posts from Substack, Ghost and Beehiiv with HTML, text and metadata; optional public comments on Substack. Paid posts contain public previews only. Ghost and Beehiiv archive coverage depends on available feeds and sitemaps. $0.003 per post.

- **URL**: https://apify.com/everyotherfriday/newsletter-archive.md
- **Developed by:** [Paul Vasquez](https://apify.com/everyotherfriday) (community)
- **Categories:** Social media, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$3.00 / 1,000 newsletter posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Newsletter Archive Scraper

Collect public newsletter posts from Substack, Ghost, and Beehiiv into a consistent Apify dataset. Supply publication homepages, including custom domains, and receive one row per matching post. The actor detects the publishing platform, discovers public posts, and optionally retrieves article content and Substack comments. It suits editorial research, newsletter inventories, content monitoring, and downstream text analysis.

The actor uses Python 3.12, the Apify SDK, and httpx. Beautiful Soup extracts article content; defusedxml parses feeds and sitemaps. No browser, account login, subscription cookie, or private API credential is required. Paywalled posts can be represented by their public metadata and available preview content. This actor does not unlock subscriber-only text.

### Input

```json
{
  "publications": [
    "https://www.lennysnewsletter.com",
    "https://newsletter.pragmaticengineer.com",
    "https://ghost.org/blog/",
    "https://thenewsletternewsletter.beehiiv.com"
  ],
  "maxPosts": 20,
  "includeBody": true,
  "includeComments": false,
  "onlyFree": false,
  "timeoutSecs": 30
}
```

`publications` is a required, nonempty array of HTTP or HTTPS publication URLs. Repeated inputs are deduplicated, query strings and fragments are removed, and redirects are followed. Preserve a publication's path when it lives under a directory. URLs containing usernames or passwords are rejected.

`maxPosts` defaults to 100 and limits successful matching posts **per publication**, not across the entire run. Errors and filtered posts do not consume this limit. `includeBody` defaults to true: Substack post JSON supplies article HTML, while Ghost and Beehiiv pages supply the largest article block, or the largest main block when no article exists. Each content row includes plain text and a computed word count. With this option false, body fields and word count are null.

`includeComments` defaults to false. Enable it to request the public Substack comments endpoint. The returned comments array flattens available nested replies, linked by `parentCommentId`. Comments are returned without commenter identities. Authentication failures and missing comment endpoints produce an empty array. Comment pagination is not implemented, so this is the endpoint's available response, not a promise of every historical comment. Other platforms return an empty array.

`publishedAfter` optionally accepts an exclusive ISO date, such as `2026-01-01`. Posts dated that day or earlier, and posts without a usable publication date, are excluded. Sitemap modification dates are deliberately not substituted for publication dates. `onlyFree` defaults to false; enable it to skip identified paywalled posts. On generic platforms, this option requests page metadata even when body output is disabled. Paywall detection uses public structured data and recognizable markup; absence of a paywall signal is not proof of unrestricted access.

`timeoutSecs` defaults to 30 and applies to each network operation. `proxyConfiguration` accepts the SDK's standard Apify Proxy or custom proxy settings. Omit it for direct connections. Configure paid proxies separately in your own environment; the supplied live-validation input uses none.

### Output and discovery

Every post includes `publication`, `platform`, `postId`, `slug`, `title`, `subtitle`, `authors`, `publishedAt`, `url`, `canonicalUrl`, `coverImage`, `paywalled`, `audience`, `wordCount`, `bodyHtml`, `bodyText`, `likes`, `commentCount`, `comments`, `tags`, and `source`. Missing scalar metadata is null; unavailable arrays are empty. Authors are publication-author names; each comment contains only `commentId`, `postId`, `parentCommentId`, `body` (text), `createdAt`, `likeCount`, `replyCount`, and `isAuthorReply` (true only when explicitly marked by the source). Word count measures extracted public text, including previews, using Unicode word tokens.

Substack discovery paginates its public archive twelve entries at a time, newest first. Details come from `/api/v1/posts/{slug}` and comments from `/api/v1/post/{id}/comments`. The source field records the archive request. Generic discovery first tries advertised RSS/Atom links and conventional feed paths, then bounded sitemap traversal. Beehiiv sitemap candidates must use `/p/`; tag, author, subscription, and landing pages are excluded where recognizable. Source records the feed or sitemap URL.

RSS is often a recent-post window. If a usable feed returns fifteen posts, this actor returns those fifteen even when the cap is twenty; it does not promise a complete historical archive. Sitemap order is publisher-defined. JavaScript-only pages, unusual themes, inaccessible feeds, and layout changes can limit extraction. Original article HTML is output as data and should be sanitized before rendering in another application.

### Reliability and pricing

HTTP 429 responses receive two retries with bounded Retry-After or exponential delays. Duplicate posts are suppressed within each publication. Individual detail failures become separate uncharged error rows; remaining posts and publications continue. SUMMARY and SUMMARY-N key-value records contain counts, paywalled counts, errors, and elapsed seconds.

The intended PPE event is `post-scraped` at **$0.003 per successful post row**. SDK charged writes enforce the available event limit. Error rows have no charged event. Configure this single custom event in Console and disable synthetic dataset/start events before publishing; the pricing JSON is a deployment specification, not proof of activated billing. Local runs do not bill.

### Local development

Create a Python 3.12 `.venv`, install `requirements.txt`, and run:

```powershell
.venv/Scripts/python.exe -m unittest discover -s tests -v
apify validate-schema .actor/input_schema.json
powershell -NoProfile -ExecutionPolicy Bypass -File validation/run_live.ps1
.venv/Scripts/python.exe validation/check_live.py
```

The live script copies INPUT.json into fresh local SDK storage and invokes `python -m src`. See VALIDATION.md and validation JSON artifacts for measured results and limitations. Docker and hosted billing require separate deployment validation. This package is prepared locally and has not been pushed or published.

### Example output

One real saved dataset row, trimmed by omitting fields only. Source: `storage/live-20260926-042940/datasets/default/000000001.json`. This is historical validation evidence, not a live response.

```json
{
  "publication": "https://www.lennysnewsletter.com",
  "platform": "substack",
  "postId": "216168140",
  "title": "Advanced evals: How to find (and fix) hidden AI failures in your product",
  "publishedAt": "2026-09-22T12:45:14.998000Z",
  "url": "https://www.lennysnewsletter.com/p/advanced-evals-how-to-find-and-fix",
  "paywalled": true,
  "audience": "only_paid",
  "wordCount": 3260
}
```

The paywall flag and audience describe the saved source response; the word count measures extracted public text, not a verified complete subscriber article. Body HTML, comments, and other fields are omitted here for readability.

### Use cases

- A newsletter sponsorship buyer can compare recent topics and publishing dates across shortlisted publications before preparing an outreach brief.
- An editorial research team can collect public article metadata into a reading queue and retain source URLs for attribution.
- A publisher operations manager can inventory discoverable posts across a Substack, Ghost, or Beehiiv portfolio before planning a content migration.
- A content strategy agency can group available public text by theme to prepare a client briefing, checking paywall flags before quoting passages.

**Pricing example:** 1,000 successful `post-scraped` events x $0.003 = **$3.00**, computed from `.actor/pay_per_event.json`. This is the declared event subtotal; it does not verify active hosted billing or include any separately applicable platform or proxy costs.

### Limitations

Discovery is limited by the public feed or sitemap and the requested per-publication cap. Preview text, missing metadata, and unavailable comments can make downstream comparisons incomplete. Review source pages before interpreting an absent post or empty comments array as evidence of inactivity.

# Actor input Schema

## `publications` (type: `array`):

Public Substack, Ghost or Beehiiv publication URLs, including custom domains.

## `maxPosts` (type: `integer`):

Maximum matching successful posts per publication.

## `includeBody` (type: `boolean`):

Fetch public post content and return HTML, plain text and word count.

## `includeComments` (type: `boolean`):

Request public Substack comments when available; other platforms return an empty array.

## `onlyFree` (type: `boolean`):

Skip posts identified as paywalled. Generic feed-only metadata may not expose paywall status.

## `publishedAfter` (type: `string`):

Optional exclusive ISO date YYYY-MM-DD. Undated posts are excluded.

## `timeoutSecs` (type: `integer`):

Per-network-operation timeout in seconds.

## `proxyConfiguration` (type: `object`):

Optional Apify Proxy or custom proxy configuration.

## Actor input object example

```json
{
  "publications": [
    "https://www.lennysnewsletter.com",
    "https://newsletter.pragmaticengineer.com"
  ],
  "maxPosts": 10,
  "includeBody": true,
  "includeComments": false,
  "onlyFree": false,
  "timeoutSecs": 30
}
```

# Actor output Schema

## `posts` (type: `string`):

All normalized post and error rows.

## `postsCsv` (type: `string`):

Dataset in CSV format.

## `summaries` (type: `string`):

SUMMARY and SUMMARY-N counts and timings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "publications": [
        "https://www.lennysnewsletter.com",
        "https://newsletter.pragmaticengineer.com"
    ],
    "maxPosts": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("everyotherfriday/newsletter-archive").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "publications": [
        "https://www.lennysnewsletter.com",
        "https://newsletter.pragmaticengineer.com",
    ],
    "maxPosts": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("everyotherfriday/newsletter-archive").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "publications": [
    "https://www.lennysnewsletter.com",
    "https://newsletter.pragmaticengineer.com"
  ],
  "maxPosts": 10
}' |
apify call everyotherfriday/newsletter-archive --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,everyotherfriday/newsletter-archive"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/eWnL3eCzmbbaEKgiw/builds/bHsW7aQexMp1hJfa4/openapi.json
