Substack & Newsletter Archive Scraper (Ghost, Beehiiv) avatar

Substack & Newsletter Archive Scraper (Ghost, Beehiiv)

Pricing

$3.00 / 1,000 newsletter posts

Go to Apify Store
Substack & Newsletter Archive Scraper (Ghost, Beehiiv)

Substack & Newsletter Archive Scraper (Ghost, Beehiiv)

Export publicly accessible newsletter posts from Substack, Ghost and Beehiiv with HTML, text and metadata; optional public comments on Substack. Paid posts contain public previews only. Ghost and Beehiiv archive coverage depends on available feeds and sitemaps. $0.003 per post.

Pricing

$3.00 / 1,000 newsletter posts

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Newsletter Archive Scraper

Collect public newsletter posts from Substack, Ghost, and Beehiiv into a consistent Apify dataset. Supply publication homepages, including custom domains, and receive one row per matching post. The actor detects the publishing platform, discovers public posts, and optionally retrieves article content and Substack comments. It suits editorial research, newsletter inventories, content monitoring, and downstream text analysis.

The actor uses Python 3.12, the Apify SDK, and httpx. Beautiful Soup extracts article content; defusedxml parses feeds and sitemaps. No browser, account login, subscription cookie, or private API credential is required. Paywalled posts can be represented by their public metadata and available preview content. This actor does not unlock subscriber-only text.

Input

{
"publications": [
"https://www.lennysnewsletter.com",
"https://newsletter.pragmaticengineer.com",
"https://ghost.org/blog/",
"https://thenewsletternewsletter.beehiiv.com"
],
"maxPosts": 20,
"includeBody": true,
"includeComments": false,
"onlyFree": false,
"timeoutSecs": 30
}

publications is a required, nonempty array of HTTP or HTTPS publication URLs. Repeated inputs are deduplicated, query strings and fragments are removed, and redirects are followed. Preserve a publication's path when it lives under a directory. URLs containing usernames or passwords are rejected.

maxPosts defaults to 100 and limits successful matching posts per publication, not across the entire run. Errors and filtered posts do not consume this limit. includeBody defaults to true: Substack post JSON supplies article HTML, while Ghost and Beehiiv pages supply the largest article block, or the largest main block when no article exists. Each content row includes plain text and a computed word count. With this option false, body fields and word count are null.

includeComments defaults to false. Enable it to request the public Substack comments endpoint. The returned comments array flattens available nested replies, linked by parentCommentId. Comments are returned without commenter identities. Authentication failures and missing comment endpoints produce an empty array. Comment pagination is not implemented, so this is the endpoint's available response, not a promise of every historical comment. Other platforms return an empty array.

publishedAfter optionally accepts an exclusive ISO date, such as 2026-01-01. Posts dated that day or earlier, and posts without a usable publication date, are excluded. Sitemap modification dates are deliberately not substituted for publication dates. onlyFree defaults to false; enable it to skip identified paywalled posts. On generic platforms, this option requests page metadata even when body output is disabled. Paywall detection uses public structured data and recognizable markup; absence of a paywall signal is not proof of unrestricted access.

timeoutSecs defaults to 30 and applies to each network operation. proxyConfiguration accepts the SDK's standard Apify Proxy or custom proxy settings. Omit it for direct connections. Configure paid proxies separately in your own environment; the supplied live-validation input uses none.

Output and discovery

Every post includes publication, platform, postId, slug, title, subtitle, authors, publishedAt, url, canonicalUrl, coverImage, paywalled, audience, wordCount, bodyHtml, bodyText, likes, commentCount, comments, tags, and source. Missing scalar metadata is null; unavailable arrays are empty. Authors are publication-author names; each comment contains only commentId, postId, parentCommentId, body (text), createdAt, likeCount, replyCount, and isAuthorReply (true only when explicitly marked by the source). Word count measures extracted public text, including previews, using Unicode word tokens.

Substack discovery paginates its public archive twelve entries at a time, newest first. Details come from /api/v1/posts/{slug} and comments from /api/v1/post/{id}/comments. The source field records the archive request. Generic discovery first tries advertised RSS/Atom links and conventional feed paths, then bounded sitemap traversal. Beehiiv sitemap candidates must use /p/; tag, author, subscription, and landing pages are excluded where recognizable. Source records the feed or sitemap URL.

RSS is often a recent-post window. If a usable feed returns fifteen posts, this actor returns those fifteen even when the cap is twenty; it does not promise a complete historical archive. Sitemap order is publisher-defined. JavaScript-only pages, unusual themes, inaccessible feeds, and layout changes can limit extraction. Original article HTML is output as data and should be sanitized before rendering in another application.

Reliability and pricing

HTTP 429 responses receive two retries with bounded Retry-After or exponential delays. Duplicate posts are suppressed within each publication. Individual detail failures become separate uncharged error rows; remaining posts and publications continue. SUMMARY and SUMMARY-N key-value records contain counts, paywalled counts, errors, and elapsed seconds.

The intended PPE event is post-scraped at $0.003 per successful post row. SDK charged writes enforce the available event limit. Error rows have no charged event. Configure this single custom event in Console and disable synthetic dataset/start events before publishing; the pricing JSON is a deployment specification, not proof of activated billing. Local runs do not bill.

Local development

Create a Python 3.12 .venv, install requirements.txt, and run:

.venv/Scripts/python.exe -m unittest discover -s tests -v
apify validate-schema .actor/input_schema.json
powershell -NoProfile -ExecutionPolicy Bypass -File validation/run_live.ps1
.venv/Scripts/python.exe validation/check_live.py

The live script copies INPUT.json into fresh local SDK storage and invokes python -m src. See VALIDATION.md and validation JSON artifacts for measured results and limitations. Docker and hosted billing require separate deployment validation. This package is prepared locally and has not been pushed or published.

Example output

One real saved dataset row, trimmed by omitting fields only. Source: storage/live-20260926-042940/datasets/default/000000001.json. This is historical validation evidence, not a live response.

{
"publication": "https://www.lennysnewsletter.com",
"platform": "substack",
"postId": "216168140",
"title": "Advanced evals: How to find (and fix) hidden AI failures in your product",
"publishedAt": "2026-09-22T12:45:14.998000Z",
"url": "https://www.lennysnewsletter.com/p/advanced-evals-how-to-find-and-fix",
"paywalled": true,
"audience": "only_paid",
"wordCount": 3260
}

The paywall flag and audience describe the saved source response; the word count measures extracted public text, not a verified complete subscriber article. Body HTML, comments, and other fields are omitted here for readability.

Use cases

  • A newsletter sponsorship buyer can compare recent topics and publishing dates across shortlisted publications before preparing an outreach brief.
  • An editorial research team can collect public article metadata into a reading queue and retain source URLs for attribution.
  • A publisher operations manager can inventory discoverable posts across a Substack, Ghost, or Beehiiv portfolio before planning a content migration.
  • A content strategy agency can group available public text by theme to prepare a client briefing, checking paywall flags before quoting passages.

Pricing example: 1,000 successful post-scraped events x $0.003 = $3.00, computed from .actor/pay_per_event.json. This is the declared event subtotal; it does not verify active hosted billing or include any separately applicable platform or proxy costs.

Limitations

Discovery is limited by the public feed or sitemap and the requested per-publication cap. Preview text, missing metadata, and unavailable comments can make downstream comparisons incomplete. Review source pages before interpreting an absent post or empty comments array as evidence of inactivity.