Substack Scraper: Newsletter Posts & Stats avatar

Substack Scraper: Newsletter Posts & Stats

Pricing

from $0.83 / 1,000 newsletter posts

Go to Apify Store
Substack Scraper: Newsletter Posts & Stats

Substack Scraper: Newsletter Posts & Stats

Export a Substack newsletter's full archive with engagement data. Returns title, subtitle, publish date, audience tier, reactions, comments, restacks, word count, tags, bylines and canonical URL per post, plus optional full post text. Accepts subdomains and custom domains.

Pricing

from $0.83 / 1,000 newsletter posts

Rating

0.0

(0)

Developer

Axiora Solutions

Axiora Solutions

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 minutes ago

Last modified

Share

Substack Scraper — export any newsletter archive with engagement data

Substack scraper that exports a publication's entire newsletter archive as structured rows: every post with its reactions, comments, restacks, word count, bylines and canonical URL. It works on both *.substack.com subdomains and custom domains (www.noahpinion.blog, www.slowboring.com), and free posts need no login or API key. The fastest way to try it: leave the prefilled publications in place and click Start.

What you get

  • title, subtitle, slug, publishedAt and canonical postUrl for every post.
  • reactionTotal plus the raw reactions breakdown by emoji.
  • commentCount, childCommentCount (replies) and restackCount.
  • audience and isFree — the exact paywall tier each post sits behind.
  • wordCount, tags, sectionName, postType and isPodcast.
  • bylines (with isGuest) and firstAuthorName, plus optional clean bodyText.

Quick start

  1. Open the Actor on Apify and list your Publications — a subdomain or a custom domain, one per line.
  2. Set Max posts per publication and Max posts for the whole run (defaults are fine), or enable Audience filter, tags, keywords and Published after to trim the run.
  3. Click Start, then read the Posts, Engagement and Content dataset tabs.
  4. Turn on Fetch full post text only if you need the bodyText field.

Minimal input:

{
"publications": ["astralcodexten.substack.com", "newsletter.pragmaticengineer.com"],
"maxPostsPerPublication": 100,
"maxPostsTotal": 500,
"audienceFilter": "all",
"includePostText": false
}

Example output

One representative dataset row:

{
"ok": true,
"errorCode": null,
"postId": "218431605",
"slug": "our-ai-midwife",
"title": "Our AI Midwife",
"subtitle": "A birth story",
"publication": "astralcodexten.com",
"publicationUrl": "https://astralcodexten.substack.com",
"canonicalUrl": "https://www.astralcodexten.com/p/our-ai-midwife",
"postUrl": "https://www.astralcodexten.com/p/our-ai-midwife",
"publishedAt": "2026-10-02T01:08:13.190Z",
"audience": "everyone",
"isFree": true,
"postType": "newsletter",
"isPodcast": false,
"coverImageUrl": "https://substackcdn.com/image/fetch/...",
"wordCount": 1556,
"commentCount": 96,
"childCommentCount": 41,
"restackCount": 18,
"reactions": { "❤": 328, "🔥": 12 },
"reactionTotal": 340,
"bylines": [
{ "id": 12345, "name": "Scott Alexander", "handle": "astralcodexten", "isGuest": false }
],
"firstAuthorName": "Scott Alexander",
"tags": [],
"truncatedBodyText": "The story of a birth with an AI doula...",
"bodyText": null,
"bodyChars": 0,
"sectionName": null,
"sectionSlug": null,
"contentHash": "7cb41a09e35d8260",
"scrapedAt": "2026-10-02T12:00:00.000Z"
}

Why engagement data is the whole point

An RSS feed gives you the title and the date. A CMS export gives you the text. Neither tells you what worked. Reaction counts, comments, replies and restacks are the signal that turns an archive into a strategy: which topics land, which formats earn a conversation, and where a publication's paywall sits.

What this Substack scraper returns

  • 📊 Engagement per post — reactionTotal plus the raw reactions breakdown by emoji, commentCount, childCommentCount (replies) and restackCount. Sort on any of them.
  • 🚪 Paywall visibility — audience is everyone for a free post or the paid tier the author chose. isFree gives you a boolean. Filter to free posts to study the free funnel, or to paid posts to see what a publication decides is worth charging for.
  • ✍️ Bylines — every author with id, name, handle and an isGuest flag, so multi-author and guest-post publications remain analysable.
  • 🎙️ Podcast posts are handled — postType and isPodcast distinguish audio issues from written ones, and wordCount plus sectionName round out the metadata.
  • 📰 Multiple publications in one run — the main input is an array, so a competitive set of twenty newsletters is one run, one dataset, with publication on every row to keep them apart.
  • 📄 Optional full text — bodyText gives clean plain text from the post page for research or embedding. Paywalled posts return the same truncated preview a non-subscriber sees: this Actor does not attempt to bypass paywalls.
  • 🎯 Filters that cut your bill — audience tier, tags, keywords and publishedAfter all run before rows are written, so filtered posts are never charged.
  • 🔁 Built for schedules — contentHash covers the post plus its engagement, so a weekly run tells you which posts gained traction after publication, not just which ones are new.
  • 🛟 Honest failures — a private or empty publication gives an ok: false row explaining that, and every other publication still lands.

Running on Apify adds scheduling, webhooks, monitoring, API and SDK access, and one-click export to JSON, CSV, Excel, Google Sheets and 20+ integrations.

How to use it

  1. Add publications to Publications — subdomain or custom domain, either works.
  2. Set Max posts per publication and Max posts for the whole run.
  3. Use Audience filter to study the free tier, the paid tier, or podcasts only.
  4. Turn on Fetch full post text if you need bodyText; leave it off if wordCount, title and truncatedBodyText are enough.
  5. Click Start, then use the Posts, Engagement and Content dataset tabs.

How do I find a publication's best-performing posts ever?

Set Max posts per publication high, leave Audience filter on All, then sort the Engagement view by reactionTotal. That is the canonical "greatest hits" analysis, and it costs one run.

How much does it cost to scrape a Substack archive?

Pricing is pay per event with one event:

EventWhat triggers itBilled
Newsletter postOne post written to the datasetper post
Actor startOnce per run, platform feeper run

Fetching full post text is included in the per-post price — it costs you run time, not money. Posts removed by filters or de-duplication are not billed, and a publication that cannot be read is not billed.

A 200-post archive is 200 billed events, and twenty publications at 200 posts each is 4,000 — that is the whole calculation. Compute, bandwidth and storage are included; there is no separate platform-usage charge on top.

Set Max cost per run in the run options for a hard ceiling. Higher Apify plans get progressively lower per-post pricing through Apify Store tier discounts.

Evaluating? Set Max posts per publication to 10 with the prefilled publications.

Example input

A fuller run with filters:

{
"publications": [
"astralcodexten.substack.com",
"newsletter.pragmaticengineer.com",
"https://newsletter.pragmaticengineer.com"
],
"maxPostsPerPublication": 150,
"maxPostsTotal": 400,
"publishedAfter": "1 year",
"audienceFilter": "all",
"tags": ["AI"],
"keywords": ["LLM"],
"includePostText": false
}

Error rows

A publication that cannot be read does not fail the run; it produces one row like this:

{
"ok": false,
"errorCode": "NOT_FOUND",
"requestedInput": "not-a-real-publication.substack.com",
"error": {
"code": "NOT_FOUND",
"message": "Archive API failed for https://not-a-real-publication.substack.com: HTTP 404 from https://not-a-real-publication.substack.com/api/v1/archive. Check that the publication is public and that the URL is correct.",
"httpStatus": 404
}
}

Use cases

  • Content strategy research — find which topics, lengths and formats actually earn reactions and comments.
  • Competitive newsletter analysis — publishing cadence, paywall placement and growth of a peer set in one dataset.
  • Creator benchmarking — compare engagement per word across publications.
  • Sponsorship and media planning — quantify a newsletter's engaged audience before buying.
  • Research corpora — bodyText plus canonicalUrl and postId is a clean, citable archive for study or embedding.
  • Trend detection — schedule weekly and diff on contentHash to catch posts that keep gaining traction.
ActorUse it for
News & RSS Feed ScraperNews coverage feeding the same research, including Google News
App Store Review ScraperVoice-of-customer data to pair with creator and product research
Domain Contact EnricherContact and company profiles for the operators behind the publications

Frequently asked questions

Do custom domains work, or only substack.com?

Both. The Actor uses the publication's own origin, so newsletter.pragmaticengineer.com and astralcodexten.substack.com behave identically. This matters because many of the largest publications have moved to custom domains.

Can it read paywalled posts?

It reads what Substack serves publicly. For a paywalled post you get the metadata, the engagement counts, and the same truncated preview a non-subscriber sees. This Actor does not bypass paywalls, and the README says so rather than leaving you to discover it. Metadata and engagement numbers for paid posts are still fully available, which is usually the analysable part anyway.

Why is bodyText null?

Two reasons. Either Fetch full post text is off, in which case nothing is fetched by design, or the fetch failed for that post — in which case bodyText is null and the row still carries truncatedBodyText plus full metadata. The post is still counted and still useful.

How many posts can I export from one publication?

The whole archive. The Actor pages the publication's archive API until it reaches your Max posts per publication limit or the archive ends. Large publications with thousands of posts are bounded only by your limits and run-cost ceiling.

What is a restack?

Substack's equivalent of a share or repost: a reader rebroadcasting the post to their own subscribers. It is the strongest distribution signal a post can have, which is why it is a first-class field here.

Do I need a Substack account or API key?

No. The publication archive API is public, and the Actor reads it directly. Nothing in the input is a credential, so an autonomous agent can call this Actor without a human.

Can I run this on a schedule?

Yes. Weekly is the useful cadence: new posts appear, and existing posts accumulate reactions and comments. Because contentHash includes the reaction total, a diff between runs surfaces posts that are still gaining traction, which is a signal you cannot get from the publish date alone.

The Actor requests the same public JSON API that a browser uses when you visit a publication's archive page, and it identifies itself. Post titles, URLs and engagement counts are public data. Full post text is the author's copyrighted work — check their licensing before republishing it or using it to train a model. This is not legal advice.

Something looks wrong — how do I report it?

Open the Issues tab on this Actor page with the publication URL and what you expected to see.


Runnable examples and how-to guides for these Actors: github.com/batow133/axiora-apify-actors