Substack Scraper: Newsletter Posts & Stats
Pricing
from $0.83 / 1,000 newsletter posts
Substack Scraper: Newsletter Posts & Stats
Export a Substack newsletter's full archive with engagement data. Returns title, subtitle, publish date, audience tier, reactions, comments, restacks, word count, tags, bylines and canonical URL per post, plus optional full post text. Accepts subdomains and custom domains.
Pricing
from $0.83 / 1,000 newsletter posts
Rating
0.0
(0)
Developer
Axiora Solutions
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 minutes ago
Last modified
Categories
Share
Substack Scraper — export any newsletter archive with engagement data
Substack scraper that exports a publication's entire newsletter archive as structured rows: every post with its reactions, comments, restacks, word count, bylines and canonical URL. It works on both *.substack.com subdomains and custom domains (www.noahpinion.blog, www.slowboring.com), and free posts need no login or API key. The fastest way to try it: leave the prefilled publications in place and click Start.
What you get
title,subtitle,slug,publishedAtand canonicalpostUrlfor every post.reactionTotalplus the rawreactionsbreakdown by emoji.commentCount,childCommentCount(replies) andrestackCount.audienceandisFree— the exact paywall tier each post sits behind.wordCount,tags,sectionName,postTypeandisPodcast.bylines(withisGuest) andfirstAuthorName, plus optional cleanbodyText.
Quick start
- Open the Actor on Apify and list your Publications — a subdomain or a custom domain, one per line.
- Set Max posts per publication and Max posts for the whole run (defaults are fine), or enable Audience filter, tags, keywords and Published after to trim the run.
- Click Start, then read the Posts, Engagement and Content dataset tabs.
- Turn on Fetch full post text only if you need the
bodyTextfield.
Minimal input:
{"publications": ["astralcodexten.substack.com", "newsletter.pragmaticengineer.com"],"maxPostsPerPublication": 100,"maxPostsTotal": 500,"audienceFilter": "all","includePostText": false}
Example output
One representative dataset row:
{"ok": true,"errorCode": null,"postId": "218431605","slug": "our-ai-midwife","title": "Our AI Midwife","subtitle": "A birth story","publication": "astralcodexten.com","publicationUrl": "https://astralcodexten.substack.com","canonicalUrl": "https://www.astralcodexten.com/p/our-ai-midwife","postUrl": "https://www.astralcodexten.com/p/our-ai-midwife","publishedAt": "2026-10-02T01:08:13.190Z","audience": "everyone","isFree": true,"postType": "newsletter","isPodcast": false,"coverImageUrl": "https://substackcdn.com/image/fetch/...","wordCount": 1556,"commentCount": 96,"childCommentCount": 41,"restackCount": 18,"reactions": { "❤": 328, "🔥": 12 },"reactionTotal": 340,"bylines": [{ "id": 12345, "name": "Scott Alexander", "handle": "astralcodexten", "isGuest": false }],"firstAuthorName": "Scott Alexander","tags": [],"truncatedBodyText": "The story of a birth with an AI doula...","bodyText": null,"bodyChars": 0,"sectionName": null,"sectionSlug": null,"contentHash": "7cb41a09e35d8260","scrapedAt": "2026-10-02T12:00:00.000Z"}
Why engagement data is the whole point
An RSS feed gives you the title and the date. A CMS export gives you the text. Neither tells you what worked. Reaction counts, comments, replies and restacks are the signal that turns an archive into a strategy: which topics land, which formats earn a conversation, and where a publication's paywall sits.
What this Substack scraper returns
- 📊 Engagement per post —
reactionTotalplus the rawreactionsbreakdown by emoji,commentCount,childCommentCount(replies) andrestackCount. Sort on any of them. - 🚪 Paywall visibility —
audienceiseveryonefor a free post or the paid tier the author chose.isFreegives you a boolean. Filter to free posts to study the free funnel, or to paid posts to see what a publication decides is worth charging for. - ✍️ Bylines — every author with id, name, handle and an
isGuestflag, so multi-author and guest-post publications remain analysable. - 🎙️ Podcast posts are handled —
postTypeandisPodcastdistinguish audio issues from written ones, andwordCountplussectionNameround out the metadata. - 📰 Multiple publications in one run — the main input is an array, so a competitive set of twenty newsletters is one run, one dataset, with
publicationon every row to keep them apart. - 📄 Optional full text —
bodyTextgives clean plain text from the post page for research or embedding. Paywalled posts return the same truncated preview a non-subscriber sees: this Actor does not attempt to bypass paywalls. - 🎯 Filters that cut your bill — audience tier, tags, keywords and
publishedAfterall run before rows are written, so filtered posts are never charged. - 🔁 Built for schedules —
contentHashcovers the post plus its engagement, so a weekly run tells you which posts gained traction after publication, not just which ones are new. - 🛟 Honest failures — a private or empty publication gives an
ok: falserow explaining that, and every other publication still lands.
Running on Apify adds scheduling, webhooks, monitoring, API and SDK access, and one-click export to JSON, CSV, Excel, Google Sheets and 20+ integrations.
How to use it
- Add publications to Publications — subdomain or custom domain, either works.
- Set Max posts per publication and Max posts for the whole run.
- Use Audience filter to study the free tier, the paid tier, or podcasts only.
- Turn on Fetch full post text if you need
bodyText; leave it off ifwordCount,titleandtruncatedBodyTextare enough. - Click Start, then use the Posts, Engagement and Content dataset tabs.
How do I find a publication's best-performing posts ever?
Set Max posts per publication high, leave Audience filter on All, then sort the Engagement view by reactionTotal. That is the canonical "greatest hits" analysis, and it costs one run.
How much does it cost to scrape a Substack archive?
Pricing is pay per event with one event:
| Event | What triggers it | Billed |
|---|---|---|
| Newsletter post | One post written to the dataset | per post |
| Actor start | Once per run, platform fee | per run |
Fetching full post text is included in the per-post price — it costs you run time, not money. Posts removed by filters or de-duplication are not billed, and a publication that cannot be read is not billed.
A 200-post archive is 200 billed events, and twenty publications at 200 posts each is 4,000 — that is the whole calculation. Compute, bandwidth and storage are included; there is no separate platform-usage charge on top.
Set Max cost per run in the run options for a hard ceiling. Higher Apify plans get progressively lower per-post pricing through Apify Store tier discounts.
Evaluating? Set Max posts per publication to 10 with the prefilled publications.
Example input
A fuller run with filters:
{"publications": ["astralcodexten.substack.com","newsletter.pragmaticengineer.com","https://newsletter.pragmaticengineer.com"],"maxPostsPerPublication": 150,"maxPostsTotal": 400,"publishedAfter": "1 year","audienceFilter": "all","tags": ["AI"],"keywords": ["LLM"],"includePostText": false}
Error rows
A publication that cannot be read does not fail the run; it produces one row like this:
{"ok": false,"errorCode": "NOT_FOUND","requestedInput": "not-a-real-publication.substack.com","error": {"code": "NOT_FOUND","message": "Archive API failed for https://not-a-real-publication.substack.com: HTTP 404 from https://not-a-real-publication.substack.com/api/v1/archive. Check that the publication is public and that the URL is correct.","httpStatus": 404}}
Use cases
- Content strategy research — find which topics, lengths and formats actually earn reactions and comments.
- Competitive newsletter analysis — publishing cadence, paywall placement and growth of a peer set in one dataset.
- Creator benchmarking — compare engagement per word across publications.
- Sponsorship and media planning — quantify a newsletter's engaged audience before buying.
- Research corpora —
bodyTextpluscanonicalUrlandpostIdis a clean, citable archive for study or embedding. - Trend detection — schedule weekly and diff on
contentHashto catch posts that keep gaining traction.
Related Actors by Axiora Solutions
| Actor | Use it for |
|---|---|
| News & RSS Feed Scraper | News coverage feeding the same research, including Google News |
| App Store Review Scraper | Voice-of-customer data to pair with creator and product research |
| Domain Contact Enricher | Contact and company profiles for the operators behind the publications |
Frequently asked questions
Do custom domains work, or only substack.com?
Both. The Actor uses the publication's own origin, so newsletter.pragmaticengineer.com and astralcodexten.substack.com behave identically. This matters because many of the largest publications have moved to custom domains.
Can it read paywalled posts?
It reads what Substack serves publicly. For a paywalled post you get the metadata, the engagement counts, and the same truncated preview a non-subscriber sees. This Actor does not bypass paywalls, and the README says so rather than leaving you to discover it. Metadata and engagement numbers for paid posts are still fully available, which is usually the analysable part anyway.
Why is bodyText null?
Two reasons. Either Fetch full post text is off, in which case nothing is fetched by design, or the fetch failed for that post — in which case bodyText is null and the row still carries truncatedBodyText plus full metadata. The post is still counted and still useful.
How many posts can I export from one publication?
The whole archive. The Actor pages the publication's archive API until it reaches your Max posts per publication limit or the archive ends. Large publications with thousands of posts are bounded only by your limits and run-cost ceiling.
What is a restack?
Substack's equivalent of a share or repost: a reader rebroadcasting the post to their own subscribers. It is the strongest distribution signal a post can have, which is why it is a first-class field here.
Do I need a Substack account or API key?
No. The publication archive API is public, and the Actor reads it directly. Nothing in the input is a credential, so an autonomous agent can call this Actor without a human.
Can I run this on a schedule?
Yes. Weekly is the useful cadence: new posts appear, and existing posts accumulate reactions and comments. Because contentHash includes the reaction total, a diff between runs surfaces posts that are still gaining traction, which is a signal you cannot get from the publish date alone.
Is scraping Substack archives legal?
The Actor requests the same public JSON API that a browser uses when you visit a publication's archive page, and it identifies itself. Post titles, URLs and engagement counts are public data. Full post text is the author's copyrighted work — check their licensing before republishing it or using it to train a model. This is not legal advice.
Something looks wrong — how do I report it?
Open the Issues tab on this Actor page with the publication URL and what you expected to see.
Runnable examples and how-to guides for these Actors: github.com/batow133/axiora-apify-actors