Substack Scraper: Posts, Full Text, Reactions & Paywall
Pricing
from $0.60 / 1,000 posts
Substack Scraper: Posts, Full Text, Reactions & Paywall
Scrape any Substack newsletter archive: title, subtitle, full post text, author, publish date, reactions, comment count and whether the post is behind the paywall. No login, no API key, pay per post.
Pricing
from $0.60 / 1,000 posts
Rating
0.0
(0)
Developer
The Mine Works
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors.
Give it Substack newsletters by name, subdomain or URL, and get their posts, newest first and as far back as you ask, as clean rows: title, subtitle, authors, publish date, word count, reaction and comment counts, whether the post is free or paid, the canonical link, and the post text. It reads each publication's own public archive, so there is nothing to log into and no key to manage.
Why choose this actor?
- Deep archives, fast. A test of this build on 2 October returned 100 Big Technology posts in 15 seconds, read from three archive pages, and our daily platform check returns posts with full text in 7 seconds. Every row carries
reactions_totalandcomment_count, so you can rank a newsletter's posts by what readers responded to. - Honest about paywalls. Every post says whether it is free (
audience: "everyone") or paid (only_paid). For paid posts you get only what Substack makes public, and a post whose text is much shorter than its word count is flaggedcontent_truncated: true. Nothing is bypassed. - Pay only for posts delivered, once. From $0.60 per 1,000 posts on Gold. A post that turns up on two archive pages, a newsletter listed twice, a publication that does not exist, and the run's report rows cost nothing beyond the $0.005 start fee.
How paging works. Substack's archive sends up to 50 posts per request, and its first answer currently holds about 23. The actor moves on by exactly the number of posts it received and keeps reading until it has maxPostsPerPublication posts or the archive has no more, so a short page never ends a run early.
Part of The Mine Works Social media and video family: Threads Scraper, Reddit Scraper, Threads Search Scraper, Instagram Profile Scraper, Instagram Followers & Following, Reddit Search Scraper.
Try it in one minute
Paste this into the JSON tab of the input page and press Start:
{"publications": ["bigtechnology"],"maxPostsPerPublication": 10,"includeBody": true}
You get Big Technology's 10 latest posts with their public text in under a minute.
Give the newsletters in publications, in any of three forms: a plain name (bigtechnology, read as bigtechnology.substack.com), a Substack subdomain (bigtechnology.substack.com), or a full URL, including newsletters on their own domain (https://www.astralcodexten.com/). Then cap each one with maxPostsPerPublication, switch the text on or off with includeBody, and keep only free or only paid posts with audienceFilter. Note that the Console form prefills platformer, which is Platformer's old Substack archive from before it moved off Substack in January 2024.
Apify's free plan includes $5 of credit every month, which covers about 4,700 posts at this actor's Free plan price ($0.001 a post plus the $0.005 start fee, in runs of 100 posts or more).
Copy to your AI assistant
themineworks/substack-scraper on Apify. Reads Substack publications' public archives and returns one row per post with title, subtitle, authors, publish date, word count, reactions, comment count, free or paid audience, canonical URL and the public post text. Call ApifyClient("TOKEN").actor("themineworks/substack-scraper").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Required: publications (string[]: a name such as "bigtechnology", a subdomain, or a full URL including custom domains). Optional: maxPostsPerPublication (1 to 2000, default 50; the actor pages through the archive, newest first, until it has that many or the archive ends), includeBody (default true; one extra request per post), audienceFilter ("" both, "everyone" free only, "only_paid" paid only). Paid posts carry only Substack's public portion; content_truncated is true when the text is much shorter than word_count. A post is delivered and billed once even if it appears on two archive pages. Rows with _type "summary" or "info" are never billed. Full spec: GET https://api.apify.com/v2/acts/themineworks~substack-scraper/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations?fpr=ymnoit&utm_source=apify-readme&utm_medium=referral
Key features
- Up to 22 fields per post:
publication,publication_host,post_id,title,subtitle,slug,url,description,authors,published_at,post_type,audience,is_paywalled,word_count,reactions_total,reactions,comment_count,cover_image,podcast_duration(podcast posts),body_text,content_truncatedandscraped_at. - Up to 2,000 posts per newsletter, newest first. The actor pages through the archive (up to 50 posts per request) until it reaches
maxPostsPerPublicationor the oldest post. Posts are deduplicated bypost_idacross pages, and the same newsletter given twice is read once. - Any publication, including custom domains. Newsletters that publish from their own domain work the same as
*.substack.comaddresses, andurlis the canonical link readers see (for Big Technology,www.bigtechnology.com/p/...). - Engagement you can sort. Substack returns reactions keyed by emoji (
{"❤": 35}); the actor keeps that map and addsreactions_totalas one number. - Text as plain prose. With
includeBodyon, each post's HTML body is fetched from Substack's post endpoint and converted to plain text (paragraphs and list bullets kept). With it off, a run reads only the archive (one request per 23 to 50 posts) and returns metadata. - Paywall status on every row.
audience,is_paywalledand, for paid posts that come back short,content_truncated.
How to use it
Basic: one newsletter
{"publications": ["bigtechnology"],"maxPostsPerPublication": 25,"includeBody": true}
Several newsletters at once
{"publications": ["bigtechnology", "https://www.astralcodexten.com/", "platformer.substack.com"],"maxPostsPerPublication": 25,"includeBody": false}
Metadata only: one or two archive requests per newsletter for 25 posts (Substack's first answer currently holds about 23), so this finishes in seconds. Each row says which newsletter it came from in publication.
Competitor research: what landed with readers
{"publications": ["bigtechnology"],"maxPostsPerPublication": 50,"includeBody": false}
Sort by reactions_total and comment_count, then compare the top posts' title, word_count and audience. In our Big Technology test, 95 of the 100 latest posts were paid, which tells you how the newsletter splits free and paid.
Free posts only, for a reading list or RAG index
{"publications": ["bigtechnology", "https://www.astralcodexten.com/"],"maxPostsPerPublication": 50,"includeBody": true,"audienceFilter": "everyone"}
Free posts come back with their whole text in body_text. Index it with title, published_at and url as metadata. The filter is applied while the archive is read, and the actor keeps paging until it has maxPostsPerPublication matching posts or the archive ends, so a newsletter with few free posts is read further back. Posts the filter removes are never charged.
New post alerts
{"publications": ["bigtechnology", "https://www.astralcodexten.com/"],"maxPostsPerPublication": 5,"includeBody": false}
Save it as a task and schedule it daily (for example 0 7 * * *). Compare post_id with your stored list to act only on new posts. The actor keeps no memory between runs, so each run is charged for every post it returns; keep maxPostsPerPublication small for alerts.
Input parameters
| Parameter | Type | Default | What it does |
|---|---|---|---|
publications | array of strings | required (form prefill: bigtechnology, platformer) | Newsletters as a plain name, a Substack subdomain or a full URL, including custom domains. |
maxPostsPerPublication | integer (1 to 2,000) | 50 (form prefill: 25) | Most posts to take from each newsletter, newest first. The actor keeps paging through the archive until it has this many or the archive ends. |
includeBody | boolean | true | Fetch each post's text, one extra request per post. Turn it off for a fast metadata pass. |
audienceFilter | string | "" (both; form prefill: everyone) | everyone for free posts only, only_paid for paid posts only, empty for both. |
"Form prefill" values fill the Console form for you but are not defaults: a run started from the form keeps only free posts unless you change the filter.
Run options. The default memory is 512 MB and the default timeout is 3,600 seconds. Our 23 post run with text took 7 seconds on the platform. With text on, every post is one extra request, so large runs (hundreds of posts per newsletter) take minutes rather than seconds.
What data do you get?
One row per post. Values Substack does not publish for a post are left out of that row rather than sent empty.
The newsletter: publication (the name you gave, or the host), publication_host.
The post: post_id, title, subtitle, description (Substack's summary line), slug, url (the canonical post link, on the newsletter's own domain when it has one), authors (bylines), published_at (ISO timestamp), post_type (such as newsletter), cover_image.
Paywall: audience (everyone or only_paid), is_paywalled (true for paid posts) and content_truncated (true when a paid post's text is under 80% of its word count).
Size and engagement: word_count (as Substack reports it for the full post), reactions (map by emoji), reactions_total, comment_count.
Text: body_text, the post as plain text, when includeBody is on.
What paid posts look like. In our Big Technology test with text (23 posts), 21 were paid. 9 of them came back with nearly their full text (Substack made it public), 10 came back as a short preview and were flagged content_truncated, and 2 had no public text at all, so those rows have no body_text. The actor never tries to get past the paywall; subscribe to the newsletter for paid text.
Each run ends with a _type: "summary" row (publications_processed, posts_delivered, paywalled_posts, archive_pages_read, charged_for) and a _type: "info" row with a short message. Neither is a post and neither is ever charged. Skip rows that have a _type field when you load posts.
Stable fields for automations
These 15 fields were present in every post row we sampled (176 rows from four runs on 2 October 2026: a platform run and a local run of the previous build, and two local runs of this build):
| Field | What it holds |
|---|---|
publication | Newsletter name as you gave it |
publication_host | The host read, such as bigtechnology.substack.com |
post_id | Substack's post ID; the key for deduplication across runs |
title | Post headline |
slug | URL slug of the post |
url | Canonical post link |
description | Summary line |
authors | List of byline names |
published_at | Publish time, ISO timestamp |
post_type | Post type, such as newsletter |
audience | everyone or only_paid |
is_paywalled | true for paid posts |
word_count | Words in the full post, per Substack |
reactions_total | All reactions summed |
comment_count | Number of comments |
reactions and scraped_at were also in every sampled row. subtitle and cover_image were in all but one (Substack leaves them blank on some posts), and body_text was in every row whose post had public text. We will not rename these fields. New fields may be added over time; existing ones keep their names.
Output examples
Real rows from our test of the previous build on 2 October 2026 (bigtechnology), with the text trimmed. This build returns the same fields.
A free post, with its full text:
{"publication": "bigtechnology","publication_host": "bigtechnology.substack.com","post_id": "217393429","title": "Here’s Everything OpenAI’s Bots (And Others) Have Hacked Or Considered Hacking","subtitle": "It’s getting hard to keep track of all the unauthorized agent activity. Here it is in one place. ","slug": "heres-everything-openais-bots-and","url": "https://www.bigtechnology.com/p/heres-everything-openais-bots-and","description": "It’s getting hard to keep track of all the unauthorized agent activity. Here it is in one place.","authors": ["Marty Swant", "Alex Kantrowitz"],"published_at": "2026-09-28T20:20:52.260Z","post_type": "newsletter","audience": "everyone","is_paywalled": false,"word_count": 950,"reactions_total": 35,"reactions": { "❤": 35 },"comment_count": 0,"body_text": "Australian Prime Minister Anthony Albanese last Friday accused OpenAI’s agents of hacking into the country’s universal health insurance system…","scraped_at": "2026-10-02T16:16:44.314Z"}
A paid post that came back as a preview:
{"publication": "bigtechnology","publication_host": "bigtechnology.substack.com","post_id": "206913880","title": "Apple’s lawsuit against OpenAI makes serious claims. Will they matter?","url": "https://www.bigtechnology.com/p/apples-lawsuit-against-openai-makes","authors": ["Marty Swant", "Alex Kantrowitz"],"published_at": "2026-07-13T22:24:43.226Z","post_type": "newsletter","audience": "only_paid","is_paywalled": true,"word_count": 1631,"reactions_total": 49,"reactions": { "❤": 49 },"comment_count": 1,"body_text": "Apple doesn’t sue often. When it has, it’s usually gone after companies like Qualcomm and Samsung that are already shipping competing products…","content_truncated": true,"scraped_at": "2026-10-02T16:17:13.529Z"}
Both examples leave out cover_image, and the second also leaves out subtitle, slug and description, for length. In real rows body_text often starts with a couple of stray bullet characters left over from Substack's page layout; they are trimmed here.
Pricing
Pay per event: you pay for each post delivered to your dataset, plus a small start fee per run. A post costs the same with or without its text.
| Event | Free | Bronze | Silver | Gold and above |
|---|---|---|---|---|
post-scraped, per post | $0.001 | $0.0009 | $0.00075 | $0.0006 |
post-scraped, per 1,000 posts | $1.00 | $0.90 | $0.75 | $0.60 |
apify-actor-start, per run | $0.005 per GB of run memory, minimum one event | same | same | same |
The start fee, exactly. Apify's apify-actor-start event is charged once when a run starts, at $0.005 for each GB of memory the run uses, with a minimum of one event. This actor runs on 512 MB by default, so a default run pays one event: $0.005. Our 2 October run shows apify-actor-start: 1 and post-scraped: 23.
Worked examples. 50 posts from one newsletter on the Free plan: $0.05 plus $0.005. Ten newsletters at 100 posts each (1,000 posts) on Gold: $0.60 plus $0.005.
Never charged: a newsletter that does not exist or returns nothing, posts removed by audienceFilter, a post that appears on two archive pages (it is delivered once), a newsletter you listed twice (read once), failed requests and their retries, and the summary and info rows. A run that delivers nothing pays only the start fee.
There is no scheduled price change for this actor. The Pricing tab on this page always shows the rate for your plan; if it and this table ever differ, the Pricing tab is right.
FAQ
What does it read? Every Substack newsletter publishes a public archive of its posts, and each post has a public page. The actor reads the archive's JSON feed for the post list and the post endpoint for the text, the same data the newsletter's own website loads for any visitor.
How many posts can I get?
Up to 2,000 per newsletter per run, set with maxPostsPerPublication (default 50), newest first. The actor pages through Substack's archive, which sends up to 50 posts per request, and stops at your number or when the archive has no older posts. In our 2 October test, Big Technology returned 100 posts from three archive pages of 23, 50 and 50 posts, reaching back to April 2025.
Do I need a Substack account or a subscription? No. The actor reads public data only. A subscription matters only if you want the full text of paid posts, which this actor does not bypass.
Why is a paid post's text short or missing?
Because Substack makes only part of it public. Rows flag it with content_truncated: true when the text is under 80% of word_count; a paid post with no public text has no body_text.
Does it work with custom domains?
Yes. Give the newsletter's own URL, and url in every row is the canonical link on that domain. Big Technology, read as bigtechnology, returned links on www.bigtechnology.com.
How fresh is the data? Every run reads the archive live; nothing is cached. Reaction and comment counts are as of the moment of the run and keep growing on new posts.
Do I need a proxy? No. The actor reads Substack's public endpoints directly, at no extra cost.
Can I run it on a schedule?
Yes. Save your input as a task, then in Apify Console go to Schedules, Create new, and pick a time or a cron expression such as 0 7 * * *. Compare post_id with your stored posts to keep only new ones.
How do I export the data?
From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API. Drop rows that have a _type field if you want posts only. Long texts are easiest to handle as JSON.
Can I use it from Claude, ChatGPT or another AI assistant?
- Connector URL:
https://mcp.apify.com/?tools=themineworks/substack-scraper. - Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
- ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
- Cursor or VS Code: add it as an HTTP MCP server with that URL.
- Claude Code:
claude mcp add -t http substack-scraper "https://mcp.apify.com/?tools=themineworks/substack-scraper".
Is it legal to scrape Substack? The actor reads only what Substack serves publicly to any visitor and never gets past a paywall or a login. Posts are their authors' copyrighted work, and bylines and comments involve people, so you are responsible for how you use the data, including Substack's terms, copyright, and data protection laws such as GDPR and CCPA. This is general information, not legal advice. This actor is independent and not affiliated with Substack.
Integrations
- Google Sheets: export a run to a sheet, or use Apify's Google Sheets integration to add each scheduled run's posts.
- Make, Zapier and n8n: use the Apify app or node to start a run and send new posts to Slack, Notion or a vector database.
- Webhooks: have Apify call your URL when a run succeeds, then read the dataset.
- API and client libraries: start runs and read datasets from Python, JavaScript or any HTTP client. The "Copy to your AI assistant" block above has the exact call.
- MCP clients: Claude, ChatGPT, Cursor, VS Code and Claude Code can call the actor as a tool through
https://mcp.apify.com.
Building a text corpus? Pair it with Podcast Episode Scraper for transcripts and LinkedIn Newsletter Scraper for LinkedIn newsletters.
More from The Mine Works
Social media and video
- Threads Scraper
- Reddit Scraper
- Threads Search Scraper
- Instagram Profile Scraper
- Instagram Followers & Following
- Reddit Search Scraper
- Twitter / X Scraper
- YouTube Transcript
- Xiaohongshu (RED) Scraper
- Telegram Channel Scraper
- Telegram Channel Finder
- Pinterest Profile Scraper
Leads and business directories
Marketing, SEO and reviews
Real estate
Science, health and government data
Jobs and hiring
Company and business data
E-commerce and marketplaces
Food and local services
Developer and AI tools
More tools
Support
Found a newsletter that returns nothing, or need a field we do not return yet? Open an issue on the Issues tab of this page with your input and the run ID, and we will reply there. To ask for a new source, email dmineworks@gmail.com. A guide for this actor also lives at themineworks.com.
Substack Scraper turns newsletters into post rows with text, reactions and honest paywall flags, at $1 or less per 1,000 posts.

