Substack Scraper — Posts, Comments & Paraphrase with AI
Pricing
from $1.00 / 1,000 results
Substack Scraper — Posts, Comments & Paraphrase with AI
Scrape Substack (substack.com) posts by keyword search, publication or post URL: title, author, date, full text, tags, reactions and comments with free sentiment, plus AI summaries. Cross-run caching returns only new posts. Export JSON, CSV, Excel or API.
Pricing
from $1.00 / 1,000 results
Rating
0.0
(0)
Developer
ActorFlow
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Substack Scraper — Newsletter Posts, Comments & Search with AI
Scrape Substack newsletter posts and comments from substack.com search results, any Substack publication, or a single post — including publications on their own custom domain. This Substack scraper extracts the title, subtitle, author, publish date, full post text, tags, cover image, reactions, comment and restack counts, and can collect every reader comment with a sentiment label at no extra cost. Optionally enrich every post with AI — a summary, keywords, sentiment or full paraphrase. Export to JSON, CSV or Excel, or call it as a Substack API from Python, JavaScript or cURL. Paste a search, publication or post URL and press Start.

🔁 Only pay for new posts. Give a run a
cacheProjectNameand this Substack scraper remembers every post it has already collected. Every later run with the same name skips those posts and returns only newly published ones, so they are never downloaded, enriched or billed twice. It makes this actor a low-cost Substack monitor for newsletters and keywords you track on a schedule.
✨ Features of this Substack newsletter scraper
- Cross-run caching: scrape only new posts — name a cache project and every later run skips posts it already collected, so scheduled runs return only fresh posts and you never pay twice for the same post
- Substack search by keyword — paste a
substack.com/search/{keyword}URL to scrape matching posts from across every publication - Full post extraction — title, subtitle, authors, publish date, body text, word count, tags, cover image and podcast URL
- Substack comments scraper — optionally collect every reader comment and reply, each tagged positive, negative, neutral or mixed at no extra charge
- AI enrichment — optional per-post summary, keywords, sentiment analysis, paraphrase, or your own custom instructions
- Engagement metrics — reaction, comment and restack counts for every post
- Paywall detection — paid posts are flagged with
isPaywalledand can be skipped entirely - Custom domain support — works with publications on their own domain, not just
*.substack.com - Pagination support — walks a whole publication archive or many pages of search results until your item limit is reached
- Proxy support — optional, and switched off by default
- No browser required — runs on plain HTTP requests, which makes it fast and cheap
🔁 Scrape only new Substack posts with cross-run caching
Most newsletter scrapers download the whole archive or search again every time they run. This one can remember what it has already scraped. Set cacheProjectName to any name, for example ai-newsletters, and the actor keeps a list of every post URL it collects under that name in your Apify account. The next run with the same name:
- skips every post already collected, whether it was found through a search, an archive or a direct post URL
- counts only new posts toward
maxItems, somaxItems: 20means 20 posts you have not seen before - never re-runs AI enrichment or comment sentiment on a cached post, so you are never charged twice for the same post
- saves the list only after a run finishes successfully. If a run fails part-way through, the next run fetches those posts again rather than missing them
How to monitor Substack for new posts:
- Enter a search URL such as
https://substack.com/search/bitcoin, or a publication archive. - Set
cacheProjectNameto a name you will reuse, such asbitcoin-watch. - Add an Apify Schedule to run it daily or hourly.
- Each run's dataset now holds only the posts published since the last run. Connect it to Slack, email, Google Sheets or a webhook to get alerts for new Substack posts.
Use a different cacheProjectName for each thing you track, or share one name across several start URLs so a post found through both a search and its publication archive is only scraped once. Leave it empty to scrape everything on every run.
🚀 How to scrape Substack newsletters in 5 steps
- Sign up for a free Apify account — includes $5 monthly credit.
- Open the actor page and click Try for free.
- Paste one or more Substack search, publication or post URLs into Start URLs.
- Click Start and wait for the run to complete.
- Download results from the Output tab in JSON, CSV, or Excel format.
You can also run this actor via the Apify API or integrate it directly into your workflows using Zapier, Make, or n8n.
💰 How much does it cost to scrape Substack?
This actor uses pay-per-result billing based on the compute units a run consumes, plus a pay-per-event charge of $0.05 for each post enriched with AI.
- New Apify accounts include $5 of free monthly credit.
- It runs on plain HTTP requests rather than a headless browser, so it costs significantly less to run than browser-based newsletter scrapers.
- Proxies are disabled by default, which keeps runs at their cheapest.
- Comment sentiment is free — comments and their sentiment labels carry no extra charge.
- AI enrichment is opt-in and only charged when it succeeds. Posts no model could enrich, and paywalled previews, are never charged.
- Caching cuts repeat-run costs. With
cacheProjectNameset, posts already collected are skipped before they are fetched or enriched, so a scheduled run only pays for new posts.
🔧 Substack scraper input configuration
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
startUrls | array | — | https://substack.com/search/bitcoin | Substack search, publication, archive, or single post URLs. Page type is detected automatically. |
maxItems | integer | — | 5 | Maximum posts to scrape per start URL. Set to 0 for no limit. |
includePaywalled | boolean | — | true | Keep paid posts (with an isPaywalled flag) or skip them entirely. |
cacheProjectName | string | — | — | Cross-run cache. Reuse the same name and later runs skip already-scraped posts, returning only new ones. |
scrapeComments | boolean | — | false | Also collect each post's comments and replies, each with a sentiment label. |
maxCommentsPerPost | integer | — | 50 | Maximum comments per post, best first. Set to 0 for all. |
aiEnabled | boolean | — | false | Turn on AI enrichment for each post. |
aiFeatures | array | — | ["summarize", "keywords", "sentiment"] | Which enrichments to generate: summary, paraphrase, keywords, sentiment, or custom. |
aiModels | array | — | ["openai/gpt-4o-mini"] | Models tried in order; the first usable result wins. |
aiCustomInstructions | string | — | — | Used only with the custom AI feature. Describe what to extract from each post. |
proxyConfiguration | object | — | {"useApifyProxy": false} | Proxy settings. Off by default. |
Supported URL types:
- Substack search —
https://substack.com/search/bitcoin?searching=all_posts - Publication archive —
https://astralcodexten.substack.com/archive - Publication home page —
https://astralcodexten.substack.com - Single post —
https://astralcodexten.substack.com/p/open-thread-450 - Custom domain —
https://www.thefp.com/p/some-post
📦 Substack scraper output data
Each result is a JSON object with the keys url, title, subtitle, slug, publication, authors, publishedAt, audience, isPaywalled, type, description, body, wordCount, tags, coverImage, podcastUrl, reactionCount, commentCount and restackCount, plus an ai object when enrichment is enabled and a comments array when comment scraping is enabled.
The dataset ships with four views: Overview, a compact table of title, authors, date and paywall status; Full post details, which adds the body text, tags and engagement counts; AI enrichment, which shows the generated summary, keywords and sentiment per post; and Comments, which lists each post's comments with their sentiment.
Sample post output:
[{"url": "https://www.astralcodexten.com/p/royce-on-san-francisco","title": "Royce On San Francisco","subtitle": "...","slug": "royce-on-san-francisco","publication": "https://astralcodexten.substack.com","authors": ["Scott Alexander"],"publishedAt": "2026-09-10T11:10:18.859Z","audience": "everyone","isPaywalled": false,"type": "newsletter","description": "...","body": "On housing prices:\n\nThe San Franciscans at this moment, living in their rag palaces, or renting them at figures that would have sounded possible in the Arabian Nights . . . seemed more like madmen than ever. A correspondent of the New York Post gives with a half-serious fury and contempt an amusing account of the landlords of San Francisco:\n\n“The people of San Francisco are mad, stark mad. A dozen times or more, duri …","wordCount": 2393,"tags": [],"coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/582979e2-4bb5-46aa-91ae-ce2ab6d05fa5_738x415.jpeg","podcastUrl": null,"reactionCount": 155,"commentCount": 76,"restackCount": 2},{"url": "https://www.astralcodexten.com/p/god-help-us-lets-try-to-learn-about","title": "God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques","subtitle": "...","slug": "god-help-us-lets-try-to-learn-about","publication": "https://astralcodexten.substack.com","authors": ["Scott Alexander"],"publishedAt": "2026-09-08T12:04:21.658Z","audience": "everyone","isPaywalled": false,"type": "newsletter","description": "...","body": "The Story So Far\n\nMechanistic interpretability is the science of “reading an AI’s mind”.\n\nLarge language models are “grown, not built”. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how.\n\nBut a neural network is just a set of simulated neurons on a computer. The person with the computer can see the neurons, the connections between them, and which one …","wordCount": 4695,"tags": [],"coverImage": "https://substackcdn.com/image/fetch/$s_!S-UV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6dfd520-f8be-4174-8bd3-428c052b53a7_400x418.png","podcastUrl": null,"reactionCount": 366,"commentCount": 208,"restackCount": 25}]
Sample comments array (when scrapeComments is on). Replies keep a parentId and depth, so threads can be rebuilt from a flat CSV:
[{"id": "277591836","parentId": null,"depth": 0,"author": "Susie","authorHandle": "susiiee","date": "2026-06-16T22:33:47.934Z","editedAt": null,"body": "I haven’t finished reading but this is so fricking relatable. I feel like I’m always watching through another lens of my life ykwim? …","isDeleted": false,"reactionCount": 188,"replyCount": 3,"sentiment": "positive","sentimentConfidence": 0.83},{"id": "278621506","parentId": "277591836","depth": 1,"author": "Vishesh Kashyap","authorHandle": "visheshkashyap","date": "2026-06-18T17:17:00.384Z","editedAt": null,"body": "It's better this way, tat u r able to act better in the situation u loved ones need u the most. …","isDeleted": false,"reactionCount": 12,"replyCount": 0,"sentiment": "positive","sentimentConfidence": 0.98}]
🐍 How to scrape Substack with Python, JavaScript or the API
Run the actor programmatically with the official Apify clients. Replace <YOUR_API_TOKEN> with the token from your Apify Console.
Python (pip install apify-client):
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("confidential_gnat/substack-newsletter-scraper").call(run_input={"startUrls": [{"url": "https://substack.com/search/artificial%20intelligence"}],"maxItems": 20,"scrapeComments": True,})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["title"], len(item.get("comments", [])))
JavaScript (npm install apify-client):
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });const run = await client.actor('confidential_gnat/substack-newsletter-scraper').call({startUrls: [{ url: 'https://astralcodexten.substack.com/archive' }],maxItems: 5,aiEnabled: false,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
cURL — start a run and wait for the dataset:
curl -X POST "https://api.apify.com/v2/acts/confidential_gnat~substack-newsletter-scraper/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \-H "Content-Type: application/json" \-d '{"startUrls": [{"url": "https://astralcodexten.substack.com/archive"}], "maxItems": 5, "aiEnabled": false}'
💡 What you can use Substack newsletter data for
- New-post monitoring — get only new posts from newsletters or keywords you follow, using cross-run caching and a schedule
- Content research — track what a writer publishes and how often
- Topic discovery — search all of Substack for a keyword and see which newsletters cover it
- Competitive analysis — monitor rival newsletters in your niche
- Audience sentiment analysis — see how readers react to each post from the tone of their comments
- Trend detection — run AI keyword extraction across an archive to see recurring themes
- Engagement benchmarking — compare reactions, comments and restacks across posts and publications
- Training and research datasets — collect long-form writing with structured metadata
Media analysts, content marketers, researchers and newsletter operators use this Substack data across publishing, market research, competitive intelligence and academic work.
⚠️ Substack scraping limitations
- Paywalled posts — paid-subscriber posts return only the short public preview, not the full text. They are flagged with
isPaywalledso you can filter them, and AI enrichment skips them so you are never charged to summarize a teaser. The actor does not log in or bypass paywalls. - Comments — collected only when
scrapeCommentsis on. On posts restricted to paid subscribers (audience: "only_paid"), Substack hides the comments from non-subscribers, socommentscomes back empty even when the post text itself is public. - Search depth — a search stops after 50 result pages (roughly 500 posts). Substack search mixes in notes and profiles; only posts are scraped.
- Rate limiting — very high-volume runs across many publications may be throttled; enable proxies if you hit limits.
❓ Frequently asked questions
Is it legal to scrape Substack?
This actor only collects data that is already publicly visible on Substack — no login, paywall bypass, or private content is accessed. Scraping publicly available data is generally considered lawful (see hiQ Labs v. LinkedIn as precedent). You remain responsible for complying with Substack's Terms of Service, each publication's own terms, and applicable copyright law when republishing content.
How do I search all of Substack by keyword?
Search on substack.com, copy the URL (for example https://substack.com/search/bitcoin?searching=all_posts) and paste it into Start URLs. The scraper collects matching posts from every publication, up to your maxItems limit.
Can I scrape Substack comments?
Yes. Turn on scrapeComments and every post gets a comments array with each comment's author, date, text, reactions and replies. Use maxCommentsPerPost to cap how many are collected, best comments first.
How is Substack comment sentiment calculated?
Each comment is classified by an AI model into one of four fixed labels — positive, negative, neutral or mixed — with a sentimentConfidence between 0 and 1. It is included free with comment scraping.
Does this scraper get paywalled Substack posts?
No — it collects what a logged-out visitor sees. For paid posts that is the public preview, which the actor flags with isPaywalled: true. Set includePaywalled to false to skip them entirely.
How many posts can I scrape from one Substack newsletter?
As many as the archive holds. maxItems limits results per start URL, so scraping three publications with maxItems: 100 returns up to 300 posts. Set it to 0 for the entire archive.
How do I scrape only new Substack posts since the last run?
Set cacheProjectName and reuse the same name every run. The actor remembers every post it has collected under that name and skips them next time, so each run returns only newly published posts and never charges you twice for the same one. Pair it with Apify Schedules to monitor a newsletter or a Substack keyword search daily or weekly.
🔗 Other actors you may find useful
- 🏠 University Living Housing Scraper — Scrapes student housing listings and property details from universityliving.com.
- 🍷 Total Wine Scraper — Scrape Total Wine & More (totalwine.com) wine, liquor and beer prices, sizes, ratings, reviews, badges, ABV, origin and taste profile from search, category or product URLs.
- 🇩🇪 German Imprint (Impressum) Scraper with AI Extraction — Finds the Impressum page on any German website and extracts the company's decision makers, legal name, address, email addresses, phone numbers, commercial register number and VAT ID as structured data using AI.
- ⭐ Google Play Store Reviews Scraper — Scrapes user reviews from Google Play Store apps (play.google.com) including review text, star rating, author, date, replies and optional sentiment tagging, with app name, developer and overall rating attached to every review.
- 📜 Google Patents Scraper — Scrapes patent data from Google Patents (patents.google.com) by keyword or URL, including title, abstract, inventors, assignee, filing and publication dates, citations, figures and PDF links.
📰 Other news and article website scrapers
- 🗞️ Google News AI Scraper — Search Google News by keyword, optionally extract full article text and AI-generated summaries, and never re-scrape the same article twice across runs.
- 🦘 Sydney Morning Herald (SMH) News Scraper — Scrape news articles from The Sydney Morning Herald (smh.com.au) — headline, author, publish date, section, keywords, images and full public article text, with a paywall flag.
- 🇮🇩 Detik News Scraper — Scrapes news articles from Detik.com, including headline, author, publish date, category, images and full article text.
💬 Support & Contact
If you encounter any issues or have questions, please open an issue
You can also find more of our actors on the Actor Flow .