Substack Scraper — Newsletter Posts & Authors API | $0.50/1k
Pricing
from $0.50 / 1,000 results
Substack Scraper — Newsletter Posts & Authors API | $0.50/1k
Scrape any Substack newsletter, custom domains too: posts with title, authors, date, free/paid audience, word count, reactions, comment count and full text of free posts, newsletter profiles and keyword search. Filters by date, audience and keyword. No login, no API key.
Pricing
from $0.50 / 1,000 results
Rating
0.0
(0)
Developer
Raffy
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
21 hours ago
Last modified
Categories
Share
What does Substack Scraper do?
Substack Scraper turns Substack newsletters into structured data: every post with its title, subtitle, author names, publish date, free or paid audience, word count, reading time, reactions, comment count, tags, cover image and the post text, exported as JSON, CSV or Excel. Enter newsletters as a subdomain (lenny), a Substack address (https://simonw.substack.com) or the newsletter's own domain (https://www.lennysnewsletter.com), paste single post URLs, or find newsletters by keyword. No Substack account, no login, no API key and no browser.
It works as a Substack API for newsletters you do not own. Every Substack page loads its data from public JSON endpoints; this Actor reads those endpoints page by page, follows newsletters that moved to a custom domain, and returns one tidy row per post (or per newsletter).
Typical use cases:
- Content and competitor research: pull the full archive of the newsletters in your niche and see what they publish, how often, and which posts get the most reactions and comments.
- Newsletter discovery and outreach lists: search Substack by keyword and get newsletters with their description, public authors, subscriber band and first post date.
- Monitoring: schedule a run with Posted on or after
1 dayto collect every new post of a list of newsletters each morning. - Datasets for AI and analytics: plain-text post bodies of free posts, ready for summarising, classification or RAG, from a Substack scraper an AI agent can call with just a newsletter name.
It reads only what Substack shows to a logged-out visitor. Paid posts come with the free preview only (isPreviewOnly: true); the Actor never tries to get around a paywall. It does not collect comments, commenters, subscriber lists or e-mail addresses.
Why use Substack Scraper?
- Posts, newsletters and search in one Actor. Archive scraping, single post URLs, a one-row-per-newsletter mode and publication search by keyword.
- Custom domains just work.
www.lennysnewsletter.com,www.thefp.comor a subdomain that redirects to one: the same Substack data is read from all of them. - Honest paywall handling.
audiencesaysfreeorpaid,wordCountis the full length Substack states, andisPreviewOnlytells you whenbodyTextis only the public preview. - Filters that save money. Free posts only, posted after / before, a keyword (Substack's own archive search), newest or most popular first, and a per-newsletter limit. Posts your filters drop are never written or billed.
- Honest results. Every row has a
status. A site that is not a Substack newsletter, an unknown post or a search with no match gives one freenot_foundrow that says why; you are billed only forokrows. - Cheap. $0.50 per 1,000 posts, Apify platform usage included.
What data can Substack Scraper extract?
One row per post (default):
| Field | Type | Description |
|---|---|---|
recordType | string | post (or publication for newsletter rows) |
title | string | Post title |
subtitle | string | Post subtitle |
publicationName | string | Newsletter name |
publicationUrl | string | Newsletter home page (custom domain or <name>.substack.com) |
publicationSubdomain | string | Newsletter subdomain on substack.com |
authors | array | Names on the post's public byline |
publishedAt | string | Publish time (ISO 8601) |
updatedAt | string | Last update time (with Include post text) |
audience | string | free (everyone) or paid (paying subscribers only) |
isPreviewOnly | boolean | true for paid posts: bodyText holds only the free preview |
wordCount | integer | Length of the whole post as Substack states it |
readingTimeMinutes | integer | Estimated reading time (250 words per minute) |
reactions | integer | Reactions (likes) |
commentCount | integer | Number of comments (a count only) |
restacks | integer | Number of restacks |
tags | array | Post tags |
section | string | Newsletter section, if the post belongs to one |
postType | string | newsletter, podcast, thread, ... |
coverImage | string | Cover image URL |
description | string | Search / SEO description of the post |
previewText | string | Short opening text from the archive list |
bodyText | string | Post body as plain text (paid posts: the free preview only) |
bodyHtml | string | Post body as HTML (with Include post HTML) |
bodyError | string | Why the full post could not be read, when the row carries only the archive data |
postId, slug, publicationId | integer / string | Substack's own ids |
url | string | Canonical post URL |
status | string | ok, not_found or error (see below) |
error | string | Reason when status is not ok |
scrapedAt | string | ISO 8601 time of extraction |
Newsletter rows (What to return = publication, and every Find newsletters by keyword result) carry name, url, subdomain, customDomain, description (the tagline), authors (the newsletter's owner, its public author), logo, subscriberText (e.g. "Hundreds of thousands of subscribers"), paidSubscriberText, subscriberCount (when the newsletter shows it publicly), hasPaidPlans, language, firstPostAt and, for search results, searchQuery.
Result status (tri-state output)
status | Meaning | Billed? |
|---|---|---|
ok | A post or newsletter row. | Yes |
not_found | Not a Substack newsletter, unknown newsletter or post, a newsletter with no public posts, a search with no match, or nothing matched your filters. error says which. | No |
error | Substack could not be read after several retries (network error, rate limit). error says why. | No |
How to scrape Substack newsletters
- Open the Actor in Apify Console and click Try for free.
- Put newsletters into Newsletters (a name like
lenny, a Substack URL or a custom domain), post links into Post URLs, or keywords into Find newsletters by keyword. - Optionally set Maximum posts per newsletter, Posted on or after, Audience or Only posts about.
- Click Start. The default run (10 posts from each of two newsletters, with full text) takes about 15 seconds.
- Open the Output tab or Export the dataset as JSON, CSV, Excel, XML or HTML.
To automate it, use the API tab (Node.js, Python, curl examples) or add a Schedule.
How much does it cost to scrape Substack?
This Actor uses pay-per-event pricing. Apify platform usage is included in these prices.
| Event | Price |
|---|---|
| Actor start | $0.005 per run |
Result (status: ok row) | $0.0005 per post (or per newsletter row) |
Examples: the default run (20 posts) costs $0.015. 1,000 posts cost $0.505, 10,000 posts $5.005. Rows with status not_found or error are never billed, and posts your filters drop are never written or billed. You can cap spending with Maximum results, Maximum posts per newsletter and the run's Max total charge option.
Input
| Field | Type | Default | Description |
|---|---|---|---|
startUrls | array of strings | simonw.substack.com, www.lennysnewsletter.com | Newsletters: subdomain, Substack URL or custom domain (post URLs also work) |
postUrls | array of strings | - | Single posts to read |
searchQueries | array of strings | - | Find newsletters by keyword (one row per newsletter) |
mode | posts or publication | posts | One row per post, or one row per newsletter |
maxItems | integer | 20 | Stop after this many rows in total |
maxPostsPerPublication | integer | - (prefill 10) | Stop reading a newsletter after this many posts |
maxPublicationsPerSearch | integer | 20 | Newsletters per search term |
fetchFullPost | boolean | true | Add bodyText and updatedAt (one extra request per post) |
includeHtml | boolean | false | Also add bodyHtml |
audience | all or free | all | free skips subscriber-only posts |
postedAfter, postedBefore | string | - | YYYY-MM-DD or relative such as 7 days |
postKeyword | string | - | Only posts Substack's archive search matches |
sort | new or top | new | Newest or most popular posts first |
proxyConfiguration | object | Apify Proxy (datacenter) | Spreads requests; residential proxies are not needed |
Example: the free posts of two newsletters published in August 2026, without the post text:
{"startUrls": ["https://www.lennysnewsletter.com", "simonw"],"audience": "free","postedAfter": "2026-08-01","postedBefore": "2026-08-31","fetchFullPost": false,"maxItems": 200}
Example: find newsletters about history:
{"searchQueries": ["history"],"maxPublicationsPerSearch": 50,"maxItems": 50}
Output
Real rows from runs on Apify (the default input, and publication mode with one unknown newsletter; long texts shortened with ...):
[{"url": "https://simonw.substack.com/p/navierstokes-rubygems-attacked-gis","status": "ok","scrapedAt": "2026-09-30T13:49:12.450Z","recordType": "post","postId": 215727411,"slug": "navierstokes-rubygems-attacked-gis","title": "Navier–Stokes, RubyGems attacked, GIS and Blender with GPT-6 Astra","subtitle": "Plus a Datasette security release, and a Pluribus Fabergé egg","description": "Plus a Datasette security release, and a Pluribus Fabergé egg","authors": ["Simon Willison"],"publishedAt": "2026-09-14T20:47:59.989Z","audience": "free","isPreviewOnly": false,"wordCount": 5230,"readingTimeMinutes": 21,"reactions": 54,"commentCount": 7,"restacks": 7,"postType": "newsletter","tags": ["openai","llms","ai"],"coverImage": "https://substackcdn.com/image/fetch/$s_!XTAQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c99ef2b-8373-4ac4-b95b-3fce518d3733_1327x934.webp","previewText": "In this newsletter:","publicationName": "Simon Willison’s Newsletter","publicationSubdomain": "simonw","publicationUrl": "https://simonw.substack.com","publicationId": 1173386,"updatedAt": "2026-09-14T20:51:58.963Z","bodyText": "In this newsletter:\nSome thoughts on the Navier–Stokes Millennium Prize Problem\nGenerating running routes with GPT-6 Astra and ChatGPT Work\nOpenAI agents attacked RubyGems back in May\nPlus 7 links, 7 quotations, 1 note, ..."},{"url": "https://www.lennysnewsletter.com/p/60-creative-growth-ideas","status": "ok","scrapedAt": "2026-09-30T13:49:14.559Z","recordType": "post","postId": 214360366,"slug": "60-creative-growth-ideas","title": "60+ new creative growth ideas","subtitle": "How to stand out when everyone is running the same playbook","description": "How to stand out when everyone is running the same playbook","authors": ["Tom Orbach"],"publishedAt": "2026-09-15T12:45:28.445Z","audience": "paid","isPreviewOnly": true,"wordCount": 5674,"readingTimeMinutes": 23,"reactions": 414,"commentCount": 6,"restacks": 34,"postType": "newsletter","tags": ["Growth"],"coverImage": "https://substackcdn.com/image/fetch/$s_!NVVg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20cbbe62-5c82-4d4f-a316-de9b885e56f2_1456x970.png","previewText": "👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs ...","publicationName": "Lenny's Newsletter","publicationSubdomain": "lenny","publicationUrl": "https://www.lennysnewsletter.com","publicationId": 10845,"updatedAt": "2026-09-15T12:50:53.552Z","bodyText": "👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs | Lenny’s Podcast | Lennybot | How I AI | Become an AI-Native Builder and my other favorite AI/PM co ..."},{"url": "https://www.thefp.com","status": "ok","scrapedAt": "2026-09-30T13:49:00.629Z","recordType": "publication","name": "The Free Press","publicationName": "The Free Press","publicationUrl": "https://www.thefp.com","publicationSubdomain": "bariweiss","publicationId": 260347,"subdomain": "bariweiss","customDomain": "www.thefp.com","description": "A new media company built on the ideals that were once the bedrock of American journalism.","authors": ["Bari Weiss"],"logo": "https://substackcdn.com/image/fetch/$s_!XTc7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F9cb7f208-a15c-46a8-a040-7e7a2150def9_1280x1280.png","hasPaidPlans": true,"language": "en","firstPostAt": "2021-01-12T20:43:07.912Z"},{"url": "https://nosuchpubxyz123456.substack.com/","status": "not_found","error": "HTTP 404: page does not exist","scrapedAt": "2026-09-30T13:49:00.035Z"}]
The paid post has audience: paid and isPreviewOnly: true: its bodyText is the free preview Substack shows to every visitor, while wordCount is the length of the whole post.
Tips
- Put many newsletters in one run: you pay the start fee once, and newsletters are read in parallel.
- Use Maximum posts per newsletter so one long archive cannot use up the run.
- For daily monitoring, schedule the run with Posted on or after
1 day: the archive is newest-first, so the scan stops as soon as it reaches older posts. - Turn Include post text off when you only need titles, dates and engagement numbers: one request per 20 posts instead of one per post.
- To get the posts of newsletters you found with a search, copy their
urlvalues into Newsletters.
Limitations
- Paid posts: only the free preview that Substack shows to logged-out visitors (
isPreviewOnly: true). The Actor does not log in and never tries to get around a paywall. - Comments, commenters, likers, subscriber lists and e-mail addresses are not collected;
commentCountis a number only. - Profile pages (
substack.com/@name) and Substack Notes are not newsletters and are not supported; enter the newsletter's address instead. - A keyword (Only posts about) uses Substack's own archive search, which also matches words inside the post, and returns matches in its relevance order.
subscriberCountis filled only when Substack shows the number publicly; otherwisesubscriberTextcarries Substack's public band ("Thousands of subscribers").
FAQ
Do I need a Substack account or API key?
No. The Actor reads the public data that Substack serves to every logged-out visitor.
Does it work with custom domains?
Yes. Enter the newsletter's own domain (https://www.thefp.com) or its Substack subdomain; a subdomain that redirects to a custom domain is followed automatically.
Can I get the full text of paid posts?
No. For paid posts you get the title, dates, author names, word count, engagement numbers and the free preview, marked with isPreviewOnly: true.
How do I find newsletters about a topic?
Put the topic into Find newsletters by keyword. Each result is one newsletter row with its name, URL, description, author and subscriber band.
Can I use this Actor from an AI agent or MCP client?
Yes. It runs with no input at all (it then reads two example newsletters), accepts plain newsletter names, and every row explains itself with status and error.
Why did I get fewer rows than maxItems?
The newsletters have fewer matching posts, a per-newsletter limit stopped them, or your run hit its Max total charge. A not_found row says when a filter matched nothing.
Related Actors
- Google News Scraper - use it to follow the same topics across news sites, not only newsletters.
- Remote Jobs Scraper - use it to collect remote job listings for the companies and industries you follow.
Legal and data-protection notice
This Actor extracts only data that Substack publishes publicly for logged-out visitors: posts, their public byline (author names) and newsletter profiles. It does not extract comments, commenter or subscriber data, e-mail addresses or phone numbers, and it does not log in or get around paywalls or other access controls. Post texts are the authors' copyrighted work: use them for analysis, research or indexing, and do not republish them without the author's permission. Personal data is protected by the GDPR in the European Union and by other regulations around the world; do not use the output to process personal data without a legitimate reason. You are responsible for complying with Substack's terms of service and applicable law when using the extracted data.
This Actor is an independent tool and is not affiliated with, endorsed by or sponsored by Substack Inc. Substack is a trademark of its owner.