Substack Scraper, Posts, Full Text and Comments, No Login
Pricing
from $2.00 / 1,000 post delivereds
Substack Scraper, Posts, Full Text and Comments, No Login
Collect public Substack newsletter posts, full text and comments through public JSON endpoints. Support custom domains and post URLs without login, cookies, a proxy or a browser.
Pricing
from $2.00 / 1,000 post delivereds
Rating
0.0
(0)
Developer
George Kioko
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Substack Scraper is an Apify Actor that collects structured posts from public Substack newsletters through their public JSON endpoints.
No login. No cookies. No proxy. No browser. Collect publication metadata, titles, dates, authors, reaction counts, public body HTML, plain text and optional comments.
Supply newsletter hosts, custom domains or individual post URLs. Each dataset row represents one post. Comments and replies stay inside that row. Post IDs are deduplicated across the run.
Pricing
You pay for each distinct post row successfully delivered and each public comment actually included in that row. Empty results, failed fetches, duplicates and skipped paid posts have no result event charge.
| Event | Price | You pay when |
|---|---|---|
apify-actor-start | $0.00005 | The platform records its synthetic start event |
post-result | $0.002 per post | A distinct post row is delivered |
comment-result | $0.0003 per comment | A comment or reply is included in a delivered row |
The price is $2 per 1,000 delivered posts and $0.30 per 1,000 included comments. Ten posts without comments have $0.02 in result charges plus the platform start event.
The SDK checks the post budget, persists the row and charges the post event. The Actor charges included comments after the post has been persisted. Before fetching comments it reserves the next post price and trims the comment allowance to the remaining budget. It stops when another post would exceed the maximum charge limit. Local and free runs write rows without paid event charges. Actor code never charges the synthetic start event.
Collect newsletter posts without login
Provide one or more publication URLs and a result limit for each publication.
{"publications": ["https://www.lennysnewsletter.com"],"maxPostsPerPublication": 10,"includeFullText": true,"includeComments": false}
How it works
- Normalize newsletter hosts, homepage URLs and individual post URLs.
- Fetch archive pages sequentially with a page limit of 50 and a desktop Chrome User-Agent.
- Optionally request public post detail and comments, then normalize the result.
- Push each distinct post and charge result events on paid runs.
- Stop at the per publication result limit, an empty or repeated page, the date cutoff, the charge limit or one minute before the run timeout.
Input -> Publication archive or single post|vPublic detail and comments|vNormalize and deduplicate|vDataset -> Event charges
Request starts are spaced at least one second apart. Each request has a 20 second timeout, shortened near the timeout reserve. Transient failures receive up to three retries with increasing delays of 1, 2 and 4 seconds. Retry-After seconds and HTTP dates take precedence when supplied.
A custom domain returning HTTP 404 triggers publication metadata and feed self link discovery. If those do not identify the Substack host, the Actor tries the custom domain's base label as a subdomain once. Arbitrary Substack links inside articles are ignored. This fallback worked for Platformer during the spike but does not guarantee a mapping for every custom domain.
What data does it extract?
Each row contains type, publication, publicationHost, postId, title, subtitle, slug, url, publishedAt, audience, isPaid, authors, authorHandles, wordCount, reactionCount, commentCount, coverImage, description, postType, bodyHtml, bodyText, isTruncated and scrapedAt.
restacks and podcastUrl appear when provided. Missing scalar metadata is null and missing list metadata is an empty array. Publication names come from publication metadata when available, otherwise the host is used. publicationHost is the API host used, which can change after fallback. url retains the canonical post URL supplied by Substack.
With includeComments, the comments array contains id, parentId, author, body, date, reactionCount and depth. Replies follow their parent in depth first order. Deleted and suppressed comments are omitted. The included array can be shorter than the reported commentCount because of visibility, the comment limit and the charge budget.
bodyText is HTML stripped of tags, scripts and styles with entity decoding and block breaks. It is plain text rather than a faithful rendering. bodyMarkdown is not produced.
What is NOT returned
- Private content, authenticated subscriber content or material beyond the public preview of paid posts.
- A global search across Substack publications.
- A guaranteed complete archive or every comment on a post.
- Markdown, media downloads or private author profile information.
Start a run
- Enter up to 50 newsletter URLs, hosts or post URLs.
- Set the maximum posts per publication and optionally add publication search or a date cutoff.
- Choose whether to include public full text, comments or only free posts.
- Run the Actor and export the dataset as JSON, CSV or Excel.
The prefilled input requests 10 posts from Lenny's Newsletter with public full text and no comments. A live local SDK run on October 6, 2026 delivered exactly 10 rows in 11.621 seconds with body HTML and plain text verified. The normal result default is 50 per publication. Default platform resources are 256 MB and a 900 second timeout.
Input
| Field | Type | Default | Meaning |
|---|---|---|---|
publications | string array | Required, prefill Lenny's Newsletter | 1 to 50 newsletter URLs, hosts or individual post URLs |
maxPostsPerPublication | integer | 50, prefill 10 | Delivered posts per supplied publication, from 1 to 5,000 |
sortBy | enum | newest | newest or top |
searchQuery | string | None | Search phrase within each publication archive |
publishedAfter | ISO date string | None | Stop at older posts, only valid with newest order |
includeFullText | boolean | true | Add public body HTML and plain text |
includeComments | boolean | false | Include public comments inside each post row |
maxCommentsPerPost | integer | 100 | Limit included comments and replies, zero disables fetching |
freeOnly | boolean | false | Skip paid and founding audience posts |
{"publications": ["platformer.substack.com", "https://www.lennysnewsletter.com"],"maxPostsPerPublication": 50,"sortBy": "newest","searchQuery": "AI","publishedAfter": "2026-01-01","includeFullText": true,"includeComments": true,"maxCommentsPerPost": 100,"freeOnly": false}
A /p/post-slug URL returns just that post and does not fetch its archive. Search applies to archives, so it does not filter a directly supplied post URL. The date and free audience filters still apply. Overlapping publication and post inputs share a deduplication set.
Invalid publication entries are logged and skipped when at least one valid input remains. Input with zero valid publications fails before fetching. A valid hostname whose endpoint fails is skipped so later publications can still deliver rows.
Output
This row comes from the live local prefill dataset on October 6, 2026. The body fields are shortened excerpts.
{"type": "post","publication": "Lenny's Newsletter","publicationHost": "www.lennysnewsletter.com","postId": 215694124,"title": "All of the Lenny & Friends Summit talks are now online!","subtitle": "Plus, some reflections and takeaways from the day","slug": "all-of-the-lenny-and-friends-summit","url": "https://www.lennysnewsletter.com/p/all-of-the-lenny-and-friends-summit","publishedAt": "2026-09-29T13:15:57.512Z","audience": "everyone","isPaid": false,"authors": ["Lenny Rachitsky"],"authorHandles": ["lenny"],"wordCount": 1276,"reactionCount": 279,"commentCount": 6,"coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/9fed347c-0e02-4190-a590-1236719b99de_9213x6142.jpeg","description": "Plus, some reflections and takeaways from the day","postType": "newsletter","bodyHtml": "<p><em>...","bodyText": "Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For m...","isTruncated": false,"scrapedAt": "2026-10-06T20:33:54.942Z","restacks": 4}
Use from MCP agents and API clients
After deployment and publication, add this Actor through the Apify MCP server to let an agent collect newsletter posts.
https://mcp.apify.com/?tools=george.the.developer/substack-scraper
From Node.js with the Apify client and an Apify token.
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('george.the.developer/substack-scraper').call({publications: ['https://www.lennysnewsletter.com'],maxPostsPerPublication: 10,includeFullText: true,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
With curl, wait for the run and return dataset rows.
curl -X POST "https://api.apify.com/v2/acts/george.the.developer~substack-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"publications":["https://www.lennysnewsletter.com"],"maxPostsPerPublication":10}'
Apify authentication runs the Actor. No Substack authentication is required.
Integrations
Clay. Read dataset items as an HTTP data source and map publication, title, author and post URL to columns.
n8n and Make. Select the published Actor, supply newsletter input and process dataset rows in the next step.
Google Sheets. Export CSV or use the Apify Google Sheets integration. JSON preserves the full nested comment arrays.
Use cases
- Researchers collect public newsletter coverage and publication dates.
- Readers organize links and publicly available text from their chosen newsletters.
- Editorial teams monitor topics, reactions and public discussion.
- Analysts compare posts across known publications without relying on global search.
Limits
Paid posts return only the public response available without authentication and are marked isTruncated. A long public body does not prove that the full paid article is available. The Actor does not log in, send cookies or bypass paywalls.
There is no global Substack publication search. searchQuery searches within the supplied publication archives only. Public endpoints and search behavior can change.
The maximum measured archive limit was 50. A short initial page still had further pages, so the Actor advances by the number of returned items and continues until an empty or repeated page. Paging was verified at offset 1000 on Noahpinion, but no unlimited archive guarantee was established. A result cap of 5,000 does not make more posts available.
Newest date cutoff stopping assumes the endpoint returns newest order. Feed changes, pinned posts and changing archives can affect completeness. Publication names may fall back to the hostname when public bylines do not identify the publication.
One request per second succeeded for the tested detail calls without HTTP 429. Rate limits can vary by publication and network. No real 429 was observed during the spike; Retry-After handling was verified with offline responses.
Custom domain discovery can fail when the publication name differs from its Substack subdomain or when the custom domain has moved to another publishing system. Failure after the fallback is logged and skipped.
Public comments can be incomplete, unavailable or restricted. A failed detail call still delivers archive metadata with null body fields and fullTextError. Failed comment calls deliver an empty comment array and commentsError. Storage and billing failures stop the run as visible errors.
FAQ
Do I need a Substack account or proxy?
No. The Actor uses unauthenticated public endpoints with a desktop User-Agent.
Can I retrieve complete paid articles?
Only the publicly available preview is returned. Paid audience posts are marked as truncated, and no paywall bypass is attempted.
Are duplicate posts charged twice?
No. Each post ID is delivered once per run, even across overlapping inputs.
What happens when a publication is empty or fails?
An empty publication has no result charges. Exhausted publication fetch failures are logged and skipped. The run exits successfully with whatever it delivered unless input, storage or billing fails.
Can I run this locally?
Yes. Install Node 22 or newer, run npm install --omit=dev --omit=optional, place input in storage/key_value_stores/default/INPUT.json and run npm start. Read rows in storage/datasets/default. Run offline checks with npm test and reproduce the live prefill check with node test/live-prefill.mjs.