Substack Scraper, Posts, Full Text and Comments, No Login avatar

Substack Scraper, Posts, Full Text and Comments, No Login

Pricing

from $2.00 / 1,000 post delivereds

Go to Apify Store
Substack Scraper, Posts, Full Text and Comments, No Login

Substack Scraper, Posts, Full Text and Comments, No Login

Collect public Substack newsletter posts, full text and comments through public JSON endpoints. Support custom domains and post URLs without login, cookies, a proxy or a browser.

Pricing

from $2.00 / 1,000 post delivereds

Rating

0.0

(0)

Developer

George Kioko

George Kioko

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Substack Scraper is an Apify Actor that collects structured posts from public Substack newsletters through their public JSON endpoints.

No login. No cookies. No proxy. No browser. Collect publication metadata, titles, dates, authors, reaction counts, public body HTML, plain text and optional comments.

Supply newsletter hosts, custom domains or individual post URLs. Each dataset row represents one post. Comments and replies stay inside that row. Post IDs are deduplicated across the run.

Pricing

You pay for each distinct post row successfully delivered and each public comment actually included in that row. Empty results, failed fetches, duplicates and skipped paid posts have no result event charge.

EventPriceYou pay when
apify-actor-start$0.00005The platform records its synthetic start event
post-result$0.002 per postA distinct post row is delivered
comment-result$0.0003 per commentA comment or reply is included in a delivered row

The price is $2 per 1,000 delivered posts and $0.30 per 1,000 included comments. Ten posts without comments have $0.02 in result charges plus the platform start event.

The SDK checks the post budget, persists the row and charges the post event. The Actor charges included comments after the post has been persisted. Before fetching comments it reserves the next post price and trims the comment allowance to the remaining budget. It stops when another post would exceed the maximum charge limit. Local and free runs write rows without paid event charges. Actor code never charges the synthetic start event.

Collect newsletter posts without login

Provide one or more publication URLs and a result limit for each publication.

{
"publications": ["https://www.lennysnewsletter.com"],
"maxPostsPerPublication": 10,
"includeFullText": true,
"includeComments": false
}

How it works

  1. Normalize newsletter hosts, homepage URLs and individual post URLs.
  2. Fetch archive pages sequentially with a page limit of 50 and a desktop Chrome User-Agent.
  3. Optionally request public post detail and comments, then normalize the result.
  4. Push each distinct post and charge result events on paid runs.
  5. Stop at the per publication result limit, an empty or repeated page, the date cutoff, the charge limit or one minute before the run timeout.
Input -> Publication archive or single post
|
v
Public detail and comments
|
v
Normalize and deduplicate
|
v
Dataset -> Event charges

Request starts are spaced at least one second apart. Each request has a 20 second timeout, shortened near the timeout reserve. Transient failures receive up to three retries with increasing delays of 1, 2 and 4 seconds. Retry-After seconds and HTTP dates take precedence when supplied.

A custom domain returning HTTP 404 triggers publication metadata and feed self link discovery. If those do not identify the Substack host, the Actor tries the custom domain's base label as a subdomain once. Arbitrary Substack links inside articles are ignored. This fallback worked for Platformer during the spike but does not guarantee a mapping for every custom domain.

What data does it extract?

Each row contains type, publication, publicationHost, postId, title, subtitle, slug, url, publishedAt, audience, isPaid, authors, authorHandles, wordCount, reactionCount, commentCount, coverImage, description, postType, bodyHtml, bodyText, isTruncated and scrapedAt.

restacks and podcastUrl appear when provided. Missing scalar metadata is null and missing list metadata is an empty array. Publication names come from publication metadata when available, otherwise the host is used. publicationHost is the API host used, which can change after fallback. url retains the canonical post URL supplied by Substack.

With includeComments, the comments array contains id, parentId, author, body, date, reactionCount and depth. Replies follow their parent in depth first order. Deleted and suppressed comments are omitted. The included array can be shorter than the reported commentCount because of visibility, the comment limit and the charge budget.

bodyText is HTML stripped of tags, scripts and styles with entity decoding and block breaks. It is plain text rather than a faithful rendering. bodyMarkdown is not produced.

What is NOT returned

  • Private content, authenticated subscriber content or material beyond the public preview of paid posts.
  • A global search across Substack publications.
  • A guaranteed complete archive or every comment on a post.
  • Markdown, media downloads or private author profile information.

Start a run

  1. Enter up to 50 newsletter URLs, hosts or post URLs.
  2. Set the maximum posts per publication and optionally add publication search or a date cutoff.
  3. Choose whether to include public full text, comments or only free posts.
  4. Run the Actor and export the dataset as JSON, CSV or Excel.

The prefilled input requests 10 posts from Lenny's Newsletter with public full text and no comments. A live local SDK run on October 6, 2026 delivered exactly 10 rows in 11.621 seconds with body HTML and plain text verified. The normal result default is 50 per publication. Default platform resources are 256 MB and a 900 second timeout.

Input

FieldTypeDefaultMeaning
publicationsstring arrayRequired, prefill Lenny's Newsletter1 to 50 newsletter URLs, hosts or individual post URLs
maxPostsPerPublicationinteger50, prefill 10Delivered posts per supplied publication, from 1 to 5,000
sortByenumnewestnewest or top
searchQuerystringNoneSearch phrase within each publication archive
publishedAfterISO date stringNoneStop at older posts, only valid with newest order
includeFullTextbooleantrueAdd public body HTML and plain text
includeCommentsbooleanfalseInclude public comments inside each post row
maxCommentsPerPostinteger100Limit included comments and replies, zero disables fetching
freeOnlybooleanfalseSkip paid and founding audience posts
{
"publications": ["platformer.substack.com", "https://www.lennysnewsletter.com"],
"maxPostsPerPublication": 50,
"sortBy": "newest",
"searchQuery": "AI",
"publishedAfter": "2026-01-01",
"includeFullText": true,
"includeComments": true,
"maxCommentsPerPost": 100,
"freeOnly": false
}

A /p/post-slug URL returns just that post and does not fetch its archive. Search applies to archives, so it does not filter a directly supplied post URL. The date and free audience filters still apply. Overlapping publication and post inputs share a deduplication set.

Invalid publication entries are logged and skipped when at least one valid input remains. Input with zero valid publications fails before fetching. A valid hostname whose endpoint fails is skipped so later publications can still deliver rows.

Output

This row comes from the live local prefill dataset on October 6, 2026. The body fields are shortened excerpts.

{
"type": "post",
"publication": "Lenny's Newsletter",
"publicationHost": "www.lennysnewsletter.com",
"postId": 215694124,
"title": "All of the Lenny & Friends Summit talks are now online!",
"subtitle": "Plus, some reflections and takeaways from the day",
"slug": "all-of-the-lenny-and-friends-summit",
"url": "https://www.lennysnewsletter.com/p/all-of-the-lenny-and-friends-summit",
"publishedAt": "2026-09-29T13:15:57.512Z",
"audience": "everyone",
"isPaid": false,
"authors": [
"Lenny Rachitsky"
],
"authorHandles": [
"lenny"
],
"wordCount": 1276,
"reactionCount": 279,
"commentCount": 6,
"coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/9fed347c-0e02-4190-a590-1236719b99de_9213x6142.jpeg",
"description": "Plus, some reflections and takeaways from the day",
"postType": "newsletter",
"bodyHtml": "<p><em>...",
"bodyText": "Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For m...",
"isTruncated": false,
"scrapedAt": "2026-10-06T20:33:54.942Z",
"restacks": 4
}

Use from MCP agents and API clients

After deployment and publication, add this Actor through the Apify MCP server to let an agent collect newsletter posts.

https://mcp.apify.com/?tools=george.the.developer/substack-scraper

From Node.js with the Apify client and an Apify token.

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('george.the.developer/substack-scraper').call({
publications: ['https://www.lennysnewsletter.com'],
maxPostsPerPublication: 10,
includeFullText: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

With curl, wait for the run and return dataset rows.

curl -X POST "https://api.apify.com/v2/acts/george.the.developer~substack-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"publications":["https://www.lennysnewsletter.com"],"maxPostsPerPublication":10}'

Apify authentication runs the Actor. No Substack authentication is required.

Integrations

Clay. Read dataset items as an HTTP data source and map publication, title, author and post URL to columns.

n8n and Make. Select the published Actor, supply newsletter input and process dataset rows in the next step.

Google Sheets. Export CSV or use the Apify Google Sheets integration. JSON preserves the full nested comment arrays.

Use cases

  • Researchers collect public newsletter coverage and publication dates.
  • Readers organize links and publicly available text from their chosen newsletters.
  • Editorial teams monitor topics, reactions and public discussion.
  • Analysts compare posts across known publications without relying on global search.

Limits

Paid posts return only the public response available without authentication and are marked isTruncated. A long public body does not prove that the full paid article is available. The Actor does not log in, send cookies or bypass paywalls.

There is no global Substack publication search. searchQuery searches within the supplied publication archives only. Public endpoints and search behavior can change.

The maximum measured archive limit was 50. A short initial page still had further pages, so the Actor advances by the number of returned items and continues until an empty or repeated page. Paging was verified at offset 1000 on Noahpinion, but no unlimited archive guarantee was established. A result cap of 5,000 does not make more posts available.

Newest date cutoff stopping assumes the endpoint returns newest order. Feed changes, pinned posts and changing archives can affect completeness. Publication names may fall back to the hostname when public bylines do not identify the publication.

One request per second succeeded for the tested detail calls without HTTP 429. Rate limits can vary by publication and network. No real 429 was observed during the spike; Retry-After handling was verified with offline responses.

Custom domain discovery can fail when the publication name differs from its Substack subdomain or when the custom domain has moved to another publishing system. Failure after the fallback is logged and skipped.

Public comments can be incomplete, unavailable or restricted. A failed detail call still delivers archive metadata with null body fields and fullTextError. Failed comment calls deliver an empty comment array and commentsError. Storage and billing failures stop the run as visible errors.

FAQ

Do I need a Substack account or proxy?

No. The Actor uses unauthenticated public endpoints with a desktop User-Agent.

Can I retrieve complete paid articles?

Only the publicly available preview is returned. Paid audience posts are marked as truncated, and no paywall bypass is attempted.

Are duplicate posts charged twice?

No. Each post ID is delivered once per run, even across overlapping inputs.

What happens when a publication is empty or fails?

An empty publication has no result charges. Exhausted publication fetch failures are logged and skipped. The run exits successfully with whatever it delivered unless input, storage or billing fails.

Can I run this locally?

Yes. Install Node 22 or newer, run npm install --omit=dev --omit=optional, place input in storage/key_value_stores/default/INPUT.json and run npm start. Read rows in storage/datasets/default. Run offline checks with npm test and reproduce the live prefill check with node test/live-prefill.mjs.