Substack Scraper | Public Full Text, Archives & Comments avatar

Substack Scraper | Public Full Text, Archives & Comments

Pricing

from $0.40 / 1,000 posts

Go to Apify Store
Substack Scraper | Public Full Text, Archives & Comments

Substack Scraper | Public Full Text, Archives & Comments

Extract public Substack posts, readable text, author and publication metadata, nested comments and archive date filters. Track content access and source coverage explicitly. Search articles by keyword through the public Posts search interface.

Pricing

from $0.40 / 1,000 posts

Rating

0.0

(0)

Developer

tingyou333 zhuang

tingyou333 zhuang

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Collect public Substack posts as structured data for publication monitoring, research datasets, editorial workflows, and content-change tracking. Get original public HTML, clean text, author and publication metadata, engagement, podcasts, and nested comments.

This independent tool is not affiliated with Substack. No Substack login, personal cookies, paid subscription, or proxy purchase is required. Paid and restricted article bodies are not unlocked.

What you can collect

  • A single article, a publication archive, or keyword search results, including custom publication domains.
  • Public newsletter, podcast, and discussion-thread posts, with date and access filters applied before the result cap.
  • Original public HTML plus readable text and a body hash for downstream change detection.
  • Public comments and available replies, preserving the source's default order and nested structure.
  • Explicit content availability, comment visibility gaps, request limits, and pagination termination reasons.

Quick start

One public article:

{
"urls": ["https://www.oneusefulthing.org/p/the-overhang"],
"maxPostsPerNewsletter": 1,
"includeContent": true,
"includePublicationInfo": true,
"includeComments": true,
"maxCommentsPerPost": 20,
"maxRequests": 8
}

Find articles about a topic:

{
"keywords": ["artificial intelligence"],
"maxSearchResultsPerKeyword": 5,
"includeContent": true,
"onlyFree": true,
"maxRequests": 20
}

Run the Actor and open its default dataset. JSON preserves nested comments and publication objects; CSV and Excel are convenient for flat article fields. Read the SUMMARY key in the run's key-value store for source errors, warnings, and coverage. A partial result keeps useful rows and explains what was unavailable; a source failure never becomes a fake article row.

Inputs

All 12 input names from automation-lab/substack-scraper are supported. Supply at least one URL or keyword. Each of the urls and keywords lists accepts at most 500 entries per run.

InputDefaultMeaning
urls[]HTTPS publication homepages, /archive, or /p/slug; custom domains accepted.
keywords[]Search public articles, preserving source relevance order.
maxSearchResultsPerKeyword20Maximum emitted matching posts per keyword; 1-100.
maxPostsPerNewsletter100Maximum emitted matching posts per archive. 0 removes this output cap.
includeContenttrueReturn public HTML and readable text. Restricted requested bodies remain empty.
includeCommentsfalseInclude public comment trees and available replies.
maxCommentsPerPost20Count every emitted parent and reply toward this cap; 0 removes the output cap.
includePublicationInfotrueInclude public publication details and subscriber labels when exposed.
contentType"all"all, newsletter, podcast, or thread.
startDateunsetInclusive UTC calendar date in YYYY-MM-DD format.
endDateunsetInclusive UTC calendar date, including the entire day.
onlyFreefalseKeep public audience=everyone posts without hidden/unlock requirements.

Additional controls help keep jobs bounded:

InputDefaultMeaning
pageSize201-50 requested archive items per page.
maxRequests1000Maximum source attempts including retries and redirects; 1-100,000. DNS resolution may make separate requests.
requestDelaySeconds0.4Minimum delay between source requests; 0.1-60 seconds.
requestTimeoutSeconds30Per-request timeout; 1-120 seconds.
maxRunSeconds600Collection time budget; 1-86,400 seconds. Set the platform run timeout higher to allow the summary to finish.

Filters apply before result caps and are rechecked against current post details. Search relevance is not chronological, so date-filtered searches may scan multiple pages. Archive offsets advance by the actual response length, including short nonterminal pages. 0 removes an output cap but does not remove request, time, source, or billing limits.

Unknown input fields are rejected rather than silently ignored. Reader-profile URLs, proxy settings, and concurrency controls are not part of the reference's public input contract. Requests are sequential. Tracking query parameters on article URLs are removed.

Output and migration

Each dataset row is one post. These reference fields are preserved:

  • Identity and dates: postId, title, subtitle, slug, url, publishedAt, updatedAt, postType.
  • Access and content: audience, isPaid, wordcount, coverImage, description, tags, bodyHtml, truncatedBodyText.
  • Engagement and media: reactionCount, commentCount, childCommentCount, restacks, hasVoiceover, podcastUrl, podcastDuration.
  • People and publication: authorName, authorHandle, authorPhotoUrl, authorBio, authorId, publicationId, publicationName, publicationUrl, subscriberCount.
  • comments and the actual observation timestamp scrapedAt.

Migration details matter: bodyHtml is null when content is not requested; publicationUrl is null when publication information is not requested; comments is null when comments are not requested. Requested but empty comments use an empty array. Requested restricted bodies use empty strings with an explicit access status.

Comments contain id, body, date, editedAt, name, handle, photoUrl, reactionCount, restacks, isAuthor, isPinned, and recursive replies. The source's default ordering is retained before applying the nested cap; no guessed ranking formula is applied.

Additional fields provide source context without changing the reference fields:

FieldsPurpose
bodyText, bodyTextLength, bodySha256Readable content and a hash of the exact returned public HTML.
contentStatuspublic_full, restricted, access_unknown, unavailable, or not_requested.
sourceUrl, discoveryPublic source URL and keyword discovery context.
publication, languageAdditional public publication details and source-reported language.
commentsStatusnot_requested, limit_reached, public_snapshot, source_partial, or unavailable.
commentsReturned, commentsAvailable, commentsTruncatedEmitted count, visible count, and whether your cap removed visible comments.
commentsReportedCount, commentsReportedTopLevelCount, commentsCountGapSource totals and gaps against the visible public snapshot.
commentsSourceHasMissingRepliesWhether returned reply counts indicate unavailable child comments.

Unknown metrics remain null instead of becoming zero. Subscriber labels retain the source's strings or rounded values; they are not estimates of paying subscribers. Source commentCount already includes replies in the observed API and must not be added to childCommentCount.

Coverage and limits

Public sources have been checked with single posts, keyword pagination, a 110-post archive ending in a real empty page, public podcasts and threads, and nested comments. Bounded private cloud checks include 100 archive rows across three pages, 20 public article bodies, and a 90-comment snapshot. These samples describe tested coverage, not a guarantee for every publication or future source response. Detailed version-specific evidence is retained in the maintainer's validation records.

Deleted, suppressed, and moderated-hidden comments are excluded. One observed post reported 91 comments but exposed 90 public comments; the output preserved the gap and a partial summary. A reported total is not a promise that hidden material can be retrieved. Very large threads may encounter source-side limits.

No paid body, subscriber-only content, login challenge, or paywall is bypassed. Source changes, custom-domain failures, restricted comments, and unavailable publication metadata are visible in SUMMARY. Requests retain TLS verification and public-address checks; local/private destinations, credentials in URLs, and nonstandard ports are rejected.

Pricing and cost controls

One saved post is one billable result. Requested public HTML, readable text, publication information and available comments are included in the post price. There is no separate per-comment or full-text event. The Pricing tab is the source of truth for the effective price.

Apify plan / discountPrice per 1,000 saved posts
Free / no discount$0.80
Starter / Bronze$0.65
Scale / Silver$0.50
Business / Gold$0.40

The startup event costs $0.0002 at the default 256 MB memory. Apify charges one startup event up to 1 GB, then one per additional GB. Platform usage is included in the event prices. For example, 100 posts at the no-discount price and default memory cost $0.0802, including public content and requested comments. Even a run with no saved posts may incur the startup event.

Set Max cost per run to control event spending; its minimum is $0.001. The Actor stops collecting when the next post no longer fits the remaining budget. The default platform timeout is 660 seconds, leaving time for the default 600-second collection budget and final summary. If you raise maxRunSeconds, also raise the platform timeout. An API client's own request timeout is separate.

Content jobs require additional article requests, and comments can increase response size. Start with a small cap and inspect the run's actual usage. For repeat monitoring, use a date window and deduplicate results in your downstream store. bodySha256 helps identify article changes.

API and scheduling

With an Apify API token in your environment:

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("tGHdx3RdqGDv2BbaR").call(run_input={
"urls": ["https://www.oneusefulthing.org/p/the-overhang"],
"includeContent": True,
"maxRequests": 8,
})
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())
summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("SUMMARY")

Save an input as an Apify task to schedule repeat runs, or use a webhook/integration to consume the dataset. Deduplication within a run uses postId. Cross-run deduplication belongs in your downstream store; summary offsets are diagnostics, not an automatic resume-token API.

FAQ

Can I collect an entire archive? Set maxPostsPerNewsletter to 0 and allow sufficient request/time budgets. Collection stops on a real empty page or a stated limit/error. Only the public listing is accessible.

Why is my run marked partial? Inspect SUMMARY: visible comments may differ from reported totals, a requested source may be unavailable, or a safety budget may have stopped collection.

Can I retrieve paid articles? Public metadata and previews may be returned. Restricted article HTML and text remain empty; no subscription access is supplied.

Do I need a Substack account or paid proxy? No. The implementation uses anonymous public sources. Source availability can still change.

Why are some fields null? They were not requested or not exposed by the source. Nulls preserve that distinction instead of inventing values.

Substack branding identifies the supported source; all brand rights remain with their owners. This Actor is an independent integration.