Substack Scraper | Public Full Text, Archives & Comments
Pricing
from $0.40 / 1,000 posts
Substack Scraper | Public Full Text, Archives & Comments
Extract public Substack posts, readable text, author and publication metadata, nested comments and archive date filters. Track content access and source coverage explicitly. Search articles by keyword through the public Posts search interface.
Pricing
from $0.40 / 1,000 posts
Rating
0.0
(0)
Developer
tingyou333 zhuang
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Collect public Substack posts as structured data for publication monitoring, research datasets, editorial workflows, and content-change tracking. Get original public HTML, clean text, author and publication metadata, engagement, podcasts, and nested comments.
This independent tool is not affiliated with Substack. No Substack login, personal cookies, paid subscription, or proxy purchase is required. Paid and restricted article bodies are not unlocked.
What you can collect
- A single article, a publication archive, or keyword search results, including custom publication domains.
- Public newsletter, podcast, and discussion-thread posts, with date and access filters applied before the result cap.
- Original public HTML plus readable text and a body hash for downstream change detection.
- Public comments and available replies, preserving the source's default order and nested structure.
- Explicit content availability, comment visibility gaps, request limits, and pagination termination reasons.
Quick start
One public article:
{"urls": ["https://www.oneusefulthing.org/p/the-overhang"],"maxPostsPerNewsletter": 1,"includeContent": true,"includePublicationInfo": true,"includeComments": true,"maxCommentsPerPost": 20,"maxRequests": 8}
Find articles about a topic:
{"keywords": ["artificial intelligence"],"maxSearchResultsPerKeyword": 5,"includeContent": true,"onlyFree": true,"maxRequests": 20}
Run the Actor and open its default dataset. JSON preserves nested comments and publication objects; CSV and Excel are convenient for flat article fields. Read the SUMMARY key in the run's key-value store for source errors, warnings, and coverage. A partial result keeps useful rows and explains what was unavailable; a source failure never becomes a fake article row.
Inputs
All 12 input names from automation-lab/substack-scraper are supported. Supply at least one URL or keyword. Each of the urls and keywords lists accepts at most 500 entries per run.
| Input | Default | Meaning |
|---|---|---|
urls | [] | HTTPS publication homepages, /archive, or /p/slug; custom domains accepted. |
keywords | [] | Search public articles, preserving source relevance order. |
maxSearchResultsPerKeyword | 20 | Maximum emitted matching posts per keyword; 1-100. |
maxPostsPerNewsletter | 100 | Maximum emitted matching posts per archive. 0 removes this output cap. |
includeContent | true | Return public HTML and readable text. Restricted requested bodies remain empty. |
includeComments | false | Include public comment trees and available replies. |
maxCommentsPerPost | 20 | Count every emitted parent and reply toward this cap; 0 removes the output cap. |
includePublicationInfo | true | Include public publication details and subscriber labels when exposed. |
contentType | "all" | all, newsletter, podcast, or thread. |
startDate | unset | Inclusive UTC calendar date in YYYY-MM-DD format. |
endDate | unset | Inclusive UTC calendar date, including the entire day. |
onlyFree | false | Keep public audience=everyone posts without hidden/unlock requirements. |
Additional controls help keep jobs bounded:
| Input | Default | Meaning |
|---|---|---|
pageSize | 20 | 1-50 requested archive items per page. |
maxRequests | 1000 | Maximum source attempts including retries and redirects; 1-100,000. DNS resolution may make separate requests. |
requestDelaySeconds | 0.4 | Minimum delay between source requests; 0.1-60 seconds. |
requestTimeoutSeconds | 30 | Per-request timeout; 1-120 seconds. |
maxRunSeconds | 600 | Collection time budget; 1-86,400 seconds. Set the platform run timeout higher to allow the summary to finish. |
Filters apply before result caps and are rechecked against current post details. Search relevance is not chronological, so date-filtered searches may scan multiple pages. Archive offsets advance by the actual response length, including short nonterminal pages. 0 removes an output cap but does not remove request, time, source, or billing limits.
Unknown input fields are rejected rather than silently ignored. Reader-profile URLs, proxy settings, and concurrency controls are not part of the reference's public input contract. Requests are sequential. Tracking query parameters on article URLs are removed.
Output and migration
Each dataset row is one post. These reference fields are preserved:
- Identity and dates:
postId,title,subtitle,slug,url,publishedAt,updatedAt,postType. - Access and content:
audience,isPaid,wordcount,coverImage,description,tags,bodyHtml,truncatedBodyText. - Engagement and media:
reactionCount,commentCount,childCommentCount,restacks,hasVoiceover,podcastUrl,podcastDuration. - People and publication:
authorName,authorHandle,authorPhotoUrl,authorBio,authorId,publicationId,publicationName,publicationUrl,subscriberCount. commentsand the actual observation timestampscrapedAt.
Migration details matter: bodyHtml is null when content is not requested; publicationUrl is null when publication information is not requested; comments is null when comments are not requested. Requested but empty comments use an empty array. Requested restricted bodies use empty strings with an explicit access status.
Comments contain id, body, date, editedAt, name, handle, photoUrl, reactionCount, restacks, isAuthor, isPinned, and recursive replies. The source's default ordering is retained before applying the nested cap; no guessed ranking formula is applied.
Additional fields provide source context without changing the reference fields:
| Fields | Purpose |
|---|---|
bodyText, bodyTextLength, bodySha256 | Readable content and a hash of the exact returned public HTML. |
contentStatus | public_full, restricted, access_unknown, unavailable, or not_requested. |
sourceUrl, discovery | Public source URL and keyword discovery context. |
publication, language | Additional public publication details and source-reported language. |
commentsStatus | not_requested, limit_reached, public_snapshot, source_partial, or unavailable. |
commentsReturned, commentsAvailable, commentsTruncated | Emitted count, visible count, and whether your cap removed visible comments. |
commentsReportedCount, commentsReportedTopLevelCount, commentsCountGap | Source totals and gaps against the visible public snapshot. |
commentsSourceHasMissingReplies | Whether returned reply counts indicate unavailable child comments. |
Unknown metrics remain null instead of becoming zero. Subscriber labels retain the source's strings or rounded values; they are not estimates of paying subscribers. Source commentCount already includes replies in the observed API and must not be added to childCommentCount.
Coverage and limits
Public sources have been checked with single posts, keyword pagination, a 110-post archive ending in a real empty page, public podcasts and threads, and nested comments. Bounded private cloud checks include 100 archive rows across three pages, 20 public article bodies, and a 90-comment snapshot. These samples describe tested coverage, not a guarantee for every publication or future source response. Detailed version-specific evidence is retained in the maintainer's validation records.
Deleted, suppressed, and moderated-hidden comments are excluded. One observed post reported 91 comments but exposed 90 public comments; the output preserved the gap and a partial summary. A reported total is not a promise that hidden material can be retrieved. Very large threads may encounter source-side limits.
No paid body, subscriber-only content, login challenge, or paywall is bypassed. Source changes, custom-domain failures, restricted comments, and unavailable publication metadata are visible in SUMMARY. Requests retain TLS verification and public-address checks; local/private destinations, credentials in URLs, and nonstandard ports are rejected.
Pricing and cost controls
One saved post is one billable result. Requested public HTML, readable text, publication information and available comments are included in the post price. There is no separate per-comment or full-text event. The Pricing tab is the source of truth for the effective price.
| Apify plan / discount | Price per 1,000 saved posts |
|---|---|
| Free / no discount | $0.80 |
| Starter / Bronze | $0.65 |
| Scale / Silver | $0.50 |
| Business / Gold | $0.40 |
The startup event costs $0.0002 at the default 256 MB memory. Apify charges one startup event up to 1 GB, then one per additional GB. Platform usage is included in the event prices. For example, 100 posts at the no-discount price and default memory cost $0.0802, including public content and requested comments. Even a run with no saved posts may incur the startup event.
Set Max cost per run to control event spending; its minimum is $0.001. The Actor stops collecting when the next post no longer fits the remaining budget. The default platform timeout is 660 seconds, leaving time for the default 600-second collection budget and final summary. If you raise maxRunSeconds, also raise the platform timeout. An API client's own request timeout is separate.
Content jobs require additional article requests, and comments can increase response size. Start with a small cap and inspect the run's actual usage. For repeat monitoring, use a date window and deduplicate results in your downstream store. bodySha256 helps identify article changes.
API and scheduling
With an Apify API token in your environment:
import osfrom apify_client import ApifyClientclient = ApifyClient(os.environ["APIFY_TOKEN"])run = client.actor("tGHdx3RdqGDv2BbaR").call(run_input={"urls": ["https://www.oneusefulthing.org/p/the-overhang"],"includeContent": True,"maxRequests": 8,})rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())summary = client.key_value_store(run["defaultKeyValueStoreId"]).get_record("SUMMARY")
Save an input as an Apify task to schedule repeat runs, or use a webhook/integration to consume the dataset. Deduplication within a run uses postId. Cross-run deduplication belongs in your downstream store; summary offsets are diagnostics, not an automatic resume-token API.
FAQ
Can I collect an entire archive? Set maxPostsPerNewsletter to 0 and allow sufficient request/time budgets. Collection stops on a real empty page or a stated limit/error. Only the public listing is accessible.
Why is my run marked partial? Inspect SUMMARY: visible comments may differ from reported totals, a requested source may be unavailable, or a safety budget may have stopped collection.
Can I retrieve paid articles? Public metadata and previews may be returned. Restricted article HTML and text remain empty; no subscription access is supplied.
Do I need a Substack account or paid proxy? No. The implementation uses anonymous public sources. Source availability can still change.
Why are some fields null? They were not requested or not exposed by the source. Nulls preserve that distinction instead of inventing values.
Substack branding identifies the supported source; all brand rights remain with their owners. This Actor is an independent integration.