Substack Scraper (posts, publications, likes, full text)
Pricing
from $0.80 / 1,000 result items
Substack Scraper (posts, publications, likes, full text)
Substack scraper with no login and no proxy: every post of any publication as a flat row — title, subtitle, authors, date, free/paid audience, likes, comments, restacks, word count, cover image and tags, plus the full text and HTML on request. Search publications by topic with subscriber counts.
Pricing
from $0.80 / 1,000 result items
Rating
0.0
(0)
Developer
Viktor Dubnytskiy
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Scrape every post of any Substack publication as a flat row — title, subtitle, authors, date, free/paid audience, likes, comments, restacks, word count, cover image and tags — with the full text and HTML available on request. Find publications by topic too. No login, no cookies, no proxy.
What you get
One row per post. Real row from the example dataset (publication lennysnewsletter.com, body truncated here):
| title | authors | likes / comments / restacks | wordCount |
|---|---|---|---|
How to turn your AI into a world-class designer | ["Anshu Chimala"] | 586 / 17 / 39 | 4604 |
Publication search (searches) returns type: publication rows instead: title, url, subdomain, subscriberCountText, authorName, logoUrl.
Full post row: postId, url, title, subtitle, slug, publication, publicationUrl, authors[], postedAt, audience, postType, likes, comments, restacks, wordCount, coverImage, description, bodyText, bodyHtml (the last two with includeBody), tags[], scrapedAt.
Use cases
- Newsletter research — pull a publication's full archive with engagement numbers and see what actually lands.
- Competitor and topic tracking — monitor mode returns only new posts or posts whose like/comment counts moved.
- Discovery — search Substack by topic and get publications with their subscriber-count text before you pitch or subscribe.
Try it in 10 seconds
Hit Start/Try it — the input already works: publications: ["lennysnewsletter.com"], maxPosts: 50, maxItems: 20, nothing required.
To track a publication over time: save the input as a Task, set mode to monitor, and put it on a schedule (Apify → Schedules → cron 0 8 * * * for daily). Each run then returns only posts that are new or whose like/comment/restack count changed, and can post the diff to webhookUrl or Telegram.
Related actors
- Medium Articles Scraper (tags, authors, publications, claps) — the same long-form content pattern for Medium.
- Reddit Scraper (subreddit posts, search, comments, no login) — see community discussion instead of newsletter posts.
- Google News Scraper (search, headlines, alerts, no login) — for headline-level monitoring instead of full posts.
How it works
- Publications → the publication's archive is paged (
sort=newortop) up tomaxPosts, straight from Substack's own public JSON. includeBody→ each post is fetched forbodyTextandbodyHtml(one extra request per post).- Searches → Substack's publication search,
searchPages× about 20 publications, returned astype: publicationrows. - No proxy is needed — these endpoints answer directly.
Input
| Field | Meaning | Default |
|---|---|---|
publications | Subdomains, domains or URLs, e.g. lennysnewsletter.com | ["lennysnewsletter.com"] |
searches | Topics → matching publications with subscriber counts | empty |
sort | Archive order: new or top | new |
maxPosts | Archive depth per publication | 50 |
includeBody | Fetch each post for full text and HTML | false |
searchPages | Search pages per topic (~20 publications each) | 2 |
maxItems | Stop after this many rows | 20 |
mode | scrape or monitor (only new/changed since last run) | scrape |
monitorStateId, webhookUrl, telegramBotToken, telegramChatId | Monitor-mode state key and alert targets | empty |
Pricing
| Event | Price |
|---|---|
| result | $0.001 per row ($1 per 1,000) |
| monitor-check | $0.005 per monitor run |
| change | $0.001 per new/changed row |
Charged only for rows actually pushed. No proxy needed.
Found it useful? A short review on the Store page helps other people find this actor and tells us what to improve. If something is wrong, open an issue on the actor page — issues are answered within a day.
Why this actor
- No login, no cookies and no proxy cost — your bill is the result events plus compute.
- Engagement is in every row: likes, comments, restacks and word count, not just title and date.
- Posts and publication discovery in one actor: search a topic, then read the archives you found.
- Full text and HTML are optional, so you do not pay the extra request when you only need metadata.
- Monitor mode with webhook and Telegram alerts on new posts and changed like counts.
Limits
- Paywalled posts return the public preview only.
- Comments and subscriber lists are not included.
subscriberCountTextis whatever the publication chooses to display (e.g. "50,000+ subscribers"), not an exact number, and many publications display nothing.
FAQ
Does it need a Substack login or a paid subscription? No. There is no login field. Paid posts appear with audience: only_paid and, with includeBody on, return the free preview — never the paywalled remainder.
Can I get the full article text? Yes, turn on includeBody: each post then carries bodyText and bodyHtml, at the cost of one extra request per post.
What happens when there are no results? No rows are pushed and no result events are charged. The RUN_SUMMARY record in the run's key-value store carries emptyReason, so a misspelled publication is distinguishable from a real block.
Changelog
- 0.1: initial release.
If this actor saved you time, a short review on its Store page genuinely helps other people find it. Found a bug or need a field that is missing? Open a ticket on the Issues tab.