Substack Scraper (posts, publications, likes, full text) avatar

Substack Scraper (posts, publications, likes, full text)

Pricing

from $0.80 / 1,000 result items

Go to Apify Store
Substack Scraper (posts, publications, likes, full text)

Substack Scraper (posts, publications, likes, full text)

Substack scraper with no login and no proxy: every post of any publication as a flat row — title, subtitle, authors, date, free/paid audience, likes, comments, restacks, word count, cover image and tags, plus the full text and HTML on request. Search publications by topic with subscriber counts.

Pricing

from $0.80 / 1,000 result items

Rating

0.0

(0)

Developer

Viktor Dubnytskiy

Viktor Dubnytskiy

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Scrape every post of any Substack publication as a flat row — title, subtitle, authors, date, free/paid audience, likes, comments, restacks, word count, cover image and tags — with the full text and HTML available on request. Find publications by topic too. No login, no cookies, no proxy.

What you get

One row per post. Real row from the example dataset (publication lennysnewsletter.com, body truncated here):

titleauthorslikes / comments / restackswordCount
How to turn your AI into a world-class designer["Anshu Chimala"]586 / 17 / 394604

Publication search (searches) returns type: publication rows instead: title, url, subdomain, subscriberCountText, authorName, logoUrl.

Full post row: postId, url, title, subtitle, slug, publication, publicationUrl, authors[], postedAt, audience, postType, likes, comments, restacks, wordCount, coverImage, description, bodyText, bodyHtml (the last two with includeBody), tags[], scrapedAt.

Use cases

  • Newsletter research — pull a publication's full archive with engagement numbers and see what actually lands.
  • Competitor and topic tracking — monitor mode returns only new posts or posts whose like/comment counts moved.
  • Discovery — search Substack by topic and get publications with their subscriber-count text before you pitch or subscribe.

Try it in 10 seconds

Hit Start/Try it — the input already works: publications: ["lennysnewsletter.com"], maxPosts: 50, maxItems: 20, nothing required.

To track a publication over time: save the input as a Task, set mode to monitor, and put it on a schedule (Apify → Schedules → cron 0 8 * * * for daily). Each run then returns only posts that are new or whose like/comment/restack count changed, and can post the diff to webhookUrl or Telegram.

How it works

  1. Publications → the publication's archive is paged (sort = new or top) up to maxPosts, straight from Substack's own public JSON.
  2. includeBody → each post is fetched for bodyText and bodyHtml (one extra request per post).
  3. Searches → Substack's publication search, searchPages × about 20 publications, returned as type: publication rows.
  4. No proxy is needed — these endpoints answer directly.

Input

FieldMeaningDefault
publicationsSubdomains, domains or URLs, e.g. lennysnewsletter.com["lennysnewsletter.com"]
searchesTopics → matching publications with subscriber countsempty
sortArchive order: new or topnew
maxPostsArchive depth per publication50
includeBodyFetch each post for full text and HTMLfalse
searchPagesSearch pages per topic (~20 publications each)2
maxItemsStop after this many rows20
modescrape or monitor (only new/changed since last run)scrape
monitorStateId, webhookUrl, telegramBotToken, telegramChatIdMonitor-mode state key and alert targetsempty

Pricing

EventPrice
result$0.001 per row ($1 per 1,000)
monitor-check$0.005 per monitor run
change$0.001 per new/changed row

Charged only for rows actually pushed. No proxy needed.

Found it useful? A short review on the Store page helps other people find this actor and tells us what to improve. If something is wrong, open an issue on the actor page — issues are answered within a day.

Why this actor

  • No login, no cookies and no proxy cost — your bill is the result events plus compute.
  • Engagement is in every row: likes, comments, restacks and word count, not just title and date.
  • Posts and publication discovery in one actor: search a topic, then read the archives you found.
  • Full text and HTML are optional, so you do not pay the extra request when you only need metadata.
  • Monitor mode with webhook and Telegram alerts on new posts and changed like counts.

Limits

  • Paywalled posts return the public preview only.
  • Comments and subscriber lists are not included.
  • subscriberCountText is whatever the publication chooses to display (e.g. "50,000+ subscribers"), not an exact number, and many publications display nothing.

FAQ

Does it need a Substack login or a paid subscription? No. There is no login field. Paid posts appear with audience: only_paid and, with includeBody on, return the free preview — never the paywalled remainder.

Can I get the full article text? Yes, turn on includeBody: each post then carries bodyText and bodyHtml, at the cost of one extra request per post.

What happens when there are no results? No rows are pushed and no result events are charged. The RUN_SUMMARY record in the run's key-value store carries emptyReason, so a misspelled publication is distinguishable from a real block.

Changelog

  • 0.1: initial release.

If this actor saved you time, a short review on its Store page genuinely helps other people find it. Found a bug or need a field that is missing? Open a ticket on the Issues tab.