Substack Publications & Posts Scraper
Pricing
from $3.00 / 1,000 results
Substack Publications & Posts Scraper
Discover Substack publications by category or name and pull their recent posts with author, subscriber count and recommended sister publications. Filter by topic, cap results, pay only per result returned.
What does Substack Publications & Posts Scraper do?
It discovers Substack publications, either by browsing a public topic category (Technology, Business, Culture, Crypto and 28 more) or from a list of publications you name directly, then pulls each one's most recent posts with the author, subscriber count, engagement numbers and the publication's recommended sister publications. Set a category and post count and you get a clean, deduplicated table of newsletters and their latest content in one run. Pricing is per result: you only pay for posts actually returned.
Data source
The Actor reads Substack's own public JSON API (the same endpoints substack.com and every *.substack.com site call from the browser): the category leaderboard for discovery, /api/v1/posts for each publication's post archive, and /api/v1/recommendations for cross-promotions. No login, no scraping of rendered HTML, data as fresh as the live site.
Why use it?
- Lead generation: find active newsletter operators in your niche along with their subscriber tier and contact handle.
- Content monitoring: track what competitors or influencers in a topic are publishing, how often, and how it performs (comments, reactions, restacks).
- Market research: measure which publications in a category are growing (subscriber counts) and who they cross-promote.
- Replaces manually browsing Substack's leaderboard and each publication's archive page by hand.
How to use it
- Open the Actor and pick one or more categories, or list specific publications by subdomain or URL.
- Click Start. Download results as JSON, CSV or Excel, or read them through the API.
- Add a Schedule and an Integration (Slack, Google Sheets, webhook) to get new posts automatically.
Input
| Field | Description |
|---|---|
categories | Discover publications from Substack's public category leaderboards (Technology, Business, Culture, and 29 more). Default: ["technology"]. |
publicationSubdomains | Optional list of exact publications to include, as a subdomain, bare name, or full substack.com URL (e.g. platformer, garbageday.substack.com). Runs in addition to any categories. |
postsPerPublication | How many of each publication's most recent posts to return. Default 3, max 20. |
maxPublications | Cap on distinct publications crawled across all selected categories. Default 15, max 100. |
includeRecommendations | Attach each publication's publicly recommended sister publications to its post records. Default true. |
maxResults | Hard stop for the run. You are billed per result, so this caps the cost. |
proxyConfiguration | Apify proxy is recommended for larger runs. |
Example input:
{"categories": ["technology"],"postsPerPublication": 3,"maxPublications": 15,"includeRecommendations": true,"maxResults": 60}
Output
Every record is one post, enriched with its publication and author:
{"id": "careerbrew:216893901","title": "Career Brew - 53 Hottest Early to Mid Career Jobs","subtitle": "$147K at Mastercard; $195K at NVIDIA; $220K at HighLevel and many more","url": "https://careerbrew.substack.com/p/career-brew-28th-sep-53-hottest-early","postDate": "2026-09-28T17:56:05.255Z","publicationName": "Career Brew","publicationSubdomain": "careerbrew","authorName": "Career Brew","category": "technology","postId": 216893901,"publicationId": 2333426,"publicationUrl": "https://careerbrew.substack.com","authorHandle": "careerbrew","type": "newsletter","audience": "everyone","wordcount": 612,"commentCount": 0,"reactionCount": 8,"restacks": 1,"subscriberTier": null,"freeSubscriberCount": 355000,"freeSubscriberCountDisplay": "355K+","publicationRank": 5,"heroText": "Increasing accessibility to education and career opportunities...","coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/...","logoUrl": "https://substackcdn.com/image/fetch/...","recommendedPublications": [{ "name": "The Founders Corner®", "subdomain": "thefoundercorner" },{ "name": "The VC Corner", "subdomain": "thevccorner" }],"scrapedAt": "2026-09-30T08:18:47.916Z"}
| Field | Meaning |
|---|---|
id | Stable key (subdomain:postId), safe for deduplication between runs |
title, subtitle, url | Post headline, summary and canonical link |
postDate | ISO date the post was published |
publicationName, publicationSubdomain, publicationUrl | The newsletter it belongs to |
authorName, authorHandle | Byline of the post |
category | The category it was discovered under, or null if it came from publicationSubdomains |
type, audience | newsletter/podcast/thread; who can read it (everyone, only_paid, founding) |
wordcount, commentCount, reactionCount, restacks | Engagement numbers |
freeSubscriberCount, freeSubscriberCountDisplay | Publication's free subscriber count, exact and rounded; only available for publications found via categories, since Substack does not expose it on a per-publication lookup |
subscriberTier | Author's public subscriber milestone badge, when Substack has assigned one |
recommendedPublications | Up to 5 publications this one recommends to its readers |
Pricing
Pay per event: a small fee when a run starts, then a fixed price per result returned. You only pay for posts that pass your filters. Set maxResults to cap any run.
Limitations
- Only covers publications still hosted on
*.substack.comor a Substack-managed custom domain; publications that migrated off Substack (their own site, their own API) are not covered. freeSubscriberCountandpublicationRankare only populated for publications discovered throughcategories; publications you name directly inpublicationSubdomainsdo not expose subscriber counts through a public per-publication lookup.- Only public, published posts are returned. Paywalled post bodies are not fetched, only their public metadata (title, engagement, audience).
FAQ
Is this legal? The Actor reads the same public JSON API Substack's own pages call for any visitor without a login, and stores only publication and post metadata the site publishes for that purpose. You are responsible for how you use the data.
Something is missing or broken? Open an issue in the Issues tab; response time is usually under a day.