Substack Scraper: Posts, Comments, Search & Sponsors avatar

Substack Scraper: Posts, Comments, Search & Sponsors

Pricing

from $3.00 / 1,000 post scrapeds

Go to Apify Store
Substack Scraper: Posts, Comments, Search & Sponsors

Substack Scraper: Posts, Comments, Search & Sponsors

Scrape any Substack newsletter or search all of Substack by keyword. Get full post text, likes, comments, restacks, author and publication data, plus the sponsors in each issue. Export to JSON, CSV or Excel.

Pricing

from $3.00 / 1,000 post scrapeds

Rating

0.0

(0)

Developer

Matias De Pascuale

Matias De Pascuale

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Substack Scraper & Newsletter Sponsor Finder

Scrape posts from any Substack newsletter — including custom domains — or search all of Substack by keyword, and see who sponsors them.

For every post you get the title, date, likes, comments, restacks, paywall status and the companies sponsoring it ("Brought to you by…", "Sponsored by…"). A summary ranks every sponsor across all the newsletters you scraped.

Who uses it

  • Brands & agencies — find newsletters your competitors sponsor, and how often.
  • Newsletter writers — build a sponsor prospect list from newsletters in your niche.
  • Researchers & investors — track engagement and posting cadence of top newsletters.
  • Content teams — export full archives, or every post on a topic, for analysis or AI summarization.
  • Trend research — search a keyword to see which newsletters cover it and which posts get the most engagement.

Example

Scraping the latest 8 posts of three tech newsletters (September 2026):

NewsletterSponsored postsSponsors found
Not Boring63%Reducto, Jack & Jill, Arena Magazine
The Pragmatic Engineer38%turbopuffer, Antithesis, Entire, Linear, WorkOS
Lenny's Newsletter0%—

Input

{
"publications": ["https://www.notboring.co", "platformer", "newsletter.pragmaticengineer.com"],
"maxPostsPerPublication": 50,
"since": "2026-01-01",
"detectSponsors": true,
"includeContent": false
}

Newsletters can be given as a URL, a name.substack.com subdomain or just the handle.

Search by keyword

Don't know which newsletters to scrape? Search all of Substack instead:

{
"searchQueries": ["AI agents", "personal finance"],
"maxPostsPerQuery": 50
}

Each result includes matchedQuery, plus the same metrics and sponsor detection. You can combine publications and searchQueries in one run; duplicate posts are skipped.

Output

Dataset — one row per post:

{
"publication": "Not Boring by Packy McCormick",
"title": "Weekly Dose of Optimism #212",
"url": "https://www.notboring.co/p/weekly-dose-of-optimism-212",
"postDate": "2026-09-25T12:36:02.277Z",
"audience": "only_paid",
"isPaywalled": true,
"wordcount": 3190,
"reactions": 68,
"comments": 0,
"restacks": 1,
"hasSponsor": true,
"sponsors": [
{
"name": "Jack & Jill",
"domain": "jackandjill.ai",
"url": "https://jackandjill.ai/jill?utm_source=notboring"
}
]
}

Set Include full post text to add contentText (for paywalled posts, only the free preview).

SUMMARY record — per newsletter: posts scraped, average likes and comments, share of paywalled and sponsored posts, sponsor list. Plus a sponsor leaderboard ranked by how many newsletters and posts each sponsor appears in.

Pricing

$0.003 per post ($3 per 1,000 posts). Posts that fail to load are saved with the error and not charged. Set a maximum cost per run and the Actor stops before exceeding it.

How sponsor detection works

The Actor looks for sponsor phrases ("Brought to you by", "Sponsored by", "Today's sponsor", "Presented by", "In partnership with"…) in headings or bold text, then collects the external links in that section, ignoring Substack, social networks and image hosts. Links from the same company are merged (pages.acme.com + acme.com).

It's accurate for classic sponsor blocks. It can miss sponsors mentioned only in plain prose without a link, and occasionally flag a cross-promotion as a sponsor.

FAQ

Does it need a login or cookies? No. It uses Substack's public endpoints, the same ones the website uses.

Can it get paid posts? It returns metadata for all posts. Full text is only available for free posts (paid posts include the free preview). Sponsor blocks usually sit at the top, so they're detected on paid posts too.

Custom domains? Yes — e.g. www.lennysnewsletter.com works the same as lenny.substack.com.