Substack Scraper - Newsletter Posts, Authors & Engagement avatar

Substack Scraper - Newsletter Posts, Authors & Engagement

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Substack Scraper - Newsletter Posts, Authors & Engagement

Substack Scraper - Newsletter Posts, Authors & Engagement

Scrape Substack newsletter archives. Post titles, authors, publish dates, likes, comments and paywall status as JSON or CSV.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Dominik

Dominik

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Extract the full post archive of any Substack publication, including authors, publish dates, engagement and paywall status.

What you can do with it

  • Content research - see what a publication covers and how often.
  • Competitor monitoring - track a rival newsletter's output and topics.
  • Newsletter discovery - build a dataset of publications in a niche.
  • Engagement analysis - compare comment counts across posts and topics.
  • Feed an AI agent or RAG pipeline - clean post metadata and descriptions.

Input

Every field is optional. Run it with no configuration and it returns useful data straight away.

FieldTypeDescription
publicationsarraySubstack URLs, e.g. https://newsletter.pragmaticengineer.com. Leave empty to use a default sample.
maxItemsintegerHow many posts to return across all publications. Default 200.
postedWithinDaysintegerOnly posts published in the last N days.

Output

Flat JSON records, exportable as JSON, CSV, Excel or XML. Fields the source does not publish come back as null rather than being omitted, so your schema stays stable between runs.

Pricing

Pay per result. You are charged only for records actually delivered - a run that matches nothing costs nothing.

Keeping it fresh

Use Apify Schedules to run this on a timer and your dataset stays current without you touching it. Add a webhook to push new records straight into your database, Google Sheet or Slack channel.

Reliability

This reads the source's own public API - no logins, no bot circumvention. That is a deliberate choice: actors that depend on defeating bot protection break without warning. This one does not.