Substack Publication Scraper avatar

Substack Publication Scraper

Pricing

from $8.25 / 1,000 items

Go to Apify Store
Substack Publication Scraper

Substack Publication Scraper

Pull every public post from any Substack publication with title, subtitle, body preview, author, publish date, podcast URL, audience type, comment count, and reactions. Filter by post type and date range. Export to JSON, CSV, or Excel for newsletter research and competitive intelligence.

Pricing

from $8.25 / 1,000 items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

2 days ago

Last modified

Share

ParseForge Banner

📰 Substack Publication Scraper

🚀 Pull every public post from any Substack publication. Title, body preview, author, podcast, paywall flag, comment count, reactions. No login, no API key, no manual scrolling.

The Substack Publication Scraper queries the public Substack archive endpoints for any publication and returns every post in the feed. Each record includes the post title, social title, subtitle, description, slug, canonical URL, publish date, post type, audience flag, paywall status, cover image, podcast duration, word count, reaction count, comment count, restack count, section info, and a truncated body preview.

Substack hosts millions of newsletters and is the largest creator-operated publishing platform on the internet. Top publications cross hundreds of thousands of paid subscribers and rival traditional media in influence. This Actor exports the full archive of any publication in a single run, letting you research content cadence, audience signals, and editorial mix without a manual subscribe-and-scroll workflow.

🎯 Target Audience💡 Primary Use Cases
Newsletter writers, content marketers, ghost writers, journalists, podcasters, researchersContent research, cadence analysis, audience mining, podcast discovery, competitive benchmarking

📋 What the Substack Publication Scraper does

Five filtering workflows in a single run:

  • 📰 Full archive export. Submit one publication subdomain or custom domain and pull its entire post archive.
  • 📅 Date range filter. Pin to a specific year, quarter, or month using minDate and maxDate.
  • 🎙️ Type filter. Restrict to newsletter, podcast, or thread posts.
  • 💎 Paywall awareness. Each record flags whether the post is everyone (free) or only_paid (subscriber-only).
  • 🔍 Engagement signals. Comment count, reaction count, restack count, and word count surface engagement patterns.

Each row reports the publication slug, post ID, full title and subtitle, slug, canonical URL, publish timestamp, type, audience, cover image URL, podcast duration when present, word count, engagement counters, and a 200-character body preview.

💡 Why it matters: Substack publications are time-machines for content strategy. Cadence, average word count, paywall ratio, and reaction-to-comment ratios all reveal what resonates. Researchers cite Substack archives in studies of opinion journalism. Ghost writers reverse-engineer voice from existing posts. Content marketers benchmark themselves against the best operators in their niche.

📊 Data fields

Each record includes: canonicalUrl, description, postDate, postId, publication, publicationDescription, publicationName, publicationSubscribersFormatted, slug, socialTitle, subtitle, title, type, url. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

🚀 How to use

  1. 🆓 Create a free Apify account. Sign up here and get $5 in free credit.
  2. 🔍 Open the Actor. Search for "Substack Publication" in the Apify Store.
  3. ⚙️ Set the publication. Enter the subdomain or custom domain and any filters.
  4. ▶️ Click Start. A 100-post run finishes in under 15 seconds.
  5. 📥 Download. Export as CSV, Excel, JSON, or XML.

⏱️ Total time from sign-up to first dataset: under five minutes.

💡 Pro Tip: browse the complete ParseForge collection for more pre-built scrapers and data tools.

Substack is a registered trademark of Substack Inc. This Actor is not affiliated with or endorsed by Substack. It reads only publicly accessible archive endpoints and respects per-publication terms of service.

🆘 Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.