Substack Scraper — Newsletter Posts & Authors API | $0.50/1k avatar

Substack Scraper — Newsletter Posts & Authors API | $0.50/1k

Pricing

from $0.50 / 1,000 results

Go to Apify Store
Substack Scraper — Newsletter Posts & Authors API | $0.50/1k

Substack Scraper — Newsletter Posts & Authors API | $0.50/1k

Scrape any Substack newsletter, custom domains too: posts with title, authors, date, free/paid audience, word count, reactions, comment count and full text of free posts, newsletter profiles and keyword search. Filters by date, audience and keyword. No login, no API key.

Pricing

from $0.50 / 1,000 results

Rating

0.0

(0)

Developer

Raffy

Raffy

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

21 hours ago

Last modified

Share

What does Substack Scraper do?

Substack Scraper turns Substack newsletters into structured data: every post with its title, subtitle, author names, publish date, free or paid audience, word count, reading time, reactions, comment count, tags, cover image and the post text, exported as JSON, CSV or Excel. Enter newsletters as a subdomain (lenny), a Substack address (https://simonw.substack.com) or the newsletter's own domain (https://www.lennysnewsletter.com), paste single post URLs, or find newsletters by keyword. No Substack account, no login, no API key and no browser.

It works as a Substack API for newsletters you do not own. Every Substack page loads its data from public JSON endpoints; this Actor reads those endpoints page by page, follows newsletters that moved to a custom domain, and returns one tidy row per post (or per newsletter).

Typical use cases:

  • Content and competitor research: pull the full archive of the newsletters in your niche and see what they publish, how often, and which posts get the most reactions and comments.
  • Newsletter discovery and outreach lists: search Substack by keyword and get newsletters with their description, public authors, subscriber band and first post date.
  • Monitoring: schedule a run with Posted on or after 1 day to collect every new post of a list of newsletters each morning.
  • Datasets for AI and analytics: plain-text post bodies of free posts, ready for summarising, classification or RAG, from a Substack scraper an AI agent can call with just a newsletter name.

It reads only what Substack shows to a logged-out visitor. Paid posts come with the free preview only (isPreviewOnly: true); the Actor never tries to get around a paywall. It does not collect comments, commenters, subscriber lists or e-mail addresses.

Why use Substack Scraper?

  • Posts, newsletters and search in one Actor. Archive scraping, single post URLs, a one-row-per-newsletter mode and publication search by keyword.
  • Custom domains just work. www.lennysnewsletter.com, www.thefp.com or a subdomain that redirects to one: the same Substack data is read from all of them.
  • Honest paywall handling. audience says free or paid, wordCount is the full length Substack states, and isPreviewOnly tells you when bodyText is only the public preview.
  • Filters that save money. Free posts only, posted after / before, a keyword (Substack's own archive search), newest or most popular first, and a per-newsletter limit. Posts your filters drop are never written or billed.
  • Honest results. Every row has a status. A site that is not a Substack newsletter, an unknown post or a search with no match gives one free not_found row that says why; you are billed only for ok rows.
  • Cheap. $0.50 per 1,000 posts, Apify platform usage included.

What data can Substack Scraper extract?

One row per post (default):

FieldTypeDescription
recordTypestringpost (or publication for newsletter rows)
titlestringPost title
subtitlestringPost subtitle
publicationNamestringNewsletter name
publicationUrlstringNewsletter home page (custom domain or <name>.substack.com)
publicationSubdomainstringNewsletter subdomain on substack.com
authorsarrayNames on the post's public byline
publishedAtstringPublish time (ISO 8601)
updatedAtstringLast update time (with Include post text)
audiencestringfree (everyone) or paid (paying subscribers only)
isPreviewOnlybooleantrue for paid posts: bodyText holds only the free preview
wordCountintegerLength of the whole post as Substack states it
readingTimeMinutesintegerEstimated reading time (250 words per minute)
reactionsintegerReactions (likes)
commentCountintegerNumber of comments (a count only)
restacksintegerNumber of restacks
tagsarrayPost tags
sectionstringNewsletter section, if the post belongs to one
postTypestringnewsletter, podcast, thread, ...
coverImagestringCover image URL
descriptionstringSearch / SEO description of the post
previewTextstringShort opening text from the archive list
bodyTextstringPost body as plain text (paid posts: the free preview only)
bodyHtmlstringPost body as HTML (with Include post HTML)
bodyErrorstringWhy the full post could not be read, when the row carries only the archive data
postId, slug, publicationIdinteger / stringSubstack's own ids
urlstringCanonical post URL
statusstringok, not_found or error (see below)
errorstringReason when status is not ok
scrapedAtstringISO 8601 time of extraction

Newsletter rows (What to return = publication, and every Find newsletters by keyword result) carry name, url, subdomain, customDomain, description (the tagline), authors (the newsletter's owner, its public author), logo, subscriberText (e.g. "Hundreds of thousands of subscribers"), paidSubscriberText, subscriberCount (when the newsletter shows it publicly), hasPaidPlans, language, firstPostAt and, for search results, searchQuery.

Result status (tri-state output)

statusMeaningBilled?
okA post or newsletter row.Yes
not_foundNot a Substack newsletter, unknown newsletter or post, a newsletter with no public posts, a search with no match, or nothing matched your filters. error says which.No
errorSubstack could not be read after several retries (network error, rate limit). error says why.No

How to scrape Substack newsletters

  1. Open the Actor in Apify Console and click Try for free.
  2. Put newsletters into Newsletters (a name like lenny, a Substack URL or a custom domain), post links into Post URLs, or keywords into Find newsletters by keyword.
  3. Optionally set Maximum posts per newsletter, Posted on or after, Audience or Only posts about.
  4. Click Start. The default run (10 posts from each of two newsletters, with full text) takes about 15 seconds.
  5. Open the Output tab or Export the dataset as JSON, CSV, Excel, XML or HTML.

To automate it, use the API tab (Node.js, Python, curl examples) or add a Schedule.

How much does it cost to scrape Substack?

This Actor uses pay-per-event pricing. Apify platform usage is included in these prices.

EventPrice
Actor start$0.005 per run
Result (status: ok row)$0.0005 per post (or per newsletter row)

Examples: the default run (20 posts) costs $0.015. 1,000 posts cost $0.505, 10,000 posts $5.005. Rows with status not_found or error are never billed, and posts your filters drop are never written or billed. You can cap spending with Maximum results, Maximum posts per newsletter and the run's Max total charge option.

Input

FieldTypeDefaultDescription
startUrlsarray of stringssimonw.substack.com, www.lennysnewsletter.comNewsletters: subdomain, Substack URL or custom domain (post URLs also work)
postUrlsarray of strings-Single posts to read
searchQueriesarray of strings-Find newsletters by keyword (one row per newsletter)
modeposts or publicationpostsOne row per post, or one row per newsletter
maxItemsinteger20Stop after this many rows in total
maxPostsPerPublicationinteger- (prefill 10)Stop reading a newsletter after this many posts
maxPublicationsPerSearchinteger20Newsletters per search term
fetchFullPostbooleantrueAdd bodyText and updatedAt (one extra request per post)
includeHtmlbooleanfalseAlso add bodyHtml
audienceall or freeallfree skips subscriber-only posts
postedAfter, postedBeforestring-YYYY-MM-DD or relative such as 7 days
postKeywordstring-Only posts Substack's archive search matches
sortnew or topnewNewest or most popular posts first
proxyConfigurationobjectApify Proxy (datacenter)Spreads requests; residential proxies are not needed

Example: the free posts of two newsletters published in August 2026, without the post text:

{
"startUrls": ["https://www.lennysnewsletter.com", "simonw"],
"audience": "free",
"postedAfter": "2026-08-01",
"postedBefore": "2026-08-31",
"fetchFullPost": false,
"maxItems": 200
}

Example: find newsletters about history:

{
"searchQueries": ["history"],
"maxPublicationsPerSearch": 50,
"maxItems": 50
}

Output

Real rows from runs on Apify (the default input, and publication mode with one unknown newsletter; long texts shortened with ...):

[
{
"url": "https://simonw.substack.com/p/navierstokes-rubygems-attacked-gis",
"status": "ok",
"scrapedAt": "2026-09-30T13:49:12.450Z",
"recordType": "post",
"postId": 215727411,
"slug": "navierstokes-rubygems-attacked-gis",
"title": "Navier–Stokes, RubyGems attacked, GIS and Blender with GPT-6 Astra",
"subtitle": "Plus a Datasette security release, and a Pluribus Fabergé egg",
"description": "Plus a Datasette security release, and a Pluribus Fabergé egg",
"authors": [
"Simon Willison"
],
"publishedAt": "2026-09-14T20:47:59.989Z",
"audience": "free",
"isPreviewOnly": false,
"wordCount": 5230,
"readingTimeMinutes": 21,
"reactions": 54,
"commentCount": 7,
"restacks": 7,
"postType": "newsletter",
"tags": [
"openai",
"llms",
"ai"
],
"coverImage": "https://substackcdn.com/image/fetch/$s_!XTAQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c99ef2b-8373-4ac4-b95b-3fce518d3733_1327x934.webp",
"previewText": "In this newsletter:",
"publicationName": "Simon Willison’s Newsletter",
"publicationSubdomain": "simonw",
"publicationUrl": "https://simonw.substack.com",
"publicationId": 1173386,
"updatedAt": "2026-09-14T20:51:58.963Z",
"bodyText": "In this newsletter:\nSome thoughts on the Navier–Stokes Millennium Prize Problem\nGenerating running routes with GPT-6 Astra and ChatGPT Work\nOpenAI agents attacked RubyGems back in May\nPlus 7 links, 7 quotations, 1 note, ..."
},
{
"url": "https://www.lennysnewsletter.com/p/60-creative-growth-ideas",
"status": "ok",
"scrapedAt": "2026-09-30T13:49:14.559Z",
"recordType": "post",
"postId": 214360366,
"slug": "60-creative-growth-ideas",
"title": "60+ new creative growth ideas",
"subtitle": "How to stand out when everyone is running the same playbook",
"description": "How to stand out when everyone is running the same playbook",
"authors": [
"Tom Orbach"
],
"publishedAt": "2026-09-15T12:45:28.445Z",
"audience": "paid",
"isPreviewOnly": true,
"wordCount": 5674,
"readingTimeMinutes": 23,
"reactions": 414,
"commentCount": 6,
"restacks": 34,
"postType": "newsletter",
"tags": [
"Growth"
],
"coverImage": "https://substackcdn.com/image/fetch/$s_!NVVg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20cbbe62-5c82-4d4f-a316-de9b885e56f2_1456x970.png",
"previewText": "👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs ...",
"publicationName": "Lenny's Newsletter",
"publicationSubdomain": "lenny",
"publicationUrl": "https://www.lennysnewsletter.com",
"publicationId": 10845,
"updatedAt": "2026-09-15T12:50:53.552Z",
"bodyText": "👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs | Lenny’s Podcast | Lennybot | How I AI | Become an AI-Native Builder and my other favorite AI/PM co ..."
},
{
"url": "https://www.thefp.com",
"status": "ok",
"scrapedAt": "2026-09-30T13:49:00.629Z",
"recordType": "publication",
"name": "The Free Press",
"publicationName": "The Free Press",
"publicationUrl": "https://www.thefp.com",
"publicationSubdomain": "bariweiss",
"publicationId": 260347,
"subdomain": "bariweiss",
"customDomain": "www.thefp.com",
"description": "A new media company built on the ideals that were once the bedrock of American journalism.",
"authors": [
"Bari Weiss"
],
"logo": "https://substackcdn.com/image/fetch/$s_!XTc7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F9cb7f208-a15c-46a8-a040-7e7a2150def9_1280x1280.png",
"hasPaidPlans": true,
"language": "en",
"firstPostAt": "2021-01-12T20:43:07.912Z"
},
{
"url": "https://nosuchpubxyz123456.substack.com/",
"status": "not_found",
"error": "HTTP 404: page does not exist",
"scrapedAt": "2026-09-30T13:49:00.035Z"
}
]

The paid post has audience: paid and isPreviewOnly: true: its bodyText is the free preview Substack shows to every visitor, while wordCount is the length of the whole post.

Tips

  • Put many newsletters in one run: you pay the start fee once, and newsletters are read in parallel.
  • Use Maximum posts per newsletter so one long archive cannot use up the run.
  • For daily monitoring, schedule the run with Posted on or after 1 day: the archive is newest-first, so the scan stops as soon as it reaches older posts.
  • Turn Include post text off when you only need titles, dates and engagement numbers: one request per 20 posts instead of one per post.
  • To get the posts of newsletters you found with a search, copy their url values into Newsletters.

Limitations

  • Paid posts: only the free preview that Substack shows to logged-out visitors (isPreviewOnly: true). The Actor does not log in and never tries to get around a paywall.
  • Comments, commenters, likers, subscriber lists and e-mail addresses are not collected; commentCount is a number only.
  • Profile pages (substack.com/@name) and Substack Notes are not newsletters and are not supported; enter the newsletter's address instead.
  • A keyword (Only posts about) uses Substack's own archive search, which also matches words inside the post, and returns matches in its relevance order.
  • subscriberCount is filled only when Substack shows the number publicly; otherwise subscriberText carries Substack's public band ("Thousands of subscribers").

FAQ

Do I need a Substack account or API key?

No. The Actor reads the public data that Substack serves to every logged-out visitor.

Does it work with custom domains?

Yes. Enter the newsletter's own domain (https://www.thefp.com) or its Substack subdomain; a subdomain that redirects to a custom domain is followed automatically.

Can I get the full text of paid posts?

No. For paid posts you get the title, dates, author names, word count, engagement numbers and the free preview, marked with isPreviewOnly: true.

How do I find newsletters about a topic?

Put the topic into Find newsletters by keyword. Each result is one newsletter row with its name, URL, description, author and subscriber band.

Can I use this Actor from an AI agent or MCP client?

Yes. It runs with no input at all (it then reads two example newsletters), accepts plain newsletter names, and every row explains itself with status and error.

Why did I get fewer rows than maxItems?

The newsletters have fewer matching posts, a per-newsletter limit stopped them, or your run hit its Max total charge. A not_found row says when a filter matched nothing.

  • Google News Scraper - use it to follow the same topics across news sites, not only newsletters.
  • Remote Jobs Scraper - use it to collect remote job listings for the companies and industries you follow.

This Actor extracts only data that Substack publishes publicly for logged-out visitors: posts, their public byline (author names) and newsletter profiles. It does not extract comments, commenter or subscriber data, e-mail addresses or phone numbers, and it does not log in or get around paywalls or other access controls. Post texts are the authors' copyrighted work: use them for analysis, research or indexing, and do not republish them without the author's permission. Personal data is protected by the GDPR in the European Union and by other regulations around the world; do not use the output to process personal data without a legitimate reason. You are responsible for complying with Substack's terms of service and applicable law when using the extracted data.

This Actor is an independent tool and is not affiliated with, endorsed by or sponsored by Substack Inc. Substack is a trademark of its owner.