Substack Scraper — Posts, Paywall & Bestseller Signals avatar

Substack Scraper — Posts, Paywall & Bestseller Signals

Pricing

from $1.31 / 1,000 post scrapeds

Go to Apify Store
Substack Scraper — Posts, Paywall & Bestseller Signals

Substack Scraper — Posts, Paywall & Bestseller Signals

Extract posts, paywall status, engagement and bestseller signals from any public Substack publication. Pure HTTP via the public archive API — no login, no API key, no browser, no residential proxy. See which competitors monetize and how much sits behind the paywall.

Pricing

from $1.31 / 1,000 post scrapeds

Rating

0.0

(0)

Developer

viralanalyzer

viralanalyzer

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

Extract posts, paywall status, engagement and bestseller signals from any public Substack publication. Pure HTTP against the public archive API — no login, no API key, no headless browser, no residential proxy.

Most Substack scrapers tell you what was published. This one tells you what is behind the paywall — which is the question that decides whether a newsletter is a business or a hobby.

🇧🇷 Versão em português no final deste README.


Features

  • Paywall intelligence — every post carries audience (everyone / only_paid) and a derived isPaywalled, plus a per-publication paywallRatio.
  • Monetization signalspaymentsEnabled, foundingPlanName and the author's bestsellerTier straight from the source.
  • Engagement — reactions, restacks, comments and word count per post.
  • Works on custom domains — pass bigtechnology or https://www.bigtechnology.com; both resolve.
  • Audience filter — keep only free posts, only paywalled posts, or everything.
  • Publication summary — one extra row per publication with paywall ratio, average reactions, average word count and top bestseller tier.
  • Honest failures — when the source blocks us or answers with an empty archive, you get a labeled diagnostic and you are not charged. Schema drift still fails loudly instead of returning silence.

Input

FieldTypeDefaultDescription
publicationsarrayHandles (bigtechnology) or URLs (https://bigtechnology.substack.com, custom domains work). 1–50 items.
maxPostsPerPublicationinteger50Max posts per publication (1–1000). Paginated in pages of 50.
audienceFilterselectallall · free · paid
includePublicationSummarybooleantrueAdds one summary row per publication.
proxyConfigurationobjectdatacenterOptional. The archive API is public — residential is not needed.

Output

Two item types share one dataset, distinguished by itemType.

itemType: "post"

publication · publicationName · customDomain · postId · title · subtitle · slug · url · postDate · postType · audience · isPaywalled · wordCount · reactionCount · restacks · commentCount · language · coverImage · description · hasAudio · podcastDuration · authors[] (name, handle, bio, bestsellerTier, twitter) · paymentsEnabled · foundingPlanName · scrapedAt

itemType: "publication_summary"

totalPosts · paidPosts · paywallRatio · avgReactions · avgWordCount · maxBestsellerTier · paymentsEnabled · foundingPlanName

Use cases

  • Competitive newsletter analysis — how much of a rival's output is paywalled, and what the free posts look like.
  • Creator economy research — find publications with payments enabled and a high bestseller tier.
  • Content strategy — correlate word count and format with reactions and restacks.
  • Partnership and sponsorship sourcing — rank publications by engagement, not by follower vanity.
  • Media monitoring — track a set of newsletters on a schedule and diff the archive.

✅ Capabilities & Limits

Stated up front, so you do not pay a run to find out.

Built here: the monetization fields - publicationName, customDomain, paymentsEnabled, foundingPlanName - are not fetched from a publication endpoint. They are mined from the first post of the first archive page, walking publishedBylines[].publicationUsers[].publication. If that newest post carries no byline with an attached publication object, all four come back null for the whole publication, even though the post data itself is complete.

Input / featureSupportedNotes
publicationsrequiredSubstack publication handles (e.g. 'bigtechnology') or full URLs (e.g. 'https://bigtechnology.substack.com' or a custom domain like…
Result volume (maxPostsPerPublication)⚠️Maximum number of posts to extract per publication. The archive API is paginated in pages of up to 50.
audienceFilterKeep only free posts, only paywalled posts, or everything. Useful to measure how much of a competitor's output sits behind the paywall.
includePublicationSummaryPush one extra summary item per publication with paywall ratio, average engagement, average word count and bestseller tier.
Proxy⚠️Optional. The Substack archive API is public and works on a direct connection — datacenter proxy is only useful to spread requests across IPs on large…

FAQ

Do I need a Substack account or API key? No. The archive endpoint is public.

Does it read paywalled article text? No. It reports that a post is paywalled and its metadata. Paid body text is not extracted — that content is not public.

Does it work with custom domains? Yes. https://www.bigtechnology.com resolves the same as bigtechnology.substack.com.

Do I need a residential proxy? No. That is exactly what makes this Actor cheap to run.

What happens when a publication does not exist? You get a SUCCEEDED run with one diagnostic item explaining why, and no charge.

How fast is it? One HTTP request returns up to 50 posts. No browser to boot.

Pricing

$0.0025 per post scraped, plus the Apify platform usage of the run, billed to you at Apify's standard rates and shown on the run page.

Pay-per-event on posts actually delivered. Diagnostic and summary rows are not charged. Platform usage (compute) is billed to your account at Apify cost — this Actor is pure HTTP, so it is a fraction of a cent per run. See the pricing panel on this page for the current rate and tiers.

  • viralanalyzer/telegram-channel-scraper — the same pure-HTTP approach for Telegram channels
  • viralanalyzer/rss-news-intelligence — feed monitoring at scale
  • viralanalyzer/linkedin-ads-library — B2B ad intelligence

Changelog

1.0 — First release. Archive API pagination, audience filter, publication summary, block-gated diagnostics, owner-skip PPE.



🇧🇷 Substack Scraper — Posts, Paywall e Sinais de Bestseller

Extraia posts, status de paywall, engajamento e sinais de bestseller de qualquer publicação pública do Substack. HTTP puro contra a API pública de arquivo — sem login, sem API key, sem navegador headless, sem proxy residencial.

A maioria dos scrapers de Substack diz o que foi publicado. Este diz o que está atrás do paywall — que é a pergunta que separa uma newsletter que é negócio de uma que é hobby.

Funcionalidades

  • Inteligência de paywall — cada post traz audience (everyone / only_paid) e o campo derivado isPaywalled, mais o paywallRatio por publicação.
  • Sinais de monetizaçãopaymentsEnabled, foundingPlanName e o bestsellerTier do autor, direto da fonte.
  • Engajamento — reações, restacks, comentários e contagem de palavras por post.
  • Funciona com domínio próprio — aceita bigtechnology ou https://www.bigtechnology.com.
  • Filtro de audiência — só posts gratuitos, só pagos, ou tudo.
  • Resumo por publicação — uma linha extra com razão de paywall, média de reações e de palavras.
  • Falha honesta — se a fonte bloquear ou devolver arquivo vazio, você recebe um diagnóstico rotulado e não é cobrado. Mudança de schema continua falhando alto, em vez de devolver silêncio.

Entrada

CampoTipoPadrãoDescrição
publicationsarrayHandles (bigtechnology) ou URLs. 1–50 itens.
maxPostsPerPublicationinteiro50Máximo de posts por publicação (1–1000).
audienceFilterseleçãoallall · free · paid
includePublicationSummarybooleanotrueAdiciona uma linha de resumo por publicação.
proxyConfigurationobjetodatacenterOpcional. A API é pública — residencial não é necessário.

Saída

Dois tipos de item no mesmo dataset, distinguidos por itemType: post e publication_summary (campos listados na seção em inglês acima).

Casos de uso

  • Análise competitiva de newsletters — quanto da produção do concorrente está no paywall.
  • Pesquisa de creator economy — achar publicações com pagamentos ativos e bestseller alto.
  • Estratégia de conteúdo — correlacionar tamanho e formato com reações e restacks.
  • Prospecção de patrocínio — ranquear por engajamento real, não por vaidade de seguidores.
  • Monitoramento — acompanhar um conjunto de newsletters em agendamento e comparar o arquivo.

FAQ

Preciso de conta ou API key do Substack? Não. O endpoint de arquivo é público.

Ele lê o texto de posts pagos? Não. Ele informa que o post é pago e os metadados dele. O conteúdo pago não é público e não é extraído.

Funciona com domínio próprio? Sim.

Preciso de proxy residencial? Não — e é justamente isso que torna este Actor barato de rodar.

E se a publicação não existir? O run termina SUCCEEDED com um item de diagnóstico explicando o motivo, e sem cobrança.

Preço

$0.0025 por post extraído, mais o uso de plataforma da Apify da execução, cobrado de você nas tarifas padrão e visível na página da run.

Cobrança por evento, sobre posts efetivamente entregues. Itens de diagnóstico e de resumo não são cobrados. O custo de plataforma é debitado na sua conta ao custo da Apify — como é HTTP puro, é uma fração de centavo por run. Veja o painel de preço nesta página para a tarifa e os tiers atuais.

Changelog

1.0 — Primeira versão. Paginação da API de arquivo, filtro de audiência, resumo por publicação, diagnóstico com gate de bloqueio, owner-skip no PPE.