Substack Email Scraper avatar

Substack Email Scraper

Pricing

from $2.49 / 1,000 results

Go to Apify Store
Substack Email Scraper

Substack Email Scraper

Substack Email Scraper SD - Substack Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Substack results by keyword, location and email domain - Substack email extractor.

Pricing

from $2.49 / 1,000 results

Rating

0.0

(0)

Developer

Leads Scraper

Leads Scraper

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

14 days ago

Last modified

Categories

Share

Substack Email Scraper

The Substack Email Scraper extracts publicly indexed contact emails from Substack newsletters, publication pages and posts, and returns them as a structured dataset ready for JSON or CSV export.

Substack is where independent writing became a business: macro and finance analysts, political and culture commentators, local news operators, fiction serialists, health and fitness coaches, niche B2B newsletters with tiny lists and enormous influence.

Almost all of them want to hear from sponsors. The Substack Email Scraper finds the address they published to make that happen — building keyword-driven Google searches against substack.com, parsing each result block structurally, and keeping only emails on the domains you specify.

This is contact discovery through public search results. The Substack Email Scraper does not log in to Substack, does not use a Substack API, and does not open the Substack website — no browser, no JavaScript rendering, no authentication, no cookies.

Who needs a Substack newsletter email finder

Media buyers and newsletter ad networks, sponsorship teams, PR and book publicists, cross-promotion partners, recruiters hiring writers, and researchers mapping the independent media landscape.

Newsletter sponsorship is a manual, relationship-driven market. The Substack Email Scraper removes the part that was never worth doing by hand.


Key Features of the Substack Email Scraper

Everything below is genuinely implemented in the Substack Email Scraper. No roadmap items.

FeatureWhat it does
Google SERP data collectionBuilds site:substack.com queries and fetches result pages through the Apify GOOGLE_SERP proxy
Subdomain coverageBecause publications live at name.substack.com, the site: operator reaches individual newsletters, not just the root domain
Query expansionBase, quoted and intitle: variants plus one variant per modifier, to get past Google's per-query result ceiling
Domain filteringOnly emails ending in your listed domains survive (@gmail.com, @yahoo.com by default)
Global deduplicationOne row per unique email across every query, page and keyword in the run
Email normalisationUnderstands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @
Junk filterRejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals
Boundary-correct matching@gmail.com never matches inside @gmail.company or @gmail.com.br
Soft-wrap repairDiscards a hit that is only the tail of another email in the same result block
Concurrency controlAn asyncio worker pool runs queries in parallel with a shared stop signal on maxEmails
Retry logicUp to 3 attempts per page with exponential backoff and a fresh proxy session per request
Block detectionCAPTCHA, "unusual traffic" and consent pages are detected and retried, never counted as empty
Resumable stateProgress persists in the key-value store, saved on PERSIST_STATE, MIGRATING and ABORTING
Structural parserLocates the <h3> title then the smallest surrounding block, rather than trusting Google's CSS class names
Whole-page fallbackA markup change degrades the crawler to "emails without account details", never to "no emails"
Streaming dataset writesEvery lead is pushed to the Apify dataset the moment it is found

How the Substack Email Scraper Works

The Substack Email Scraper pipeline has six stages and no hidden state.

  1. Read input. Keywords, optional location, email domains, limits and concurrency.
  2. Build queries. Each keyword pairs with each domain into a site:substack.com search, e.g. site:substack.com fintech newsletter "@gmail.com" "London".
  3. Expand queries. With expandQueries on, quoted and intitle: phrasings also run, plus one variant per modifier. Base queries always go first.
  4. Fetch result pages. Asynchronous aiohttp requests through the Apify GOOGLE_SERP proxy, with pagination up to maxPagesPerQuery.
  5. Parse blocks. The parser reads the result title, the Substack site label and the post snippet as metadata for each candidate lead.
  6. Extract, deduplicate, store. A domain-filtered regex pulls addresses from the block text; new ones stream into the dataset, duplicates are dropped.

Everything the Substack Email Scraper returns was already visible to anyone running the same search manually.


What Data Does the Substack Email Scraper Extract

Each item carries the same 14 fields, so the output schema of the Substack Email Scraper never shifts between runs.

FieldMeaning
networkPlatform name
keywordThe keyword that produced the lead
queryThe exact Google query used
titleRaw result title
accountNameAccount label Google prints
fullNameDisplay name parsed from a profile-style title; empty for post snippets
usernameURL-safe publication handle when one is exposed; otherwise null
profileUrlCanonical newsletter URL when a handle is known; otherwise empty
urlDirect Substack link when exposed, else the profile URL
descriptionBio or post snippet, cleaned of labels and engagement counters
emailLower-cased email address
emailDomainThe matched domain, e.g. @gmail.com
possiblyTruncatedtrue when Google's snippet ellipsis touched the email
foundAtISO 8601 UTC timestamp

Substack publications are subdomains, so a populated username produces a profileUrl of the form https://handle.substack.com/.


Input Fields

Nine fields drive the Substack Email Scraper. Only keywords is required.

FieldTypeDefaultDescription
keywordsarray (required)["newsletter", "writer"]Search terms: niche, job title, industry
locationstring""Optional location phrase added to every query
customDomainsarray["@gmail.com", "@yahoo.com"]Only emails on these domains are kept; the @ is optional
maxEmailsinteger 1–1000020Stop after this many unique emails
countryCodestring""Two-letter country for the search proxy (US, GB, DE…)
expandQueriesbooleantrueSearch each keyword × domain pair in several phrasings
queryModifiersarray["email", "contact", "sponsor", "advertise", "partnership"]Extra words combined with each keyword when expansion is on
maxPagesPerQueryinteger 1–5030Pagination cap per query
maxConcurrencyinteger 1–205Parallel queries

JSON input example

{
"keywords": ["fintech newsletter", "climate newsletter", "local news"],
"location": "London",
"customDomains": ["@gmail.com", "@protonmail.com"],
"maxEmails": 300,
"countryCode": "GB",
"expandQueries": true,
"queryModifiers": ["email", "contact", "sponsor", "advertise", "partnership"],
"maxPagesPerQuery": 30,
"maxConcurrency": 5
}

Output Schema and Dataset Export

The Substack Email Scraper writes leads to the Apify dataset as they are discovered, so results accumulate visibly during the run.

JSON output example

{
"network": "Substack",
"keyword": "fintech newsletter",
"query": "site:substack.com fintech newsletter \"@gmail.com\" \"London\"",
"title": "Macro Margins | Daniel Okereke | Substack",
"accountName": "macromargins",
"fullName": "Daniel Okereke",
"username": "macromargins",
"profileUrl": "https://macromargins.substack.com/",
"url": "https://macromargins.substack.com/",
"description": "Weekly fintech and payments analysis from London. Sponsorship and partnership enquiries: macromargins.ads@gmail.com",
"email": "macromargins.ads@gmail.com",
"emailDomain": "@gmail.com",
"possiblyTruncated": false,
"foundAt": "2026-08-31T13:05:41.118Z"
}

Export the Substack Email Scraper dataset as JSON, CSV, XLSX or XML from the Apify console, or read it over the API into your CRM.


How to Run It

  1. Open the Substack Email Scraper on Apify and click Try for free.
  2. Replace the default keywords with the newsletter category you are targeting — crypto newsletter, parenting newsletter, B2B SaaS newsletter, film criticism.
  3. Set the email domains you want. A narrow list produces a cleaner dataset than a broad one.
  4. Optionally add a location and a countryCode — essential for local news and regional newsletters.
  5. Choose maxEmails and start the run.
  6. Follow the live log: it reports pages fetched, blocked pages, retries and emails per page.
  7. Export the dataset, or chain the Substack Email Scraper into a workflow through the Apify API.

Tips for a better yield

Feed the Substack Email Scraper a topic plus the word "newsletter". writer is broad; supply chain newsletter is a query that converts.

Leave query expansion on. The sponsor, advertise and partnership modifiers are tuned for exactly how Substack writers phrase their ad pages.

If a niche runs dry, change customDomains before you change the keyword — many publishers run a dedicated ads inbox on a different provider.


Use Cases for the Substack Email Scraper

Use caseHow the Substack Email Scraper helps
Newsletter sponsorship buyingBuild a category-specific list of publishers open to advertising
Ad network supply sourcingRecruit independent newsletters into a network or marketplace
PR and book publicityPitch reviews and interviews to writers covering your subject
Cross-promotion and swapsFind newsletters of comparable size in an adjacent niche
Recruiting writersSource proven independent writers for contract or staff roles
Media and market researchMap who publishes in a vertical and how they describe their audience
CRM enrichmentAppend newsletter URLs and bio snippets to existing contact records
Podcast and event bookingInvite domain experts who already publish regularly

Why Choose This Actor

Substack has no exportable publisher directory with contact details, and its own discovery pages surface recommendations rather than emails.

The Substack Email Scraper automates the tedious half of newsletter lead generation — searching, paginating, copying, deduplicating — and leaves the qualification to you.

Because the Substack Email Scraper parses structurally instead of by CSS class name, and because a whole-page fallback exists, a Google layout change degrades result quality rather than breaking the run.

It also complements its siblings. Independent writers rarely stay on one platform: pair the Substack Email Scraper with the Medium Email Scraper for the ones still publishing there, and with the Patreon Email Scraper for creators monetising through memberships instead.

For the audience layer around a newsletter, the X Email Scraper covers where most Substack writers promote, the Reddit Email Scraper reaches the communities that share their posts, and the LinkedIn Email Scraper is stronger for B2B newsletter operators.


Limitations

These constraints on the Substack Email Scraper are real. Read them before your first run.

  • Only publicly indexed emails. The Substack Email Scraper finds addresses Google has already crawled. A publisher who never published an email will not appear.
  • Google's result cap. A single query returns roughly 300 results at most. That is precisely why query expansion exists — leave it enabled.
  • possiblyTruncated. When Google's snippet ellipsis touches an address, the row is flagged true. Verify those before outreach.
  • Apify GOOGLE_SERP proxy required. The Substack Email Scraper cannot run without Apify proxy credentials.
  • Free-plan cap. Free Apify plans are limited to 100 emails per run; paid plans are uncapped.
  • username and profileUrl availability. These are populated only when Substack exposes a publication handle in Google's result. Individual post results sometimes print only a display name, so those rows keep accountName and fullName but have an empty username and profileUrl. That is a Google limitation, not a bug.
  • Variable yield. Output depends on keywords, domains and location. No volume is guaranteed.

Example: a full run

A typical Substack Email Scraper run with three keywords, two domains, expansion on and maxEmails: 300.

The Substack Email Scraper issues base queries first (site:substack.com climate newsletter "@gmail.com"), then quoted and intitle: variants, then one variant per modifier — email, contact, sponsor, advertise, partnership.

Five queries run concurrently, each paginating until it runs dry or reaches maxPagesPerQuery. Addresses are normalised, domain-filtered, deduplicated globally and streamed to the dataset until the 300th unique email trips the shared stop signal.

Blocked or failed queries are re-queued once at the end of the run, so a transient CAPTCHA does not quietly cost you an entire keyword.


The Substack Email Scraper belongs to a family of contact-discovery Actors that share this engine and output schema.

ActorWhat it collects
Substack Email and Phone Number ScraperEmails and phone numbers from Substack
Substack Phone Number ScraperPublic phone numbers from Substack
Behance Email ScraperPublic contact emails from Behance
Bigo Live Email ScraperPublic contact emails from Bigo Live
Bluesky Email ScraperPublic contact emails from Bluesky
Bumble Email ScraperPublic contact emails from Bumble
Clubhouse Email ScraperPublic contact emails from Clubhouse
Dailymotion Email ScraperPublic contact emails from Dailymotion
DeviantArt Email ScraperPublic contact emails from DeviantArt
Discord Email ScraperPublic contact emails from Discord
Dribbble Email ScraperPublic contact emails from Dribbble
Facebook Email ScraperPublic contact emails from Facebook
Goodreads Email ScraperPublic contact emails from Goodreads
Hinge Email ScraperPublic contact emails from Hinge
Instagram Email ScraperPublic contact emails from Instagram
KakaoTalk Email ScraperPublic contact emails from KakaoTalk
Kick Email ScraperPublic contact emails from Kick
Lemon8 Email ScraperPublic contact emails from Lemon8
Likee Email ScraperPublic contact emails from Likee
LINE Email ScraperPublic contact emails from LINE
LinkedIn Email ScraperPublic contact emails from LinkedIn
Mastodon Email ScraperPublic contact emails from Mastodon
Medium Email ScraperPublic contact emails from Medium
Mixcloud Email ScraperPublic contact emails from Mixcloud
Patreon Email ScraperPublic contact emails from Patreon
Pinterest Email ScraperPublic contact emails from Pinterest
Quora Email ScraperPublic contact emails from Quora
Reddit Email ScraperPublic contact emails from Reddit
Rumble Email ScraperPublic contact emails from Rumble
Snapchat Email ScraperPublic contact emails from Snapchat
SoundCloud Email ScraperPublic contact emails from SoundCloud
Telegram Email ScraperPublic contact emails from Telegram
Threads Email ScraperPublic contact emails from Threads
TikTok Email ScraperPublic contact emails from TikTok
Tinder Email ScraperPublic contact emails from Tinder
Tumblr Email ScraperPublic contact emails from Tumblr
Twitch Email ScraperPublic contact emails from Twitch
Vimeo Email ScraperPublic contact emails from Vimeo
VK Email ScraperPublic contact emails from VK
WeChat Email ScraperPublic contact emails from WeChat
Weibo Email ScraperPublic contact emails from Weibo
X Email ScraperPublic contact emails from X

FAQ

What is the Substack Email Scraper?

An Apify Actor that extracts publicly indexed contact emails from Substack newsletters, publication pages and posts via Google search results, and writes them to a structured dataset.

Does the Substack Email Scraper log in to Substack?

No. The Substack Email Scraper does not log in, does not call a Substack API, and does not open substack.com. It reads Google search results only.

Does it find subscriber emails?

No, and it never could. The Substack Email Scraper only sees addresses a publisher chose to publish on a public page — typically a sponsorship or contact inbox.

Where do the emails come from?

Titles, snippets and site labels in Google's public index, usually from an about page, an advertise page, or the footer of a post.

Can I choose which email providers are returned?

Yes. customDomains filters the Substack Email Scraper's extraction step. Defaults are @gmail.com and @yahoo.com, and the leading @ is optional.

Why is username empty on some rows?

Google sometimes prints only a display name for an individual post. Where no publication handle is exposed, username is null and profileUrl is empty, while accountName and fullName remain populated.

What does possiblyTruncated: true mean?

Google's snippet ellipsis touched the address, so it may be cut off. Treat those rows as unverified before sending anything.

How many Substack newsletter emails will one run return?

It depends on keywords, domains and location — no volume is guaranteed. Free Apify plans stop at 100 emails per run; paid plans are uncapped.

Do I need a proxy?

Yes. The Substack Email Scraper requires the Apify GOOGLE_SERP proxy and will not run without Apify proxy credentials.

Should I leave query expansion on?

Almost always. Google caps a single query near 300 results, and expansion is how the Substack Email Scraper works around that ceiling.

Does the output schema ever change?

No. Every item the Substack Email Scraper writes carries the same 14 fields, which keeps downstream mapping stable.

Can I export to CSV or schedule recurring runs?

Yes. Export as JSON, CSV, XLSX or XML, and use Apify Schedules to run the Substack Email Scraper on a recurring basis into the same dataset.


Leave a review

If the Substack Email Scraper saved you time, please leave a star rating and a short review on the Actor page.

Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.

If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.

Support

Questions about the Substack Email Scraper, feature requests, or need a custom build? Email neurodata.apify@gmail.com.

Not affiliated with, endorsed by, or officially connected to Substack. Use collected data in line with applicable privacy and anti-spam law.