Substack Email Scraper
Pricing
from $2.49 / 1,000 results
Substack Email Scraper
Substack Email Scraper SD - Substack Email Scraper is a lead generation tool that extracts leads with public contact emails, account names and profile URLs from Substack results by keyword, location and email domain - Substack email extractor.
Pricing
from $2.49 / 1,000 results
Rating
0.0
(0)
Developer
Leads Scraper
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
14 days ago
Last modified
Categories
Share
Substack Email Scraper
The Substack Email Scraper extracts publicly indexed contact emails from Substack newsletters, publication pages and posts, and returns them as a structured dataset ready for JSON or CSV export.
Substack is where independent writing became a business: macro and finance analysts, political and culture commentators, local news operators, fiction serialists, health and fitness coaches, niche B2B newsletters with tiny lists and enormous influence.
Almost all of them want to hear from sponsors. The Substack Email Scraper finds the address they published to make that happen — building keyword-driven Google searches against substack.com, parsing each result block structurally, and keeping only emails on the domains you specify.
This is contact discovery through public search results. The Substack Email Scraper does not log in to Substack, does not use a Substack API, and does not open the Substack website — no browser, no JavaScript rendering, no authentication, no cookies.
Who needs a Substack newsletter email finder
Media buyers and newsletter ad networks, sponsorship teams, PR and book publicists, cross-promotion partners, recruiters hiring writers, and researchers mapping the independent media landscape.
Newsletter sponsorship is a manual, relationship-driven market. The Substack Email Scraper removes the part that was never worth doing by hand.
Key Features of the Substack Email Scraper
Everything below is genuinely implemented in the Substack Email Scraper. No roadmap items.
| Feature | What it does |
|---|---|
| Google SERP data collection | Builds site:substack.com queries and fetches result pages through the Apify GOOGLE_SERP proxy |
| Subdomain coverage | Because publications live at name.substack.com, the site: operator reaches individual newsletters, not just the root domain |
| Query expansion | Base, quoted and intitle: variants plus one variant per modifier, to get past Google's per-query result ceiling |
| Domain filtering | Only emails ending in your listed domains survive (@gmail.com, @yahoo.com by default) |
| Global deduplication | One row per unique email across every query, page and keyword in the run |
| Email normalisation | Understands name [at] domain [dot] com, name (at) domain, name @ domain.com, domain .com, zero-width characters and the full-width @ |
| Junk filter | Rejects placeholders such as email@, yourname@, test@, xxx@ and single-character locals |
| Boundary-correct matching | @gmail.com never matches inside @gmail.company or @gmail.com.br |
| Soft-wrap repair | Discards a hit that is only the tail of another email in the same result block |
| Concurrency control | An asyncio worker pool runs queries in parallel with a shared stop signal on maxEmails |
| Retry logic | Up to 3 attempts per page with exponential backoff and a fresh proxy session per request |
| Block detection | CAPTCHA, "unusual traffic" and consent pages are detected and retried, never counted as empty |
| Resumable state | Progress persists in the key-value store, saved on PERSIST_STATE, MIGRATING and ABORTING |
| Structural parser | Locates the <h3> title then the smallest surrounding block, rather than trusting Google's CSS class names |
| Whole-page fallback | A markup change degrades the crawler to "emails without account details", never to "no emails" |
| Streaming dataset writes | Every lead is pushed to the Apify dataset the moment it is found |
How the Substack Email Scraper Works
The Substack Email Scraper pipeline has six stages and no hidden state.
- Read input. Keywords, optional location, email domains, limits and concurrency.
- Build queries. Each keyword pairs with each domain into a
site:substack.comsearch, e.g.site:substack.com fintech newsletter "@gmail.com" "London". - Expand queries. With
expandQuerieson, quoted andintitle:phrasings also run, plus one variant per modifier. Base queries always go first. - Fetch result pages. Asynchronous
aiohttprequests through the Apify GOOGLE_SERP proxy, with pagination up tomaxPagesPerQuery. - Parse blocks. The parser reads the result title, the Substack site label and the post snippet as metadata for each candidate lead.
- Extract, deduplicate, store. A domain-filtered regex pulls addresses from the block text; new ones stream into the dataset, duplicates are dropped.
Everything the Substack Email Scraper returns was already visible to anyone running the same search manually.
What Data Does the Substack Email Scraper Extract
Each item carries the same 14 fields, so the output schema of the Substack Email Scraper never shifts between runs.
| Field | Meaning |
|---|---|
network | Platform name |
keyword | The keyword that produced the lead |
query | The exact Google query used |
title | Raw result title |
accountName | Account label Google prints |
fullName | Display name parsed from a profile-style title; empty for post snippets |
username | URL-safe publication handle when one is exposed; otherwise null |
profileUrl | Canonical newsletter URL when a handle is known; otherwise empty |
url | Direct Substack link when exposed, else the profile URL |
description | Bio or post snippet, cleaned of labels and engagement counters |
email | Lower-cased email address |
emailDomain | The matched domain, e.g. @gmail.com |
possiblyTruncated | true when Google's snippet ellipsis touched the email |
foundAt | ISO 8601 UTC timestamp |
Substack publications are subdomains, so a populated username produces a profileUrl of the form https://handle.substack.com/.
Input Fields
Nine fields drive the Substack Email Scraper. Only keywords is required.
| Field | Type | Default | Description |
|---|---|---|---|
keywords | array (required) | ["newsletter", "writer"] | Search terms: niche, job title, industry |
location | string | "" | Optional location phrase added to every query |
customDomains | array | ["@gmail.com", "@yahoo.com"] | Only emails on these domains are kept; the @ is optional |
maxEmails | integer 1–10000 | 20 | Stop after this many unique emails |
countryCode | string | "" | Two-letter country for the search proxy (US, GB, DE…) |
expandQueries | boolean | true | Search each keyword × domain pair in several phrasings |
queryModifiers | array | ["email", "contact", "sponsor", "advertise", "partnership"] | Extra words combined with each keyword when expansion is on |
maxPagesPerQuery | integer 1–50 | 30 | Pagination cap per query |
maxConcurrency | integer 1–20 | 5 | Parallel queries |
JSON input example
{"keywords": ["fintech newsletter", "climate newsletter", "local news"],"location": "London","customDomains": ["@gmail.com", "@protonmail.com"],"maxEmails": 300,"countryCode": "GB","expandQueries": true,"queryModifiers": ["email", "contact", "sponsor", "advertise", "partnership"],"maxPagesPerQuery": 30,"maxConcurrency": 5}
Output Schema and Dataset Export
The Substack Email Scraper writes leads to the Apify dataset as they are discovered, so results accumulate visibly during the run.
JSON output example
{"network": "Substack","keyword": "fintech newsletter","query": "site:substack.com fintech newsletter \"@gmail.com\" \"London\"","title": "Macro Margins | Daniel Okereke | Substack","accountName": "macromargins","fullName": "Daniel Okereke","username": "macromargins","profileUrl": "https://macromargins.substack.com/","url": "https://macromargins.substack.com/","description": "Weekly fintech and payments analysis from London. Sponsorship and partnership enquiries: macromargins.ads@gmail.com","email": "macromargins.ads@gmail.com","emailDomain": "@gmail.com","possiblyTruncated": false,"foundAt": "2026-08-31T13:05:41.118Z"}
Export the Substack Email Scraper dataset as JSON, CSV, XLSX or XML from the Apify console, or read it over the API into your CRM.
How to Run It
- Open the Substack Email Scraper on Apify and click Try for free.
- Replace the default keywords with the newsletter category you are targeting —
crypto newsletter,parenting newsletter,B2B SaaS newsletter,film criticism. - Set the email domains you want. A narrow list produces a cleaner dataset than a broad one.
- Optionally add a
locationand acountryCode— essential for local news and regional newsletters. - Choose
maxEmailsand start the run. - Follow the live log: it reports pages fetched, blocked pages, retries and emails per page.
- Export the dataset, or chain the Substack Email Scraper into a workflow through the Apify API.
Tips for a better yield
Feed the Substack Email Scraper a topic plus the word "newsletter". writer is broad; supply chain newsletter is a query that converts.
Leave query expansion on. The sponsor, advertise and partnership modifiers are tuned for exactly how Substack writers phrase their ad pages.
If a niche runs dry, change customDomains before you change the keyword — many publishers run a dedicated ads inbox on a different provider.
Use Cases for the Substack Email Scraper
| Use case | How the Substack Email Scraper helps |
|---|---|
| Newsletter sponsorship buying | Build a category-specific list of publishers open to advertising |
| Ad network supply sourcing | Recruit independent newsletters into a network or marketplace |
| PR and book publicity | Pitch reviews and interviews to writers covering your subject |
| Cross-promotion and swaps | Find newsletters of comparable size in an adjacent niche |
| Recruiting writers | Source proven independent writers for contract or staff roles |
| Media and market research | Map who publishes in a vertical and how they describe their audience |
| CRM enrichment | Append newsletter URLs and bio snippets to existing contact records |
| Podcast and event booking | Invite domain experts who already publish regularly |
Why Choose This Actor
Substack has no exportable publisher directory with contact details, and its own discovery pages surface recommendations rather than emails.
The Substack Email Scraper automates the tedious half of newsletter lead generation — searching, paginating, copying, deduplicating — and leaves the qualification to you.
Because the Substack Email Scraper parses structurally instead of by CSS class name, and because a whole-page fallback exists, a Google layout change degrades result quality rather than breaking the run.
It also complements its siblings. Independent writers rarely stay on one platform: pair the Substack Email Scraper with the Medium Email Scraper for the ones still publishing there, and with the Patreon Email Scraper for creators monetising through memberships instead.
For the audience layer around a newsletter, the X Email Scraper covers where most Substack writers promote, the Reddit Email Scraper reaches the communities that share their posts, and the LinkedIn Email Scraper is stronger for B2B newsletter operators.
Limitations
These constraints on the Substack Email Scraper are real. Read them before your first run.
- Only publicly indexed emails. The Substack Email Scraper finds addresses Google has already crawled. A publisher who never published an email will not appear.
- Google's result cap. A single query returns roughly 300 results at most. That is precisely why query expansion exists — leave it enabled.
possiblyTruncated. When Google's snippet ellipsis touches an address, the row is flaggedtrue. Verify those before outreach.- Apify GOOGLE_SERP proxy required. The Substack Email Scraper cannot run without Apify proxy credentials.
- Free-plan cap. Free Apify plans are limited to 100 emails per run; paid plans are uncapped.
usernameandprofileUrlavailability. These are populated only when Substack exposes a publication handle in Google's result. Individual post results sometimes print only a display name, so those rows keepaccountNameandfullNamebut have an emptyusernameandprofileUrl. That is a Google limitation, not a bug.- Variable yield. Output depends on keywords, domains and location. No volume is guaranteed.
Example: a full run
A typical Substack Email Scraper run with three keywords, two domains, expansion on and maxEmails: 300.
The Substack Email Scraper issues base queries first (site:substack.com climate newsletter "@gmail.com"), then quoted and intitle: variants, then one variant per modifier — email, contact, sponsor, advertise, partnership.
Five queries run concurrently, each paginating until it runs dry or reaches maxPagesPerQuery. Addresses are normalised, domain-filtered, deduplicated globally and streamed to the dataset until the 300th unique email trips the shared stop signal.
Blocked or failed queries are re-queued once at the end of the run, so a transient CAPTCHA does not quietly cost you an entire keyword.
Related Actors
The Substack Email Scraper belongs to a family of contact-discovery Actors that share this engine and output schema.
| Actor | What it collects |
|---|---|
| Substack Email and Phone Number Scraper | Emails and phone numbers from Substack |
| Substack Phone Number Scraper | Public phone numbers from Substack |
| Behance Email Scraper | Public contact emails from Behance |
| Bigo Live Email Scraper | Public contact emails from Bigo Live |
| Bluesky Email Scraper | Public contact emails from Bluesky |
| Bumble Email Scraper | Public contact emails from Bumble |
| Clubhouse Email Scraper | Public contact emails from Clubhouse |
| Dailymotion Email Scraper | Public contact emails from Dailymotion |
| DeviantArt Email Scraper | Public contact emails from DeviantArt |
| Discord Email Scraper | Public contact emails from Discord |
| Dribbble Email Scraper | Public contact emails from Dribbble |
| Facebook Email Scraper | Public contact emails from Facebook |
| Goodreads Email Scraper | Public contact emails from Goodreads |
| Hinge Email Scraper | Public contact emails from Hinge |
| Instagram Email Scraper | Public contact emails from Instagram |
| KakaoTalk Email Scraper | Public contact emails from KakaoTalk |
| Kick Email Scraper | Public contact emails from Kick |
| Lemon8 Email Scraper | Public contact emails from Lemon8 |
| Likee Email Scraper | Public contact emails from Likee |
| LINE Email Scraper | Public contact emails from LINE |
| LinkedIn Email Scraper | Public contact emails from LinkedIn |
| Mastodon Email Scraper | Public contact emails from Mastodon |
| Medium Email Scraper | Public contact emails from Medium |
| Mixcloud Email Scraper | Public contact emails from Mixcloud |
| Patreon Email Scraper | Public contact emails from Patreon |
| Pinterest Email Scraper | Public contact emails from Pinterest |
| Quora Email Scraper | Public contact emails from Quora |
| Reddit Email Scraper | Public contact emails from Reddit |
| Rumble Email Scraper | Public contact emails from Rumble |
| Snapchat Email Scraper | Public contact emails from Snapchat |
| SoundCloud Email Scraper | Public contact emails from SoundCloud |
| Telegram Email Scraper | Public contact emails from Telegram |
| Threads Email Scraper | Public contact emails from Threads |
| TikTok Email Scraper | Public contact emails from TikTok |
| Tinder Email Scraper | Public contact emails from Tinder |
| Tumblr Email Scraper | Public contact emails from Tumblr |
| Twitch Email Scraper | Public contact emails from Twitch |
| Vimeo Email Scraper | Public contact emails from Vimeo |
| VK Email Scraper | Public contact emails from VK |
| WeChat Email Scraper | Public contact emails from WeChat |
| Weibo Email Scraper | Public contact emails from Weibo |
| X Email Scraper | Public contact emails from X |
FAQ
What is the Substack Email Scraper?
An Apify Actor that extracts publicly indexed contact emails from Substack newsletters, publication pages and posts via Google search results, and writes them to a structured dataset.
Does the Substack Email Scraper log in to Substack?
No. The Substack Email Scraper does not log in, does not call a Substack API, and does not open substack.com. It reads Google search results only.
Does it find subscriber emails?
No, and it never could. The Substack Email Scraper only sees addresses a publisher chose to publish on a public page — typically a sponsorship or contact inbox.
Where do the emails come from?
Titles, snippets and site labels in Google's public index, usually from an about page, an advertise page, or the footer of a post.
Can I choose which email providers are returned?
Yes. customDomains filters the Substack Email Scraper's extraction step. Defaults are @gmail.com and @yahoo.com, and the leading @ is optional.
Why is username empty on some rows?
Google sometimes prints only a display name for an individual post. Where no publication handle is exposed, username is null and profileUrl is empty, while accountName and fullName remain populated.
What does possiblyTruncated: true mean?
Google's snippet ellipsis touched the address, so it may be cut off. Treat those rows as unverified before sending anything.
How many Substack newsletter emails will one run return?
It depends on keywords, domains and location — no volume is guaranteed. Free Apify plans stop at 100 emails per run; paid plans are uncapped.
Do I need a proxy?
Yes. The Substack Email Scraper requires the Apify GOOGLE_SERP proxy and will not run without Apify proxy credentials.
Should I leave query expansion on?
Almost always. Google caps a single query near 300 results, and expansion is how the Substack Email Scraper works around that ceiling.
Does the output schema ever change?
No. Every item the Substack Email Scraper writes carries the same 14 fields, which keeps downstream mapping stable.
Can I export to CSV or schedule recurring runs?
Yes. Export as JSON, CSV, XLSX or XML, and use Apify Schedules to run the Substack Email Scraper on a recurring basis into the same dataset.
Leave a review
If the Substack Email Scraper saved you time, please leave a star rating and a short review on the Actor page.
Reviews are how other buyers judge whether a tool works, and they tell us which features to build next.
If something did not work, email neurodata.apify@gmail.com instead - bugs get fixed faster than they get complained about.
Support
Questions about the Substack Email Scraper, feature requests, or need a custom build? Email neurodata.apify@gmail.com.
Not affiliated with, endorsed by, or officially connected to Substack. Use collected data in line with applicable privacy and anti-spam law.