# Substack Publications & Posts Scraper (`talcoe/substack-publications`) Actor

Discover Substack publications by category or name and pull their recent posts with author, subscriber count and recommended sister publications. Filter by topic, cap results, pay only per result returned.

- **URL**: https://apify.com/talcoe/substack-publications.md
- **Developed by:** [M Usama](https://apify.com/talcoe) (community)
- **Categories:** News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Substack Publications & Posts Scraper do?

It discovers Substack publications, either by browsing a public topic category (Technology, Business, Culture, Crypto and 28 more) or from a list of publications you name directly, then pulls each one's most recent posts with the author, subscriber count, engagement numbers and the publication's recommended sister publications. Set a category and post count and you get a clean, deduplicated table of newsletters and their latest content in one run. Pricing is per result: you only pay for posts actually returned.

### Data source

The Actor reads Substack's own public JSON API (the same endpoints substack.com and every `*.substack.com` site call from the browser): the category leaderboard for discovery, `/api/v1/posts` for each publication's post archive, and `/api/v1/recommendations` for cross-promotions. No login, no scraping of rendered HTML, data as fresh as the live site.

### Why use it?

- Lead generation: find active newsletter operators in your niche along with their subscriber tier and contact handle.
- Content monitoring: track what competitors or influencers in a topic are publishing, how often, and how it performs (comments, reactions, restacks).
- Market research: measure which publications in a category are growing (subscriber counts) and who they cross-promote.
- Replaces manually browsing Substack's leaderboard and each publication's archive page by hand.

### How to use it

1. Open the Actor and pick one or more categories, or list specific publications by subdomain or URL.
2. Click Start. Download results as JSON, CSV or Excel, or read them through the API.
3. Add a Schedule and an Integration (Slack, Google Sheets, webhook) to get new posts automatically.

### Input

| Field | Description |
|---|---|
| `categories` | Discover publications from Substack's public category leaderboards (Technology, Business, Culture, and 29 more). Default: `["technology"]`. |
| `publicationSubdomains` | Optional list of exact publications to include, as a subdomain, bare name, or full substack.com URL (e.g. `platformer`, `garbageday.substack.com`). Runs in addition to any categories. |
| `postsPerPublication` | How many of each publication's most recent posts to return. Default 3, max 20. |
| `maxPublications` | Cap on distinct publications crawled across all selected categories. Default 15, max 100. |
| `includeRecommendations` | Attach each publication's publicly recommended sister publications to its post records. Default true. |
| `maxResults` | Hard stop for the run. You are billed per result, so this caps the cost. |
| `proxyConfiguration` | Apify proxy is recommended for larger runs. |

Example input:

```json
{
    "categories": ["technology"],
    "postsPerPublication": 3,
    "maxPublications": 15,
    "includeRecommendations": true,
    "maxResults": 60
}
```

### Output

Every record is one post, enriched with its publication and author:

```json
{
    "id": "careerbrew:216893901",
    "title": "Career Brew - 53 Hottest Early to Mid Career Jobs",
    "subtitle": "$147K at Mastercard; $195K at NVIDIA; $220K at HighLevel and many more",
    "url": "https://careerbrew.substack.com/p/career-brew-28th-sep-53-hottest-early",
    "postDate": "2026-09-28T17:56:05.255Z",
    "publicationName": "Career Brew",
    "publicationSubdomain": "careerbrew",
    "authorName": "Career Brew",
    "category": "technology",
    "postId": 216893901,
    "publicationId": 2333426,
    "publicationUrl": "https://careerbrew.substack.com",
    "authorHandle": "careerbrew",
    "type": "newsletter",
    "audience": "everyone",
    "wordcount": 612,
    "commentCount": 0,
    "reactionCount": 8,
    "restacks": 1,
    "subscriberTier": null,
    "freeSubscriberCount": 355000,
    "freeSubscriberCountDisplay": "355K+",
    "publicationRank": 5,
    "heroText": "Increasing accessibility to education and career opportunities...",
    "coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/...",
    "logoUrl": "https://substackcdn.com/image/fetch/...",
    "recommendedPublications": [
        { "name": "The Founders Corner®", "subdomain": "thefoundercorner" },
        { "name": "The VC Corner", "subdomain": "thevccorner" }
    ],
    "scrapedAt": "2026-09-30T08:18:47.916Z"
}
```

| Field | Meaning |
|---|---|
| `id` | Stable key (`subdomain:postId`), safe for deduplication between runs |
| `title`, `subtitle`, `url` | Post headline, summary and canonical link |
| `postDate` | ISO date the post was published |
| `publicationName`, `publicationSubdomain`, `publicationUrl` | The newsletter it belongs to |
| `authorName`, `authorHandle` | Byline of the post |
| `category` | The category it was discovered under, or `null` if it came from `publicationSubdomains` |
| `type`, `audience` | `newsletter`/`podcast`/`thread`; who can read it (`everyone`, `only_paid`, `founding`) |
| `wordcount`, `commentCount`, `reactionCount`, `restacks` | Engagement numbers |
| `freeSubscriberCount`, `freeSubscriberCountDisplay` | Publication's free subscriber count, exact and rounded; only available for publications found via `categories`, since Substack does not expose it on a per-publication lookup |
| `subscriberTier` | Author's public subscriber milestone badge, when Substack has assigned one |
| `recommendedPublications` | Up to 5 publications this one recommends to its readers |

### Pricing

Pay per event: a small fee when a run starts, then a fixed price per result returned. You only pay for posts that pass your filters. Set `maxResults` to cap any run.

### Limitations

- Only covers publications still hosted on `*.substack.com` or a Substack-managed custom domain; publications that migrated off Substack (their own site, their own API) are not covered.
- `freeSubscriberCount` and `publicationRank` are only populated for publications discovered through `categories`; publications you name directly in `publicationSubdomains` do not expose subscriber counts through a public per-publication lookup.
- Only public, published posts are returned. Paywalled post bodies are not fetched, only their public metadata (title, engagement, audience).

### FAQ

**Is this legal?** The Actor reads the same public JSON API Substack's own pages call for any visitor without a login, and stores only publication and post metadata the site publishes for that purpose. You are responsible for how you use the data.

**Something is missing or broken?** Open an issue in the Issues tab; response time is usually under a day.

# Actor input Schema

## `categories` (type: `array`):

Discover publications from Substack's public category leaderboards. Leave a category out to skip it.

## `publicationSubdomains` (type: `array`):

Optional. Add exact Substack publications to scrape (subdomain, full substack.com URL, or bare name like "platformer"). These run in addition to any categories above.

## `postsPerPublication` (type: `integer`):

How many of each publication's most recent posts to return.

## `maxPublications` (type: `integer`):

Cap on how many distinct publications to crawl across all selected categories.

## `includeRecommendations` (type: `boolean`):

Attach each publication's publicly recommended sister publications to its post records.

## `maxResults` (type: `integer`):

Hard stop for the run. You are billed per result, so this caps the cost.

## `proxyConfiguration` (type: `object`):

Apify proxy is recommended for larger runs.

## Actor input object example

```json
{
  "categories": [
    "technology"
  ],
  "publicationSubdomains": [],
  "postsPerPublication": 3,
  "maxPublications": 15,
  "includeRecommendations": true,
  "maxResults": 60,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "categories": [
        "technology"
    ],
    "publicationSubdomains": [],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("talcoe/substack-publications").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "categories": ["technology"],
    "publicationSubdomains": [],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("talcoe/substack-publications").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "categories": [
    "technology"
  ],
  "publicationSubdomains": [],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call talcoe/substack-publications --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,talcoe/substack-publications"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xnIFOp0MOPuDbImcR/builds/6et1sFSBisnEWSD9R/openapi.json
