# Tech Jobs Newsletter Aggregator (`parseforge/tech-jobs-newsletter-substack-aggregator-scraper`) Actor

Aggregate hand-curated tech job postings from top Substack newsletters into one structured dataset. Pulls role, company, apply URL, ATS source, newsletter, post title, post URL and date. Export CSV, Excel, JSON or XML for talent pipelines and research.

- **URL**: https://apify.com/parseforge/tech-jobs-newsletter-substack-aggregator-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Jobs, Business, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $19.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are a software tools running on the Apify platform, for all kinds of web data extraction and automation use cases.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js/docs.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python/docs.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

![ParseForge Banner](https://github.com/ParseForge/apify-assets/blob/ad35ccc13ddd068b9d6cba33f323962e39aed5b2/banner.jpg?raw=true)

## 💼 Tech Jobs Newsletter Aggregator

> 🚀 **Aggregate hand-curated tech job postings from top Substack newsletters.**

> 🕒 **Last updated** 2026-07-17 · **📊 9 fields** per record · **Apify residential proxy**

Pulls structured job rows (company, role, link, newsletter) from publications like job hunting sux, Nonlinear News, Beyond Bay Street. CSV, Excel, JSON or XML output. Every record is mapped to a consistent schema - role, company, apply url, ats, newsletter name, post title and more, so you can filter, join and export without any glue code. No account or API key required to try it.

| 🎯 Target Audience | 💡 Primary Use Cases |
|---|---|
| Analysts, developers, researchers and operators who need this data as clean rows | Data collection, enrichment, monitoring, and feeding apps, sheets or models |

### 📋 What the Tech Jobs Newsletter Aggregator does

- Collects the records you ask for as clean, structured rows
- Returns 9 fields per record including role, company, apply url, ats, newsletter name, post title
- Consistent, flat schema - ready for a spreadsheet, database or app
- Exports to CSV, Excel, JSON, or XML

> 💡 **Why it matters:** the raw source returns nested, paged responses. This Actor flattens them into one clean row per record, so the data drops straight into your workflow.

### 🎬 Full Demo (_🚧 Coming soon_)

### ⚙️ Input

| Field | Type | Description |
|---|---|---|
| `newsletters` | array | Substack-powered newsletter feed URLs to aggregate. Each URL should be the RSS endpoint (typically /feed). e.g. `https://www.jobhuntingsux.com/feed, https://nonlinearnews.com/feed, https://beyondbayst.substack.com/feed, https://aitidbits.substack.com/feed, https://thepragmaticengineer.substack.com/feed, https://lennysnewsletter.substack.com/feed`. |
| `lookbackDays` | integer | Only include posts published within the last N days. Set 0 to skip date filtering. e.g. `30`. |
| `keyword` | string | Filter by job title or company keyword (case-insensitive substring match). Leave empty for all. |
| `maxItems` | integer | Cap on the number of results. Free: 10. Paid: up to 1,000,000. |
| `proxyConfiguration` | object | Apify Proxy configuration. Defaults to Apify Proxy. e.g. `[object Object]`. |

````

{"newsletters":\["https://www.jobhuntingsux.com/feed","https://nonlinearnews.com/feed","https://beyondbayst.substack.com/feed","https://aitidbits.substack.com/feed","https://thepragmaticengineer.substack.com/feed","https://lennysnewsletter.substack.com/feed"],"lookbackDays":30,"maxItems":25,"proxyConfiguration":{"useApifyProxy":true}}

````

### 📊 Output

Each item is one record with 9 fields:

| Field | Type | Description |
|---|---|---|
| `role` | string | Role |
| `company` | string | Company |
| `applyUrl` | string | Apply URL |
| `ats` | string | ATS / Source |
| `newsletterName` | string | Newsletter |
| `postTitle` | string | Post Title |
| `postUrl` | string | Post URL |
| `publishedAt` | string | Published |
| `scrapedAt` | string | Scraped |

### ✨ Why choose this Actor

- Clean, flat schema - ready to use
- Residential proxy support for reliable access
- Export to CSV, Excel, JSON, or XML
- Pay-per-event billing - only pay for records returned

### 🚀 How to use

1. [Create a free account w/ $5 credit](https://console.apify.com/sign-up?fpr=vmoqkp)
2. Open the Actor on Apify Console
3. Fill in the input fields
4. Click Start
5. Download the results as CSV, Excel, JSON, or XML

### 💼 Use cases

**Research and analysis** - study the data at scale in your own tools.

**Enrichment** - add structured records to a database, sheet or app.

**Monitoring** - track changes over time with scheduled runs.

**Datasets** - build clean datasets for reporting, dashboards or models.

### 🔌 Automating Tech Jobs Newsletter Aggregator

Trigger the Actor via Make, Zapier, Slack, Airbyte, GitHub Actions, Google Drive, n8n or any HTTP webhook on Apify's integration platform.

### 🤖 Ask an AI assistant about this scraper

Paste this README into ChatGPT, Claude, Perplexity or Copilot and ask how to integrate the data into your workflow.

### ❓ Frequently Asked Questions

**❓ Do I need an API key?** No - just set the input and run.

**❓ Free users?** Capped at 10 items per run - perfect for testing.

**❓ What output formats?** CSV, Excel, JSON, or XML - pick on the Storage tab.

**❓ Is the data live?** Yes - fetched at run time.

**❓ Can I schedule it?** Yes - use Apify's scheduler or any webhook integration.

### 🔌 Integrate with any app

Apify provides webhooks, REST API, scheduled runs, dataset downloads, and integrations with Make, Zapier, n8n, Slack, Google Drive, Airbyte, GitHub Actions, and more.

> 💡 **Pro Tip:** browse the complete [ParseForge collection](https://apify.com/parseforge) for more scrapers and all-in-one combos.

**🆘 Need Help?** [Open our contact form](https://tally.so/r/BzdKgA)

> **⚠️ Disclaimer:** independent tool, not affiliated with any data source named above. Only publicly available data is collected.

# Actor input Schema

## `newsletters` (type: `array`):

Substack-powered newsletter feed URLs to aggregate. Each URL should be the RSS endpoint (typically /feed).
## `lookbackDays` (type: `integer`):

Only include posts published within the last N days. Set 0 to skip date filtering.
## `keyword` (type: `string`):

Filter by job title or company keyword (case-insensitive substring match). Leave empty for all.
## `maxItems` (type: `integer`):

Free users: Limited to 10 newsletter posts (preview). Paid users: Optional, max 1,000,000

## Actor input object example

```json
{
  "newsletters": [
    "https://www.jobhuntingsux.com/feed",
    "https://nonlinearnews.com/feed",
    "https://beyondbayst.substack.com/feed",
    "https://aitidbits.substack.com/feed",
    "https://thepragmaticengineer.substack.com/feed",
    "https://lennysnewsletter.substack.com/feed"
  ],
  "lookbackDays": 30,
  "maxItems": 10
}
````

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "newsletters": [
        "https://www.jobhuntingsux.com/feed",
        "https://nonlinearnews.com/feed",
        "https://beyondbayst.substack.com/feed",
        "https://aitidbits.substack.com/feed",
        "https://thepragmaticengineer.substack.com/feed",
        "https://lennysnewsletter.substack.com/feed"
    ],
    "lookbackDays": 30,
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/tech-jobs-newsletter-substack-aggregator-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "newsletters": [
        "https://www.jobhuntingsux.com/feed",
        "https://nonlinearnews.com/feed",
        "https://beyondbayst.substack.com/feed",
        "https://aitidbits.substack.com/feed",
        "https://thepragmaticengineer.substack.com/feed",
        "https://lennysnewsletter.substack.com/feed",
    ],
    "lookbackDays": 30,
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/tech-jobs-newsletter-substack-aggregator-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "newsletters": [
    "https://www.jobhuntingsux.com/feed",
    "https://nonlinearnews.com/feed",
    "https://beyondbayst.substack.com/feed",
    "https://aitidbits.substack.com/feed",
    "https://thepragmaticengineer.substack.com/feed",
    "https://lennysnewsletter.substack.com/feed"
  ],
  "lookbackDays": 30,
  "maxItems": 10
}' |
apify call parseforge/tech-jobs-newsletter-substack-aggregator-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=parseforge/tech-jobs-newsletter-substack-aggregator-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "Tech Jobs Newsletter Aggregator",
        "description": "Aggregate hand-curated tech job postings from top Substack newsletters into one structured dataset. Pulls role, company, apply URL, ATS source, newsletter, post title, post URL and date. Export CSV, Excel, JSON or XML for talent pipelines and research.",
        "version": "0.1",
        "x-build-id": "PGtX3k1oed3E3dITK"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/parseforge~tech-jobs-newsletter-substack-aggregator-scraper/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-parseforge-tech-jobs-newsletter-substack-aggregator-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/parseforge~tech-jobs-newsletter-substack-aggregator-scraper/runs": {
            "post": {
                "operationId": "runs-sync-parseforge-tech-jobs-newsletter-substack-aggregator-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/parseforge~tech-jobs-newsletter-substack-aggregator-scraper/run-sync": {
            "post": {
                "operationId": "run-sync-parseforge-tech-jobs-newsletter-substack-aggregator-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "properties": {
                    "newsletters": {
                        "title": "Newsletters",
                        "type": "array",
                        "description": "Substack-powered newsletter feed URLs to aggregate. Each URL should be the RSS endpoint (typically /feed).",
                        "default": [
                            "https://www.jobhuntingsux.com/feed",
                            "https://nonlinearnews.com/feed",
                            "https://beyondbayst.substack.com/feed",
                            "https://aitidbits.substack.com/feed",
                            "https://thepragmaticengineer.substack.com/feed",
                            "https://lennysnewsletter.substack.com/feed"
                        ],
                        "items": {
                            "type": "string"
                        }
                    },
                    "lookbackDays": {
                        "title": "Lookback Days",
                        "minimum": 0,
                        "maximum": 3650,
                        "type": "integer",
                        "description": "Only include posts published within the last N days. Set 0 to skip date filtering.",
                        "default": 30
                    },
                    "keyword": {
                        "title": "Job or company keyword",
                        "type": "string",
                        "description": "Filter by job title or company keyword (case-insensitive substring match). Leave empty for all."
                    },
                    "maxItems": {
                        "title": "Maximum newsletter posts",
                        "minimum": 1,
                        "maximum": 1000000,
                        "type": "integer",
                        "description": "Free users: Limited to 10 newsletter posts (preview). Paid users: Optional, max 1,000,000"
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
