# Y Combinator Jobs Scraper - YC Startup Jobs & New Postings (`neverempty/yc-startup-jobs-monitor`) Actor

Job postings at Y Combinator startups from ycombinator.com/jobs pages: title, company, batch, location, salary, equity, experience, visa and apply link. Read role, location or company job pages, or monitor them and get only new postings.

- **URL**: https://apify.com/neverempty/yc-startup-jobs-monitor.md
- **Developed by:** [NeverEmpty](https://apify.com/neverempty) (community)
- **Categories:** Jobs, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.75 / 1,000 job returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Y Combinator Jobs Scraper - YC Startup Jobs & New Postings

Job postings at Y Combinator startups, read from the public **jobs pages on ycombinator.com**, one row per posting: title, company, batch, location, remote flag, job type, role, salary range, equity range, minimum experience, visa policy, skills, how long ago it was posted, and the job page and apply links.

Three kinds of pages can be read, in any mix:

1. **Role pages** - `ycombinator.com/jobs/role/<role>` (software-engineer, designer, product-manager, ...).
2. **Location pages** - `ycombinator.com/jobs/location/<location>` (remote, san-francisco, new-york, ...), or role + location pages.
3. **Company jobs pages** - `ycombinator.com/companies/<slug>/jobs`, the open postings of one company.

Turn on **monitoring** and run it on a schedule to get **only the postings that were not there before**.

No login, no cookies. Only what the public page shows is returned. The apply link is passed through as a column; it is never opened.

### What you get

| Column | Meaning |
|---|---|
| `jobId`, `title`, `jobUrl` | YC's numeric id of the posting, its title and its page on ycombinator.com |
| `companyName`, `companySlug`, `companyUrl`, `companyBatch`, `companyOneLiner`, `companyLogoUrl` | The company as shown on the posting, its YC page and batch (`S22`) |
| `location`, `isRemote` | Location text as written (`New York, NY, US / Remote (US)`); `isRemote` is true when that text contains "Remote" |
| `jobType` | `Full-time`, `Contract`, `Internship` |
| `role`, `roleName`, `roleSpecialty` | Role code (`eng`), its display name (`Engineering`) and specialty (`Backend`) |
| `salaryRange`, `equityRange` | As written on the page (`$120K - $180K`, `0.10% - 0.50%`), not converted |
| `minExperience`, `minSchoolYear` | As written (`3+ years`, `Any (new grads ok)`) |
| `visa` | As written: `Will sponsor`, `US citizen/visa only`, `US citizenship/visa not required` |
| `skills` | Skills listed on the posting |
| `hiringManager` | The hiring manager's name when the page shows one. It was empty on every posting sampled on 2026-10-05. No contact details are added |
| `postedAgo`, `lastActiveAgo` | The page's own relative text (`2 months`, `6 days`). The page gives no exact date, so none is invented |
| `applyUrl` | The page's apply link (a Y Combinator sign-in address) |
| `listedOn`, `postingsOnPage` | The page this posting was read from, and how many postings that page carried |
| `change`, `previousCheckedAt` | Monitoring runs only: `new-job`, and when jobs were last recorded for these filters |
| `input`, `scrapedAt`, `status` | What you typed, when the page was read, `ok` for a job row |

A value that is not on the page is `null`, never an empty string.

**A posting that appears on several pages is returned once and charged once** (matched by `jobId`); `listedOn` is the first page it was read from.

### How much of the job board this covers

Each page carries a limited set of postings, and that set is all the Actor can read from it:

- A **company jobs page** lists that company's open postings (3 for Stripe, 0 for Airbnb on 2026-10-05).
- A **role or location page** shows a selection, not every open job: 39 postings on the software-engineer page, 25 on designer, 35 on remote, 50 on boston on 2026-10-05. Reading the same page three times a few seconds apart gave the same set of postings.
- The pages have no "next page" link, and addresses with a query (`?...`) are not used, so there is no way to page further. The site does not publish a total number of open jobs.

Every row has `postingsOnPage`, and the log states the count for each page. To cover more, give several roles and locations (or turn on role + location pages, which are separate selections), and add the companies you care about by slug. This Actor does not claim to return every job at every YC company.

### Input

- **`roles`** - role slugs, one per line. On 2026-10-05 the site listed: `software-engineer`, `designer`, `product-manager`, `recruiting-hr`, `sales-manager`, `marketing`, `support`, `operations`, `science`.
- **`locations`** - location slugs, one per line: `remote`, `san-francisco`, `new-york`, `los-angeles`, `seattle`, `boston`, `austin`, `chicago`, `india`.
- **`combineRolesAndLocations`** - when both are filled in, read one page per role + location pair (`/jobs/role/software-engineer/remote`) instead of the separate role and location pages.
- **`companies`** - company slugs (`stripe`) or YC company page addresses, one per line.
- **`remoteOnly`**, **`visaSponsorshipOnly`**, **`withSalaryOnly`**, **`keywords`** - filters, applied after a page has been read. All filters that are set must match; any one keyword is enough. A posting with no location does not match `remoteOnly`, and one with no salary range does not match `withSalaryOnly`.
- **`maxJobs`** - with monitoring off, stop when this many rows have been returned (default 200).
- **`monitoringMode`**, **`resetMonitoringState`** - see below.
- **`useProxy`** - requests go directly to www.ycombinator.com; a residential proxy is tried once per refused request, at most three times in a run, only when the site answers HTTP 403 or 429.

At most 200 pages are read in one run.

The site answers an unknown role with its Software Engineer page and an unknown location with a general list, both with HTTP 200. Before any row is returned, the role, location or company the page itself reports is compared with what was asked for; if they differ, nothing from that page is returned (a free `no-such-page` or `page-mismatch` row instead).

### Monitoring mode

Turn **monitoringMode** on and schedule the Actor with the same input.

- The **first time a page is read** with a given set of filters, the postings on it are only remembered. Nothing is returned and nothing is charged for that page (a free `baseline-saved` row). To get the postings that are listed now, run once with monitoring off.
- On **later runs**, each page that is read and compared is charged one monitoring check (`page-checked`), and only postings whose `jobId` has not been seen before come back, with `change: "new-job"`. A page with nothing new returns one free `no-new-jobs` row and costs only the check. The check applies to every page that is read and compared, including a company page with no open posting and a page whose postings were all returned from another page in the same run.
- `maxJobs` is not applied in monitoring mode: every new posting on the pages you gave is returned.
- A posting that could not be delivered because of a charge limit is not remembered; it still counts as new. A page that could not be read is not charged and nothing about it is remembered.
- A posting that did not match your filters when it was first seen is remembered and is not returned later, even if the posting is edited.

"New" means new on the pages you watch. A company jobs page lists all of that company's open postings, so a new row there is a newly listed posting. A role or location page shows a selection, so an older posting that enters the selection is also returned once; `postedAgo` tells them apart.

Job ids are remembered per set of filters, so two schedules with different filters do not affect each other. Do not run two schedules with the same filters at the same moment: the platform's key-value store has no atomic update, so two runs finishing together can overwrite each other's record and return the same posting twice.

### What is charged

- **`job-returned`** - one per job row.
- **`page-checked`** - monitoring mode only: one per page that was read and compared with the remembered job ids. The first read of a page (which only remembers its postings) is not charged, and runs with monitoring off never charge it.

There is no start fee and no monthly fee. Postings dropped by your filters are not charged. Rows that explain why nothing was returned are free: `no-such-company`, `no-such-page`, `page-mismatch`, `no-jobs-on-page`, `invalid-input`, `duplicate-input`, `unreadable`, `blocked`, `robots-disallowed`, `no-match`, `no-new-jobs`, `already-returned`, `baseline-saved`, `limit-reached`, `not-read`, `budget-reached`.

If you set a maximum total charge for a run, the Actor stops at the first page for which the limit has no room for one more job row (in monitoring mode: one check plus one row) and adds a free `not-read` row naming the pages it did not read. The first monitoring read of a page costs nothing, so it is read even when the limit is used up, unless the run has already stopped.

### How it reads

- `robots.txt` of www.ycombinator.com is read at the start of every run and followed (rules for `User-agent: *`). Only `/jobs/role/...`, `/jobs/location/...` and `/companies/<slug>/jobs` are requested, never an address with a query.
- One request at a time with a pause between them; a failed request is retried once. Redirects are not followed.
- If a page is refused (HTTP 403 / 429 or a verification page), or three pages in a row cannot be read, the run stops sending requests and reports what it did not read.

### Limits

- Coverage is what the pages carry (see above), not the whole job board.
- The full job description text is not included; `jobUrl` leads to it.
- Dates are relative text as shown on the page.
- This Actor is not affiliated with or endorsed by Y Combinator. You are responsible for how you use the data, including Y Combinator's terms of use.

### Thanks for using this Actor

We build these tools for people who run them every day, and we improve them from what users tell us.

- **Missing a field, or need another filter?** Tell us in the **Issues** tab. If the data is there, we add it.
- **Found a bug or a wrong value?** Post the run ID in the **Issues** tab. Wrong data is the thing we fix first.

If this Actor saved you time, a short review helps other people find it.

# Actor input Schema

## `roles` (type: `array`):

Role pages to read, as the slug used on ycombinator.com/jobs/role/<slug>. On 2026-10-05 the site listed these roles: software-engineer, designer, product-manager, recruiting-hr, sales-manager, marketing, support, operations, science. One per line. Each role page carries a limited set of postings (39 on the software-engineer page and 25 on the designer page on 2026-10-05), not every open job. A role the site does not list comes back as a "no-such-page" row that is not charged.

## `locations` (type: `array`):

Location pages to read, as the slug used on ycombinator.com/jobs/location/<slug>: remote, san-francisco, new-york, los-angeles, seattle, boston, austin, chicago, india. One per line. Each location page carries a limited set of postings (35 on the remote page and 50 on the boston page on 2026-10-05). A location the site has no page for comes back as a "no-such-page" row that is not charged.

## `combineRolesAndLocations` (type: `boolean`):

When both Roles and Locations are filled in: off reads each role page and each location page separately. On reads one page per pair instead (ycombinator.com/jobs/role/<role>/<location>, for example software-engineer + remote) and does not read the separate role and location pages.

## `companies` (type: `array`):

Companies whose jobs page is read (ycombinator.com/companies/<slug>/jobs): the company slug (stripe) or its Y Combinator page address (https://www.ycombinator.com/companies/stripe). One per line. Roles, Locations and Companies together: at most 200 pages per run. A slug with no company comes back as a "no-such-company" row and a company with no open posting as a "no-jobs-on-page" row; these rows are free (in monitoring mode, a page that was read and compared is still charged its monitoring check).

## `remoteOnly` (type: `boolean`):

Keep only postings whose location text on the page contains "Remote" (for example "New York, NY, US / Remote (US)"). Postings with no location do not match.

## `visaSponsorshipOnly` (type: `boolean`):

Keep only postings whose visa field on the page is "Will sponsor". The other values seen on 2026-10-05 are "US citizen/visa only" and "US citizenship/visa not required".

## `withSalaryOnly` (type: `boolean`):

Keep only postings that show a salary range on the page. The range is returned as written ("$120K - $180K", "€75 - €95 EUR"), not converted.

## `keywords` (type: `array`):

Keep only postings where at least one of these words appears in the job title, company name, company one-liner, role, role specialty or skills. Not case-sensitive. One per line.

## `maxJobs` (type: `integer`):

With monitoring off, the run stops when this many job rows have been returned; pages after that are not read. Not applied in monitoring mode, where every new posting on the pages you gave is returned.

## `monitoringMode` (type: `boolean`):

Remembers job ids between runs. The first time a page is read with a given set of filters, the postings on it are only remembered: nothing is returned and nothing is charged for that page. On later runs each page read and compared is charged one monitoring check (page-checked), and only postings whose id has not been seen before are returned (job-returned). "New" means new on the pages you watch: a company's jobs page lists all of its open postings, while a role or location page shows a limited selection, so an older posting that enters that selection is also returned once (see postedAgo). A posting that did not match your filters when first seen is remembered and not returned later. Job ids are remembered per set of filters. Do not run two schedules with the same filters at the same moment.

## `resetMonitoringState` (type: `boolean`):

Clears everything remembered for this Actor on your account (for every set of filters) before this run reads anything, so with monitoring on every page is a first read again: its postings are remembered and nothing is returned. Use it for one run and turn it off again.

## `useProxy` (type: `boolean`):

Requests go directly to www.ycombinator.com, one at a time with a pause between them. If it answers HTTP 403 or 429, the request is sent one more time through an Apify residential proxy, at most three times in a run; if it is refused again the run stops sending requests and says what it did not read. Turn this off to never use a proxy.

## Actor input object example

```json
{
  "roles": [
    "software-engineer",
    "designer"
  ],
  "combineRolesAndLocations": false,
  "remoteOnly": false,
  "visaSponsorshipOnly": false,
  "withSalaryOnly": false,
  "maxJobs": 200,
  "monitoringMode": false,
  "resetMonitoringState": false,
  "useProxy": true
}
```

# Actor output Schema

## `results` (type: `string`):

One row per job posting found on the Y Combinator jobs pages that were read: job id, title, company, batch, location, remote flag, job type, role, salary and equity range as shown, minimum experience, visa policy, skills, how long ago it was posted, job page and apply links, and the page it was found on. A posting that appears on several pages is returned once. Pages that do not exist, pages with no posting, pages that could not be read, runs with no match or no new posting, and runs that hit a limit come back as their own rows and are not charged.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "roles": [
        "software-engineer",
        "designer"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neverempty/yc-startup-jobs-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "roles": [
        "software-engineer",
        "designer",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("neverempty/yc-startup-jobs-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "roles": [
    "software-engineer",
    "designer"
  ]
}' |
apify call neverempty/yc-startup-jobs-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neverempty/yc-startup-jobs-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lMVXkbobydUCkyRl0/builds/dXDzwwIFQaRVduzZN/openapi.json
