# Y Combinator Companies Scraper & API - YC Startup Directory (`rel8ble/yc-companies-scraper`) Actor

List Y Combinator startups from ycombinator.com/companies. Input: optional search text, batches (e.g. W24), status, industries, regions, tags, team size, hiring flag. One result = one company: name, website, one-liner, batch, status, industry, team size, location, tags, founders. $1 per 1,000.

- **URL**: https://apify.com/rel8ble/yc-companies-scraper.md
- **Developed by:** [Giovanni Rich](https://apify.com/rel8ble) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Y Combinator Companies Scraper & API - YC Startup Directory, Founders and Batches

**Y Combinator Companies Scraper** is a YC startup directory scraper and unofficial Y Combinator API: filter ycombinator.com/companies by batch (W24, S25, ...), status, industry, region, tag, team size or hiring flag, and export every company's name, website, one-liner, description, batch, status, industry, team size, location, tags and launch date as JSON, CSV or Excel. Turn on `includeDetails` to add founders (name, title, LinkedIn), the company's LinkedIn, Twitter/X, Crunchbase and GitHub links, year founded, city, country and open jobs.

It reads the same public search index the YC website uses, over plain HTTP, with **no headless browser and no login**. The whole directory (about 6,300 companies) downloads in about 10 seconds; a 20-company test costs a fraction of a cent.

### How to use

1. Pick your filters: **Batches** (e.g. `W24`, `Summer 2024`), **Company status**, **Industries**, **Regions**, **Tags**, **team size**, **Hiring only**. Or type a **Search text** like `AI agents`. Leave everything empty to get the whole directory.
2. Turn on **Include founders, social links and jobs** if you need founder names, LinkedIn/Twitter/Crunchbase/GitHub URLs, year founded or open roles. This opens one YC page per company (still fast: 5-10 companies per second).
3. Click **Start**. Download the results as JSON, CSV, Excel or HTML from the **Output** tab, or pull them through the Apify API, Make, Zapier, n8n or an MCP client.

### What you get

- **Company basics**: name, slug, YC profile URL, website, one-liner, long description, logo, former names.
- **YC data**: batch (long and short form), status (Active / Acquired / Public / Inactive), stage, YC Top Company flag, nonprofit flag, launch date, hiring flag.
- **Firmographics**: industry, sub-industry, all industries, tags, regions, location string, team size.
- **With `includeDetails`**: founders (name, title, bio, LinkedIn URL, Twitter URL, active flag), company LinkedIn / Twitter / Facebook / Crunchbase / GitHub URLs, year founded, city, country, open job count and titles, number of Launch YC posts, latest launch title, news count, Demo Day video URL.
- **Clean and deduplicated**: one row per company, stable numeric `id`, so you can diff runs or join with other data.

### Use cases

- **Lead generation**: every active B2B company from the last four batches with 5-50 employees, with founder names and LinkedIn URLs, straight into your CRM or outreach tool.
- **Investor and competitor research**: all YC companies in Fintech in Latin America, or every company tagged "Developer Tools" that is hiring.
- **Market maps and newsletters**: the newest launches (`sortBy: "launch_date"`), a batch overview the day after Demo Day, or Top Companies by industry.
- **Recruiting and job search**: companies that are hiring, with their open job titles.
- **Datasets for AI agents**: a clean, filterable company list that an agent can query through MCP with plain-English filters.

### Input example

| Field | Default | Description |
|---|---|---|
| `query` | - | Free-text search, like the search box on YC's site: `"AI agents"`, `"fintech Brazil"`, `"Stripe"` |
| `batches` | all | `["W24", "Summer 2024"]`. Short form: W = Winter, S = Summer, F = Fall, X = Spring |
| `statuses` | all | Any of `Active`, `Acquired`, `Public`, `Inactive` |
| `industries` | all | YC industry names, e.g. `["B2B", "Fintech"]` (list below) |
| `regions` | all | YC region/country names, e.g. `["Europe", "India"]` (list below) |
| `tags` | all | YC tags, e.g. `["SaaS", "Developer Tools"]` |
| `hiringOnly` / `topCompaniesOnly` / `nonprofitOnly` | false | Flags from the YC directory |
| `teamSizeMin` / `teamSizeMax` | - | Self-reported team size range |
| `sortBy` | `default` | `default` (YC order) or `launch_date` (newest launch first) |
| `maxResults` | 100 | Companies to save; `0` = all matching |
| `includeDetails` | false | Adds founders, social links, year founded, city, country, open jobs |

Example: active B2B companies from the two most recent full batches, small teams, with founders:

```json
{
    "batches": ["W25", "S25"],
    "statuses": ["Active"],
    "industries": ["B2B"],
    "teamSizeMin": 2,
    "teamSizeMax": 50,
    "includeDetails": true,
    "maxResults": 0
}
```

Example: the whole directory, basics only (about 6,300 rows, about 10 seconds):

```json
{ "maxResults": 0 }
```

### Output example

One dataset item per company. This one is from a real run with `includeDetails: true` (description shortened).

```json
{
    "id": 30355,
    "name": "assistant-ui",
    "slug": "assistant-ui",
    "ycUrl": "https://www.ycombinator.com/companies/assistant-ui",
    "website": "https://assistant-ui.com",
    "oneLiner": "Open Source React.js Library for AI Chat",
    "longDescription": "assistant-ui helps frontend developers add AI chat to their apps. ...",
    "batch": "Winter 2025",
    "batchShort": "W25",
    "status": "Active",
    "stage": "Early",
    "industry": "B2B",
    "subindustry": "Infrastructure",
    "industries": ["B2B", "Infrastructure"],
    "tags": ["Developer Tools", "Generative AI", "Chat", "Web Development", "AI Assistant"],
    "regions": ["United States of America", "America / Canada"],
    "locations": "San Francisco, CA, USA",
    "teamSize": 3,
    "launchedAt": "2025-02-06T20:25:28.000Z",
    "isHiring": false,
    "topCompany": false,
    "nonprofit": false,
    "formerNames": [],
    "logoUrl": "https://bookface-images.s3.amazonaws.com/small_logos/11f41b527f94c63e7d0ee1e19db038344f713a8f.png",
    "scrapedAt": "2026-10-08T00:25:34.092Z",
    "yearFounded": 2024,
    "city": "San Francisco",
    "country": "US",
    "linkedinUrl": "https://linkedin.com/company/assistant-ui",
    "twitterUrl": "https://x.com/assistantui",
    "facebookUrl": null,
    "crunchbaseUrl": null,
    "githubUrl": "https://github.com/assistant-ui",
    "founders": [
        {
            "name": "Simon Farshid",
            "title": "Founder/CEO",
            "bio": "Simon is the Founder and CEO of assistant-ui. ...",
            "linkedinUrl": "https://linkedin.com/in/simon-farshid",
            "twitterUrl": "https://twitter.com/simonfarshid",
            "isActive": true
        }
    ],
    "founderCount": 1,
    "openJobsCount": 0,
    "openJobTitles": [],
    "launchesCount": 1,
    "latestLaunchTitle": "💬 assistant-ui - Open Source Typescript/React Library for AI Chat",
    "newsCount": 0,
    "demoDayVideoUrl": null
}
```

Without `includeDetails`, the item stops at `scrapedAt`. Fill rates on the full directory (6,279 companies, Oct 2026): website 99%, one-liner 97%, description 93%, team size 98%, location 98%, tags 86%; batch, status, industry, regions, launch date and logo 100%. With details (30-company sample): LinkedIn 97%, Twitter/X 70%, GitHub 33%, Crunchbase 10%, founders, year founded, city and country 100%.

### Pricing

Pay per result: **$1.00 per 1,000 results** (one result = one company saved to the dataset, with or without details).

- 1,000 companies = $1.00
- The whole directory (about 6,300 companies) = about $6.30

You're never charged for failed requests or duplicates. If you set a maximum cost per run, the scraper stops cleanly when it reaches it. Apify's free plan includes $5 of monthly platform credit, enough to download most of the directory.

### Integrations

- **Make, Zapier and n8n**: start runs and push new YC companies into your CRM, Slack or Airtable.
- **Google Sheets**: export the dataset straight into a spreadsheet, or refresh it on a schedule.
- **Apify API**: run the actor and fetch results over REST, or with the official JavaScript and Python clients.
- **Webhooks**: get notified when a run finishes and process the data right away.
- **Schedules**: run it weekly with `sortBy: "launch_date"` to catch new launches, or after each Demo Day for the new batch.
- **MCP for AI agents**: through the Apify MCP server (https://mcp.apify.com), Claude, ChatGPT, Cursor and other agents can call this Y Combinator API directly and read the results.

### Filter values

**Statuses**: Active, Acquired, Public, Inactive.

**Industries** (company count in Oct 2026): B2B (3199), Consumer (885), Healthcare (703), Fintech (663), Engineering, Product and Design (625), Industrials (480), Infrastructure (328), Productivity (232), Marketing (170), Real Estate and Construction (163), Manufacturing and Robotics (160), Operations (154), Supply Chain and Logistics (144), Healthcare IT (141), Finance and Accounting (140), Sales (137), Retail (127), Analytics (125), Education (125), Home and Personal (123), Security (123), Payments (122), Consumer Health and Wellness (119), Content (118), Social (111), Food and Beverage (94), Consumer Finance (89), Housing and Real Estate (85), Human Resources (83), Recruiting and Talent (76), Credit and Lending (75), Banking and Exchange (73), Healthcare Services (73), Insurance (72), Gaming (70), Aviation and Space (67), Therapeutics (65), Legal (60), Drug Discovery and Delivery (59), Asset Management (58), Diagnostics (58), Energy (54), Climate (53), Construction (51), Apparel and Cosmetics (49), Consumer Electronics (49), Medical Devices (44), Government (43), Travel, Leisure and Tourism (36), Industrial Bio (34), Agriculture (31), Transportation Services (27), Defense (25), Office Management (25), Drones (24), Automotive (21), Virtual and Augmented Reality (20), Job and Career Services (19).

**Regions** (most common): America / Canada, United States of America, Remote, Partly Remote, Fully Remote, Europe, South Asia, United Kingdom, India, Latin America, Canada, Southeast Asia, Middle East and North Africa, France, Mexico, Africa, Germany, Singapore, Nigeria, Brazil, Israel, Colombia, East Asia, Indonesia, Oceania, Sweden, Argentina, Spain, Australia, Chile, United Arab Emirates, Denmark, Netherlands, Egypt, South Korea, Switzerland, Pakistan, Kenya, Peru, Philippines, Norway, Hong Kong, Ireland, Malaysia, Turkey, Vietnam, Poland, Portugal, China, and about 50 more countries.

**Tags** (most common of 337): B2B, SaaS, Artificial Intelligence, AI, Fintech, Developer Tools, Marketplace, Generative AI, Consumer, Machine Learning, E-commerce, Healthcare, Analytics, Health Tech, Hardware, Open Source, Productivity, Education, AI Assistant, Robotics, Biotech, API, Payments, Infrastructure, Climate, Hard Tech, Logistics, Enterprise Software, Sales, Digital Health, Marketing, Manufacturing, Data Engineering, Finance, Automation, Supply Chain, Crypto / Web3, Security, Video, Insurance, Gaming, Proptech, Real Estate, Computer Vision, Workflow Automation, Compliance, HR Tech, Construction, Social, Recruiting, Medical Devices, LegalTech, Energy, Delivery.

**Batches**: every batch from Summer 2005 to the current one. The biggest are Winter 2022 (398), Summer 2021 (391) and Winter 2021 (335). Both `"Winter 2024"` and `"W24"` work.

### Limits (read before large runs)

- **About 6,300 companies in total.** That's the whole public directory; there's nothing more to page through.
- **Filters are exact-match on YC's own labels.** `"Fintech"` works, `"fin tech"` doesn't. Use the lists above or copy labels from the YC website. Free-text `query` is fuzzy and searches names, descriptions, tags and locations.
- **The search index serves at most 1,000 results per filter combination.** The scraper handles this automatically by splitting big result sets by batch (no batch has more than ~400 companies). You'll only hit the cap if you combine a huge free-text query with `maxResults: 0`; narrow it by batch or industry if the log says so.
- **Founder and contact data is what YC publishes.** Names, titles and LinkedIn/Twitter URLs come from the public company page. Emails, phone numbers and founder photos are not collected.
- **Team size and location are self-reported** by the companies and can be out of date.
- **Jobs**: `openJobTitles` lists the roles shown on the YC company page (up to 50). Full job descriptions are not included.

### FAQ

**Is it legal to scrape Y Combinator's company directory?**
This actor collects publicly available information about companies, plus founder names and public professional profile links that YC publishes on each company page. You're responsible for using it in line with YC's terms and privacy laws such as GDPR. This is not legal advice: if you're unsure, check with a lawyer.

**Does it need a browser, cookies or a login?**
No. It uses the public search index behind ycombinator.com/companies and reads company pages over plain HTTP. That's why the whole directory takes about 10 seconds.

**How do I get only new companies since my last run?**
Save the `id` values from your previous run and filter them out, or run with `sortBy: "launch_date"` and a small `maxResults` on a schedule. Each company's `id` is stable.

**Why did my run return fewer companies than I expected?**
Check the exact spelling of batch, industry, region and tag names (see **Filter values**). Filters are AND-ed across fields and OR-ed within a field: `industries: ["B2B"]` + `regions: ["Europe"]` means B2B companies in Europe. The `RUN_SUMMARY` record in the key-value store shows how many companies matched your filters.

**What does `includeDetails` cost?**
The same per result. It just takes a bit longer (one extra request per company).

**Can I get the jobs themselves?**
Only the titles shown on each company page. A dedicated jobs scraper is a different tool.

### How it works (for developers)

ycombinator.com/companies is a search UI on top of an Algolia index. The actor loads that page once to read the current public search key, then queries the index directly with the same facet filters the website uses (`batch`, `status`, `industries`, `regions`, `tags`, `isHiring`, `top_company`, `nonprofit`, `team_size`). Because Algolia returns at most 1,000 hits per query, result sets bigger than that are split into one query per batch. With `includeDetails`, each company's YC page is fetched and the embedded page JSON (founders, social links, jobs, launches) is parsed. Every field is read defensively: a malformed record just has fewer fields.

Run it locally:

```bash
npm install
npm test                                  # parser tests on saved fixtures
APIFY_LOCAL_STORAGE_DIR=./storage node src/main.js   # input in storage/key_value_stores/default/INPUT.json
```

# Actor input Schema

## `query` (type: `string`):

Optional. Free-text search over company name, one-liner, description, tags and location, exactly like the search box on ycombinator.com/companies. Examples: "AI agents", "fintech Brazil", "Stripe". Leave empty to list every company that matches the filters below.

## `batches` (type: `array`):

Optional. YC batches to include, one per line. Accepts the long form "Winter 2024" or the short form "W24" (W = Winter, S = Summer, F = Fall, X = Spring). Empty = all batches since 2005 (about 6,300 companies).

## `statuses` (type: `array`):

Optional. Keep only companies with these statuses. Empty = all.

## `industries` (type: `array`):

Optional. Industry or sub-industry names as YC spells them, one per line, e.g. "B2B", "Fintech", "Healthcare", "Consumer", "Developer Tools", "Sales", "Infrastructure", "Real Estate and Construction". A company matches if any of its industries match. Empty = all. The full list is in the README.

## `regions` (type: `array`):

Optional. Region or country names as YC lists them, e.g. "United States of America", "Europe", "India", "Latin America", "Remote", "United Kingdom", "Canada". Empty = all.

## `tags` (type: `array`):

Optional. YC tags, e.g. "SaaS", "Artificial Intelligence", "Marketplace", "Developer Tools", "Climate", "Open Source". A company matches if it carries any of them. Empty = all.

## `hiringOnly` (type: `boolean`):

Optional boolean. true = only companies currently marked as hiring on YC. Default false.

## `topCompaniesOnly` (type: `boolean`):

Optional boolean. true = only companies on YC's Top Companies list (the highest-valued alumni). Default false.

## `nonprofitOnly` (type: `boolean`):

Optional boolean. true = only nonprofit companies. Default false.

## `teamSizeMin` (type: `integer`):

Optional. Keep companies with at least this many employees (YC's self-reported team size). Example: 10.

## `teamSizeMax` (type: `integer`):

Optional. Keep companies with at most this many employees. Example: 50.

## `sortBy` (type: `string`):

Optional. "default" = YC's directory order (relevance for a search text, otherwise top companies first). "launch_date" = most recently launched first. Default "default".

## `maxResults` (type: `integer`):

Optional. Maximum number of companies to save, integer >= 0. Default 100. 0 = no limit (the whole directory is about 6,300 companies). You pay per company saved.

## `includeDetails` (type: `boolean`):

Optional boolean. true = also open each company's YC page and add founders (name, title, LinkedIn/Twitter URL), company LinkedIn/Twitter/Facebook/Crunchbase/GitHub URLs, year founded, city, country, open job titles and launch count. One extra request per company, about 5-10 companies per second. Same price per result. Default false.

## `maxConcurrency` (type: `integer`):

Optional, advanced. Parallel requests, integer 1-20. Default 5. Only matters with includeDetails.

## `maxRequestRetries` (type: `integer`):

Optional, advanced. Retries per failed request, integer 0-20; each retry uses a new session. Default 5.

## `proxyConfiguration` (type: `object`):

Optional, advanced. Apify Proxy settings object, e.g. {"useApifyProxy": true}. Default: Apify Proxy on (datacenter), which is enough for this site. Leave unset.

## Actor input object example

```json
{
  "batches": [
    "Summer 2024"
  ],
  "hiringOnly": false,
  "topCompaniesOnly": false,
  "nonprofitOnly": false,
  "sortBy": "default",
  "maxResults": 20,
  "includeDetails": false,
  "maxConcurrency": 5,
  "maxRequestRetries": 5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `companies` (type: `string`):

All YC companies found, one item each, with batch, status, industry, team size, website, location, tags and optional founders and social links.

## `summary` (type: `string`):

Matching company count, searches made, filters used and the stop reason.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "",
    "batches": [
        "Summer 2024"
    ],
    "maxResults": 20,
    "includeDetails": false,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("rel8ble/yc-companies-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "",
    "batches": ["Summer 2024"],
    "maxResults": 20,
    "includeDetails": False,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("rel8ble/yc-companies-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "",
  "batches": [
    "Summer 2024"
  ],
  "maxResults": 20,
  "includeDetails": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call rel8ble/yc-companies-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rel8ble/yc-companies-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/2YDP1LC6bB2J4reC2/builds/O5FdH8vpzX2ZRW00V/openapi.json
