# Y Combinator Companies Scraper (`scrapyx/ycombinator-companies-scraper`) Actor

Scrapes the Y Combinator startup directory by industry. Each row carries company name, description, YC batch, location, logo and social links, with optional founders, team size, tags and open job postings.

- **URL**: https://apify.com/scrapyx/ycombinator-companies-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Y Combinator Companies Scraper

Scrapes the **Y Combinator** startup directory by industry. YC is the most
prominent startup accelerator in the world — its company directory is
startup/business intelligence data that has not existed anywhere in this
portfolio before.

***

### What each row contains

- **Name, one-liner and long description**
- **YC batch** (e.g. W24, S23) and **year founded**
- **Location** — city and country
- **Logo** and social links (website, LinkedIn, Twitter/X, GitHub, Crunchbase)
- With `includeCompanyDetails`: **founders** (names, roles, LinkedIn),
  **team size**, **tags**, and **open job postings**

#### Record types

| `recordType` | One per | Purpose |
| --- | --- | --- |
| `SEARCH_SUMMARY` | industry | upstream's own page count, requests spent |
| `COMPANY` | company | the company itself, upstream shape preserved |
| `ERROR` | failed input | so every input maps to at least one row |

***

### Industry is the only filter offered, and it's a real, current list

`industry` is validated against Y Combinator's **own sitemap**, fetched once
per run — not a hardcoded list that could go stale as YC adds categories. An
unrecognised value is refused before any crawl, with the real list included
in the error.

A **location** filter also exists on the site, but is deliberately **not**
offered here: an unresolvable location silently substitutes an unrelated
result set rather than erroring, and — unlike industry — there is no
authoritative list of valid location slugs to validate against. Rather than
offer a filter that can quietly return the wrong data, it's left out
entirely; every company's own `city`/`country`/`location` fields are still
in the output for filtering downstream.

***

### Input

| Field | Notes |
| --- | --- |
| `industry` | One YC industry slug, e.g. `developer-tools`, `fintech`, `generative-ai` |
| `industries` | Optional array to run several industries in one call |
| `includeCompanyDetails` | One extra request per company, for founders/team/tags/jobs |
| `maxItems` | Per industry |

***

### Example

```json
{
  "industries": ["developer-tools", "fintech", "generative-ai"],
  "includeCompanyDetails": true,
  "maxItems": 200
}
```

### Known limits

- No location filter (see above).
- Free-text search (`?query=`) was tested and found completely inert on this
  surface — not offered.
- A company can appear under more than one industry if YC tags it that way;
  running several industries in one call may produce the same company more
  than once across different `SEARCH_SUMMARY` groups (each is scoped to its
  own industry, by design — the same way a multi-location real-estate crawl
  in this portfolio can return the same listing from an overlapping area).

### Anti-bot

None. Seven TLS fingerprints tested on both list and detail pages returned
clean data every time. The Actor still rotates fingerprints on transient
failure and defaults to residential proxy, since cloud egress can be
fingerprinted differently from a local test — an earlier actor in this
portfolio (eFinancialCareers) found exactly that gap on its first cloud run.
`robots.txt`'s `/companies?*` (query-string) path is never used by this
Actor — it works entirely through the unrestricted
`/companies/industry/{slug}` and `/companies/{slug}` path segments.

# Actor input Schema

## `industry` (type: `string`):

One YC industry category slug, e.g. `developer-tools`, `fintech`, `generative-ai`, `machine-learning`, `healthcare`. There are 107 in total — the Actor reads the current, authoritative list from Y Combinator's own sitemap on every run, so it never goes stale, and refuses an unrecognised slug up front with the real list rather than silently returning nothing.

For several industries in one run, use `industries` below instead.

## `industries` (type: `array`):

Optional. Run several industries in one Actor call, each producing its own SEARCH\_SUMMARY row. When set, the `industry` field above is ignored.

## `includeCompanyDetails` (type: `boolean`):

The company list already carries name, one-liner description, location, YC batch, logo and social links. This fetches each company's own page for founders (names, roles, LinkedIn), team size, year founded, tags, and open job postings — one extra request per company.

## `maxItems` (type: `integer`):

Stop after this many companies per industry. Set to 0 for everything in that industry.

## `maxConcurrency` (type: `integer`):

Upper bound on requests in flight at once, across all industries and detail fetches.

## `minRequestInterval` (type: `integer`):

Paces how often requests START, without tying up a concurrency slot. No rate limiting was observed, so this defaults to 0.

## `proxyConfiguration` (type: `object`):

Residential by default. No bot challenge appeared on any of the 7 TLS fingerprints tested across list and detail pages, but datacentre egress from a cloud platform can be fingerprinted differently from a local test.

## Actor input object example

```json
{
  "industry": "developer-tools",
  "industries": [],
  "includeCompanyDetails": false,
  "maxItems": 200,
  "maxConcurrency": 4,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "industry": "developer-tools"
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/ycombinator-companies-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "industry": "developer-tools" }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/ycombinator-companies-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "industry": "developer-tools"
}' |
apify call scrapyx/ycombinator-companies-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/ycombinator-companies-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aXc8aTJaouXywVB9e/builds/yBdDdRVH2bPFTyQPx/openapi.json
