Y Combinator Startups Scraper avatar

Y Combinator Startups Scraper

Pricing

from $1.20 / 1,000 company scrapeds

Go to Apify Store
Y Combinator Startups Scraper

Y Combinator Startups Scraper

Extract YC company data using a public JSON endpoint; coverage and availability vary. Filter by batch, status, industry, region and team size.

Pricing

from $1.20 / 1,000 company scrapeds

Rating

0.0

(0)

Developer

Automation Lab

Automation Lab

Maintained by Community

Actor stats

2

Bookmarked

87

Total users

5

Monthly active users

18 days ago

Last modified

Categories

Share

What does Y Combinator Startups Scraper do?

Y Combinator Startups Scraper extracts structured data from the Y Combinator startup directory. It uses a publicly accessible YC JSON endpoint to collect company profiles including names, websites, descriptions, team sizes, batch info, industries, funding status, and hiring data.

Directory size and available batches change over time; results reflect the YC endpoint at run time. Filter by batch (W24, S23), status (Active, Acquired, Public), industry, region, team size, tags, or hiring status. It makes direct HTTP requests without login or a configured proxy; upstream availability can vary.

Try it now on Apify with the "Start" button — the prefilled input requests up to 20 AI startups.

Who is it for?

Venture Capital & Angel Investors

  • Track new YC batches as they launch to discover investment opportunities early
  • Filter by industry + status to find active startups in your investment thesis
  • Monitor acquired/public companies for exit pattern analysis

Sales & Business Development Teams

  • Build targeted lead lists of YC-backed companies by industry and team size
  • Identify companies currently hiring (growing = budget for new tools)
  • Enrich with website URLs for outbound prospecting campaigns

Market Researchers & Analysts

  • Analyze YC batch composition trends across industries and regions
  • Track startup survival rates by batch vintage
  • Study which industries YC is betting on each season

Recruiters & Talent Teams

  • Find YC startups that are actively hiring
  • Target companies by team size (early-stage vs growth-stage)
  • Build lists of potential employer partners by industry

Why use Y Combinator Startups Scraper?

  • Structured data source — requests a public YC JSON endpoint rather than parsing directory HTML; the endpoint can still change or be unavailable
  • Directory coverage — paginate the companies currently returned by YC, including historical batches where available
  • Rich filtering — search by keyword, batch, status, industry, region, inclusive team-size range, tags, hiring status, and top company designation
  • No login required — direct HTTP access is used; temporary errors are retried up to twice before failing visibly
  • Pay per run and company — a start charge applies once per run, plus a charge per extracted company
  • API-first — integrate with any workflow via the Apify API. Schedule runs, export to Google Sheets, trigger webhooks
  • Lightweight — HTTP-only Actor configured at 256 MB; runtime depends on filters, pagination and upstream response times

What data can you extract?

Each Y Combinator company profile includes:

FieldDescription
nameCompany name
slugURL slug on YC
websiteCompany website URL
oneLinerShort company description
longDescriptionFull company description
teamSizeNumber of employees
batchYC batch (e.g., W24, S23, F25)
statusActive, Acquired, Inactive, or Public
industriesArray of industries (e.g., B2B, Healthcare, Fintech)
regionsGeographic regions
locationsSpecific city/state/country
tagsTags (e.g., Artificial Intelligence, SaaS)
isHiringWhether the company is currently hiring
isTopCompanyYC top company designation
smallLogoUrlCompany logo URL
ycUrlFull YC profile URL
scrapedAtTimestamp of data extraction

How much does it cost to scrape Y Combinator startups?

This Actor uses pay-per-event pricing — each run has a start charge and each extracted company has a separate charge. No monthly subscription. All platform costs are included.

FreeStarter ($29/mo)Scale ($199/mo)Business ($999/mo)
Per company$0.0023$0.002$0.00156$0.0012
Run start$0.005$0.005$0.005$0.005
For a single run, multiply the emitted company count by the relevant per-company rate and add the one-time run start charge.

Higher-tier plans get additional volume discounts.

Real-world cost examples:

For example, 25 results in one Free-tier run cost 25 × the Free per-company rate plus the one-time start fee. Actual counts and runtime depend on current YC data and filters; available Apify credit and other charges vary.

How to scrape Y Combinator startups

  1. Go to the Y Combinator Startups Scraper page on Apify Store
  2. Click "Start" to open the actor in Apify Console
  3. Configure your search filters:
    • Enter a search query (e.g., "AI", "fintech") or leave empty for all companies
    • Select a YC batch (e.g., "W24") to focus on a specific cohort
    • Choose a status filter (Active, Acquired, Inactive, Public)
    • Set an optional minimum and/or maximum team size to target an early- or growth-stage segment
    • Toggle Currently hiring only to find growing companies
  4. Set the Max companies limit (start small with 25 to preview results)
  5. Click "Start" to run the scraper
  6. Download results in JSON, CSV, Excel, or connect via API

Example input — scrape all W24 AI startups:

{
"searchQuery": "AI",
"batch": "W24",
"maxCompanies": 100
}

Example input — find hiring healthcare startups with teams of 10–50:

{
"industry": "Healthcare",
"isHiring": true,
"minTeamSize": 10,
"maxTeamSize": 50,
"maxCompanies": 200
}

Team-size bounds are inclusive. When either bound is set, companies whose team size is missing or non-numeric are excluded rather than treated as zero.

Input parameters

ParameterTypeDefaultDescription
searchQuerystring""Search by company name or description
batchstring""YC batch filter (e.g., W24, S23, F25)
statusstring""Company status: Active, Acquired, Inactive, Public
industrystring""Industry filter (e.g., Healthcare, Fintech, B2B)
regionstring""Region filter (e.g., United States, Europe)
minTeamSizeintegerunsetInclusive minimum team size (0 or greater)
maxTeamSizeintegerunsetInclusive maximum team size (0 or greater)
isHiringbooleanfalseOnly companies currently hiring
tagsstring""Tag filter (e.g., Artificial Intelligence, SaaS)
isTopCompanybooleanfalseOnly YC top companies
maxCompaniesinteger100Max companies to scrape (0 = unlimited)

Output example

{
"id": 30837,
"name": "AirCaps",
"slug": "aircaps",
"website": "https://aircaps.com",
"smallLogoUrl": "https://bookface-images.s3.amazonaws.com/small_logos/839111803eb4ccce6e6e411847617a96d8d7d880.png",
"oneLiner": "The AI copilot for in-person conversations.",
"longDescription": "AirCaps is bringing AI assistance to in-person conversations...",
"teamSize": 2,
"ycUrl": "https://www.ycombinator.com/companies/aircaps",
"batch": "F25",
"status": "Active",
"industries": ["Consumer"],
"regions": ["United States of America", "America / Canada"],
"locations": ["San Francisco, CA, USA"],
"tags": ["Artificial Intelligence", "Productivity", "AI", "Conversational AI"],
"badges": [],
"isHiring": false,
"isTopCompany": false,
"scrapedAt": "2026-03-30T12:00:00.000Z"
}

Tips for best results

  • Start small — use maxCompanies: 25 for your first run to preview the data and estimate costs
  • Use server-side filters first — searchQuery, batch, status, tags, and isHiring are filtered server-side and run faster than industry, region, or team-size bounds (which require scanning all pages)
  • Combine filters — narrow down results by combining batch + status + tags for precise targeting
  • Schedule weekly runs — set up a scheduled run to track new YC companies as batches launch
  • Export to Google Sheets — use the Apify Google Sheets integration for automatic updates to your CRM or deal flow tracker
  • Full available directory — set maxCompanies: 0 to paginate all companies currently returned by the endpoint; allow more runtime for large scans and client-side filters

Integrations

  • Y Combinator Scraper → Google Sheets — automatically update your deal flow spreadsheet with new YC startups each week
  • Y Combinator Scraper → Slack — get notified when new companies match your investment criteria (e.g., Healthcare + hiring)
  • Y Combinator Scraper → Zapier/Make — trigger outbound email sequences when new startups appear in your target industry
  • Y Combinator Scraper → CRM (HubSpot, Salesforce) — enrich your pipeline with YC company data including website, team size, and description
  • Scheduled runs — run weekly to catch new batch announcements and company status changes
  • Webhooks — trigger downstream processing as soon as a scrape completes

API usage

You can access Y Combinator Startups Scraper programmatically using the Apify API.

Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('automation-lab/ycombinator-scraper').call({
searchQuery: 'AI',
batch: 'W24',
maxCompanies: 50,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

Python

from apify_client import ApifyClient
client = ApifyClient('YOUR_APIFY_TOKEN')
run = client.actor('automation-lab/ycombinator-scraper').call(run_input={
'searchQuery': 'AI',
'batch': 'W24',
'maxCompanies': 50,
})
items = client.dataset(run['defaultDatasetId']).list_items().items
print(items)

cURL

curl -X POST "https://api.apify.com/v2/acts/automation-lab~ycombinator-scraper/runs?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"searchQuery": "AI", "batch": "W24", "maxCompanies": 50}'

Use with AI agents via MCP

Y Combinator Startups Scraper is available as a tool for AI assistants that support the Model Context Protocol (MCP).

Add the Apify MCP server to your AI client — this gives you access to all Apify actors, including this one:

Setup for Claude Code

$claude mcp add --transport http apify "https://mcp.apify.com?tools=automation-lab/ycombinator-scraper"

Setup for Claude Desktop, Cursor, or VS Code

Add this to your MCP config file:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=automation-lab/ycombinator-scraper"
}
}
}

Your AI assistant will use OAuth to authenticate with your Apify account on first use.

Example prompts

Once connected, try asking your AI assistant:

  • "Use automation-lab/ycombinator-scraper to find all AI startups from the W24 YC batch"
  • "Scrape Y Combinator companies that are currently hiring in healthcare"
  • "Get the full list of YC top companies with their websites and team sizes"

Learn more in the Apify MCP documentation.

Data handling, dependencies, and support

This Actor makes HTTPS requests to a publicly accessible Y Combinator companies JSON endpoint. It uses no AI provider, paid data service, proxy, login, or user-supplied credential. Search and filter inputs are sent only to that API as query parameters where supported; client-side filters remain inside the Actor run. The Actor keeps no separate cache or external copy. Results and operational logs use Apify storage and follow the retention and deletion settings of your Apify account.

Y Combinator is named only to identify the public data source. This Actor is independently operated and is not affiliated with or endorsed by Y Combinator. For help, use the Issues tab on the Actor's Apify Store page.

Legality

This actor accesses publicly available data from the Y Combinator company directory through a publicly accessible JSON endpoint without authentication.

We follow ethical scraping practices:

  • Only access publicly available information
  • Use the public endpoint without bypassing access controls
  • Respect rate limits and server resources
  • Do not collect personal data beyond what companies voluntarily publish

For more information, see the Apify ethical web scraping guide.

FAQ

How fast is the scraper? It uses direct JSON requests without browser rendering. Runtime varies with the number of pages, filters and YC availability; transient errors may require retries.

How much does it cost to scrape all YC companies? The current directory size varies. Multiply actual emitted results by the per-company rate in the table and add the one-time start charge; your plan's rate may differ.

Is this better than scraping the YC website directly? This Actor requests structured JSON rather than parsing directory HTML. The endpoint may change or become unavailable; terminal failures are reported rather than treated as empty results.

Why do some companies have empty descriptions? Some YC companies don't fill in their longDescription on the YC directory. The oneLiner field is almost always populated, but longDescription may be empty for newer or less active companies.

Why does the industry filter take longer? The YC API doesn't support industry filtering server-side, so the scraper must scan all pages and filter locally. Use searchQuery, batch, status, or tags filters for faster results — those are processed server-side.