Y Combinator Scraper — YC Companies, Founders & Batches
Pricing
from $1.00 / 1,000 result items
Y Combinator Scraper — YC Companies, Founders & Batches
Scrape the Y Combinator directory: 6,000+ startups with batch, industry, tags, status, team size, website and one-liner, or 13,000+ founders with role, batches and current company. Filter by batch, industry, region, tag or hiring status. No login or API key, and no 1,000-result cap.
Pricing
from $1.00 / 1,000 result items
Rating
0.0
(0)
Developer
R.L.
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
2 days ago
Last modified
Categories
Share
Extract structured data from the official Y Combinator startup directory and YC founders directory - no browser rendering required. This Actor talks directly to the same Algolia search API that powers both directory pages, so it's fast, reliable, and returns every field the website itself uses (name, description, batch, industry, location, team size, founder titles, and more). Run it on the Apify platform to get scheduled runs, API access, webhooks, and proxy support out of the box.
What can you use YC directory data for?
- Lead generation & investor research - build a list of startups by batch, industry, or region to find investment targets or partnership leads.
- Recruiting & sales prospecting - pull founder names, titles, and current companies to build outreach lists.
- Market research - analyze YC's startup portfolio by industry, stage, or funding batch over time.
- Competitive intelligence - track which companies are hiring, which have shut down, and how industries trend across batches.
How do I scrape the YC directory?
- Click Try for free.
- On the Input tab, choose a Mode:
companiesorfounders. - Optionally set a Search query (same as the search box on the directory page), a Max items limit, and pick from any of the narrowing dropdowns (Batch, Region, Industry, Tags, Status, Is Hiring, Role) - each is populated with the same real options shown in the filter panel on the site, so there's nothing to type.
- Click Start and wait for the run to finish.
- Open the Dataset tab to browse, export (JSON, CSV, Excel, HTML, and more), or download your results via the API.
What input does it take?
| Field | Type | Description |
|---|---|---|
mode | String | companies to scrape ycombinator.com/companies, or founders to scrape ycombinator.com/founders. |
query | String | Free-text search query. Leave empty to fetch the whole directory. |
maxItems | Integer | Maximum number of records to scrape. 0 = no limit. |
companyBatch | Array (dropdown) | Companies only - narrow by YC batch, full name (e.g. Summer 2024). |
founderBatch | Array (dropdown) | Founders only - narrow by YC batch, short code (e.g. S24). |
companyRegions | Array (dropdown) | Companies only - narrow by region/location (e.g. United States of America). |
founderRegions | Array (dropdown) | Founders only - narrow by current region/location. |
industries | Array (dropdown) | Companies only - narrow by industry (e.g. Fintech). |
tags | Array (dropdown) | Companies only - narrow by tag (e.g. B2B, AI). |
status | Array (dropdown) | Companies only - narrow by status: Active, Acquired, Public, or Inactive. |
isHiring | Boolean | Companies only - only return companies that are currently hiring. |
ycTitles | Array (dropdown) | Founders only - narrow by role (e.g. Founder, CEO, CTO). Lists the most common roles. |
Filters combine as an AND across fields and an OR within one field, so industries: ["Fintech"]
with companyBatch: ["Winter 2024", "Summer 2024"] means "fintech companies from either 2024
batch". Each dropdown is populated with the actual values from the live directory (fetched at
schema-build time), so you pick from a list instead of typing free text. Filters that belong to the
other mode are ignored rather than rejected.
Which YC batches and filters can you pick?
| Filter | Options | Examples |
|---|---|---|
| Batch (companies) | 50 | Summer 2005 … Winter 2027, plus Spring 2025/Spring 2026, Fall 2024–Fall 2026, and Unspecified |
| Batch (founders) | 49 | the same batches as short codes: S05 … W27, P25/P26 for Spring, F24–F26 for Fall |
| Industry (companies) | 59 | B2B, Consumer, Healthcare, Fintech, Infrastructure, Productivity, Real Estate and Construction, Healthcare IT |
| Tags (companies) | 333 | SaaS, Artificial Intelligence, AI, Generative AI, Developer Tools, Marketplace, E-commerce, Machine Learning, Health Tech |
| Status (companies) | 4 | Active, Inactive, Acquired, Public |
| Region (companies) | 101 | United States of America, Europe, India, Latin America, Canada, United Kingdom, South Asia, Remote, Fully Remote |
| Region (founders) | 105 | as above, for the founder's current location |
| Role (founders) | 30 | Founder, CEO, CTO, Co-Founder, COO, CPO, CMO, President, Head of Product |
YC's batch calendar changed over the years: two batches a year (Winter and Summer) up to 2024, then
four — Winter, Spring, Summer and Fall. Both are in the dropdowns, so a run can target
Fall 2026 as easily as Summer 2013.
What does each row contain?
Example item in companies mode:
{"id": 531,"name": "DoorDash","slug": "doordash","url": "https://www.ycombinator.com/companies/doordash","website": "http://doordash.com","one_liner": "Restaurant delivery.","batch": "Summer 2013","status": "Public","team_size": 8600,"industry": "Consumer","industries": ["Consumer", "Food and Beverage"],"tags": ["Marketplace", "E-commerce"],"all_locations": "San Francisco, CA, USA","is_hiring": true}
Example item in founders mode:
{"id": 14644,"full_name": "Tony Xu","url": "https://www.ycombinator.com/founders/tony-xu","current_company": "DoorDash","current_title": "CEO","yc_titles": ["CEO"],"batches": ["S13"],"current_region": "United States of America"}
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
Which fields does each row have?
Companies
| Field | Description |
|---|---|
id, object_id | YC's numeric company id and the Algolia record id |
name, slug, url | Company name and links to its YC profile page |
website | Company's own website |
one_liner, long_description | Short and full descriptions |
batch, status, stage | YC batch, current company status (Active/Acquired/Public/Inactive), and growth stage |
team_size | Reported employee count |
industry, subindustry, industries, tags | Industry classification and tags |
all_locations, regions | Headquarters location(s) |
is_hiring, top_company, nonprofit | Flags |
former_names, launched_at, small_logo_thumb_url | Previous names, launch timestamp, logo |
Founders
| Field | Description |
|---|---|
id, object_id | YC's numeric founder id and the Algolia record id |
full_name, first_name, last_name, url_slug, url, hnid | Identity and profile link |
current_company, current_title, company_slug, all_companies_text | Current role |
yc_titles, batches | Titles and YC batches across their companies |
yc_industries, yc_parent_industries, yc_subindustries | Industry classification |
current_region, top_company, avatar_thumb | Current location, flag, photo |
How does it fetch a whole directory past Algolia's 1,000-hit limit?
Algolia only lets a single query walk through its first 1,000 hits, even though it reports the true
total in nbHits — which is why scrapers built on the plain search endpoint stop at 1,000 rows. This
Actor shards each search instead: it splits the directory by YC batch, and then bisects the numeric
id range of any shard that still reports more than 1,000 hits, recursing until every slice fits
under the cap. The slices together cover 100% of the matching records, and results are de-duplicated
by Algolia record id, so a full directory run really does return the full directory.
What do the YC terms mean?
- Batch — the cohort a company went through YC in. Companies carry the full name (
Summer 2013); founders carry the short code (S13). Winter isW, Summer isS, Spring isP(P25= Spring 2025, becauseSwas already taken) and Fall isF. A founder can list several batches if they went through YC more than once. - Status —
Active(operating),Acquired(bought by another company),Public(listed on a stock exchange) orInactive(shut down or dormant). This is YC's own label on the profile, not a live company-registry check. - Stage — the company's stage as YC records it on the profile, separate from
status. - Industry vs. subindustry vs. tags —
industryis the single top-level category on the profile,industriesthe full list including the parent category,subindustrythe narrower one, andtagsthe free-form keyword labels (SaaS,Generative AI) the site's Tags filter uses. Tags are much more numerous and much less structured than industries. - One-liner vs. long description —
one_lineris the single-sentence pitch shown in the directory list;long_descriptionis the full profile write-up. - Top company (
top_company) — YC's own flag marking its most successful alumni; it is a boolean on the record, not a ranking. hnid— the founder's Hacker News username, where YC has one on file. Useful for joining to Hacker News activity.all_locationsvs.regions—all_locationsis the free-text office location string (San Francisco, CA, USA);regionsis the normalised list the Region filter searches, which is where values likeRemoteandFully Remotelive.- Bookface — YC's private, login-only founder network. This Actor never touches it; everything here comes from the two public directory pages.
How much does it cost to scrape the YC directory?
Pay per result — $0.001 per row ($1.00 per 1,000 rows). Apify platform usage (compute) is billed on top at standard rates, and it is small: the Actor makes plain JSON API calls, needs no browser and no proxy, and a full directory run finishes in about a minute.
Verified against the live indexes on 2026-09-27, a whole-directory run is:
| Run | Rows | Result cost |
|---|---|---|
| Every YC company | 6,253 | ~$6.25 |
| Every YC founder | 13,905 | ~$13.91 |
One recent batch of companies (Summer 2024 is 248, Winter 2025 is 165, Fall 2026 is 95) | 95–250 | $0.10–$0.25 |
Use maxItems to cap a run, and the Batch/Industry/Region filters to pay only for the slice you
need.
How do I scope a run sensibly?
- Leave
maxItemsat0to get the complete directory; set it lower for quick tests. - Use the Batch/Region/Industry/Tags/Status/Role filters to scope a run instead of scraping everything and filtering afterwards.
- Schedule runs (e.g. weekly) via Apify's built-in Scheduler to track new YC batches and founder updates over time.
FAQ
Do I need a Y Combinator account, login, or API key? No. The Actor uses the public, search-only Algolia key that YC embeds in its own directory pages, and it reads that key live from the page each run (falling back to a bundled copy) so a YC key rotation does not break your runs. Nothing authenticated is touched — in particular not Bookface.
How many companies and founders are there? Checked live on 2026-09-27: 6,253 companies and 13,905 founders. Both grow with every batch, so a scheduled run is the easy way to keep a copy current.
What does one row contain? One company or one founder, flattened — no nesting beyond the list fields (industries, tags, batches, yc_titles). Companies and founders are separate runs, chosen with mode; there is no combined row.
Does it include founder emails or phone numbers? No. The founders directory exposes a name, profile URL, current company and title, YC batches, industry tags, region, photo and Hacker News username — no email addresses, phone numbers or any other contact details. If you need contact data you will have to source it elsewhere.
Can I really get more than 1,000 results? Yes. Algolia caps a single query at 1,000 explorable hits; the Actor shards by batch and bisects the id range until each slice is under that cap, then de-duplicates. That is how a 6,253-row companies run is possible at all.
How do I filter to one batch? Use companyBatch in companies mode (full names like Winter 2025) or founderBatch in founders mode (short codes like W25). Filters for the other mode are ignored, so a companies-only filter left in a founders run does nothing rather than erroring.
Can I schedule it to catch each new batch? Yes, with Apify Schedules. A weekly companies run with no filters keeps a full mirror; a run filtered to the current batch is the cheap way to watch just the newest cohort.
How is the cost bounded? Rows are $0.001 each, so the ceiling of any run is maxItems × $0.001 when you set a cap, and about $6.25 / $13.91 for the entire companies / founders directory. Narrowing by batch, industry or region is the cheapest way to get only what you need.
Why might a run return fewer rows than the directory page shows? Records are de-duplicated by Algolia record id, and the directory's own counts include facets that your filter combination may exclude. If a number looks off, run without filters and compare totals.
Disclaimers and support
This Actor only reads publicly available data exposed by YC's own directory search API and does not access any authenticated Bookface content. It is provided for research and lead-generation purposes; please review Y Combinator's Terms of Use before using the data. YC may change its underlying API at any time without notice, which could require an update to this Actor. If you hit an issue or need custom fields, open a ticket on the Issues tab.
Did you find this useful?
⭐ Rate this actor on Apify! Your feedback helps other users find it and helps us keep improving it.