Y Combinator Companies Scraper & API - YC Startup Directory avatar

Y Combinator Companies Scraper & API - YC Startup Directory

Pricing

from $1.00 / 1,000 results

Go to Apify Store
Y Combinator Companies Scraper & API - YC Startup Directory

Y Combinator Companies Scraper & API - YC Startup Directory

List Y Combinator startups from ycombinator.com/companies. Input: optional search text, batches (e.g. W24), status, industries, regions, tags, team size, hiring flag. One result = one company: name, website, one-liner, batch, status, industry, team size, location, tags, founders. $1 per 1,000.

Pricing

from $1.00 / 1,000 results

Rating

0.0

(0)

Developer

Giovanni Rich

Giovanni Rich

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Y Combinator Companies Scraper & API - YC Startup Directory, Founders and Batches

Y Combinator Companies Scraper is a YC startup directory scraper and unofficial Y Combinator API: filter ycombinator.com/companies by batch (W24, S25, ...), status, industry, region, tag, team size or hiring flag, and export every company's name, website, one-liner, description, batch, status, industry, team size, location, tags and launch date as JSON, CSV or Excel. Turn on includeDetails to add founders (name, title, LinkedIn), the company's LinkedIn, Twitter/X, Crunchbase and GitHub links, year founded, city, country and open jobs.

It reads the same public search index the YC website uses, over plain HTTP, with no headless browser and no login. The whole directory (about 6,300 companies) downloads in about 10 seconds; a 20-company test costs a fraction of a cent.

How to use

  1. Pick your filters: Batches (e.g. W24, Summer 2024), Company status, Industries, Regions, Tags, team size, Hiring only. Or type a Search text like AI agents. Leave everything empty to get the whole directory.
  2. Turn on Include founders, social links and jobs if you need founder names, LinkedIn/Twitter/Crunchbase/GitHub URLs, year founded or open roles. This opens one YC page per company (still fast: 5-10 companies per second).
  3. Click Start. Download the results as JSON, CSV, Excel or HTML from the Output tab, or pull them through the Apify API, Make, Zapier, n8n or an MCP client.

What you get

  • Company basics: name, slug, YC profile URL, website, one-liner, long description, logo, former names.
  • YC data: batch (long and short form), status (Active / Acquired / Public / Inactive), stage, YC Top Company flag, nonprofit flag, launch date, hiring flag.
  • Firmographics: industry, sub-industry, all industries, tags, regions, location string, team size.
  • With includeDetails: founders (name, title, bio, LinkedIn URL, Twitter URL, active flag), company LinkedIn / Twitter / Facebook / Crunchbase / GitHub URLs, year founded, city, country, open job count and titles, number of Launch YC posts, latest launch title, news count, Demo Day video URL.
  • Clean and deduplicated: one row per company, stable numeric id, so you can diff runs or join with other data.

Use cases

  • Lead generation: every active B2B company from the last four batches with 5-50 employees, with founder names and LinkedIn URLs, straight into your CRM or outreach tool.
  • Investor and competitor research: all YC companies in Fintech in Latin America, or every company tagged "Developer Tools" that is hiring.
  • Market maps and newsletters: the newest launches (sortBy: "launch_date"), a batch overview the day after Demo Day, or Top Companies by industry.
  • Recruiting and job search: companies that are hiring, with their open job titles.
  • Datasets for AI agents: a clean, filterable company list that an agent can query through MCP with plain-English filters.

Input example

FieldDefaultDescription
query-Free-text search, like the search box on YC's site: "AI agents", "fintech Brazil", "Stripe"
batchesall["W24", "Summer 2024"]. Short form: W = Winter, S = Summer, F = Fall, X = Spring
statusesallAny of Active, Acquired, Public, Inactive
industriesallYC industry names, e.g. ["B2B", "Fintech"] (list below)
regionsallYC region/country names, e.g. ["Europe", "India"] (list below)
tagsallYC tags, e.g. ["SaaS", "Developer Tools"]
hiringOnly / topCompaniesOnly / nonprofitOnlyfalseFlags from the YC directory
teamSizeMin / teamSizeMax-Self-reported team size range
sortBydefaultdefault (YC order) or launch_date (newest launch first)
maxResults100Companies to save; 0 = all matching
includeDetailsfalseAdds founders, social links, year founded, city, country, open jobs

Example: active B2B companies from the two most recent full batches, small teams, with founders:

{
"batches": ["W25", "S25"],
"statuses": ["Active"],
"industries": ["B2B"],
"teamSizeMin": 2,
"teamSizeMax": 50,
"includeDetails": true,
"maxResults": 0
}

Example: the whole directory, basics only (about 6,300 rows, about 10 seconds):

{ "maxResults": 0 }

Output example

One dataset item per company. This one is from a real run with includeDetails: true (description shortened).

{
"id": 30355,
"name": "assistant-ui",
"slug": "assistant-ui",
"ycUrl": "https://www.ycombinator.com/companies/assistant-ui",
"website": "https://assistant-ui.com",
"oneLiner": "Open Source React.js Library for AI Chat",
"longDescription": "assistant-ui helps frontend developers add AI chat to their apps. ...",
"batch": "Winter 2025",
"batchShort": "W25",
"status": "Active",
"stage": "Early",
"industry": "B2B",
"subindustry": "Infrastructure",
"industries": ["B2B", "Infrastructure"],
"tags": ["Developer Tools", "Generative AI", "Chat", "Web Development", "AI Assistant"],
"regions": ["United States of America", "America / Canada"],
"locations": "San Francisco, CA, USA",
"teamSize": 3,
"launchedAt": "2025-02-06T20:25:28.000Z",
"isHiring": false,
"topCompany": false,
"nonprofit": false,
"formerNames": [],
"logoUrl": "https://bookface-images.s3.amazonaws.com/small_logos/11f41b527f94c63e7d0ee1e19db038344f713a8f.png",
"scrapedAt": "2026-10-08T00:25:34.092Z",
"yearFounded": 2024,
"city": "San Francisco",
"country": "US",
"linkedinUrl": "https://linkedin.com/company/assistant-ui",
"twitterUrl": "https://x.com/assistantui",
"facebookUrl": null,
"crunchbaseUrl": null,
"githubUrl": "https://github.com/assistant-ui",
"founders": [
{
"name": "Simon Farshid",
"title": "Founder/CEO",
"bio": "Simon is the Founder and CEO of assistant-ui. ...",
"linkedinUrl": "https://linkedin.com/in/simon-farshid",
"twitterUrl": "https://twitter.com/simonfarshid",
"isActive": true
}
],
"founderCount": 1,
"openJobsCount": 0,
"openJobTitles": [],
"launchesCount": 1,
"latestLaunchTitle": "💬 assistant-ui - Open Source Typescript/React Library for AI Chat",
"newsCount": 0,
"demoDayVideoUrl": null
}

Without includeDetails, the item stops at scrapedAt. Fill rates on the full directory (6,279 companies, Oct 2026): website 99%, one-liner 97%, description 93%, team size 98%, location 98%, tags 86%; batch, status, industry, regions, launch date and logo 100%. With details (30-company sample): LinkedIn 97%, Twitter/X 70%, GitHub 33%, Crunchbase 10%, founders, year founded, city and country 100%.

Pricing

Pay per result: $1.00 per 1,000 results (one result = one company saved to the dataset, with or without details).

  • 1,000 companies = $1.00
  • The whole directory (about 6,300 companies) = about $6.30

You're never charged for failed requests or duplicates. If you set a maximum cost per run, the scraper stops cleanly when it reaches it. Apify's free plan includes $5 of monthly platform credit, enough to download most of the directory.

Integrations

  • Make, Zapier and n8n: start runs and push new YC companies into your CRM, Slack or Airtable.
  • Google Sheets: export the dataset straight into a spreadsheet, or refresh it on a schedule.
  • Apify API: run the actor and fetch results over REST, or with the official JavaScript and Python clients.
  • Webhooks: get notified when a run finishes and process the data right away.
  • Schedules: run it weekly with sortBy: "launch_date" to catch new launches, or after each Demo Day for the new batch.
  • MCP for AI agents: through the Apify MCP server (https://mcp.apify.com), Claude, ChatGPT, Cursor and other agents can call this Y Combinator API directly and read the results.

Filter values

Statuses: Active, Acquired, Public, Inactive.

Industries (company count in Oct 2026): B2B (3199), Consumer (885), Healthcare (703), Fintech (663), Engineering, Product and Design (625), Industrials (480), Infrastructure (328), Productivity (232), Marketing (170), Real Estate and Construction (163), Manufacturing and Robotics (160), Operations (154), Supply Chain and Logistics (144), Healthcare IT (141), Finance and Accounting (140), Sales (137), Retail (127), Analytics (125), Education (125), Home and Personal (123), Security (123), Payments (122), Consumer Health and Wellness (119), Content (118), Social (111), Food and Beverage (94), Consumer Finance (89), Housing and Real Estate (85), Human Resources (83), Recruiting and Talent (76), Credit and Lending (75), Banking and Exchange (73), Healthcare Services (73), Insurance (72), Gaming (70), Aviation and Space (67), Therapeutics (65), Legal (60), Drug Discovery and Delivery (59), Asset Management (58), Diagnostics (58), Energy (54), Climate (53), Construction (51), Apparel and Cosmetics (49), Consumer Electronics (49), Medical Devices (44), Government (43), Travel, Leisure and Tourism (36), Industrial Bio (34), Agriculture (31), Transportation Services (27), Defense (25), Office Management (25), Drones (24), Automotive (21), Virtual and Augmented Reality (20), Job and Career Services (19).

Regions (most common): America / Canada, United States of America, Remote, Partly Remote, Fully Remote, Europe, South Asia, United Kingdom, India, Latin America, Canada, Southeast Asia, Middle East and North Africa, France, Mexico, Africa, Germany, Singapore, Nigeria, Brazil, Israel, Colombia, East Asia, Indonesia, Oceania, Sweden, Argentina, Spain, Australia, Chile, United Arab Emirates, Denmark, Netherlands, Egypt, South Korea, Switzerland, Pakistan, Kenya, Peru, Philippines, Norway, Hong Kong, Ireland, Malaysia, Turkey, Vietnam, Poland, Portugal, China, and about 50 more countries.

Tags (most common of 337): B2B, SaaS, Artificial Intelligence, AI, Fintech, Developer Tools, Marketplace, Generative AI, Consumer, Machine Learning, E-commerce, Healthcare, Analytics, Health Tech, Hardware, Open Source, Productivity, Education, AI Assistant, Robotics, Biotech, API, Payments, Infrastructure, Climate, Hard Tech, Logistics, Enterprise Software, Sales, Digital Health, Marketing, Manufacturing, Data Engineering, Finance, Automation, Supply Chain, Crypto / Web3, Security, Video, Insurance, Gaming, Proptech, Real Estate, Computer Vision, Workflow Automation, Compliance, HR Tech, Construction, Social, Recruiting, Medical Devices, LegalTech, Energy, Delivery.

Batches: every batch from Summer 2005 to the current one. The biggest are Winter 2022 (398), Summer 2021 (391) and Winter 2021 (335). Both "Winter 2024" and "W24" work.

Limits (read before large runs)

  • About 6,300 companies in total. That's the whole public directory; there's nothing more to page through.
  • Filters are exact-match on YC's own labels. "Fintech" works, "fin tech" doesn't. Use the lists above or copy labels from the YC website. Free-text query is fuzzy and searches names, descriptions, tags and locations.
  • The search index serves at most 1,000 results per filter combination. The scraper handles this automatically by splitting big result sets by batch (no batch has more than ~400 companies). You'll only hit the cap if you combine a huge free-text query with maxResults: 0; narrow it by batch or industry if the log says so.
  • Founder and contact data is what YC publishes. Names, titles and LinkedIn/Twitter URLs come from the public company page. Emails, phone numbers and founder photos are not collected.
  • Team size and location are self-reported by the companies and can be out of date.
  • Jobs: openJobTitles lists the roles shown on the YC company page (up to 50). Full job descriptions are not included.

FAQ

Is it legal to scrape Y Combinator's company directory? This actor collects publicly available information about companies, plus founder names and public professional profile links that YC publishes on each company page. You're responsible for using it in line with YC's terms and privacy laws such as GDPR. This is not legal advice: if you're unsure, check with a lawyer.

Does it need a browser, cookies or a login? No. It uses the public search index behind ycombinator.com/companies and reads company pages over plain HTTP. That's why the whole directory takes about 10 seconds.

How do I get only new companies since my last run? Save the id values from your previous run and filter them out, or run with sortBy: "launch_date" and a small maxResults on a schedule. Each company's id is stable.

Why did my run return fewer companies than I expected? Check the exact spelling of batch, industry, region and tag names (see Filter values). Filters are AND-ed across fields and OR-ed within a field: industries: ["B2B"] + regions: ["Europe"] means B2B companies in Europe. The RUN_SUMMARY record in the key-value store shows how many companies matched your filters.

What does includeDetails cost? The same per result. It just takes a bit longer (one extra request per company).

Can I get the jobs themselves? Only the titles shown on each company page. A dedicated jobs scraper is a different tool.

How it works (for developers)

ycombinator.com/companies is a search UI on top of an Algolia index. The actor loads that page once to read the current public search key, then queries the index directly with the same facet filters the website uses (batch, status, industries, regions, tags, isHiring, top_company, nonprofit, team_size). Because Algolia returns at most 1,000 hits per query, result sets bigger than that are split into one query per batch. With includeDetails, each company's YC page is fetched and the embedded page JSON (founders, social links, jobs, launches) is parsed. Every field is read defensively: a malformed record just has fewer fields.

Run it locally:

npm install
npm test # parser tests on saved fixtures
APIFY_LOCAL_STORAGE_DIR=./storage node src/main.js # input in storage/key_value_stores/default/INPUT.json