Ycombinator $0.8๐Ÿ’ฐ  Companies Scraper avatar

Ycombinator $0.8๐Ÿ’ฐ Companies Scraper

Pricing

from $0.80 / 1,000 results

Go to Apify Store
Ycombinator $0.8๐Ÿ’ฐ  Companies Scraper

Ycombinator $0.8๐Ÿ’ฐ Companies Scraper

Pull every ycombinator.com company across every batch, with founders, social URLs, application Q&A, demo-day video, photos, and partner. Filter by 14 dimensions or paste any /companies search URL straight from your browser.

Pricing

from $0.80 / 1,000 results

Rating

3.0

(1)

Developer

Abot API

Abot API

Maintained by Community

Actor stats

1

Bookmarked

18

Total users

5

Monthly active users

8 days ago

Last modified

Share

Y Combinator Scraper: Startup Directory, Founders, Jobs & Launches

Y Combinator Scraper turns the YC startup directory into structured data. Get 50+ fields per company, including founders, social links, application details, open job postings, news mentions and Launch YC product launches. Filter across 14 dimensions in a form, or paste any /companies search or company URL straight from your browser, then export to JSON, CSV or Excel, or pull it through the API.

Why This Scraper?

  • Deep company profiles. 50+ fields per company when detail fetching is on: founders with bio, title and social links, LinkedIn, Twitter, Crunchbase, GitHub and Facebook URLs, demo-day and application videos, application answers, photos, and the primary YC partner.
  • Open roles and coverage. Job postings with title, location, salary range, equity range, visa notes and an apply URL, plus news mentions and Launch YC product launches.
  • 14 filter dimensions. Free-text query, batch, industry, subindustry, region, tag, status, funding stage, hiring-only, top-ranked, nonprofit, and public-media flags, all combinable in one search.
  • Two ways in. Build filters in the input form, or paste any /companies search or company URL copied from a browser (multiple URLs supported in one run).
  • Streaming output. Each company is written to the dataset as soon as it's collected, so a stopped run still keeps everything found so far.
  • Resumable at two levels. A killed or migrated run picks itself back up automatically without re-charging for companies already saved, and a brand-new run can continue a previous one by ID.
  • Built for monitoring. Incremental mode tracks a saved search over time and labels each row new, updated, unchanged, reappeared or expired.

Use Cases

  • Recruiting and talent sourcing: filter hiring companies by industry, region or stage and pull their open roles in one pass.
  • Investor and market research: scan a batch or stage for company counts, industries and funding stage mix.
  • Sales and partnership prospecting: build lead lists by industry, region or tag, with founder names and company sites.
  • Founder benchmarking and competitive tracking: compare peer companies in the same subindustry or batch.
  • Journalism and content: track Launch YC announcements and news mentions as they appear.
  • Recurring monitoring: schedule a saved search and get only the companies that are new, changed, or newly listed since the last run.

Data You Get

Sample shape: values are illustrative placeholders, not from a live company.

FieldExample
id12345
slug"example-co"
name"Example Co"
url"https://www.ycombinator.com/companies/example-co"
oneLiner"One sentence company pitch."
website"https://example.com"
batch"W24"
batchName"Winter 2024"
status"Active"
stage"Seed"
industry"B2B"
tags["SaaS", "AI"]
regions["United States of America", "Remote"]
teamSize8
yearFounded2024
isHiringtrue
topCompanyfalse
founders[{fullName: "Founder Name", title: "CEO", bio: "Sample bio.", linkedinUrl: "...", twitterUrl: "...", isActive: true}]
founderCount2
jobPostings[{title: "Senior Engineer", location: "Remote", salaryRange: "...", equityRange: "...", visa: "Open to all", applyUrl: "..."}]
jobPostingCount9
newsItems[{title: "Article title", url: "...", date: "Jan 01, 2026"}]
launches[{title: "Launch announcement title", body: "Body of the Launch YC post."}]
scrapedAt"2026-01-01T00:00:00.000Z"

Every company record also carries objectID, subindustry, industries (a list), allLocations, launchedAt, nonprofit, formerNames, logoUrl, smallLogoUrl, location, city, cityTag, country, linkedinUrl, twitterUrl, facebookUrl, crunchbaseUrl, githubUrl, demoDayVideoUrl, appVideoUrl, appAnswers, freeResponseAnswers, primaryGroupPartner, companyPhotos, hasAppVideo, hasDemoDayVideo, hasQuestionAnswers, jobsUrl, newsUrl, newsCount and launchCount. For companies found through a search or pasted search URL, a run with detail fetching turned off still fills in the search-facing fields (name, url, one-liner, batch, status, stage, industry, tags, regions, website), but leaves the social and video links, application answers, the primary partner, year founded and city as null, and founders, job postings, news items, launches and photos as empty lists (counts 0); turn detail fetching on to populate them. Pasted company URLs are loaded from the profile page only, so they never get stage, regions, industries or launchedAt, their isHiring and topCompany flags always read false, and with detail fetching off they carry little more than the slug.

How to Use

  1. Pick a mode: search (build filters below) or url (paste one or more /companies links).
  2. In search mode, set the filters you need. In URL mode, paste search links, company links, or a mix.
  3. Turn Fetch full detail page on for founders, social links, jobs, news and launches, or off for a faster, search-only pull.
  4. Set Max companies (and optionally Max search pages) to control run size and cost, then click Start.

Most-recent batches with full enrichment:

{
"mode": "search",
"batches": ["Winter 2026", "Spring 2026"],
"fetchDetails": true,
"maxListings": 100
}

AI agent companies in B2B, hiring only:

{
"mode": "search",
"query": "ai agents",
"industries": ["B2B"],
"tags": ["Artificial Intelligence", "AI"],
"isHiring": true,
"fetchDetails": true,
"maxListings": 50
}

Top-ranked companies that are public or acquired, across all batches:

{
"mode": "search",
"topCompany": true,
"statuses": ["Public", "Acquired"],
"fetchDetails": true,
"maxListings": 0
}

URL mode, multiple links pasted at once:

{
"mode": "url",
"urls": [
"https://www.ycombinator.com/companies?batch=Winter%202024",
"https://www.ycombinator.com/companies?batch=Summer%202024",
"https://www.ycombinator.com/companies/example-co"
],
"fetchDetails": true
}

Run it from your code

Python:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("abotapi/ycombinator-com-scraper").call(run_input={
"mode": "search",
"isHiring": True,
"fetchDetails": True,
"maxListings": 25,
})
for company in client.dataset(run["defaultDatasetId"]).iterate_items():
print(company["name"], company["url"])

JavaScript:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('abotapi/ycombinator-com-scraper').call({
mode: 'search',
isHiring: true,
fetchDetails: true,
maxListings: 25,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

Or connect it to Make, Zapier, n8n, Google Sheets or webhooks from the Integrations tab.

Resume and recurring updates

  • Resume from a previous run (resumeFromRunId) continues one interrupted run in a brand-new run: paste its run or dataset ID and this run skips every company already saved there (matched by objectID), so it only appends new companies. A killed or migrated run also checkpoints itself automatically and picks back up in the same run without any input, and without re-charging for companies it already collected.
  • Incremental mode (incrementalMode) is for a search or URL set you run on a schedule. Each company is classified NEW, UPDATED (with changedFields), UNCHANGED (suppressed and not billed unless emitUnchanged is on), EXPIRED, or REAPPEARED. EXPIRED is only produced after a run that scanned the entire tracked search or URL set with no cap, no page limit and no resumeFromRunId cutting it short, and only with emitExpired on; a capped or partial run leaves the previous state untouched instead of guessing. REAPPEARED can only happen to a company that a prior complete scan already marked EXPIRED, and it shows up again on a later run, so it never appears unless emitExpired has been used at least once for that search. stateKey names or shares the stored baseline; leave it empty and one is derived automatically from your filters or URLs. Do not combine resumeFromRunId with an incrementalMode search that already has saved state, the two solve different problems and the run stops with an error if they conflict. With incremental mode off, output is exactly as before.

Send results into your apps (MCP connectors)

Optionally pipe the scraped results into the apps you already use, via Model Context Protocol (MCP) connectors. This is an extra delivery step after the scrape: the Apify dataset is never changed.

What gets written to the connector: a condensed, human-readable summary of each company, not the full JSON. Each item becomes one entry with a title and its key fields flattened to plain text. The complete record always stays in the Apify dataset.

  1. Authorize a connector once under Apify โ†’ Settings โ†’ Integrations (Notion, Linear, Airtable, or Apify).
  2. Select it in the "Pipe results into your apps" input field. If the picker is empty, you haven't authorized a connector yet.
  3. For Notion, also set notionParentPageUrl to the page where items should be created.

The connection is mediated by Apify's MCP proxy, so this actor never sees your third-party credentials. Leave the field empty to skip.

Input Parameters

ParameterTypeDefaultDescription
modestringsearchsearch (filters) or url (paste links). The other mode's fields are ignored.
querystringnoneFree-text search across name, one-liner and description. Leave blank to match all companies.
batchesarraynone (prefill: ["Winter 2026", "Spring 2026"])YC batch names. Accepts long form (Winter 2024) or short codes (W24). Leave empty for all batches.
industriesarraynoneTop-level industry, e.g. B2B, Consumer, Healthcare, Fintech. Multiple values are OR'd.
subindustriesarraynoneRefined industry tag, e.g. Fintech -> Payments. Multiple values OR'd.
regionsarraynoneHQ region or country, e.g. United States of America, Europe, Remote. Multiple values OR'd.
tagsarraynoneTopic tags, e.g. SaaS, Artificial Intelligence, Developer Tools. Multiple values OR'd.
statusesarraynoneOne or more of Active, Inactive, Acquired, Public.
stagesarraynoneOne or more of Seed, Early, Growth.
isHiringbooleanfalseOnly companies with active job postings.
topCompanybooleanfalseOnly YC top-ranked companies.
nonprofitbooleanfalseOnly nonprofit companies.
hasDemoDayVideobooleanfalseOnly companies with a public demo-day video.
hasAppVideobooleanfalseOnly companies whose application video is public.
hasAppAnswersbooleanfalseOnly companies whose application answers are public.
sortBystringrelevancerelevance or launch-date-desc (newest launched first).
urlsarraynone (prefill: one sample search URL)URL mode only. Any /companies search URL or company detail URL, one or more.
fetchDetailsbooleantrueAdds founders, social URLs, application Q&A, video URLs, photos, partner, jobs, news, launches, year founded and city.
maxListingsinteger20Cap on total companies. 0 = unlimited (the entire matching search or URL set).
maxPagesintegernone (prefill 0; both mean unlimited)Optional bound on search pages walked per query. Leave empty or 0 to walk every result page until maxListings is hit or the results end.
resumeFromRunIdstringnonePaste a previous run ID or dataset ID to continue a full walk across separate runs, skipping companies already saved there.
incrementalModebooleanfalseTrack a recurring search or URL set and label each row's changeType.
stateKeystringnoneNames the saved baseline this run compares against. Leave empty to auto-derive one from your filters or URLs.
emitUnchangedbooleanfalseAlso return, and bill, companies with no detected change.
emitExpiredbooleanfalseAlso return, and bill, companies no longer found, but only after a run that scanned the full tracked search with nothing cutting it short.
proxyobjectApify Proxy (datacenter)Connection settings. Residential is supported as an upgrade if you ever see throttling.
mcpConnectorsarraynoneOptional: send a summary of each company to apps you authorized under Integrations.
notionParentPageUrlstringnoneNotion connector only: page under which company pages are created.
maxNotifyListingsinteger50Cap on items written to each connector per run. Does not affect the dataset.

Output Example

Sample record: values are illustrative placeholders, not from a live company.

{
"id": 12345,
"slug": "example-co",
"objectID": "12345",
"name": "Example Co",
"url": "https://www.ycombinator.com/companies/example-co",
"oneLiner": "One sentence company pitch.",
"longDescription": "Longer multi-paragraph description from the application.",
"website": "https://example.com",
"batch": "W24",
"batchName": "Winter 2024",
"status": "Active",
"stage": "Seed",
"industry": "B2B",
"subindustry": "B2B -> Productivity",
"industries": ["B2B", "Productivity"],
"tags": ["SaaS", "AI"],
"regions": ["United States of America", "Remote"],
"allLocations": "San Francisco, CA, USA; Remote",
"teamSize": 8,
"yearFounded": 2024,
"launchedAt": 1700000000,
"isHiring": true,
"topCompany": false,
"nonprofit": false,
"logoUrl": "https://bookface-images.s3.us-west-2.amazonaws.com/logos/0000000000000000000000000000000000000000.png",
"location": "San Francisco, CA, USA",
"city": "San Francisco",
"country": "United States",
"linkedinUrl": "https://www.linkedin.com/company/example-co",
"twitterUrl": "https://twitter.com/example_co",
"crunchbaseUrl": "https://www.crunchbase.com/organization/example-co",
"githubUrl": "https://github.com/example-co",
"primaryGroupPartner": {
"id": 1,
"fullName": "Partner Name",
"url": "https://www.ycombinator.com/people/partner-name",
"avatarUrl": "https://example-cdn.ycombinator.com/avatars/partner.jpg"
},
"founders": [
{
"fullName": "Founder Name",
"title": "CEO",
"bio": "Founder bio text.",
"linkedinUrl": "https://www.linkedin.com/in/founder-name",
"twitterUrl": "https://twitter.com/founder_name",
"isActive": true
}
],
"founderCount": 1,
"jobPostings": [
{
"title": "Senior Engineer",
"location": "Remote",
"type": "Full-time",
"role": "eng",
"salaryRange": "$140K - $180K",
"equityRange": "0.10% - 0.50%",
"visa": "Open to all",
"applyUrl": "https://www.workatastartup.com/jobs/0000"
}
],
"jobPostingCount": 1,
"newsItems": [
{ "title": "Article title", "url": "https://example.com/article", "date": "Jan 01, 2026" }
],
"newsCount": 1,
"launches": [
{ "title": "Launch announcement title", "body": "Body of the Launch YC post." }
],
"launchCount": 1,
"scrapedAt": "2026-01-01T00:00:00.000Z"
}

Plan Requirement

The default proxy setting works out of the box for this actor. For large or frequent runs, or if you ever see requests being refused, switch to the residential proxy group under Proxy. Cost scales with the number of rows returned, so Max companies is your cost cap. Fetch full detail page adds one extra page fetch per company, which makes runs slower but does not change the per-row price.

FAQ

How much does it cost?

You pay per company returned. The Pricing tab shows the current rates. Use Max companies to cap the cost of any run, and turn Fetch full detail page off for a faster, lighter, search-only pull.

This actor collects only publicly available startup-directory data. You are responsible for how you use it: follow Y Combinator's terms and the laws that apply to you, and get legal advice if you plan commercial redistribution. Company facts are generally not protected, but logos, photos and founder bios may carry their own rights.

Can I get only new or changed companies on a schedule?

Yes. Schedule the actor from the Schedules tab with the same filters or URLs each time, and turn on Incremental mode. Each run then labels every company NEW, UPDATED, UNCHANGED (not returned unless you turn that on), and, only after a run that scans the whole tracked search with nothing cutting it short, EXPIRED. A company only comes back as REAPPEARED if an earlier complete scan had already marked it EXPIRED and it shows up again later.

What's the difference between "Resume from a previous run" and "Incremental mode"?

Resume from a previous run continues one specific interrupted or capped run: paste its run or dataset ID and this run appends only the companies it hasn't already saved. Incremental mode is for monitoring the same search or URL set over time on a schedule; it remembers state itself and classifies every row. Don't combine the two for a search that already has saved incremental state, the run stops with an error rather than guessing which baseline to use.

Why did my run fail instead of returning an empty dataset?

An empty dataset usually means your filters or URLs matched nothing, or requests were quietly refused for part of the run; this actor does not fail the whole run for that, it saves whatever it managed to collect, which can be nothing. Try loosening your filters or running again. An actual failed run happens only for a handful of clear input problems: a Resume from a previous run ID that doesn't match any run or dataset this account can read, combining that field with an incremental search that already has saved state, or URL mode with an empty URL list or any entry that isn't a complete web address starting with https:// (one malformed entry fails the whole run).

Can I use it with AI agents or MCP?

Yes. Call it from any Apify integration or MCP client, and use the connector field to push results into Notion, Linear or Airtable.

๐Ÿ”— Want more leads data?

Pair this actor with these related scrapers from the same team:

๐Ÿ“‡ BetaList Scraper
Scrape BetaList.com startup profiles with founder and contact enrichment. Extract startup...
๐Ÿ’ผ Gupy Jobs Scraper
Scrape public jobs from Gupy.io by keyword, state, city, workplace type, job type...
๐Ÿ“‡ GoWork FR & DE Company Reviews and Profile Scraper
Extract company profiles from GoWork France and Germany, including contact details...
๐Ÿจ Comparably
From $1/1K. Extract rich company data from comparably.com using company names or profile...
๐Ÿ“‡ JobStreet
From $1/1K. Pull JobStreet company profiles and employee reviews across Malaysia...
๐Ÿ“ฑ Clutch Scraper
Scrape Clutch company directories and profiles. Get names, ratings, review counts, hourly...

๐Ÿ‘‰ Browse all abotapi scrapers

๐Ÿ’ฌ Support & custom scrapers

  • ๐Ÿž Found a bug or a missing field? Open a ticket on the Issues tab. We usually reply within hours.
  • ๐Ÿ› ๏ธ Need another site, extra fields or a private build? Email abotapi@proton.me or message Telegram @abotapi.
  • โญ Enjoying it? A quick review on the actor page helps other users find it.