Y Combinator Companies & Jobs Scraper (all batches) avatar

Y Combinator Companies & Jobs Scraper (all batches)

Pricing

from $1.00 / 1,000 company returneds

Go to Apify Store
Y Combinator Companies & Jobs Scraper (all batches)

Y Combinator Companies & Jobs Scraper (all batches)

Export YC companies filtered by batch, industry, tag, hiring status or keyword, with founders and social links on request. Get current job postings with salary and equity ranges for recruiting, investment research and market mapping.

Pricing

from $1.00 / 1,000 company returneds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Y Combinator Companies and Jobs

Build a focused company directory or collect public startup vacancies from Y Combinator companies. This actor combines the anonymous yc-oss directory with public Work at a Startup company and job pages. It needs no source API key, account, browser, or login. Use it for batch research, hiring dashboards, startup discovery, and regular exports into your own analysis tools.

Quick start

Run the supplied INPUT.json to select up to 40 Summer 2025 companies and attempt up to five job details per company. It requests both company and job rows. This is also the small automated smoke input; measured timings and counts are recorded in VALIDATION.md. Source availability and network conditions can change its duration.

For a larger company export, choose mode companies and increase maxCompanies. With no input, the actor selects up to 500 companies across all batches. Python 3.12 is the runtime. Create a virtual environment, install requirements.txt, then run python -m src. To repeat the supplied local validation on Windows, run powershell -File validation/run_live.ps1 after installing dependencies into .venv. The script creates separate local storage for each execution and saves a report under validation.

Filters and limits

Mode accepts companies, jobs, or both. Batches accept short names such as W25 and S25, or full names such as Winter 2025 and Summer 2025. Industries and tags accept display names or normalized slugs. Values within one list are alternatives; different filters are combined with AND. Empty lists impose no restriction.

Status accepts active, acquired, public, inactive, or all. Query performs a case-insensitive substring search across the company name and one-liner. Omit isHiring to include both hiring and non-hiring businesses; true selects hiring businesses and false selects non-hiring businesses. topCompanyOnly selects the source's top-company flag. Hiring flags reflect the directory snapshot and can lag behind job pages.

maxCompanies defaults to 500 and caps selected unique companies after filtering. Selection preserves directory order; it does not rank companies by size or freshness. maxJobsPerCompany defaults to 50 and caps unique detail attempts, including failed requests. includeFounders defaults to false and exposes founders only when present in the yc-oss company record. It does not enrich founders from other pages. timeoutSecs defaults to 20 per HTTP operation.

Output fields

Company rows include id, name, slug, oneLiner, longDescription, website, batch, status, industry, industries, tags, teamSize, location, regions, foundedYear, isHiring, isTopCompany, launchedAt, ycUrl, logoUrl, socials, founders, and source. Socials contains linkedin, twitter, github, facebook, and crunchbase keys. Missing scalar fields remain null, while missing lists are empty. launchedAt retains the source Unix timestamp; it is not used to invent a founding year.

Job rows include companySlug, companyName, jobId, title, role, location, remote, salaryMin, salaryMax, currency, equity, experience, engagement, postedAt, url, descriptionText, and source. Descriptions are converted to plain text and truncated to 2,000 characters. Employment labels are lowercase, with underscores changed to hyphens. Salary amounts retain source units; no currency conversion or annualization occurs. An ambiguous dollar sign alone does not establish USD. Missing publication dates and occupational categories remain null. Remote is inferred from explicit telecommuting data or a location containing “remote”; unknown locations produce null.

Every dataset row has rowType: company, job, error, or summary. Errors identify the public source URL and a concise failure reason. Companies with a recognized empty job list receive a free summary. A filter matching no companies also receives a free summary. The SUMMARY key-value record reports selected companies, successful company and job counts, errors, empty job lists, charge-limit status, and elapsed seconds.

Sources and reliability

The actor downloads yc-oss companies/all.json once and filters locally. The project's metadata describes its daily refresh and available batch, industry, and tag endpoints. Local filtering avoids intersecting multiple remote lists. The live source currently names batches with full season names, so short input codes are normalized before matching.

Work at a Startup job IDs come from company-page embedded data or public links. Detail parsing prefers JobPosting JSON-LD and supports the current public embedded page data when JSON-LD is absent. Embedded details must match both requested job ID and company slug. Recommended jobs are not harvested. HTTP requests explicitly accept HTML; this avoids the observed default-header 406 response. No proxy is enabled or required by the observed sources.

HTTP 429 and server errors receive two retries with exponential backoff; transport failures also receive two retries. Other HTTP failures are reported without retries. Partial successes survive later failures. Algolia's public index on the YC company directory is a documented manual fallback if yc-oss becomes unavailable; this version does not automatically discover or call Algolia credentials or endpoints.

Pricing and deployment

Each successful company row requests one company-returned event at $0.001; each successful job row requests one job-returned event at $0.002 through Actor.charge. Errors and zero-result summaries are uncharged. Charge limits stop further successful output. Charging precedes persistence, so a storage failure after charging cannot be rolled back locally. Configure both custom events and disable synthetic start/dataset events before publication. Local runs do not bill. This task does not push or publish the actor.

Example output

Recorded local validation output from storage/live-20260926-060735/datasets/default/000000001.json, dataset row 1. Fields are omitted for brevity; retained values are unchanged. This is a historical example, not a current-source claim.

{
"rowType": "company",
"id": 29697,
"name": "Stormy",
"slug": "stormy",
"website": "https://stormy.ai",
"batch": "Summer 2025",
"status": "Active",
"industry": "B2B",
"teamSize": 5,
"foundedYear": null,
"isHiring": false,
"source": "https://yc-oss.github.io/api/companies/all.json"
}

Use cases

  • A venture capital analyst filters a YC batch and industry to prepare a company research shortlist with source links.
  • A recruiting agency exports public jobs for selected companies and reviews location and salary fields before matching candidates.
  • A startup ecosystem researcher compares saved company exports across batches using IDs, tags, and directory status.
  • A sales operations team selects relevant B2B companies for manual account qualification using websites and company descriptions.

Pricing example

1,000 company rows and 500 job rows cost (1,000 x $0.001) + (500 x $0.002) = $2.00 in declared events. Rates come from .actor/pay_per_event.json. This calculation is an event subtotal, not a measured invoice; local validation does not bill.

Limitations

Directory records and job pages can disagree or change between runs. A missing job is not evidence that a company has stopped hiring. Job-attempt caps and partial failures bound coverage; review error and summary rows before treating an export as complete.