# Y Combinator Companies & Jobs Scraper (all batches) (`everyotherfriday/yc-companies`) Actor

Export YC companies filtered by batch, industry, tag, hiring status or keyword, with founders and social links on request. Get current job postings with salary and equity ranges for recruiting, investment research and market mapping.

- **URL**: https://apify.com/everyotherfriday/yc-companies.md
- **Developed by:** [Paul Vasquez](https://apify.com/everyotherfriday) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 company returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Y Combinator Companies and Jobs

Build a focused company directory or collect public startup vacancies from Y Combinator companies. This actor combines the anonymous yc-oss directory with public Work at a Startup company and job pages. It needs no source API key, account, browser, or login. Use it for batch research, hiring dashboards, startup discovery, and regular exports into your own analysis tools.

### Quick start

Run the supplied INPUT.json to select up to 40 Summer 2025 companies and attempt up to five job details per company. It requests both company and job rows. This is also the small automated smoke input; measured timings and counts are recorded in VALIDATION.md. Source availability and network conditions can change its duration.

For a larger company export, choose mode `companies` and increase maxCompanies. With no input, the actor selects up to 500 companies across all batches. Python 3.12 is the runtime. Create a virtual environment, install requirements.txt, then run `python -m src`. To repeat the supplied local validation on Windows, run `powershell -File validation/run_live.ps1` after installing dependencies into `.venv`. The script creates separate local storage for each execution and saves a report under validation.

### Filters and limits

Mode accepts `companies`, `jobs`, or `both`. Batches accept short names such as W25 and S25, or full names such as Winter 2025 and Summer 2025. Industries and tags accept display names or normalized slugs. Values within one list are alternatives; different filters are combined with AND. Empty lists impose no restriction.

Status accepts active, acquired, public, inactive, or all. Query performs a case-insensitive substring search across the company name and one-liner. Omit isHiring to include both hiring and non-hiring businesses; true selects hiring businesses and false selects non-hiring businesses. topCompanyOnly selects the source's top-company flag. Hiring flags reflect the directory snapshot and can lag behind job pages.

maxCompanies defaults to 500 and caps selected unique companies after filtering. Selection preserves directory order; it does not rank companies by size or freshness. maxJobsPerCompany defaults to 50 and caps unique detail attempts, including failed requests. includeFounders defaults to false and exposes founders only when present in the yc-oss company record. It does not enrich founders from other pages. timeoutSecs defaults to 20 per HTTP operation.

### Output fields

Company rows include id, name, slug, oneLiner, longDescription, website, batch, status, industry, industries, tags, teamSize, location, regions, foundedYear, isHiring, isTopCompany, launchedAt, ycUrl, logoUrl, socials, founders, and source. Socials contains linkedin, twitter, github, facebook, and crunchbase keys. Missing scalar fields remain null, while missing lists are empty. launchedAt retains the source Unix timestamp; it is not used to invent a founding year.

Job rows include companySlug, companyName, jobId, title, role, location, remote, salaryMin, salaryMax, currency, equity, experience, engagement, postedAt, url, descriptionText, and source. Descriptions are converted to plain text and truncated to 2,000 characters. Employment labels are lowercase, with underscores changed to hyphens. Salary amounts retain source units; no currency conversion or annualization occurs. An ambiguous dollar sign alone does not establish USD. Missing publication dates and occupational categories remain null. Remote is inferred from explicit telecommuting data or a location containing “remote”; unknown locations produce null.

Every dataset row has rowType: company, job, error, or summary. Errors identify the public source URL and a concise failure reason. Companies with a recognized empty job list receive a free summary. A filter matching no companies also receives a free summary. The SUMMARY key-value record reports selected companies, successful company and job counts, errors, empty job lists, charge-limit status, and elapsed seconds.

### Sources and reliability

The actor downloads [yc-oss companies/all.json](https://yc-oss.github.io/api/companies/all.json) once and filters locally. The project's [metadata](https://yc-oss.github.io/api/meta.json) describes its daily refresh and available batch, industry, and tag endpoints. Local filtering avoids intersecting multiple remote lists. The live source currently names batches with full season names, so short input codes are normalized before matching.

Work at a Startup job IDs come from company-page embedded data or public links. Detail parsing prefers JobPosting JSON-LD and supports the current public embedded page data when JSON-LD is absent. Embedded details must match both requested job ID and company slug. Recommended jobs are not harvested. HTTP requests explicitly accept HTML; this avoids the observed default-header 406 response. No proxy is enabled or required by the observed sources.

HTTP 429 and server errors receive two retries with exponential backoff; transport failures also receive two retries. Other HTTP failures are reported without retries. Partial successes survive later failures. Algolia's public index on the YC company directory is a documented manual fallback if yc-oss becomes unavailable; this version does not automatically discover or call Algolia credentials or endpoints.

### Pricing and deployment

Each successful company row requests one company-returned event at $0.001; each successful job row requests one job-returned event at $0.002 through Actor.charge. Errors and zero-result summaries are uncharged. Charge limits stop further successful output. Charging precedes persistence, so a storage failure after charging cannot be rolled back locally. Configure both custom events and disable synthetic start/dataset events before publication. Local runs do not bill. This task does not push or publish the actor.

### Example output

Recorded local validation output from [storage/live-20260926-060735/datasets/default/000000001.json](storage/live-20260926-060735/datasets/default/000000001.json), dataset row 1. Fields are omitted for brevity; retained values are unchanged. This is a historical example, not a current-source claim.

```json
{
    "rowType":  "company",
    "id":  29697,
    "name":  "Stormy",
    "slug":  "stormy",
    "website":  "https://stormy.ai",
    "batch":  "Summer 2025",
    "status":  "Active",
    "industry":  "B2B",
    "teamSize":  5,
    "foundedYear":  null,
    "isHiring":  false,
    "source":  "https://yc-oss.github.io/api/companies/all.json"
}
```

### Use cases

- A venture capital analyst filters a YC batch and industry to prepare a company research shortlist with source links.
- A recruiting agency exports public jobs for selected companies and reviews location and salary fields before matching candidates.
- A startup ecosystem researcher compares saved company exports across batches using IDs, tags, and directory status.
- A sales operations team selects relevant B2B companies for manual account qualification using websites and company descriptions.

### Pricing example

1,000 company rows and 500 job rows cost (1,000 x $0.001) + (500 x $0.002) = **$2.00** in declared events. Rates come from [the local event declaration](.actor/pay_per_event.json). This calculation is an event subtotal, not a measured invoice; local validation does not bill.

### Limitations

Directory records and job pages can disagree or change between runs. A missing job is not evidence that a company has stopped hiring. Job-attempt caps and partial failures bound coverage; review error and summary rows before treating an export as complete.

# Actor input Schema

## `mode` (type: `string`):

Choose company rows, job rows, or both.

## `batches` (type: `array`):

Match any value in this list; combine different filters with AND. Batches accept S25 or Summer 2025; tags and industries accept names or slugs.

## `industries` (type: `array`):

Match any value in this list; combine different filters with AND. Batches accept S25 or Summer 2025; tags and industries accept names or slugs.

## `tags` (type: `array`):

Match any value in this list; combine different filters with AND. Batches accept S25 or Summer 2025; tags and industries accept names or slugs.

## `status` (type: `string`):

Filter the directory company status.

## `query` (type: `string`):

Case-insensitive substring in name or one-liner.

## `isHiring` (type: `boolean`):

Omit to include both hiring and non-hiring companies; false selects non-hiring only.

## `topCompanyOnly` (type: `boolean`):

Only return companies marked top\_company.

## `includeFounders` (type: `boolean`):

Include founders from the company record when present; no founder enrichment requests.

## `maxCompanies` (type: `integer`):

Maximum selected companies.

## `maxJobsPerCompany` (type: `integer`):

Maximum unique job detail attempts per company.

## `timeoutSecs` (type: `integer`):

Timeout per HTTP operation in seconds.

## Actor input object example

```json
{
  "mode": "companies",
  "batches": [],
  "industries": [],
  "tags": [],
  "status": "all",
  "topCompanyOnly": false,
  "includeFounders": false,
  "maxCompanies": 500,
  "maxJobsPerCompany": 50,
  "timeoutSecs": 20
}
```

# Actor output Schema

## `rows` (type: `string`):

Dataset JSON

## `csv` (type: `string`):

Dataset CSV

## `summary` (type: `string`):

Run counts and elapsed time

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("everyotherfriday/yc-companies").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("everyotherfriday/yc-companies").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call everyotherfriday/yc-companies --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,everyotherfriday/yc-companies"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0bv6Qbcar3quQnQvh/builds/TCWzFz9atPvF5NfgS/openapi.json
