# Changelog of Y Combinator Jobs Scraper | YC Startup Hiring Data (`tqm/ycombinator-scraper`) Actor

- **URL**: https://apify.com/tqm/ycombinator-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/tqm/ycombinator-scraper.md

## Changelog

### 2026-09-10 (later) — `roles` failed OPEN

```ts
if (roles.length > 0 && role) { ... }   // before
```

A card with no detected role **bypassed the filter** instead of being excluded. This repo's own
`README.md` recorded it as *"`roles` leaks"*; the Store listing never mentioned it, and it was never
fixed.

**Now fails closed**, with `roles_unknown` as its own rejection reason.

**It does not currently bite** — `role` was populated on 20 of 20 cards measured 2026-09-10, so no
row reaches the bypass. That is a property of today's board, not of the code. It is the same failure
mode that **inverted `categories` on `remoteok-scraper`**, found the same day and biting there
because 14 of 99 rows were unclassifiable.

> ⚠️ **A filter that skips rows with missing data is a latent inversion.** It is invisible while the
> field happens to be well-populated, and it surfaces on the day a buyer pays for the run.

### 2026-09-10 — YC-EQUITY-01: `equityOnly` and `hasEquity` removed

`hasEquity` was derived by searching `` `${title} ${companyDescription} ${role}` `` for the word
"equity". That text never contains it, so `hasEquity` was **`false` on every row ever returned** and
`equityOnly: true` filtered out the **entire** dataset, always. Confirmed against the live Store
Actor 2026-09-10: `equityOnly: true`, `maxJobs: 50` → 0 rows, FAILED.

**Both are now gone** — from `src/main.ts`, `.actor/INPUT_SCHEMA.json`, `.actor/dataset_schema.json`
and both READMEs.

**Why removal rather than a louder warning.** The listing already warned — the input was titled
*"(rarely matches)"* and its description opened with *"⚠️ Almost always returns nothing"*. That was
not good enough. **An input whose documented behaviour is "does nothing" is worse than an absent
one**, and a field that is 100% populated with the same wrong value is worse than an honest null,
because nobody audits a field that looks full.

**And "rarely" was too generous.** Measured against the live board 2026-09-10, **YC's list page
carries no equity element at all** — a card's `.whitespace-nowrap` details are
`[jobType, department, role, salary?]` and nothing more. Equity is on the individual job page, i.e.
one extra page-load per row: the cost structure that already killed the Upwork detail tier (N1) and
`indiehackers` (N3). Do not reintroduce this without pricing that fetch first.

**TQM's nightly is unaffected** — `apify_ingest.py` never sent `equityOnly` (its config carries a
comment saying why) and nothing in the pipeline reads `hasEquity`. Checked before removing, not
assumed.

### Unreleased

#### 2026-09-08 — a short result set now says whose limit it was

Measured against the live Store listing: asking for the advertised maximum of **50** returned
**20** (40%), and the count reproduced exactly across separate runs — so this is the board's real
inventory, not a flaky scrape. The status message said only `Returned 20 YCombinator postings.`, which a
buyer who asked for 50 cannot tell apart from a broken Actor.

`finishRun()` now receives `requested: maxJobs` and appends the reason:

```
Returned 20 YCombinator postings. That is every posting on the board matching your
filters — asking for 50 will not return more today.
```

`RUN_STATS` also carries `requested` and `boardExhausted`. Logic and tests live in
`apify-actor-kit` (`inventory.ts`); `src/kit/` here is generated.

### \[1.1.0] - 2026-08-27

#### Added — a zero-row run now says why, and fails

Adopts the shared `src/kit/`, generated from [`apify-actor-kit`](https://github.com/Atredies/apify-actor-kit).

Every filter here was previously a bare `continue`, so a run that returned nothing
gave **no clue which filter emptied it**. Rejections are now counted per filter and
surfaced in the message and in `RUN_STATS.filteredOut`. A refused listing page is
counted too, so `blocked` is distinguishable from "no postings found".
`failOnZeroResults` (default `true`).

#### Fixed — two inputs advertised more than they deliver

The breakdown made both measurable in a single run, on 2026-08-27:

| input | measured | now |
|---|---|---|
| `equityOnly: true` | **rejected 20 of 20** | titled *rarely matches*; description states that YC's list page does not state equity for most postings, so the filter nearly always empties the result |
| `techStackFilter` (devops/kubernetes/aws) | **rejected 19 of 20** | description states the stack is inferred from a very short title and blurb, and suggests filtering client-side for recall |

Neither is silently broken any more: an emptied run now fails and names the filter
that did it. **A filter that silently does not apply is a refund request.**

#### Fixed — `maxJobs` advertised a ceiling it cannot reach

The schema offered up to **1,000**. YCombinator's jobs page is a single
un-paginated list — measured 2026-08-27 it carried **20** postings, and this actor
does not paginate. Maximum corrected to **50**.

#### Changed

- One batched `Dataset.pushData()` per run instead of one call per row.

### Unreleased — 2026-08-13

Documentation only. **No behavior changes**; build `0.0.3` is unmodified.

#### Added

- Full documentation suite: `docs/ARCHITECTURE.md`, `docs/DEVELOPMENT.md`, `docs/YC-DOM.md`,
  `docs/IMPROVEMENT-PLAN.md`, and `CLAUDE.md`.
- `.actor/dataset_schema.json` — the output surface was previously undeclared.
- README rewritten as a buyer-facing listing with an explicit known-limitations table.

#### Verified against live YC

Two platform runs on 2026-08-13:

| Input | Rows | Cost | Duration |
|---|---:|---:|---:|
| `maxJobs: 10` | 10 | $0.0010 | 5 s |
| `maxJobs: 100` | 20 | $0.0052 | 43 s |

20 rows is source exhaustion: one server-rendered page, no pagination. Runs on `CheerioCrawler` over
plain HTTP with no browser, which is why it is among the cheapest actors in the portfolio.

Data quality is the best of the TypeScript actors — `ycBatch`, `role`, `companyDescription` and a
genuine multi-currency salary parser (`USD`/`GBP`/`EUR`/`INR`) all populate reliably.

#### Findings recorded (not fixed)

- A zero-row run exits **SUCCEEDED** — total scrape failure is indistinguishable from success.
- **`equityOnly: true` returns zero rows, always.** `hasEquity` is derived from title + company
  blurb + role, which never mention equity, so it is `false` on every row.
- Every selector is a **Tailwind utility class** (`li.my-2.flex.h-auto.w-full`, …). Any YC restyle
  breaks extraction silently. Most fragile actor in the portfolio.
- Detail chips (`jobType`, department, `role`, salary) are read **by array index** — a missing chip
  shifts every later field into the wrong chip.
- `maxJobs` advertises a maximum of 1,000; ~20 is achievable.
- `postedAt` is relative prose (`"2 days ago"`), not a timestamp.
- `techStack` is matched against title + company blurb + role, never a job description — empty on
  \~8 of 10 rows.
- An unparseable salary is dropped entirely rather than kept as raw text.
- The `K`/`M` multiplier is applied string-wide rather than per number.
- The `roles` filter is bypassed for jobs with no detected role.
- The `jobId` fallback embeds `Date.now()`.
- `department` is extracted into a local variable and never emitted.

All are itemized and prioritized in `docs/IMPROVEMENT-PLAN.md`, which flags one investigation to run
**before** any of the DOM work: check whether YC ships its listings as embedded JSON, since that
would replace the fragile selectors and solve pagination in a single change.

***

### 0.0.3 and earlier

No changelog was kept. Build `0.0.3` is the deployed build as of 2026-08-13; the last source commit
was 2025-11-16.
