# Changelog of Hacker News Jobs Scraper | Who Is Hiring Threads (`tqm/hackernews-job-scraper`) Actor

- **URL**: https://apify.com/tqm/hackernews-job-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/tqm/hackernews-job-scraper.md

## Changelog

### 2026-09-10 — HN-ZERO-01: a zero-posting run now fails and says why

This was the **last of the seven listed actors with no zero-row contract at all**, and the seventh
silent zero in the portfolio's history. Measured on the live Store actor 2026-09-10: an impossible
`techStackFilter` returned **SUCCEEDED, `[]`, 7.0s** — a buyer's only signal being an empty Output
tab, with no way to tell too-narrow filters from a markup change.

- **Wired up `apify-actor-kit`** (`src/kit/`, vendored) and added the `failOnZeroResults` input,
  default `true` — the same contract the other five board actors already had. This actor is now a
  registered consumer in `apify-actor-kit/sync-to-actors.sh`.
- **Instrumented the causes**: `commentsSeen` (what HN served), `postingsParsed` (what survived
  validation), a per-filter rejection breakdown, and a `blocked` count in `failedRequestHandler`,
  which previously only logged. A zero run now reports `filtered_out` / `blocked` / `none_usable` /
  `source_empty` rather than one green lie.
- **The "no threads found" path was a second silent zero**, one level up: `threads.length === 0`
  called `Actor.exit()` bare. It now goes through the same contract — HN publishes a hiring thread
  every month, so finding none for every month requested is a real failure.

**Caught by running it, not by reading it.** The first cut passed
`Dataset.getInfo().itemCount` as the pushed count. That value is **eventually consistent**: a run
that had just pushed 40 rows read back **0**, so a perfectly healthy run reported `parsed_not_pushed`
and FAILED. It was harmless while it only fed a log line and became a bug the moment it decided
whether the run failed. The count is now `itemsPushed`, the same synchronously-reserved counter that
enforces the spend cap.

> ⚠️ **A value that was only ever logged has no track record.** Promoting one to a control decision
> is a change of contract, not a change of caller.

Two deliberate non-changes, both of which would have been wrong:

- **The reply skip is NOT counted as a buyer filter.** `includeReplies` is off by default, so
  counting it would make nearly every zero-row run blame a filter the buyer never set — and
  `diagnoseZeroRows` tests the filter breakdown before every other cause. A skipped reply is simply
  not a parsed posting.
- **`requested` is not wired to `maxItems`.** On the five board actors it names a *board* ceiling, so
  a short set means the board is exhausted. Here the ceiling is however many postings N months of
  threads hold — 3 months returns ~750 against a default cap of 1000 every time, which is not a
  shortfall worth reporting.

### Unreleased — 2026-09-04 — N4: one dataset item is now one job posting

**Breaking output change.** The actor emitted one item per *thread*, with every posting nested in
`jobPostings[]`. It now emits one item per posting, with the thread fields denormalized onto each row.

The old shape was defensible on the merits — the thread *is* the unit HN publishes — and it stopped
being defensible the moment the actor had a price. **Apify bills per dataset item.** The verified
August 2026 thread delivered 230 postings as **one billable result**. There is no price that works
in both directions: cover support at 230-postings-per-result and the listing reads as absurd; price
per posting and the whole run bills as $0.001.

At the portfolio price of $1.50/1,000 that same run now bills as **230 results = $0.345**, against
$0.0018 of measured cost.

> **The billing unit is part of the data model, not a decision downstream of it.** This shape was
> chosen before the actor had a price, and it was correct until the price existed.

#### Also changed, all for the same reason

| | |
|---|---|
| `includeReplies` | default **true → false**. A reply is a discussion comment, not a vacancy, and it would now bill as a result. It also bypasses `techStackFilter` and `remoteOnly` entirely (HN-DOM), so a filtered run was returning unfiltered replies |
| `maxItems` | **new**, default 1000. Per-result pricing means the row count is the bill, and a 12-month run over ~230-posting threads is a few thousand rows. Enforced by slicing **before** the push - Apify bills what reaches the dataset, so a cap applied afterwards is a cap the buyer already paid for |
| `title` | **removed**. Hard-coded null on every row since the actor was written. Advertising a field that is never populated is the same defect as a filter that never narrows |
| `rawText` | **removed**. It was `description.substring(0, 500)` - no information the full field lacks, on every row of a 230-row payload |

#### Output fields

Added `threadTotalComments` (was `totalComments`, at thread level). Every row now carries `source`,
`commentId`, `commentUrl`, `company`, `location`, `remote`, `salary`, `skills`, `description`,
`depth`, `parentCommentId`, the four `thread*` fields and `scrapedAt`.

Two dataset **views** added - `overview` and `full` - which the flat shape makes possible; a nested
array cannot be rendered as a table.

#### Not affected

TQM's own nightly pipeline does **not** consume this actor. It scrapes the same threads through its
own Python `hackernews.py` via the Algolia API, and is not registered in `apify_ingest.ACTORS`.
Checked before the reshape rather than assumed.

`tsc --noEmit` clean. Closes IMPROVEMENT-PLAN items 1 and 7.

### Unreleased — 2026-08-13

Documentation only. **No behavior changes**; build `0.0.16` is unmodified.

#### Added

- Full documentation suite: `docs/ARCHITECTURE.md`, `docs/DEVELOPMENT.md`, `docs/HN-DOM.md`,
  `docs/IMPROVEMENT-PLAN.md`, and `CLAUDE.md`.
- `.actor/dataset_schema.json` — the output surface was previously undeclared, and now documents the
  nested `jobPostings[]` structure.
- README rewritten as a buyer-facing listing with an explicit nested-output warning.

#### Fixed (documentation)

- **The README documented an input that does not exist.** It described `includeDescription`; the
  actual input is `includeReplies`. Corrected.

#### Verified against live HackerNews

Platform run on 2026-08-13 (`monthsToScrape: 1`, `includeReplies: false`):

| Metric | Value |
|---|---:|
| Dataset items | **1** (a thread) |
| Job postings inside it | **230** |
| Thread comments | 350 |
| Cost | $0.0018 |
| Duration | 13 s |

Roughly **$0.008 per 1,000 job postings** — the cheapest source in the portfolio by an order of
magnitude, because one HTTP fetch yields hundreds of postings.

#### Findings recorded (not fixed)

- **Output is nested — one dataset item per thread**, unique in this portfolio. On Apify's
  pay-per-result model, which bills per dataset item, a run delivering 230 postings bills as **one**
  result. This blocks sensible pricing and breaks consumer expectations.
- **`title` is hard-coded `undefined`** — advertised, never populated, 0 of 230 postings.
- The Algolia thread lookup takes `hits[0]` without verifying the author is `whoishiring`, that the
  title matches exactly, or that the month is right.
- The `company` extractor's third pattern guesses "first three words if all capitalized" and produced
  `"Location: Charleston SC"` as a company name. Fill rate 219/230, accuracy lower.
- `techStackFilter` and `remoteOnly` are applied **only at depth 0**, so with the default
  `includeReplies: true` replies bypass filtering entirely.
- A run finding no threads exits **SUCCEEDED**.
- `rawText` is `description.substring(0, 500)` — pure duplication, inflating dataset size.
- Comment depth is derived from an indent image's pixel width (`width / 40`). If HN moves to CSS
  indentation every comment silently becomes depth 0, and `includeReplies: false` stops working.
- `getParentCommentId` is O(n²) — ~60,000 depth computations on a 350-comment thread.
- There is no result cap; run size is governed only by `monthsToScrape`.

#### What this actor gets right

- **It uses the Algolia HN Search API to locate threads** rather than scraping HN's search — a
  documented public API, and a large part of why a run costs a fifth of a cent. (Compare the sibling
  RemoteOK actor, which drives Puppeteer against a site that publishes a JSON API.)
- **`commentId` is HN's own identifier** — a genuine stable ID rather than a derived slug. The
  strongest identity story in the portfolio.
- Plain HTTP, no browser, correctly matching a server-rendered source.
- Every field is derived from the individual comment's own text — correctly scoped.

All findings are itemized and prioritized in `docs/IMPROVEMENT-PLAN.md`. The headline item is
flattening the output: it is a prerequisite for pricing the actor and it changes the shape every
other item is written against.

***

### 0.0.16 and earlier

No changelog was kept. Build `0.0.16` is the deployed build as of 2026-08-13; the last source commit
was 2025-11-14.

Note the repo and `.actor/actor.json` are named `hackernews-scraper`, while the deployed actor is
`moisecristian2/hackernews-job-scraper`.
