# Hacker News Scraper: Stories, Comments, Search & Alerts (`jtpalms/hacker-news`) Actor

Hacker News front page, new, best, Ask HN, Show HN and job lists, or keyword search with date, points and comment filters. Optional comments, flattened with depth. Schedule it with "Only new items" for brand and topic alerts. Official HN and Algolia APIs. USD 0.50 per 1,000 stories.

- **URL**: https://apify.com/jtpalms/hacker-news.md
- **Developed by:** [JT Palms](https://apify.com/jtpalms) (community)
- **Categories:** News, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 stories

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hacker News Scraper: Stories, Comments, Search & Alerts

Get Hacker News as clean JSON, CSV or Excel. Pick the lists you want (front page, newest, best, Ask HN, Show HN, jobs) or search the whole HN archive by keyword, with date, points and comment filters. Add each story's comments, flattened into rows with their depth. Turn on **Only new items** and schedule it to get an alert whenever a new story mentions your brand, product or topic.

It uses the two official ways to read Hacker News: the Hacker News API that Y Combinator publishes, and the Algolia HN Search API. No login, no page scraping, no browser.

**USD 0.50 per 1,000 stories, USD 0.20 per 1,000 comments.** Failed lists or searches are free, and monitor runs with nothing new cost nothing.

### What people use it for

- **Brand and topic monitoring to Slack.** Search for your company, product or competitors, turn on **Only new items**, and schedule it every hour. Connect the task to Slack, email, Google Sheets or a webhook and each new mention lands there. Search **Stories and comments** to catch mentions deep in threads too.
- **Launch tracking.** Follow **Show HN** and **Launch HN** posts, or watch the front page for stories from your domain. See points and comment counts climb, and pull the comments to read the feedback.
- **Research datasets.** Pull every story about a topic for a date range (`"rust"`, 2024-01-01 to 2024-12-31, at least 50 points) for trend analysis, or a front-page snapshot every day for a longitudinal dataset.
- **AI agents and LLM pipelines.** Comments come as plain text with story, parent and depth, ready for summarising a discussion, sentiment analysis or RAG. Agents can call it through the Apify API or MCP server to answer "what does HN think about X?".
- **Hiring and job signals.** The **Jobs** list gives YC company job posts; keyword search on jobs finds roles by technology.

### Sample output

Real rows from the front page with comments, trimmed:

```json
[
  {
    "type": "story",
    "kind": "story",
    "id": 49880036,
    "title": "Pirating the Pirates",
    "url": "https://mubi.com/en/notebook/posts/pirating-the-pirates",
    "domain": "mubi.com",
    "author": "piotrgrabowski",
    "points": 205,
    "commentCount": 67,
    "createdAt": "2026-09-28T15:54:15.000Z",
    "hnUrl": "https://news.ycombinator.com/item?id=49880036",
    "text": null,
    "storyId": 49880036,
    "depth": 0,
    "rank": 1,
    "source": "top"
  },
  {
    "type": "comment",
    "kind": "comment",
    "id": 49881259,
    "author": "schlauerfox",
    "createdAt": "2026-09-28T17:17:00.000Z",
    "hnUrl": "https://news.ycombinator.com/item?id=49881259",
    "text": "One thing to note, is that the library of congress has the power to create the exceptions to the DMCA. The EFF lobbies for expansion of the exact powers the article is advocating for.\nhttps://www.eff.org/issues/dmca-rulemaking",
    "storyId": 49880036,
    "storyTitle": "Pirating the Pirates",
    "parentId": 49880036,
    "depth": 1,
    "source": "top"
  }
]
```

A keyword search result (`postgres`, stories and comments):

```json
{
  "type": "story",
  "kind": "show",
  "id": 49882674,
  "title": "Show HN: iCli - Postgres health reporter and index optimizer",
  "url": "https://github.com/abhiraj-ku/iCli",
  "domain": "github.com",
  "author": "abhirajabhi312",
  "points": 2,
  "commentCount": 0,
  "createdAt": "2026-09-28T18:51:27.000Z",
  "hnUrl": "https://news.ycombinator.com/item?id=49882674",
  "source": "search: postgres"
}
```

Every row has the same fields, so CSV and Excel exports line up:

| Field | What it is |
|---|---|
| `type` | `story`, `job`, `poll` or `comment`. |
| `kind` | `story`, `ask`, `show`, `launch`, `tell`, `job`, `poll` or `comment`. Text posts without a link count as `ask`, as on HN. |
| `id`, `hnUrl` | The HN item ID and its page on news.ycombinator.com. |
| `title`, `url`, `domain` | Story title, the link it points to, and that link's domain without `www.`. |
| `author` | The HN username that posted it. |
| `points`, `commentCount` | Upvotes and total comments at the time of the run. HN does not publish comment scores. |
| `createdAt` | When it was posted, ISO 8601 in UTC. |
| `text` | Plain text of an Ask HN, Show HN or job post, or of a comment. HTML is removed, links are written out in full, paragraphs become blank lines. |
| `storyId`, `storyTitle`, `parentId`, `depth` | For comments: the story, the item it replies to, and how deep it is (1 = reply to the story). Stories have depth 0. |
| `rank` | Position in the list (1 = top of the front page). Lists only. |
| `source` | Which list or search keyword produced the row. |

### How to use it

#### Lists

1. Leave **What to get** on **Lists** and pick one or more lists. **Top** is the front page ranking.
2. Optional: **Minimum points**, **Minimum comments**, **Posted after** and **Must contain (any of)** (for example `rust`, `postgres*`).
3. Set **Max stories per list or keyword** (default 100).
4. Click **Start**. Export as JSON, CSV or Excel, or read the dataset through the Apify API.

#### Keyword search

1. Set **What to get** to **Keyword search** and add one keyword or phrase per line. Quotes make an exact phrase: `"open source"`.
2. Choose **Search in**: stories, comments, both, or only Ask HN, Show HN, Launch HN, jobs, polls or the current front page. **Ask HN** also covers Tell HN and other text posts; the `kind` field tells them apart.
3. Optional filters: **Posted after** and **Posted before** (a date like `2026-01-01` or a relative value like `7 days`), **Minimum points**, **Minimum comments**.
4. **Sort** newest first (any number of results) or by relevance (up to 1,000 per keyword).

#### Comments

Turn on **Include comments**. Each story is followed by its comments, in the order HN ranks them. **Comment depth** 1 gives only top-level comments; 2 adds the replies to them. **Max comments per story** caps the count, taking top-level comments first, so a small number gives you the best of the discussion.

#### Set it up as an alert

1. Turn on **Only new items**.
2. Choose what the first run does: **All current items** returns what matches now (for a search without **Posted after**, the last 7 days); **None, just set the starting point** returns nothing and only records what exists, so you hear only about new items.
3. Save it as a task and add a **schedule**, for example every hour.
4. In the task's **Integrations** tab, connect Slack, email, Google Sheets, Zapier, Make or a webhook.

Examples:

- New stories mentioning your product: Keyword search, `yourproduct`, Search in **Stories and comments**, Only new items, hourly.
- Front page stories about your field: Lists, **Top**, Must contain `robotics`, Only new items, every 30 minutes.
- Big launches only: Keyword search with no keyword, Search in **Show HN**, Minimum points `100`, Only new items, daily.

With **Minimum points** the alert fires when a story reaches the threshold, not when it is posted, so a slow climber still triggers once. If you run two schedules on the same list or keyword with different filters, give each a **Monitor name** (under Monitoring) so they keep separate memories.

### Pricing

| What | Price |
|---|---|
| Story, Ask HN, Show HN, job or poll in the results | USD 0.0005 (USD 0.50 per 1,000) |
| Comment in the results | USD 0.0002 (USD 0.20 per 1,000) |
| A list or search that fails | Free |
| Monitor run where nothing is new | Free |

Examples: the front page (30 stories) with 20 comments each is 30 x 0.0005 + 600 x 0.0002 = USD 0.135. A brand alert that finds 5 new mentions a day costs about USD 0.08 a month. 10,000 stories for a research dataset cost USD 5.

Set a maximum cost per run in the run options. The actor stops cleanly when it gets there, and with **Only new items** anything it did not save is kept for the next run.

### Limits

- Lists hold what HN holds: Top and New up to 500 stories, Best 200, Ask, Show and Jobs up to 200.
- Relevance search stops at 1,000 results per keyword (a limit of the search API). Newest-first search has no such limit; the actor walks back through time.
- The search index can be a few minutes behind the live site, and its points and comment counts are refreshed periodically. With **Include comments**, points and comment counts are re-read live from the official API.
- **Only new items** remembers up to 20,000 item IDs per list or keyword in your `hacker-news-monitor` key-value store. A search monitor looks back 48 hours before the newest item it has seen, so a story that crosses **Minimum points** more than two days after it was posted is not reported.
- Deleted and flagged ("dead") items are skipped. The replies under a deleted comment are kept; the replies under a flagged one are not, as on HN by default.
- The actor does not log in, vote, post, or read user profiles.

### FAQ

**Is this allowed?** Yes. It only uses the official Hacker News API, which Y Combinator publishes for exactly this kind of use, and the public Algolia HN Search API that powers the search box at the bottom of every HN page. It never scrapes news.ycombinator.com pages. The content belongs to its authors and to Hacker News: quote and link back, and check HN's terms before republishing large amounts of it.

**Does it collect personal data?** It returns what HN shows publicly on each post: the text and the author's username. It does not read user profiles, "about" fields, karma or submission histories.

**Why did the first alert run return items?** With **All current items** the first run returns what matches now. Choose **None, just set the starting point** to hear only about items posted after the first run.

**Can I start over?** Delete the matching record in the `hacker-news-monitor` key-value store (Storage, Key-value stores), or use a new **Monitor name**.

**I get HTTP 429 from search.** The search API allows 10,000 requests an hour per IP address, and cloud IPs are shared. Turn on a proxy under Advanced.

**Something missing or wrong?** Open an issue with your input.

# Actor input Schema

## `mode` (type: `string`):

Lists: the stories on Hacker News lists right now (front page, newest, best, Ask HN, Show HN, jobs), in the order HN shows them. Search: stories or comments matching keywords, with date, points and comment filters, from the whole HN archive.

## `lists` (type: `array`):

Used with "Lists". Top is the front page ranking (up to 500 stories), New the newest submissions (500), Best the highest-voted recent stories (200), plus Ask HN, Show HN and Jobs.

## `searchQueries` (type: `array`):

Used with "Keyword search". One search per line, for example your brand, a product or a topic. Put a phrase in quotes for an exact match ("open source"). Each line is searched and monitored on its own. Leave empty to match everything that passes the other filters.

## `searchType` (type: `string`):

Used with "Keyword search". Which items to search: stories, comments, or both, or only Ask HN (which also covers Tell HN and other text posts), Show HN, Launch HN, jobs, polls, or the stories on the front page right now.

## `sortBy` (type: `string`):

Used with "Keyword search". Newest first can read any number of results; relevance (best match, then points) stops at 1,000 results per keyword. With "Only new items" the search always runs newest first.

## `dateFrom` (type: `string`):

Optional. Only items posted on or after this date. A date like 2026-09-01, or a relative value like "7 days" or "24 hours" (counted back from the start of each run).

## `dateTo` (type: `string`):

Optional. Only items posted on or before this date (the whole day is included), or a relative value like "1 day".

## `minPoints` (type: `integer`):

Optional. Only stories with at least this many points (upvotes). Does not apply to comments, because Hacker News does not publish comment scores.

## `minComments` (type: `integer`):

Optional. Only stories with at least this many comments.

## `keywords` (type: `array`):

Optional extra filter on the title, text and link. Whole words, not case-sensitive: "AI" does not match "said". End a word with \* for any ending ("startup\*"), or write /pattern/ for a regular expression. Useful with Lists, for example front page stories that mention your field.

## `maxItems` (type: `integer`):

At most this many results per list or per search keyword in each run, not counting comments. With "Only new items", new results beyond this number are kept for the next run.

## `includeComments` (type: `boolean`):

Add each story's comments, in the order Hacker News ranks them, as separate rows right after the story. Every comment row has the story ID, parent ID, depth and plain text. Comments are billed per comment.

## `maxCommentDepth` (type: `integer`):

How deep into reply threads to go. 1 is only top-level comments, 2 adds their direct replies, and so on.

## `maxCommentsPerStory` (type: `integer`):

At most this many comments per story. Top-level comments are taken first, in HN's ranking, then replies, so a low number gives you the best comments.

## `onlyNew` (type: `boolean`):

Remember what earlier runs returned and output only stories or comments not seen before. Schedule the actor (every hour, every morning) to use it as an alert: runs with nothing new output nothing and cost nothing. The memory is kept in your "hacker-news-monitor" key-value store.

## `firstRunOutput` (type: `string`):

What the first monitor run outputs. All: what matches now (up to the maximum; for searches without "Posted after", the last 7 days). None: nothing; it only records what is there now, so you get only items that appear after today.

## `monitorId` (type: `string`):

Optional. Keeps a separate memory of seen items. Use a different name for each schedule that watches the same list or keyword with different filters, so they do not share what counts as new.

## `maxConcurrency` (type: `integer`):

How many API requests to run at the same time. The default is fast and polite.

## `proxyConfiguration` (type: `object`):

Optional. Only needed if the search API answers HTTP 429 (too many requests from a shared IP address).

## Actor input object example

```json
{
  "mode": "lists",
  "lists": [
    "top"
  ],
  "searchType": "story",
  "sortBy": "date",
  "maxItems": 100,
  "includeComments": false,
  "maxCommentDepth": 2,
  "maxCommentsPerStory": 30,
  "onlyNew": false,
  "firstRunOutput": "all",
  "maxConcurrency": 10
}
```

# Actor output Schema

## `results` (type: `string`):

All output rows in the default dataset (JSON, CSV, Excel via the format parameter).

## `summary` (type: `string`):

Counts and per-input status for this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("jtpalms/hacker-news").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("jtpalms/hacker-news").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call jtpalms/hacker-news --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jtpalms/hacker-news"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kS7jhvxCGz551tdzn/builds/O0gZucc6t7N49cymR/openapi.json
