# Hacker News Scraper (`pdcul/hacker-news-scraper`) Actor

Scrape Top, New, Best, Ask HN, Show HN and Job stories from Hacker News via the official keyless Firebase API — title, score, author, comment count and links, no proxy needed.

- **URL**: https://apify.com/pdcul/hacker-news-scraper.md
- **Developed by:** [Dan Cristian Podina](https://apify.com/pdcul) (community)
- **Categories:** AI, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Hacker News Scraper

Pull the latest **Hacker News** stories — Top, New, Best, Ask HN, Show HN, or Jobs — as clean, structured data: title, score, author, comment count, the outbound link, and the HN discussion link. Great for tracking the front page, feeding fresh tech headlines to an LLM, or building a daily digest.

### What it does

- **Every front-page list** — choose `top`, `new`, `best`, `ask`, `show`, or `job`.
- **Official, keyless source** — reads the public [Hacker News Firebase API](https://github.com/HackerNews/API), the same data that powers the site. No login, no key, no proxy.
- **Clean, structured output** — each story as `{ id, title, url, score, by, time, timeIso, descendants, type, text, hnUrl }`.
- **Reliable** — a public JSON API that's meant to be fetched by machines, so it just works from a datacenter IP on the free plan.

### Input

| Field | Type | Description |
|---|---|---|
| `listType` | `top` | `new` | `best` | `ask` | `show` | `job` | Which HN list to scrape. Default `top`. |
| `maxItems` | integer | How many stories to fetch from the top of the list (1–200). Default `30`. |

#### Example input

```json
{
  "listType": "top",
  "maxItems": 5
}
```

### Output

Each story becomes one dataset item:

```json
{
  "id": 42345678,
  "title": "Show HN: I built a thing",
  "url": "https://example.com/thing",
  "score": 512,
  "by": "pg",
  "time": 1754660580,
  "timeIso": "2026-08-08T14:03:00Z",
  "descendants": 137,
  "type": "story",
  "text": null,
  "hnUrl": "https://news.ycombinator.com/item?id=42345678"
}
```

Field notes:

- **`url`** is `null` for Ask HN and other text posts — the discussion *is* the content. Use `hnUrl` to open it on Hacker News, and `text` for the post body (HTML).
- **`descendants`** is the comment count.
- **`time`** is unix seconds; **`timeIso`** is the same instant as a UTC ISO-8601 string.

### Notes

- Only public data is accessed (the same JSON that powers news.ycombinator.com).
- Ordering follows Hacker News' own ranking for the chosen list at the moment of the run.
- The Actor makes **N+1 requests**: one for the id list, then one per story. Items that fail to fetch are skipped rather than aborting the run.

### Local development

```bash
apify run -i '{"listType":"top","maxItems":5}'
```

The fetch core (`src/hn.py`) has no Apify dependency and uses only the Python standard library, so it can be imported and run directly for quick testing.

# Actor input Schema

## `listType` (type: `string`):

Which Hacker News list to scrape. Ask HN, Show HN and Jobs mirror the corresponding front-page tabs.

## `maxItems` (type: `integer`):

How many stories to fetch from the top of the chosen list. The API returns one id list, then each story is fetched individually (so the run makes N+1 requests).

## Actor input object example

```json
{
  "listType": "top",
  "maxItems": 30
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "listType": "top"
};

// Run the Actor and wait for it to finish
const run = await client.actor("pdcul/hacker-news-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "listType": "top" }

# Run the Actor and wait for it to finish
run = client.actor("pdcul/hacker-news-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "listType": "top"
}' |
apify call pdcul/hacker-news-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pdcul/hacker-news-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TL5KOcK8mLuEOqMCy/builds/1ixJ03jzCaUAjgtL6/openapi.json
