# Hacker News Scraper (Cheap) (`data_api/hacker-news-scraper-cheap`) Actor

Hacker News scraper that pulls stories, jobs, Ask HN and Show HN posts from news.ycombinator.com, so developers and SEO teams can track tech trends and job listings without manual browsing.

- **URL**: https://apify.com/data\_api/hacker-news-scraper-cheap.md
- **Developed by:** [Data API](https://apify.com/data_api) (community)
- **Categories:** Developer tools, Jobs, News
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.99 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are a software tools running on the Apify platform, for all kinds of web data extraction and automation use cases.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js/docs.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python/docs.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

## Hacker News Scraper

![Hacker News Scraper](cover.jpg)

Hacker News is one of the best signals in tech, but the site gives you no easy way to pull it into a spreadsheet. You are stuck refreshing the front page, copying links by hand, and losing track of what scored well yesterday. This scraper reads any HN feed for you and hands back clean rows: the headline, the outbound link, the vote score, who submitted it, how many replies it drew, and a direct link to the thread. Point it at Top, New, Best, Ask, Show, or Jobs, or feed it your own list of HN page URLs, and export the lot as JSON, CSV, or Excel.

### What you get

One row per post, with the same fields every time, so your columns line up whether you open the results in a sheet or load them into a database. Anything Hacker News does not show (vote scores and authors on job posts, for example) comes back as `null` rather than dropping out. Each row carries:

- **The post** — `headline`, `linkUrl`, `sourceDomain`, `entryKind`, and the numeric `postId`
- **The signal** — `voteScore`, `repliesTotal`, `submitter`, and `postedAgo`
- **The links** — `threadUrl` straight to the discussion, plus its `feedPosition` and a `collectedAt` timestamp

### Quick start

1. Hit **Try for free** and open the input form.
2. Pick a **Feed** — Top, New, Best, Ask HN, Show HN, or Jobs.
3. Set a **Results limit** to control how many rows you bring back (each page holds 30).
4. Press **Start**, then export the results as JSON, CSV, Excel, or XML once the run finishes.

![How it works](how-it-works.jpg)

To crawl specific pages instead of a feed, paste their URLs into **Seed URLs** — those take over from the Feed setting.

### Use cases

- **Trend tracking** — watch which stories climb the front page and which topics keep resurfacing
- **Founder and job research** — pull the Jobs and Show HN feeds to see who is hiring and what people are launching
- **Newsletter curation** — grab the day's top stories with their scores and links in one export
- **Content and SEO research** — find the domains and headlines that get traction with a technical audience
- **Sentiment and discussion analysis** — pair each post with its reply count and thread link for deeper reading
- **Dataset building** — snapshot a feed on a schedule and stack the rows into a historical record over time

### Input

| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `feedChoice` | string | Yes | Which Hacker News feed to pull: `top`, `new`, `best`, `ask`, `show`, or `jobs`. Default `top`. |
| `seedUrls` | array of strings | No | Specific HN page URLs to crawl instead of a feed. When set, these override `feedChoice`. |
| `resultsLimit` | integer | No | Largest number of rows per run. Each feed page holds 30 entries. Default `25`. |
| `timeoutSeconds` | integer | No | Seconds to wait on each request before giving up. Default `240`; raise it if you hit timeouts. |

#### Example input

```json
{
    "feedChoice": "best",
    "seedUrls": [],
    "resultsLimit": 25,
    "timeoutSeconds": 240
}
````

### Output

Every post on the feed becomes one row, and each field is always present — values Hacker News does not expose come back as `null` so the dataset stays rectangular.

#### Example output

```json
{
    "postId": 48031684,
    "feedPosition": 1,
    "headline": "Agents can now create Cloudflare accounts, buy domains, and deploy products",
    "linkUrl": "https://blog.cloudflare.com/agents-stripe-projects/",
    "sourceDomain": "cloudflare.com",
    "voteScore": 200,
    "submitter": "rolph",
    "repliesTotal": 108,
    "threadUrl": "https://news.ycombinator.com/item?id=48031684",
    "postedAgo": "3 hours ago",
    "entryKind": "story",
    "collectedAt": "2026-06-29T12:00:00.000000+00:00"
}
```

#### Output fields

| Field | Type | Description |
|-------|------|-------------|
| `postId` | integer | Numeric Hacker News identifier for the entry |
| `feedPosition` | integer | Where the entry sits in the listing order |
| `headline` | string | Title text of the post or story |
| `linkUrl` | string | Outbound link the post points to; the HN thread link for Ask and Show entries |
| `sourceDomain` | string | Host portion of the linked URL |
| `voteScore` | integer | Upvotes on the entry; `null` for job posts |
| `submitter` | string | Account name of whoever posted it; `null` for job posts |
| `repliesTotal` | integer | Comment count on the thread; `null` for job posts |
| `threadUrl` | string | Link straight to the HN comment thread |
| `postedAgo` | string | Plain-language age of the entry |
| `entryKind` | string | Category: story, job, ask, show, or launch |
| `collectedAt` | string | ISO 8601 UTC timestamp of when the row was captured |

### Tips for best results

- **Match the feed to the question.** Use New for the freshest posts, Best for all-time winners, and Jobs or Show HN when you want a single post type.
- **Keep test runs small.** Set `resultsLimit` to 20–30 while you confirm the output fits your pipeline, then raise it for the full batch.
- **Mind the 30-per-page rule.** Each feed page holds 30 entries, so the scraper pages forward automatically until it reaches your limit.
- **Expect `null` on job posts.** Jobs carry no vote score, author, or reply count, so those fields stay empty by design — that is not an error.
- **Raise `timeoutSeconds`** if a run reports timeouts; slower network conditions sometimes need a longer window.

### How can I use Hacker News data?

**How can I use the Hacker News Scraper to track trending tech stories?**
Run the Top or Best feed on a schedule and each post comes back with its vote score, reply count, and outbound link. Sort by `voteScore` or `repliesTotal` to see what is climbing, and compare snapshots over time to spot which topics keep returning to the front page.

**How can I scrape Hacker News jobs and Show HN posts?**
Set `feedChoice` to `jobs` or `show`, or drop the matching page URLs into `seedUrls`. The scraper returns each entry's headline, link, submitter, and thread URL, so you can build a feed of who is hiring or a list of fresh launches without scrolling the site by hand.

**How can I export Hacker News data to a spreadsheet for research?**
Pick a feed, set a `resultsLimit`, and run it. Every post becomes one row with consistent columns, ready to download as JSON, CSV, Excel, or XML. From there you can pivot on domains, authors, or post types to study what resonates with a technical audience.

### Is it legal to scrape data?

Our actors are ethical and do not extract any private user data, such as email addresses or private contact information. They only extract what the user has chosen to share publicly. We therefore believe that our actors, when used for ethical purposes by Apify users, are safe.

However, you should be aware that your results could contain personal data. Personal data is protected by the GDPR in the European Union and by other regulations around the world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether your reason is legitimate, consult your lawyers.

You can also read Apify's blog post on the [legality of web scraping](https://blog.apify.com/is-web-scraping-legal/).

### Support

Questions, feature requests, or a field you'd like added? Reach out at <data.apify@proton.me> and we'll get back to you.

# Actor input Schema

## `feedChoice` (type: `string`):

Pick the Hacker News listing you want. Top is the front page, New is the latest submissions, Best is the all-time highest voted, and Ask, Show, and Jobs narrow the results to those post types.

## `seedUrls` (type: `array`):

Optional. Drop in specific Hacker News page URLs to crawl instead of a feed. Anything you list here overrides the Feed setting above.

## `resultsLimit` (type: `integer`):

Largest number of rows to bring back in a single run. Each feed page holds 30 entries.

## `timeoutSeconds` (type: `integer`):

How long to wait on each request before giving up. Bump it up if you start seeing timeout errors.

## Actor input object example

```json
{
  "feedChoice": "best",
  "seedUrls": [
    "https://news.ycombinator.com/show",
    "https://news.ycombinator.com/ask"
  ],
  "resultsLimit": 25,
  "timeoutSeconds": 240
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "feedChoice": "new",
    "seedUrls": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("data_api/hacker-news-scraper-cheap").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "feedChoice": "new",
    "seedUrls": [],
}

# Run the Actor and wait for it to finish
run = client.actor("data_api/hacker-news-scraper-cheap").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "feedChoice": "new",
  "seedUrls": []
}' |
apify call data_api/hacker-news-scraper-cheap --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=data_api/hacker-news-scraper-cheap",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "Hacker News Scraper (Cheap)",
        "description": "Hacker News scraper that pulls stories, jobs, Ask HN and Show HN posts from news.ycombinator.com, so developers and SEO teams can track tech trends and job listings without manual browsing.",
        "version": "0.0",
        "x-build-id": "QHeVTi6yWRXmkpfar"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/data_api~hacker-news-scraper-cheap/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-data_api-hacker-news-scraper-cheap",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/data_api~hacker-news-scraper-cheap/runs": {
            "post": {
                "operationId": "runs-sync-data_api-hacker-news-scraper-cheap",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/data_api~hacker-news-scraper-cheap/run-sync": {
            "post": {
                "operationId": "run-sync-data_api-hacker-news-scraper-cheap",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "required": [
                    "feedChoice"
                ],
                "properties": {
                    "feedChoice": {
                        "title": "Feed",
                        "enum": [
                            "top",
                            "new",
                            "best",
                            "ask",
                            "show",
                            "jobs"
                        ],
                        "type": "string",
                        "description": "Pick the Hacker News listing you want. Top is the front page, New is the latest submissions, Best is the all-time highest voted, and Ask, Show, and Jobs narrow the results to those post types.",
                        "default": "top"
                    },
                    "seedUrls": {
                        "title": "Seed URLs",
                        "type": "array",
                        "description": "Optional. Drop in specific Hacker News page URLs to crawl instead of a feed. Anything you list here overrides the Feed setting above.",
                        "items": {
                            "type": "string"
                        }
                    },
                    "resultsLimit": {
                        "title": "Results limit",
                        "minimum": 1,
                        "maximum": 1000,
                        "type": "integer",
                        "description": "Largest number of rows to bring back in a single run. Each feed page holds 30 entries.",
                        "default": 25
                    },
                    "timeoutSeconds": {
                        "title": "Request timeout (seconds)",
                        "minimum": 50,
                        "maximum": 1200,
                        "type": "integer",
                        "description": "How long to wait on each request before giving up. Bump it up if you start seeing timeout errors.",
                        "default": 240
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
