Hacker News Scraper — Stories, Comments & Who Is Hiring
Pricing
from $0.70 / 1,000 story scrapeds
Hacker News Scraper — Stories, Comments & Who Is Hiring
Get Hacker News stories from the official HN API: top, new, best, Ask HN, Show HN and jobs. Filter by keyword and minimum score, optionally with comments. Clean text, direct links, JSON/CSV/Excel.
Pricing
from $0.70 / 1,000 story scrapeds
Rating
0.0
(0)
Developer
KeyMan98
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Get live Hacker News stories — Top, New, Best, Ask HN, Show HN, and the monthly Who Is Hiring thread — with comments, filtered by keyword or score. Pulls data through Hacker News' own official, public Firebase API (hacker-news.firebaseio.com) — no HTML scraping of news.ycombinator.com, no third-party data provider — and exports the results to CSV, Excel, or JSON.
What you get (output fields)
For each story, one dataset row with:
id/type— the Hacker News item ID, and its type as reported by the API (e.g.story,job).title/url— the story's title and, if it links out, its external URL.text— for text-only posts (most Ask HN, Show HN, and job posts): the body as clean plain text — HTML entities decoded, paragraphs turned into line breaks, links kept as plain URLs.by— the author's public Hacker News username (public data from the official API).score— points at the time of the run.descendants— total comment count on the story, as reported by HN (independent of how many were actually fetched).time— when the story was submitted (ISO 8601, UTC).hnUrl— the story's discussion page on news.ycombinator.com.domain— the domain ofurl, withoutwww.(null for text-only posts).comments— a list of comments (id,by,text,time,parent,depth), only if "Include comments" was on; otherwise null.error— set only on rows that could not be resolved (see below); null otherwise.
Who it's for
- Tech trend / product monitoring — track what's on the Hacker News front page for a brand, technology, or company name.
- Research and datasets — pull structured HN discussions (stories + comment trees) for analysis or as training/evaluation data.
- "Who is Hiring" tracking — scan the monthly Ask HN hiring thread for a role, technology, or "remote" without reading it by hand.
- Alerts — run on a schedule and filter by keyword/score to get notified only when something relevant appears.
This Actor does not do full-text search over HN's history — see "Limitations" below.
How to use
- Story type —
top,new,best,ask(Ask HN),show(Show HN), orjob(Who is Hiring / job posts). Ignored if "Item IDs" is set. - Max items — stop after this many stories pass the filters below.
- Keywords (optional) — keep only stories whose title or text contains at least one of these words (case-insensitive).
- Minimum score (optional) — keep only stories with at least this many points.
- Include comments — fetch each story's comment tree too (breadth-first, shallow replies first).
- Max comments per story — cap on how many comments to fetch per story, if comments are included.
- Item IDs (optional) — specific HN item IDs to fetch directly instead of a story list; overrides "Story type".
Input example (JSON)
{"storyType": "top","maxItems": 30,"keywords": [],"minScore": null,"includeComments": false,"maxCommentsPerStory": 20,"itemIds": []}
Output example (JSON)
{"id": 1,"type": "story","title": "Y Combinator","url": "http://ycombinator.com","text": null,"by": "pg","score": 61,"descendants": 19,"time": "2006-10-09T18:21:51+00:00","hnUrl": "https://news.ycombinator.com/item?id=1","domain": "ycombinator.com","comments": null,"error": null}
If an item ID is not found or can't be fetched
That entry becomes one error row: error is set to a short explanation, every other field is null. The run does not fail, the rest of the list keeps running, and you are not charged for that row. A story that's been deleted or marked dead on Hacker News is instead skipped entirely — no row at all, since that's HN's own removal, not a fetch problem.
Pricing
Pay only for stories actually returned, after filters — nothing charged for an item ID that could not be resolved, and nothing extra for comments (they're included in the story's own charge). Pricing model: pay-per-event.
| Event | When it's charged | Price |
|---|---|---|
item-scraped | a story was returned in the results (after filters) | 0.0007 USD |
Limitations
- No full-text search over HN's history. This Actor only reads live lists (top/new/best/ask/show/job, each up to 500 items) and specific item IDs you already know. For searching HN's entire archive by keyword, use the separate, community-run Algolia HN Search API (
hn.algolia.com/api) — not used here. - No per-user data. Only the public
byusername on each story/comment is returned — no karma, no submission history. - Deleted or dead stories are skipped, never returned as a row.
commentsreflects what was actually fetched (bounded by "Max comments per story");descendantsis HN's own total comment count, and the two can differ on a busy story.- Job posts (
storyTypejob) have noscorefield — setting "Minimum score" always excludes them, regardless of the value.
FAQ
Am I charged if an item ID is not found?
No. You are only charged for stories actually returned in the results, after filters.
Where does the data come from?
The official, public Hacker News Firebase API (hacker-news.firebaseio.com/v0), documented at github.com/HackerNews/API. No HTML page of news.ycombinator.com is ever fetched or parsed.
Does fetching comments cost extra?
No. Comments are included in the same per-story charge — only whether "Include comments" is on changes what's inside the comments field, not the price.
How do keyword and score filters work?
keywords matches as a whole word, case-insensitively, against the story's title and text — kept if at least one keyword appears in either. "AI" won't match inside "said" or "email", and keywords like "C++" or ".NET" still match correctly at the edge of a sentence. minScore keeps only stories with at least that many points — note that job posts (storyType job) have no score at all, so setting minScore always excludes them. Both filters apply before anything is added to the dataset, so filtered-out stories are never charged.
How often is the data updated?
Every run re-fetches live from the HN API — results reflect the score/comment count at run time, not a cached snapshot.
Can I monitor Hacker News on a schedule?
Yes. Set "Story type" and a keyword or score filter, then run this Actor on a schedule (Apify's built-in scheduler).
Can I re-check specific stories later?
Yes — pass their IDs in "Item IDs" (found in a story's hnUrl, after ?id=). This ignores "Story type" and fetches each ID once, even if listed twice.
Can I use this through the Apify API or an MCP server?
Yes, like any Apify Actor — through the standard Apify API, or through the Apify MCP server if you use Claude, Cursor, or another MCP-enabled client.
Export
Results can be downloaded from the Apify dataset as JSON, CSV, or Excel, or accessed via the Apify API.