Hacker News Scraper – Stories, Comments, Search & Who's Hiring avatar

Hacker News Scraper – Stories, Comments, Search & Who's Hiring

Pricing

from $0.20 / 1,000 hn items

Go to Apify Store
Hacker News Scraper – Stories, Comments, Search & Who's Hiring

Hacker News Scraper – Stories, Comments, Search & Who's Hiring

Hacker News data from the official API and Algolia HN Search: top/new/best/Ask/Show/job stories, keyword and brand monitoring with only-new alerts, full comment trees flattened with depth, and the monthly Who is hiring thread parsed into job posts.

Pricing

from $0.20 / 1,000 hn items

Rating

0.0

(0)

Developer

Cemal Atakli

Cemal Atakli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Get Hacker News data as clean JSON: front page and new/best/Ask HN/Show HN/job stories, keyword and brand monitoring across stories and comments, full comment trees flattened with depth and parent IDs, and the monthly "Ask HN: Who is hiring?" thread parsed into job posts (company, roles, location, remote, salary, links).

It uses only the official Hacker News API (Firebase) and the public Algolia HN Search API. No browser, no proxies, no login, so it is fast and costs $0.20 per 1,000 items.

What it does

  • Story lists. top (front page ranking), new, best, ask, show, job, with title, URL, domain, points, comment count, author, time, text and rank.
  • Search & brand monitoring. Search stories, comments or both through Algolia HN Search. Filter by date range (2026-09-01 or 7 days), min points, min comments, author, Show HN / Ask HN / front page. Sort by date (no 1,000-result cap: the Actor pages by time) or relevance. Strict matching removes Algolia's typo matches (rust → trust).
  • Only-new mode for alerts. Schedule it hourly: each run returns only stories, comments or job posts that were not returned before. Works for searches, lists, comment threads and the hiring thread.
  • Full comment trees. Give item IDs or URLs and get the item plus every comment as one row each, with depth, parentId, rankAmongSiblings, replyCount, plain text (HTML converted, links kept in links). Limit depth or count. Top-level comments follow HN's own ranking.
  • "Who is hiring?" parser. Finds the thread for any month (latest, 2026-09, March 2019), splits it into job posts and extracts best-effort fields from the Company | Role | Location | Remote | Salary | URL header: company, ycBatch, roles, locations, remote, workArrangement, employmentType, salaryText / salaryMin / salaryMax / salaryCurrency, visaSponsorship, companyUrl, applyUrls, emails. The full raw text is always kept. Filter by keywords (python, Berlin) or remote only.
  • Robust: retries with backoff, per-item errors become free error rows, honest User-Agent, max 4 parallel requests to Algolia.

Use cases

  • Brand and competitor monitoring: get an alert (Slack, email, webhook) whenever your product, domain or competitor is mentioned in a story or comment.
  • Job search and recruiting data: every Who is hiring post as structured rows; track remote jobs, salaries and stacks month by month.
  • Market and trend research: what HN says about Rust, Postgres or AI agents over time; top stories by points for a date range.
  • AI / LLM datasets: clean comment threads with depth and parent IDs for summarisation, sentiment analysis or RAG.
  • Launch tracking: follow your Show HN thread and get only the new comments each hour.
  • Newsletters and dashboards: the daily front page or best stories as JSON, CSV or Excel.

Input examples

Front page (default, one click):

{ "mode": "lists", "lists": ["top"], "maxItemsPerList": 30 }

Brand monitoring, scheduled hourly:

{
"mode": "search",
"searchQueries": ["Apify", "\"web scraping\"", "crawlee"],
"searchTags": "storyOrComment",
"createdAfter": "7 days",
"onlyNew": true
}

A thread with all comments:

{ "mode": "items", "itemIds": ["https://news.ycombinator.com/item?id=8863"], "maxCommentDepth": 0 }

Remote Python jobs from this month's Who is hiring:

{ "mode": "whoIsHiring", "hiringMonth": "latest", "hiringKeywords": ["python"], "remoteOnly": true }
FieldDefaultNotes
modelistslists, search, items, whoIsHiring
lists / maxItemsPerList["top"] / 30top, new, best, ask, show, job
searchQueries–One per line; quotes for phrases. Rows get matchedQuery
searchTagsstorystory, comment, storyOrComment, showHn, askHn, poll, frontPage, any
sortBydaterelevance is capped at 1,000 results per query by Algolia
createdAfter / createdBefore–2026-09-01 or 7 days
minPoints / minComments0Stories only
author–HN username; works without a query
maxResultsPerQuery100
exactMatchtrueNo typo matches; every word must start a word in the text
itemIds–IDs or item URLs (items mode)
includeCommentsfalseAlso comment trees for list / search stories
maxCommentDepth / maxCommentsPerItem0 / 00 = all
commentSourceautoauto = Algolia tree in one request (falls back to the official API), firebase = official API only
hiringMonthlatest2026-09, September 2026
hiringKeywords / hiringKeywordsMode / remoteOnly– / any / falseFilters for job posts
onlyNew / stateKeyfalse / –Monitoring memory
maxItems0Total cap on rows
includeHtmlfalseAdds HN's original HTML as textHtml

Output examples

One dataset row per item. The Output tab has Stories & search results, Comments, Who is hiring – job posts and Errors views. Full examples are in SAMPLE_OUTPUT.json.

Story (lists mode):

{
"type": "story", "id": 49940394, "list": "top", "rank": 1,
"title": "Newgrounds.com – A community of games, music, and art",
"url": "https://www.newgrounds.com/", "domain": "newgrounds.com",
"points": 211, "numComments": 55, "author": "azhenley",
"createdAt": "2026-10-03T00:55:25Z", "createdAtUnix": 1790988925,
"parentId": null, "replyCount": 25, "text": null, "links": [],
"hnUrl": "https://news.ycombinator.com/item?id=49940394"
}

Comment (items mode):

{
"type": "comment", "id": 9272, "author": "dhouston", "createdAt": "2007-04-05T16:47:01Z",
"storyId": 8863, "storyTitle": "My YC app: Dropbox - Throw away your USB drive",
"parentId": 9224, "depth": 2, "rankAmongSiblings": 1, "replyCount": 1,
"text": "1. re: the first part, many people want something plug and play. …",
"links": [], "hnUrl": "https://news.ycombinator.com/item?id=9272", "commentSource": "algolia"
}

Job post (Who is hiring mode):

{
"type": "job-post", "id": 49922584, "month": "2026-10", "position": 2,
"threadTitle": "Ask HN: Who is hiring? (October 2026)",
"headerLine": "PrairieLearn (Remote US) — Full-Stack Software Engineer — TypeScript / Postgres / React / AI",
"company": "PrairieLearn", "companyUrl": "https://www.prairielearn.com",
"roles": ["Full-Stack Software Engineer", "AI"], "locations": ["US"],
"remote": true, "workArrangement": ["remote"], "employmentType": ["full-time"],
"salaryText": "$100k-$180k", "salaryMin": 100000, "salaryMax": 180000, "salaryCurrency": "USD", "salaryPeriod": "year",
"otherHeaderParts": ["TypeScript", "Postgres", "React"],
"applyUrls": ["https://www.prairielearn.com/jobs-ashby?..."],
"text": "PrairieLearn (Remote US) — Full-Stack Software Engineer — … (full post)",
"hnUrl": "https://news.ycombinator.com/item?id=49922584"
}
  • Search results also have matchedQuery, tags and, for comments, storyId, storyTitle, storyUrl.
  • Errors: "type": "error" with input and error. Not charged.
  • The key-value store record OUTPUT holds the run summary (rows by type, API requests, stop reason).

Pricing

Pay per event. You pay only for rows saved.

EventPrice
Hacker News item (story, comment or job post)$0.0002 ($0.20 / 1,000)
Actor start$0.00005

Compared with other Store Actors (public Store prices, 2026-10-01):

ActorPrice per 1,000 itemsSearch / monitoringComment treesWho is hiring parser
Hacker News Scraper (this Actor)$0.20yes, only-newyes, flattened with depthyes
gentle_cloud/hacker-news-scraper$0.20–––
ryanclinton/hackernews-search$5.00yes––

Set Maximum cost per run in the run options to cap spending. The Actor stops cleanly at the limit, and in only-new mode the rows it could not save come back on the next run.

FAQ

Where does the data come from? From the official Hacker News API (hacker-news.firebaseio.com, live) and Algolia's public HN Search API (hn.algolia.com), which Hacker News links to for search. No HTML is scraped.

How does only-new mode work? Returned IDs are stored per query / list / thread in a named key-value store derived from your input (hacker-news-data-…), or from stateKey if you set one. The first run returns the current results; later runs return only new ones. For a search sorted by date the Actor stops paging as soon as it reaches results from an earlier run, so scheduled runs are cheap.

Why does a search return fewer results than Algolia's nbHits? Algolia matches typos and prefixes. With exactMatch on (default) typo tolerance is off and each word must appear at the start of a word, so rust no longer matches trust, but postgres still finds PostgreSQL. Turn it off to get Algolia's raw matching. Dropped results are free.

Are comment trees complete? auto reads the whole tree from Algolia in one request (fast; Algolia can lag a few minutes behind HN on brand-new comments, and lists replies by time). When the Algolia tree is clearly incomplete, the Actor switches to the official API. Choose firebase for a live, exact HN ordering at every level. Deleted and flagged comments are skipped (their replies are kept) unless includeDeletedComments is on.

How accurate is the Who is hiring parser? It is best-effort. On the October 2026 thread, company was found for 95% of 183 posts, location for 85%, remote/onsite for 93%, roles for 71% and salary text for 42% (numeric min/max for 37%; many posts do not state a salary). headerStructured: false marks posts that do not follow the A | B | C format. The full text is always included, so you can run your own parser or an LLM on it.

Does it include personal data? Rows contain public HN usernames, which are part of the public data. Job posts may contain contact emails that companies posted for applications. Use the data in line with the GDPR and HN's guidelines; do not use it for spam.

Use with AI agents / Apify MCP

  • Apify MCP server: add gazidev/hacker-news-data to your MCP client (Claude Desktop, Cursor, VS Code) via https://mcp.apify.com?actors=gazidev/hacker-news-data. An agent can call it with {"mode":"search","searchQueries":["your product"],"createdAfter":"7 days"} and read what HN says about it.
  • API: POST https://api.apify.com/v2/acts/gazidev~hacker-news-data/run-sync-get-dataset-items?token=... with the input JSON returns the rows directly, handy as a "Hacker News tool" for LangChain or LlamaIndex agents.
  • Scheduled alerts: create a Task with onlyNew: true, schedule it, and connect the Slack, email, Google Sheets or webhook integration.

More from the same developer