Hacker News Scraper — Stories, Comments, Ask HN, Show HN avatar

Hacker News Scraper — Stories, Comments, Ask HN, Show HN

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Hacker News Scraper — Stories, Comments, Ask HN, Show HN

Hacker News Scraper — Stories, Comments, Ask HN, Show HN

Scrape Hacker News stories and full comment threads via the official Firebase API. Top stories, Ask HN, Show HN, jobs, new. Built for founder/VC sentiment monitoring, topic clustering, trend tracking. Clean JSON, comment tree depth-controlled.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Berkan Kaplan

Berkan Kaplan

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Hacker News Intelligence 🟠

foXLabs web & community series: GitHub trending · Wikipedia companies · Community listening

🎉 Turn Hacker News into clean, structured intelligence — no login, no API key, one row per item, with the title, author, points, comments and links. Built for developer-tool marketing, hiring and tech research.

🔍 What is the Hacker News Intelligence — and when should you use it?

Give this actor keywords or topics and it returns matching items from public Hacker News posts and comments — as clean, deduplicated rows you can filter, export or feed to an AI agent. Every run reads the source live.

Use it when you need: a HN item company list for outreach; formation / status monitoring; or a canonical registry record for KYB and due diligence.

Use something else when: you need private analytics — this reads public Hacker News content only.

🤖 Use with AI agents

Already on the Apify MCP server? Ask for this Actor by name: foxlabs/hackernews-intelligence.

Your agent can pay for its own runs. This Actor is pay-per-event with agentic payments, so an agent can discover it, run it and settle the bill over x402 (USDC on Base) or Skyfire — no Apify account or API token of its own. Billing is the same either way: per delivered record, never for errors.

Otherwise paste this into Claude, ChatGPT, Cursor or any MCP-enabled assistant:

I want to pull HN item company records using the Apify Actor `foxlabs/hackernews-intelligence`.
Input: `mode`, `query`, `searchType`, `sortBy` and more — see the Input table below. `maxResults` caps how many results are returned.
Start with: {"mode":"who_is_hiring","maxResults":1000}
Ask me what to look up, run the Actor, then summarise the rows as a table.

The machine-readable API, MCP config and OpenAPI definition live at apify.com/foxlabs/hackernews-intelligence.md.

📋 Overview

Everything you need to turn public Hacker News posts and comments into clean, structured data — in one actor, with no login, cookies or API key.

Why teams pick this actor:

  • Whole feed, one call — name or ID in, matching items out.
  • 🧹 No empty-promise columns — only fields this registry actually fills; degenerate columns are removed.
  • 🔗 Stable identifiers — every row carries the source's own IDs, ready to join across runs and to other Fox Labs actors.
  • 💰 Per-row pricing — a minimal price per delivered row, no subscription.
  • 🤖 Agent-ready — MCP + x402 agentic payments.

✨ Features

  • 🔍 Name or ID lookup — relevance-ranked name search or exact registry-ID lookup.
  • 🏢 Full entity profile — status, legal form, formation date, address and the registry’s own contact fields.
  • 🧹 Clean schema — deduplicated camelCase rows, ready for CSV/Excel/JSON.

🎬 Quick Start

curl -X POST "https://api.apify.com/v2/acts/foxlabs~hackernews-intelligence/runs?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mode":"who_is_hiring","maxResults":1000}'

🚀 Getting Started (3 steps)

  1. Choose your targets — keywords or topics.
  2. Set the capmaxResults limits how many results are returned.
  3. Run and export — get a clean dataset as JSON, CSV or Excel.

📥 Input

{"mode":"who_is_hiring","maxResults":1000}
FieldTypeDescription
modestringSearch = query the full HN archive (stories & comments) with points/date filters — unlimited depth, no 1,000-result wall. Who is hiring = parse the monthly 'Ask…
querystringSearch mode. What to search for — a brand, product, topic (e.g. "postgres", "Supabase", "rust async"). Leave empty to browse everything in the date window.
searchTypestringSearch mode. Stories = titles/links (launches, articles). Comments = the discussion text (where brand mentions and opinions actually live). Both = the two merged.
sortBystringDate (default) streams newest-first with UNLIMITED depth. Relevance uses the API's ranked order but is hard-capped at 1,000 results by the API.
minPointsintegerSearch mode, stories only. Keep only stories with at least this many upvotes (e.g. 100 = front-page material). Leave empty for all.
minCommentsintegerSearch mode, stories only. Keep only stories with at least this many comments. Leave empty for all.
authorstringSearch mode. Only items by this HN username (e.g. "patio11"). Leave empty for all authors.
threadTypestringWhich monthly thread to parse: companies hiring (default), people looking for work, or the freelancer thread.
monthstringWhich month's thread, e.g. 2026-07. Leave empty for the latest thread.
datePresetstringSearch mode: how far back to search. The archive goes back to 2007 — 'All time' really is all time.
dateFromstringYYYY-MM-DD. Only used when Date range = Custom.
dateTostringYYYY-MM-DD. Only used when Date range = Custom. Leave empty for 'today'.
feedstringFeed mode: which live list to pull.
includeCommentsbooleanFeed mode: also walk each story's comment tree into the row (slower).
commentDepthintegerFeed mode, with comments on: how deep to walk each thread (1-5).
maxCommentsPerStoryintegerFeed mode, with comments on: cap per story (1-500).
maxResultsintegerHard cap on dataset rows. Set 0 for unlimited (the sliding-window search really can walk the whole archive).

📤 Output

One row per result, saved to the dataset. Every row carries scrapedAt. Lookups that cannot be completed are reported in the run log rather than silently dropped.

FieldDescription
typeType
threadIdThread Id
threadTitleThread Title
companyCompany
headlineHeadline
roleLineRole Line
remoteRemote
urlsUrls
authorAuthor
hnUrlHn Url
createdAtCreated At
textText
fetchedAtFetched At

💼 Use cases

1. Launch monitoring — track HN discussion of a product. Input: product keywords. Output: items + points + comments. Use: a launch dashboard.

2. Hiring signals — read "Who is hiring" threads. Input: keywords. Output: items + text. Use: a hiring list.

3. Tech trends — spot rising topics on HN. Input: topic keywords. Output: items over time. Use: a trend report.

🔗 Integration

JavaScript / Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('foxlabs/hackernews-intelligence').call({"mode":"who_is_hiring","maxResults":1000});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0]);

Python

from apify_client import ApifyClient
client = ApifyClient('YOUR_TOKEN')
run = client.actor('foxlabs/hackernews-intelligence').call(run_input={"mode":"who_is_hiring","maxResults":1000})
for item in client.dataset(run['defaultDatasetId']).iterate_items():
print(item)

Automation (n8n / Zapier / Make): schedule or webhook → HTTP request to the actor API with your input → handle the JSON dataset → push to a sheet, CRM or dashboard.

📊 Pricing

Pay-per-event: per delivered record. Empty or failed lookups are never billed. View current pricing.

❓ FAQ

Do I need an account, login or API key? No. This reads public Hacker News posts and comments.

What do I search by? Keywords or topics.

How current is the data? Every run queries the source live, so results are as fresh as the registry.

What does each row cover? One HN item: title, author, points, comment count, URL and text where present.

Can I export to CSV / Excel / JSON? Yes — directly from the Apify dataset.

🐛 Troubleshooting

  • Fewer rows than expected — raise maxResults, or refine the input.
  • A name returns an unexpected entity — it matched a similar registered name; search the exact registry ID.
  • No rows for a name — try the entity’s exact legal name or its registry ID.

This actor reads public Hacker News posts and comments. Results can still contain personal data (e.g. a person’s name); personal data is protected by the GDPR and similar laws, so only process it with a legitimate basis. See Apify’s blog post on the legality of web scraping.

🤝 Support & contact

Changelog

0.2.10 — 2026-09-20 — README examples corrected against the real input schema

  • The README's code examples did not match this Actor. They used queries and maxResultsPerQuery — keys that do not exist in this Actor's input schema — with a placeholder value, and the input table listed those same phantom fields. Anyone who copied the AI-agent, cURL, JavaScript or Python example got a failing run. Every example now uses the real schema and matches the Console prefill: {"mode":"who_is_hiring","maxResults":1000}
  • The input table is regenerated from input_schema.json, so it lists the fields the Actor actually accepts.
  • Removed claims carried over from the same generator template where present: "formation / status monitoring", "a canonical registry record for KYB and due diligence", "every row carries query", and industry described as a NACE code.
  • No code, output field or pricing change.

0.2 — 2026-09-07

  • Dropped empty-promise columns. Removed emails — public Hacker News posts and comments does not carry them, so they were shipped as always-null columns. Only fields this source actually fills are now emitted.
  • Enabled AI-agent payments (x402) + rebuilt the README to the full standard (What-is / when, AI-agents + x402 agentic payments + MCP, Overview, Features, Use cases, Integration, FAQ, Troubleshooting, Support & contact).

0.0

  • Initial release: data from public Hacker News posts and comments by name or registry ID.