IndieHackers Scraper
Under maintenancePricing
from $10.00 / 1,000 results
IndieHackers Scraper
Under maintenanceExtract IndieHackers product data, founder stories, and revenue milestones without an API key. Perfect for market research, VC deal sourcing, and SaaS competitive analysis.
Pricing
from $10.00 / 1,000 results
Rating
0.0
(0)
Developer
Web Data Labs
Maintained by CommunityActor stats
3
Bookmarked
45
Total users
8
Monthly active users
8 days ago
Last modified
Categories
Share
Breaking change in v2.0: IndieHackers rebuilt their site and retired the search index this actor used. v2.0 scrapes the public server-rendered site instead. Lost versus v1: free-text query search, numFollowers, followers sorting, taxonomy fields (verticals, revenueModels, funding, platforms, allTags), real timestamps (startDate, createdAt, publishedAt), and founderUserIds (replaced by founders[] objects). Kept: name, tagline, description, url, websiteUrl, avatarUrl, monthlyRevenueUSD, twitterHandle, updatedAt, scrapedAt. New: revenueIsExact, revenueUpdatedAt, postsCount, facebookUrl, claimed, categories[], businessModels[], founders[] (username/name/role/avatar).
Extract the IndieHackers product catalog — self-reported monthly revenue, founders, taglines, categories — from the public site. No API key needed.
What it does
Enumerates IndieHackers products two ways:
- Listing mode (default): scrapes the public
/productsdirectory slices — category, business model, revenue range, and four sort orders. Anonymous browsing is capped at roughly 18 products per filter combination behind IndieHackers' signup wall. For the revenue sorts (highest-revenue/lowest-revenue) the site-wide slice is queried with a $10M/mo ceiling — claims above that can't be real anyway — so its 18 cards are the true revenue head rather than the usual troll-dominated mix. WhenmaxItemsneeds more candidates than the head holds, the actor walks category top-18 slices in batches and stops as soon as the pool covers the request (a substringqueryfilter still walks all of them — its match rate can't be estimated). If the delivered tail draws on a partial category walk, the run status says exactly what the ranking covers. Only products whose page carries a usable revenue figure are returned — implausible self-reported values (troll claims above $10M/mo) are skipped, not billed. Good for quick, fresh, filtered pulls. - Catalog mode (
includeFullCatalog: true): enumerates every product slug in the public IndieHackers sitemap (~30k products, ~8% dead) and scrapes each product page. This is the bulk-extraction path.
Each product page yields: name, tagline, description, website, avatar, self-reported monthly revenue (optionally the exact figure from the product's /revenue subpage), posts count, founders (username, name, role, avatar), twitter/facebook links, claimed flag, and the latest update date.
Use cases
- VCs & accelerators — source bootstrapped SaaS deals with self-reported revenue traction
- Market research — map products by category or business model
- Competitive analysis — track competitors' reported revenue and update cadence
- Lead lists — build outreach lists of founders with public revenue
Input
| Field | Type | Description |
|---|---|---|
categories | string[] | Category filter (listing mode). Each category caps at ~18 results anonymously. |
businessModel | string | free, ads, commission, transactions, or subscriptions (listing mode). |
minRevenue | integer | Minimum self-reported monthly revenue (USD). Server-side in listing mode, post-parse in catalog mode. |
maxRevenue | integer | Maximum monthly revenue (listing mode). |
sortBy | string | highest-revenue (default), lowest-revenue, recently-updated, recently-added. |
includeFullCatalog | boolean | Enumerate the full sitemap (~30k products) instead of listing slices. |
fetchExactRevenue | boolean | Also fetch each product's /revenue subpage for the exact figure + last-updated date. 2 requests per product. |
maxItems | integer | Max products to return. Default 100, max 30000. |
proxyConfiguration | object | Apify proxy settings. |
Legacy input mapping (v1 → v2)
Old inputs are accepted and mapped, with a warning in the run log:
| v1 input | v2 behavior |
|---|---|
query: "text" | Post-fetch substring filter over name/tagline/description. Public search no longer exists upstream. |
sortBy: "newest" | Mapped to recently-added. |
sortBy: "revenue" | Mapped to highest-revenue. |
sortBy: "followers" | Mapped to highest-revenue + warning — the followers metric is gone upstream. |
minRevenue, maxItems | Unchanged. |
Example input
{"categories": ["ai", "productivity"],"minRevenue": 1000,"sortBy": "highest-revenue","maxItems": 50,"fetchExactRevenue": true}
Output
Each item contains:
{"productId": "nomad-list","name": "Nomads.com","tagline": "The best cities to live and work remotely for digital nomads","description": "To accelerate the freedom of global movement enabled by remote work","url": "https://www.indiehackers.com/product/nomad-list","websiteUrl": "https://nomads.com","avatarUrl": "https://storage.googleapis.com/indie-hackers.appspot.com/product-avatars/nomad-list/200x200.webp","monthlyRevenueUSD": 35000,"revenueIsExact": true,"revenueUpdatedAt": "2024-05-07","postsCount": 2,"twitterHandle": "levelsio","facebookUrl": null,"founders": [{"username": "levelsio", "name": "Pieter Levels", "role": "Founder", "avatarUrl": "..."}],"claimed": false,"categories": ["travel"],"businessModels": [],"updatedAt": "2025-03-09","scrapedAt": "2026-09-22T12:00:00.000Z"}
Field notes:
monthlyRevenueUSD— self-reported by founders;nullwhen unreported or implausible (troll values above $10M/mo are rejected). Abbreviated on the product page ($14K→ 14000,$2.6MM→ 2600000); exact whenfetchExactRevenuehits the/revenuesubpage (revenueIsExact: true, plusrevenueUpdatedAt). In revenue-sorted listing runs, products with no usable revenue figure are omitted entirely.categories/businessModels— populated only for products discovered through a filtered listing slice; product pages carry no taxonomy markup.updatedAt— latest milestone post date, not a record-modified timestamp.claimed—falsewhen the page shows IndieHackers' "not been claimed" placeholder.- v1 fields
numFollowers,startDate,createdAt,publishedAt,founderUserIds,verticals,revenueModels,funding,platforms,allTags— removed; no public source exists.
Pricing
Pay-per-event: you only pay for product rows actually delivered to the dataset. No fixed cost, no compute charges. Dead product pages and filter misses are never billed.
Reliability
The actor fails loudly instead of returning empty success when IndieHackers is unreachable, blocked, or the markup drifts — a run that can't reach the site ends FAILED, not SUCCEEDED-0. In catalog mode a named key-value store caches the sitemap inventory for 7 days.
Notes
- Data comes from indiehackers.com public pages. Revenue is whatever founders voluntarily disclosed — many report nothing, and self-reported figures are unverified.
- Anonymous listing slices cap at ~18 products each; use
includeFullCatalogfor bulk extraction. - Respect the data: don't spam founders or scrape for malicious purposes.
⭐ Leave a Review
If this actor saved you time, a quick review on the Apify Store helps other developers find it. Takes 30 seconds and means a lot.
Found this actor useful? Leave a review - it helps a lot. Something broken? Report it on the Issues tab.