IndieHackers Scraper avatar

IndieHackers Scraper

Under maintenance

Pricing

from $10.00 / 1,000 results

Go to Apify Store
IndieHackers Scraper

IndieHackers Scraper

Under maintenance

Extract IndieHackers product data, founder stories, and revenue milestones without an API key. Perfect for market research, VC deal sourcing, and SaaS competitive analysis.

Pricing

from $10.00 / 1,000 results

Rating

0.0

(0)

Developer

Web Data Labs

Web Data Labs

Maintained by Community

Actor stats

3

Bookmarked

45

Total users

8

Monthly active users

8 days ago

Last modified

Categories

Share

Breaking change in v2.0: IndieHackers rebuilt their site and retired the search index this actor used. v2.0 scrapes the public server-rendered site instead. Lost versus v1: free-text query search, numFollowers, followers sorting, taxonomy fields (verticals, revenueModels, funding, platforms, allTags), real timestamps (startDate, createdAt, publishedAt), and founderUserIds (replaced by founders[] objects). Kept: name, tagline, description, url, websiteUrl, avatarUrl, monthlyRevenueUSD, twitterHandle, updatedAt, scrapedAt. New: revenueIsExact, revenueUpdatedAt, postsCount, facebookUrl, claimed, categories[], businessModels[], founders[] (username/name/role/avatar).

Extract the IndieHackers product catalog — self-reported monthly revenue, founders, taglines, categories — from the public site. No API key needed.

What it does

Enumerates IndieHackers products two ways:

  • Listing mode (default): scrapes the public /products directory slices — category, business model, revenue range, and four sort orders. Anonymous browsing is capped at roughly 18 products per filter combination behind IndieHackers' signup wall. For the revenue sorts (highest-revenue / lowest-revenue) the site-wide slice is queried with a $10M/mo ceiling — claims above that can't be real anyway — so its 18 cards are the true revenue head rather than the usual troll-dominated mix. When maxItems needs more candidates than the head holds, the actor walks category top-18 slices in batches and stops as soon as the pool covers the request (a substring query filter still walks all of them — its match rate can't be estimated). If the delivered tail draws on a partial category walk, the run status says exactly what the ranking covers. Only products whose page carries a usable revenue figure are returned — implausible self-reported values (troll claims above $10M/mo) are skipped, not billed. Good for quick, fresh, filtered pulls.
  • Catalog mode (includeFullCatalog: true): enumerates every product slug in the public IndieHackers sitemap (~30k products, ~8% dead) and scrapes each product page. This is the bulk-extraction path.

Each product page yields: name, tagline, description, website, avatar, self-reported monthly revenue (optionally the exact figure from the product's /revenue subpage), posts count, founders (username, name, role, avatar), twitter/facebook links, claimed flag, and the latest update date.

Use cases

  • VCs & accelerators — source bootstrapped SaaS deals with self-reported revenue traction
  • Market research — map products by category or business model
  • Competitive analysis — track competitors' reported revenue and update cadence
  • Lead lists — build outreach lists of founders with public revenue

Input

FieldTypeDescription
categoriesstring[]Category filter (listing mode). Each category caps at ~18 results anonymously.
businessModelstringfree, ads, commission, transactions, or subscriptions (listing mode).
minRevenueintegerMinimum self-reported monthly revenue (USD). Server-side in listing mode, post-parse in catalog mode.
maxRevenueintegerMaximum monthly revenue (listing mode).
sortBystringhighest-revenue (default), lowest-revenue, recently-updated, recently-added.
includeFullCatalogbooleanEnumerate the full sitemap (~30k products) instead of listing slices.
fetchExactRevenuebooleanAlso fetch each product's /revenue subpage for the exact figure + last-updated date. 2 requests per product.
maxItemsintegerMax products to return. Default 100, max 30000.
proxyConfigurationobjectApify proxy settings.

Legacy input mapping (v1 → v2)

Old inputs are accepted and mapped, with a warning in the run log:

v1 inputv2 behavior
query: "text"Post-fetch substring filter over name/tagline/description. Public search no longer exists upstream.
sortBy: "newest"Mapped to recently-added.
sortBy: "revenue"Mapped to highest-revenue.
sortBy: "followers"Mapped to highest-revenue + warning — the followers metric is gone upstream.
minRevenue, maxItemsUnchanged.

Example input

{
"categories": ["ai", "productivity"],
"minRevenue": 1000,
"sortBy": "highest-revenue",
"maxItems": 50,
"fetchExactRevenue": true
}

Output

Each item contains:

{
"productId": "nomad-list",
"name": "Nomads.com",
"tagline": "The best cities to live and work remotely for digital nomads",
"description": "To accelerate the freedom of global movement enabled by remote work",
"url": "https://www.indiehackers.com/product/nomad-list",
"websiteUrl": "https://nomads.com",
"avatarUrl": "https://storage.googleapis.com/indie-hackers.appspot.com/product-avatars/nomad-list/200x200.webp",
"monthlyRevenueUSD": 35000,
"revenueIsExact": true,
"revenueUpdatedAt": "2024-05-07",
"postsCount": 2,
"twitterHandle": "levelsio",
"facebookUrl": null,
"founders": [{"username": "levelsio", "name": "Pieter Levels", "role": "Founder", "avatarUrl": "..."}],
"claimed": false,
"categories": ["travel"],
"businessModels": [],
"updatedAt": "2025-03-09",
"scrapedAt": "2026-09-22T12:00:00.000Z"
}

Field notes:

  • monthlyRevenueUSD — self-reported by founders; null when unreported or implausible (troll values above $10M/mo are rejected). Abbreviated on the product page ($14K → 14000, $2.6MM → 2600000); exact when fetchExactRevenue hits the /revenue subpage (revenueIsExact: true, plus revenueUpdatedAt). In revenue-sorted listing runs, products with no usable revenue figure are omitted entirely.
  • categories / businessModels — populated only for products discovered through a filtered listing slice; product pages carry no taxonomy markup.
  • updatedAt — latest milestone post date, not a record-modified timestamp.
  • claimed — false when the page shows IndieHackers' "not been claimed" placeholder.
  • v1 fields numFollowers, startDate, createdAt, publishedAt, founderUserIds, verticals, revenueModels, funding, platforms, allTags — removed; no public source exists.

Pricing

Pay-per-event: you only pay for product rows actually delivered to the dataset. No fixed cost, no compute charges. Dead product pages and filter misses are never billed.

Reliability

The actor fails loudly instead of returning empty success when IndieHackers is unreachable, blocked, or the markup drifts — a run that can't reach the site ends FAILED, not SUCCEEDED-0. In catalog mode a named key-value store caches the sitemap inventory for 7 days.

Notes

  • Data comes from indiehackers.com public pages. Revenue is whatever founders voluntarily disclosed — many report nothing, and self-reported figures are unverified.
  • Anonymous listing slices cap at ~18 products each; use includeFullCatalog for bulk extraction.
  • Respect the data: don't spam founders or scrape for malicious purposes.

⭐ Leave a Review

If this actor saved you time, a quick review on the Apify Store helps other developers find it. Takes 30 seconds and means a lot.


Found this actor useful? Leave a review - it helps a lot. Something broken? Report it on the Issues tab.