# Shopify Store Scraper 🛍️ Prospector (catalogue + contacts) (`tagadanar/shopify-store-prospector`) Actor

Turn a domain list into a qualified Shopify prospecting sheet: catalogue size, price range, brands and types carried, products shipped in the last 30 days, discount depth, collections, theme, socials and role mailboxes. Read from each store's own public storefront JSON. No API key, no proxy.

- **URL**: https://apify.com/tagadanar/shopify-store-prospector.md
- **Developed by:** [Tagada Data](https://apify.com/tagadanar) (community)
- **Categories:** E-commerce, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.60 / 1,000 store analyzeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Shopify Store Prospector 🛍️ (catalogue + contacts, no key)

Paste a list of domains. Get back a qualified prospecting sheet: which of them
actually run Shopify, how big and how active each catalogue is, what the store
sells and at what price, which brands it carries, how hard it discounts, what
theme it runs, and where to contact it. Everything is read from each store's own
public storefront JSON, so there is no API key, no app install, no OAuth and no
proxy anywhere in the run.

StoreLeads, the database most agencies use for this, publishes $75 a month for
interactive use of its UI and $250 a month for the plan that lets you export
(storeleads.app, read 2026-08-26). This is the same job priced per store, with
no subscription, and you keep the raw rows.

### How it compares on the Apify Store

Users and prices read from the public Store API on 2026-08-26. Per the
2026-08-20 recon, no listing in this pool carries more than seven reviews, and
the biggest one has none at all, so nobody here has an established quality bar
yet.

| Listing | 30-day users | Price per store row | Reads a headless storefront? | Free for a miss? |
| --- | --- | --- | --- | --- |
| `clearpath/shopify-store-leads` | 176 | $5.99 / 1k (deepest $4.99) | not advertised | not advertised |
| `apivault_labs/website-leads-database` | 100 | $8.00 / 1k | not advertised | not advertised |
| `apivault_labs/shopify-store-analyzer` | 77 | $15.00 / 1k (deepest $4.00) | not advertised | not advertised |
| `webdatalabs/shopify-store-intelligence` | 10 | $10.00 / 1k (deepest $4.00) | not advertised | not advertised |
| **this actor** | new | **$8.00 / 1k** (deepest $5.60) | **yes** | **yes** |

"Not advertised" means exactly that: it is what their public listing says, not a
test of their code.

#### The headless problem, and why it matters to your list

A growing share of the best Shopify brands run a custom front end (Hydrogen, a
Next.js build, a headless CMS). Ask `brand.com/products.json` and you get a 404
HTML page, so a naive storefront scraper writes them off as "not Shopify" and
they fall out of your list, which is the opposite of what you want: those are the
biggest accounts in it.

This actor reads the homepage, recovers the store's `.myshopify.com` handle from
its own markup, and reads the catalogue there instead. Measured 2026-08-25:
`skims.com/products.json` answers 404, `skimsbody.myshopify.com/products.json`
answers 200 with the real catalogue. When even that is closed, the row still
comes back labelled `shopify-headless` with the firmographics, theme and socials
attached, instead of a wrong `not-shopify`.

### What you get

One `store` row per domain.

| Field | What it is |
| --- | --- |
| `status` | `ok`, `empty`, `shopify-headless`, `password`, `blocked`, `not-shopify`, `unreachable`, `invalid-domain` |
| `statusDetail` | A plain sentence saying what happened, per store |
| `isShopify`, `myshopifyDomain`, `canonicalDomain` | Platform verdict and the store's Shopify handle |
| `shopName`, `city`, `province`, `country`, `currency` | Firmographics from the store's own shop record |
| `productCount`, `productsAnalyzed`, `productCountIsFloor`, `catalogueSizeBand`, `catalogueTruncated` | Exact catalogue size, how many products the summary was computed over, and the size band: starter / micro / small / mid / large / enterprise |
| `priceMin`, `priceMax`, `priceAvg` | Price positioning, in the store's own currency |
| `vendorCount`, `topVendors` | Own-brand or multi-brand retailer, and which brands it carries |
| `productTypeCount`, `topProductTypes` | What it actually sells |
| `newProductsLast30d`, `newProductsLast90d`, `newestProductAt`, `oldestProductAt`, `catalogueAgeDays`, `productsUpdatedLast30d` | Merchandising velocity and catalogue age. New counts come from each product's creation date, not its published date, because Shopify rewrites the published date every time a merchant re-publishes an old product |
| `discountedProducts`, `discountRate`, `maxDiscountPercent` | How hard the store is discounting right now |
| `inStockProducts`, `outOfStockProducts`, `inStockRate` | Stock health |
| `variantCount`, `avgVariantsPerProduct`, `imageCount` | SKU depth and content investment |
| `collectionCount`, `collectionsTruncated`, `topCollections` | Merchandising structure |
| `themeName`, `themeSchemaName`, `themeSchemaVersion`, `usesFreeShopifyTheme` | Theme stack, and whether the store is still on a free Shopify theme |
| `socialLinks` | Instagram, TikTok, Facebook, X, YouTube, Pinterest, LinkedIn company page |
| `publicEmails`, `contactUrls` | Generic role mailboxes and contact / wholesale pages published on the store itself |
| `prospectFlags` | The short list you sort on, e.g. `ships-new-products-weekly`, `catalogue-stale-90d`, `heavy-discounting`, `multi-brand-retailer`, `own-brand`, `free-shopify-theme`, `headless-storefront`, `stock-thin`, `premium-pricing` |

Turn the products toggle on and you also get one `product` row per product:
title, handle, URL, vendor, product type, tags, price, compare-at price, sale
depth, stock, variant count, image count and the created / updated / published
dates.

#### On contact data

Only generic role mailboxes (`hello@`, `info@`, `wholesale@` and the like) that
the store publishes on its own homepage are ever returned. A named personal
mailbox is discarded and never stored. There is no people data of any kind in
the output.

### Who buys this

**Agencies and SaaS selling into DTC.** Filter a domain list down to live
Shopify stores, then sort by `catalogueSizeBand`, `newProductsLast30d` and
`usesFreeShopifyTheme` to find the accounts worth a call. A store on Dawn with
900 products and 40 new listings a month is a different pitch from a parked
40-product store on a custom theme.

**3PL, packaging and fulfilment.** `productCount`, `variantCount` and
`newProductsLast30d` are the closest public proxy there is for how much a brand
actually ships, and `country` plus `currency` tell you which warehouse it needs.

**Competitor and price watching.** Run the same list weekly. `priceAvg`,
`discountRate`, `maxDiscountPercent` and `newProductsLast30d` move before the
press release does.

**Sourcing and retail buying.** `topVendors` and `topProductTypes` tell you
which retailers already carry a category, and `inStockRate` tells you who is
running out.

### Input

```json
{
  "stores": ["allbirds.com", "colourpop.com"]
}
```

Products as well, capped so one large retailer cannot swamp the run:

```json
{
  "stores": [
    "https://www.allbirds.com/collections/mens",
    "deathwishcoffee.com",
    "kith.com"
  ],
  "includeProducts": true,
  "maxProductsPerStore": 250
}
```

| Input | Type | Default | Notes |
| --- | --- | --- | --- |
| `stores` | array of strings | required | A bare domain, a domain with www, or any URL on the store. Duplicates merge. |
| `includeProducts` | boolean | `false` | Adds one row per product on top of the store row. |
| `maxProductsPerStore` | integer | `100` | Cap on product rows per store. The store summary always covers the whole catalogue. |

Anything the list field can be handed works: a single string, a comma or
newline separated list, or a dataset piped in from another scraper with `url`
or `domain` keys on each item.

### Pricing

You pay per store we actually read, and per product row delivered. Domains that
turn out to be blocked by a merchant firewall, password-protected, unreachable
or not on Shopify come back labelled and cost nothing beyond the run start.

Platform usage is on us, so the price you see is the price you pay. There is no
proxy setting because none is needed: the storefront JSON surface is open.

### FAQ

**Does this need a Shopify API key or a private app?**
No. Every path it reads (`/meta.json`, `/products.json`, `/collections.json`)
is served publicly by Shopify storefronts by default.

**How do I find Shopify stores in the first place?**
Version 1 takes a domain list you already have: a CRM export, an ad-library
pull, a competitor's stockist page, a conference attendee list. It does not
discover stores on its own. Discovery through a tech-detection database needs
a paid API key, which would put a cost on this actor that you would end up
paying for, so it is deliberately out of scope.

**Why does one of my stores come back `blocked`?**
Some merchants put their own firewall in front of Shopify. `gymshark.com`
answers 403 from any datacenter and `bombas.com` answers 429 (both measured
2026-08-25). The run does not stop, the store is labelled, and you are not
charged for it.

**What does `password` mean?**
The store exists on Shopify but still has its password page on, which is what
an unlaunched or seasonal store looks like. For some prospecting jobs that is
the best signal in the sheet.

**Is `productCount` the whole catalogue?**
Almost always exactly, however big the store is. The catalogue is read in full
up to 5,000 products; past that the exact total is recovered with a handful of
tiny index requests rather than downloading the rest, `catalogueTruncated` turns
`true`, and `productsAnalyzed` tells you how many products the price, brand and
velocity figures were computed over.

There is one hard limit, and it is Shopify's, not ours: the storefront JSON
refuses to look past product number 25,000 (`page * limit` above 25,000 returns
HTTP 400, measured on kith.com 2026-08-26). A store that still answers at that
depth comes back with `productCount` 25000 and `productCountIsFloor` `true`,
meaning at least that many. No reader of this surface can do better, so treat a
tool that quotes you an exact number above 25,000 with suspicion.

**Can I get best sellers?**
No, and be careful with anyone who says they can from this surface. The JSON
endpoint ignores `sort_by=best-selling`: the response is byte identical with
and without it (measured 2026-08-25). This actor reports what it can prove.

**Can I track a store over time?**
Yes. Re-run the same list on a schedule. `newProductsLast30d`, `priceAvg`,
`discountRate` and `inStockRate` are designed to be diffed run over run.

**Does it collect personal data?**
No. Role mailboxes published by the store itself, social profile links and
company firmographics only. Named personal mailboxes are dropped.

### Related actors

- `agency-leads-scraper` and `clutch-agency-leads` for the agencies on the
  other side of this market.
- `b2b-lead-verifier` to check a company behind a domain against public
  registers before you call it.
- `saas-pricing-monitor` if the pricing pages, rather than the catalogues, are
  what you need to watch.

***

Keywords: shopify store scraper, shopify store leads, shopify products json,
dtc brand list, shopify prospecting, ecommerce lead generation, shopify store
finder, storeleads alternative, shopify catalogue export, shopify competitor
price tracking, shopify tech stack lookup, headless shopify detection.

# Actor input Schema

## `stores` (type: `array`):

One entry per store. A bare domain (<code>allbirds.com</code>), a domain with www, or any URL on the store (<code>https://www.allbirds.com/collections/mens</code>) all work, and duplicates are merged. Stores that turn out not to run Shopify, that are password-protected, or that sit behind a merchant firewall come back as a labelled row and are not charged.

## `includeProducts` (type: `boolean`):

Add one row per product on top of the store row: title, vendor, type, price, compare-at price, sale depth, stock, variants, images and dates. Leave it off for a pure prospecting list, which is cheaper because product rows are billed separately.

## `maxProductsPerStore` (type: `integer`):

Cap on product rows for each store, so one 10,000-product retailer cannot swamp the run. The store row still summarises the whole catalogue whatever this is set to. Ignored when the products toggle is off.

## Actor input object example

```json
{
  "stores": [
    "allbirds.com",
    "colourpop.com"
  ],
  "includeProducts": false,
  "maxProductsPerStore": 25
}
```

# Actor output Schema

## `stores` (type: `string`):

One store signal row per domain in the default dataset, plus optional product rows.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "stores": [
        "allbirds.com",
        "colourpop.com"
    ],
    "maxProductsPerStore": 25
};

// Run the Actor and wait for it to finish
const run = await client.actor("tagadanar/shopify-store-prospector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "stores": [
        "allbirds.com",
        "colourpop.com",
    ],
    "maxProductsPerStore": 25,
}

# Run the Actor and wait for it to finish
run = client.actor("tagadanar/shopify-store-prospector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "stores": [
    "allbirds.com",
    "colourpop.com"
  ],
  "maxProductsPerStore": 25
}' |
apify call tagadanar/shopify-store-prospector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tagadanar/shopify-store-prospector"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fht0NEFcw0WUTF3eu/builds/70Nz59Zq3Txz3aQ5X/openapi.json
