# Shopify Store Checker: Product Count and Scrape Cost Estimate (`montyburrows/shopify-stores`) Actor

Check which of your domains are readable Shopify stores, how many products each lists and what a full catalogue scrape would cost. Each miss says why: blocked, not found or not Shopify. A free dry run shows the price first, a spend cap stops the run, and a failed run costs nothing.

- **URL**: https://apify.com/montyburrows/shopify-stores.md
- **Developed by:** [Monty Burrows](https://apify.com/montyburrows) (community)
- **Categories:** E-commerce, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.56 / 1,000 domain checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Shopify Store Checker

Give it a list of domains. Get back which of them are readable Shopify storefronts, how many
products each one holds, and **what a full catalogue scrape would cost you** before you start one.

One row per domain, whatever the answer is. "That is not a Shopify store" is the answer you came
for, and it costs the same two requests to establish as "it is, and it has 1,675 products".

### Why run this before the catalogue scraper

**Roughly half of the domains anyone pastes are not readable Shopify storefronts.** Forty-three
brand-name guesses were tried on 2026-09-23, the way anyone builds a first list, and twenty
answered with a catalogue. The other twenty-three failed in six distinct ways.

Running the [Shopify Product & Variant Scraper](https://apify.com/montyburrows/shopify-products) over an unfiltered list
means paying to discover that. Running this first costs **$0.0008 a domain**: five thousand
domains is $4, against roughly $60 of catalogue scraping spent on the ones that were never going
to answer.

And for the ones that do answer, this tells you what reading them is worth before you commit:

```
www.deathwishcoffee.com   ok           145 products   2.88 variants each   $0.13 to scrape
www.hismileteeth.com      ok            66 products   3.06 variants each   $0.06 to scrape
hiutdenim.co.uk           notFound
www.gymshark.com          blocked
www.misfitsmarket.com     notShopify
```

That is a real run, over live domains, on 2026-09-23.

### What you get

![All 14 coffee roaster domains from a real run, one row each: outcome, store name, product count, variants per product and what a full catalogue scrape would cost](https://api.apify.com/v2/key-value-stores/eKO8tTWxCJLEfh79b/records/shopify-stores-example-output.png)

![A finished run's results in the Apify Console, in table view, with the Overview, The shops and The misses views](https://api.apify.com/v2/key-value-stores/eKO8tTWxCJLEfh79b/records/shopify-stores-console-dataset.png)

One row per domain.

| Field                        | Description                                                                              |
| ---------------------------- | ---------------------------------------------------------------------------------------- |
| `domain`                     | The domain you asked about                                                               |
| `resolvedDomain`             | The host that answered. Different when the apex redirects to the www                     |
| `isShopify`                  | Whether the catalogue endpoint answered with a product list                              |
| `outcome`                    | Which of the ways a domain can answer this one answered. See below                       |
| `statusCode`                 | What the catalogue endpoint answered with                                                |
| `storeName` / `storeCountry` | The shop's own name and country, as it states them                                       |
| `myshopifyDomain`            | The shop's platform identity, stable when it changes its public domain                   |
| `currency`                   | ISO-4217, read from the shop rather than guessed from its country                        |
| `productCount`               | How many products the merchant lists in their own public sitemap. Empty when not counted |
| `variantsPerProduct`         | Measured on the first page of the catalogue                                              |
| `catalogueCostUsd`           | What a full catalogue run would charge, at the free-plan rate. Empty when not counted    |
| `detail`                     | A one-line explanation, on the rows that need one                                        |

#### The outcomes, and what each one means for your list

| Outcome       | What it means                                                            | Worth retrying |
| ------------- | ------------------------------------------------------------------------ | -------------- |
| `ok`          | A readable catalogue. This is the list you want                          |                |
| `empty`       | A real Shopify shop with nothing in stock. Still a shop                  |                |
| `notFound`    | No catalogue at that address. Not Shopify, or the merchant turned it off |                |
| `notShopify`  | A live site that answered 200 and served a web page                      |                |
| `blocked`     | Edge bot protection refused us. **Not worked around**                    |                |
| `rateLimited` | We were asked to slow down                                               | **Yes**        |
| `unavailable` | A 5xx. The shop is having a bad day                                      | **Yes**        |

**`isShopify` answers "can you read it", not "what platform is it on".** A shop that is plainly
Shopify and sits behind Cloudflare comes back `false` with an outcome of `blocked`, because the
second question is not one this Actor can answer honestly from outside, and a column that
implied otherwise would be worse than no column.

#### About the product count

It comes from the merchant's **own public sitemap**, not from walking the catalogue. That is one
or two small requests instead of up to seven large ones, and it is a more honest number: measured
across six stores, the catalogue endpoint returned between zero and ten more products than the
sitemap, and every extra read was an app-generated non-product such as a shipping-protection
contract or a bundle-builder placeholder.

So `productCount` is a floor on what a catalogue run yields, and a fair estimate of what it is
worth.

A store that sells in several countries lists a separate product sitemap for each one:
`liquiddeath.com` lists 201 and `www.aloyoga.com` 1,290. The count is for the store's primary
market, which is the catalogue a run over that domain reads, and it costs one request per part of
that market's own sitemap rather than one for every country.

**With the count switched off, `productCount` and `catalogueCostUsd` are empty**, and they are
also empty when a store's sitemap cannot be read. The first page of the catalogue is not a count:
it stops at 250 products, and an estimate built on it would quote a large store a fraction of
what reading it costs.

`variantsPerProduct` is measured on the first page only, and it is the main source of error in
the cost estimate. `catalogueCostUsd` uses the **free-plan rate**, which is the dearest rung, so
the estimate is never lower than the bill.

### Input

| Setting                         | What it does                                                                     |
| ------------------------------- | -------------------------------------------------------------------------------- |
| **Domains**                     | One per line. A bare domain, or any URL on the store. Up to 2,500                |
| **Count each store's products** | On by default. Adds the sitemap count and the cost estimate. Off, both are empty |
| **Dry run**                     | Count what you would get and what it would cost, then stop                       |
| **Maximum results**             | Hard cap on rows. The run stops the moment it is reached                         |
| **Maximum spend (USD)**         | Hard cap on cost. The run stops before exceeding it                              |

**The dry run here is exact rather than estimated.** One row per domain, known before a single
request goes out, so the number you see is the number you will get.

### Pricing

**From $0.56 per 1,000 domains checked**, and nothing else. No per-run fee, no charge for compute
or retries.

| Your Apify plan | Per domain | Per 1,000 domains |
| --------------- | ---------- | ----------------- |
| Free            | $0.0008    | $0.80             |
| Bronze          | $0.00072   | $0.72             |
| Silver          | $0.00064   | $0.64             |
| Gold            | $0.00056   | $0.56             |
| Platinum        | $0.00056   | $0.56             |
| Diamond         | $0.00056   | $0.56             |

![A finished dry run in the Apify Console: its status line reads "Dry run: about 14 results for roughly $0.0112. Nothing was charged."](https://api.apify.com/v2/key-value-stores/eKO8tTWxCJLEfh79b/records/shopify-stores-console-dry-run.png)

The status line is the estimate. A dry run writes no rows, so its results table stays empty, and it charges nothing.

![A real dry run's estimate: 14 domains for $0.0112 against caps of 20 results and $0.02, with $0.00 charged, beside the three billing rules](https://api.apify.com/v2/key-value-stores/eKO8tTWxCJLEfh79b/records/shopify-stores-cost-control.png)

![How a run becomes a bill: each domain is checked once, held in a ledger, every row is checked, and only a run that passes is charged](https://api.apify.com/v2/key-value-stores/eKO8tTWxCJLEfh79b/records/shopify-stores-how-billing-works.png)

**Every domain is charged, whatever the answer is**, and this listing is the one place that is
true across this developer's Actors. Everywhere else, a source that returned nothing costs
nothing, because there were no rows. Here the row is the answer: twenty-three of the
forty-three domains measured were not readable storefronts, and knowing which twenty-three is the
whole job. Charging only for the hits would mean charging double for them, which is the same
money with a better story.

A domain listed twice is checked once and billed once.

**What you are not charged for is a wrong answer.** If most of a run comes back `blocked`, that is
a fact about our exit IP rather than about your list, and the run fails and bills nothing rather
than selling you several thousand rows about somebody else's shop.

### What happens when the source changes

Every run is checked against what a healthy run looks like, and **a run that fails a check is not
billed**:

- **Every domain got an answer.** One row per domain is this Actor's entire promise, so a run
  that silently returned 900 rows for 1,000 domains would have dropped a hundred answers and
  looked complete doing it.
- **We are not being blocked at scale**, per the paragraph above.
- **The pacing was honoured.** One request per second per domain. A run here touches up to 2,500
  other people's shops, and the run fails its own health check if two requests to one of them ever
  go out closer together.
- **The request budget held.** Two requests per domain, or up to four with the product count on. A
  number far above that means a retry storm, which costs compute rather than your bill, and you
  should still know about it.

### How it reads the sources

Through the public endpoints every Shopify storefront serves on the merchant's own domain:
`/meta.json` for the shop's own facts, `/products.json` for the catalogue, and the merchant's own
`/sitemap.xml` for the product count. No login, no token, no session, no browser, and no personal
data of any kind.

**A store that refuses us is reported, never worked around.** On the seventeen stores whose
`robots.txt` was read on 2026-09-23, every path this Actor fetches is permitted, and none of them
publishes a crawl delay that applies to a general-purpose agent. A 403 from edge bot protection is
a shop saying no, and the answer to it is a row in your dataset rather than a different exit IP.

`robots.txt` is a crawling policy rather than a contract, and every merchant has their own terms.
The claim here is the narrow one that is actually true: this Actor fetches paths the store's own
`robots.txt` permits, at one request per second, with no login and no personal data.

### Limits

- **2,500 domains per run.**
- **You supply the domains.** This Actor checks a list; it does not find one. Knowing which
  domains are worth checking is a different product, and a dearer one.
- **It cannot tell you a shop is Shopify if the shop will not talk to us.** `blocked` means
  exactly that and nothing more.
- **One currency per store.** A shop that serves different prices by geography states one
  currency, and what you get is the one it stated to the run.

### Run it on a schedule

Save your input as a task and add an Apify Schedule to run it daily, weekly or hourly. When a run
finishes, Apify's integrations can pass its rows to Google Sheets, Zapier, Make or n8n, and a
webhook can call your own endpoint. A run that fails costs nothing and does not fire an
integration or webhook set to run on success.

Most lists need one run, before a catalogue scrape. On a schedule it checks the whole list again
every time, and every domain is charged whatever the answer. `rateLimited` and `unavailable` are
the two outcomes worth checking again later.

### Use it from an AI agent

Apify's MCP server loads this Actor as a single tool:

```text
https://mcp.apify.com/?tools=montyburrows/shopify-stores
```

An agent with it can run the Actor with the same inputs as the form, dry run and spend cap
included, and read the answer for each domain, billed to your Apify account at the prices above.

### FAQ

#### How much does it cost to check whether a site is a Shopify store?

$0.80 per 1,000 domains on Apify's free plan, and $0.56 per 1,000 on Gold and above, with no start
fee. Every domain is charged whatever the answer, because "not a Shopify store" is the answer.
Apify's free plan gives $5 of usage a month, which at $0.0008 a domain is 6,250 domains. A domain
listed twice is billed once, a failed run and an empty run cost nothing, and the dry run is free
and exact.

#### Is it legal to check Shopify stores this way?

This Actor reads public endpoints every Shopify storefront serves on the merchant's own domain:
`/meta.json`, `/products.json` and the merchant's `/sitemap.xml`, with no login, no token and no
personal data. A store that refuses a request is reported, never worked around. Every merchant has
their own terms, whether a particular use is lawful depends on what you do with the data and where
you are, and nothing here is legal advice.

#### Does `isShopify: false` mean the site is not on Shopify?

Not always. It answers "can you read it", not "what platform is it on": a Shopify shop behind bot
protection comes back `false` with the outcome `blocked`.

#### How accurate is the scrape cost estimate?

`productCount` comes from the merchant's own sitemap and is a floor on what a catalogue run yields.
`catalogueCostUsd` uses the free-plan rate, the dearest rung, and variants per product measured on
the first catalogue page, which is its main source of error.

#### Can it find Shopify stores for me?

No. It checks a list you supply; it does not find one.

### Other Actors from this developer

Every one has the same free dry run and spend cap, and none charges for a failed run.

More for Shopify stores:

- [Shopify Product Scraper](https://apify.com/montyburrows/shopify-products): every product and variant from a list of Shopify stores, with SKU, price, compare-at price and stock
- [Shopify Price & Stock Tracker](https://apify.com/montyburrows/shopify-prices): price and stock for the product pages you choose, one row per variant, built to run on a schedule

And for other data:

- [Google Flights Scraper](https://apify.com/montyburrows/google-flights): live Google Flights fares for any route and date
- [RSS Feed Reader](https://apify.com/montyburrows/rss-feeds): RSS, Atom and RDF feeds in one table
- [Domain Expiry, WHOIS & DNS Lookup](https://apify.com/montyburrows/domain-rdap): expiry dates, registrar and live DNS for a list of domains

### Support and feature requests

Found a bug, need another platform checked, or want a field adding?
Email **actors@montyburrows.com**. Feature requests are welcome and usually quick.

# Actor input Schema

## `domains` (type: `array`):

The domains to check, one per line. A bare domain like beardbrand.com, or any URL on the store. Redirects are followed, so the apex and the www form of one shop give one answer. Up to 2,500 per run.

## `countProducts` (type: `boolean`):

Reads each readable store's own sitemap to count what it lists publicly, and estimates what a full catalogue run would cost. On by default, because that estimate is the reason to run this before the catalogue scraper rather than instead of it. It adds two or three small requests per store that answered. Switched off, the count and the estimate are left empty.

## `dryRun` (type: `boolean`):

Count what you would get and what it would cost, then stop. Nothing is written and you are charged nothing. The count is exact here rather than estimated: one row per domain, known before a single request.

## `maxResults` (type: `integer`):

The most domains this run may answer about. The run stops as soon as it is reached. This is a hard cap, not a target.

## `maxCostUsd` (type: `number`):

The most this run may cost you, in US dollars. The run stops before exceeding it. Leave empty to be limited only by Maximum results.

## `proxy` (type: `object`):

Apify Proxy configuration. The default (datacentre) is the cheapest option and is all this check needs.

## `proxyTier` (type: `string`):

Datacentre is cheap and fast. Residential costs considerably more, and a shop that refuses us refuses on the first request whatever the exit IP is.

## `maxConcurrency` (type: `integer`):

How many requests to run at once against a single host. Each domain is checked one at a time regardless, paced to one request per second.

## `maxRequestRetries` (type: `integer`):

How many times to retry a failed request before recording the outcome it failed with.

## `requestTimeoutSecs` (type: `integer`):

How long a single request may take before it is retried.

## `debug` (type: `boolean`):

Log every request and retry. Useful when opening a support ticket.

## Actor input object example

```json
{
  "domains": [
    "www.deathwishcoffee.com",
    "www.gymshark.com",
    "hiutdenim.co.uk"
  ],
  "countProducts": true,
  "dryRun": false,
  "maxResults": 2500,
  "maxCostUsd": 2,
  "proxy": {
    "useApifyProxy": true
  },
  "proxyTier": "datacenter",
  "maxConcurrency": 4,
  "maxRequestRetries": 4,
  "requestTimeoutSecs": 30,
  "debug": false
}
```

# Actor output Schema

## `domains` (type: `string`):

One row per domain. A dry run writes none: its estimate is in the run summary.

## `runSummary` (type: `string`):

What the run did and what it charged. After a dry run, the estimate: the count, the price and every note that qualifies them.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "www.deathwishcoffee.com",
        "www.gymshark.com",
        "hiutdenim.co.uk"
    ],
    "maxCostUsd": 2
};

// Run the Actor and wait for it to finish
const run = await client.actor("montyburrows/shopify-stores").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "www.deathwishcoffee.com",
        "www.gymshark.com",
        "hiutdenim.co.uk",
    ],
    "maxCostUsd": 2,
}

# Run the Actor and wait for it to finish
run = client.actor("montyburrows/shopify-stores").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "www.deathwishcoffee.com",
    "www.gymshark.com",
    "hiutdenim.co.uk"
  ],
  "maxCostUsd": 2
}' |
apify call montyburrows/shopify-stores --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,montyburrows/shopify-stores"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/2mP46d5PRe5HAaudg/builds/aURGNvRtAMh1w409y/openapi.json
