# Shopify Store Scraper — Emails, Phones & Apps (`b2b_leads/shopify-real-time-data-scraper`) Actor

Find Shopify stores by niche, keyword, country or direct URL. Get store identity, product counts, public emails, phone numbers, social profiles, Shopify themes and installed apps. Results stream live and can be sent to any webhook. Free Apify accounts are limited to 2 results per run.

- **URL**: https://apify.com/b2b\_leads/shopify-real-time-data-scraper.md
- **Developed by:** [Emmanuel](https://apify.com/b2b_leads) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Shopify Real-Time Data Scraper

**Turn Shopify merchants into a structured, sales-ready dataset — store identity, product counts, contact details, social profiles, themes and installed apps, streamed to you as each record is ready.**

Free Apify accounts are limited to **2 results per run**. Upgrade to a paid Apify plan to get full, unlimited data.

***

### ⚠️ Free tier / paid plan notice (please read first)

| Account type | What happens |
| --- | --- |
| **Paid Apify plan** | Full output. No caps. |
| **Free Apify plan (default `limit` mode)** | Up to **2 results** are exported, then the run finishes cleanly with an upgrade notice. |
| **Free Apify plan (`block` mode)** | Nothing is exported. The run stops immediately with an upgrade notice. |
| **Local development** | No gating at all, so you can test freely. |

This is a deliberate, transparent product restriction — not a bug. The run always
finishes gracefully (never a system failure) and the run summary always contains a
`paywall` object describing exactly what was applied. See
[Free tier & monetization](#free-tier--monetization) for the full rules.

***

### What it does

The Shopify Real-Time Data Scraper finds online stores and turns each one into a
clean, consistent JSON record. Point it at a niche, a set of keywords, a country,
or a list of store URLs, and it delivers:

- **Store identity** — brand name, website, country, niche tags.
- **Catalog signals** — product count and average price where available.
- **Lead details** — business email addresses and phone numbers.
- **Social profiles** — Instagram, Facebook, TikTok, X/Twitter, YouTube, LinkedIn and Pinterest.
- **Technology signals** — the active Shopify theme and the list of installed Shopify apps.
- **Descriptive text** — a short store description for segmentation and enrichment.

Every record is written to the run's dataset **as soon as it is ready**, so you can
watch results appear in real time and memory stays light on very large runs.

#### What you get

- A clean JSON dataset with a documented, stable field structure.
- Incremental, live output instead of one big file at the end.
- Optional real-time webhook delivery to a CRM, Slack, Zapier, Make or Google Sheets.
- A run summary that shows totals, the stop reason, and the plan status.
- The run total is the sum of `maxResults` across every search task you add, so you control volume task by task.
- A spending-limit guard that stops the run the moment your budget is exhausted.

***

### Who it is for

- **Lead-generation agencies** building prospect lists for ecommerce, marketing and fulfilment services.
- **Dropshipping and product-research sellers** who want to study niches, storefronts and product volume.
- **Shopify agencies and freelancers** prospecting brands that may need migration, theme, CRO or app services.
- **Marketing teams** mapping social presence and contact data for outreach.
- **Competitive intelligence analysts** tracking which stores are active in a niche or country.
- **Startup founders** validating a niche before committing budget.
- **Sales and CRM teams** importing clean, deduplicated merchant records.
- **Data analysts and researchers** building datasets for market studies.
- **Automation users** chaining this Actor into n8n, Make, Zapier or a custom workflow.

***

### Use cases

1. **Niche prospecting** — find stores in a niche to study assortment and positioning.
2. **Agency outreach** — identify Shopify stores that could benefit from your services.
3. **Competitor mapping** — size up who else is operating in a category or country.
4. **Theme & app signals** — spot stores running specific themes or app stacks.
5. **Social outreach lists** — collect Instagram/TikTok handles for influencer-style outreach.
6. **Cold email lists** — gather public business email addresses.
7. **Phone-based sales lists** — gather public business phone numbers.
8. **App ecosystem research** — see which app ecosystems dominate a niche.
9. **Market sizing** — estimate how crowded a niche is by counting merchants.
10. **Country-level market entry research** — evaluate a new region before launching.
11. **Agency pitch lists** — filter by product count to target high-volume merchants.
12. **Store enrichment** — take a list of known domains and fill in contact data.
13. **CRM hygiene** — keep store records fresh before a campaign.
14. **New-store monitoring** — re-run a niche regularly to spot new or growing stores.
15. **Vendor shortlisting** — compare technology choices across suppliers.
16. **Content partnership discovery** — find brands active on a given social network.
17. **Wholesale outreach** — find retailers in a product vertical.
18. **Price benchmarking research** — compare average price signals across a niche.
19. **Data-enrichment pipelines** — feed a warehouse or spreadsheet with structured records.
20. **Bulk list building** — collect thousands of records in one run.

***

### Key features

- **Enable lead details (on by default)** — adds contact and social data. Stores with
  missing contact fields are still exported: lead details are an *addition*, never a filter.
  Adds a little extra time per business, but output stays complete and predictable.
- **Search tasks** — run several niche / keyword / country combinations in a single run.
- **Direct store lookup** — supply your own store URLs instead of searching.
- **Bounded parallelism** — search tasks and lead-detail work run concurrently for speed.
- **Incremental dataset writes** — memory stays light on long runs.
- **Country filter** — restrict results to a specific market.
- **Tech-stack enrichment** — optional theme and installed-app detection.
- **Real-time webhooks** — push each record to your own service.
- **Plan-aware gating** — transparent free-tier limits, configurable from the Console.
- **Spending-limit aware** — stops as soon as your budget is reached.
- **Secret-safe errors** — failures are reported as plain, fixed sentences.
- **MCP-ready** — works with the Apify MCP Server.

***

### Quick start

#### Example 1 — a single niche in the US

```json
{
  "enableDirectoryBrowse": true,
  "searchTasks": [
    { "niche": "clothing", "country": "US", "maxResults": 10 }
  ],
  "enableLeadDetails": true
}
```

#### Example 2 — several tasks at once

```json
{
  "searchTasks": [
    { "niche": "jewelry", "country": "US", "maxResults": 20 },
    { "niche": "skincare", "country": "UK", "maxResults": 20 },
    { "niche": "pet", "maxResults": 20 }
  ],
  "enableLeadDetails": true,
  "enrichThemeApps": true
}
```

#### Example 3 — enrich a list of stores you already have

```json
{
  "enableDirectoryBrowse": false,
  "enableStoreDetails": true,
  "storeUrls": [
    "gymshark.com",
    "allbirds.com",
    "warburtons.com"
  ],
  "enableLeadDetails": true
}
```

#### Example 4 — fast run without lead details

```json
{
  "searchTasks": [{ "niche": "coffee", "maxResults": 50 }],
  "enableLeadDetails": false
}
```

In Example 4 the run total is 50, because that is the only task. In Example 2
the three tasks add up to 60. There is no separate total to configure.

***

### Full input schema

All fields are optional — sensible defaults are applied when you leave them blank.

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `enableDirectoryBrowse` | boolean | `true` | Turn on niche/category discovery. |
| `searchTasks` | array of objects | one starter task | The searches to perform. |
| `searchTasks[].niche` | string | — | Niche or category, e.g. `clothing`, `jewelry`, `skincare`. |
| `searchTasks[].keyword` | string | `""` | Optional keyword or brand term. |
| `searchTasks[].country` | string | `""` | Country name or 2-letter code, e.g. `US`, `UK`, `Germany`. |
| `searchTasks[].maxResults` | integer | `10` | Results to aim for from this task. The run total is the sum of this across all tasks. |
| `enableKeywordSearch` | boolean | `false` | Force keyword-style discovery for every task. |
| `keywords` | array of strings | `[]` | Convenience list that becomes search tasks. |
| `enableStoreDetails` | boolean | `false` | Process `storeUrls` directly. |
| `storeUrls` | array of strings | `[]` | Store websites to process, with or without a scheme. |
| `enableLeadDetails` | boolean | `true` | Add contact and social details. Free accounts are limited to 2 results. |
| `enrichSocials` | boolean | follows `enableLeadDetails` | Collect social profile links. |
| `enrichContact` | boolean | follows `enableLeadDetails` | Collect email addresses and phone numbers. |
| `enrichThemeApps` | boolean | `false` | Detect the Shopify theme and installed apps. |
| `webhookUrl` | string | `""` | Optional destination for real-time delivery. |
| `webhookFormat` | string | `json` | `json` for the full record, `slack` for a compact message. |
| `proxyConfiguration` | object | US residential | Connection settings for the run. |

#### `enableDirectoryBrowse` — boolean, default `true`

Turns on niche discovery. When `true` (and at least one search task exists), the
Actor fills its results from the niches you chose.

#### `searchTasks` — array of objects

The main way to configure a run. Add one object per search. Tasks run
concurrently, so three tasks finish in roughly the time of one.

```json
{
  "searchTasks": [
    { "niche": "clothing", "country": "US", "maxResults": 25 },
    { "niche": "watches", "maxResults": 15 }
  ]
}
```

##### `searchTasks[].niche` — string

The category to focus on. Common values include `clothing`, `fashion`,
`streetwear`, `jewelry`, `watches`, `sneakers`, `handbags`, `skincare`,
`cosmetics`, `makeup`, `soap`, `furniture`, `home-decor`, `candles`, `coffee`,
`tea`, `supplements`, `fitness-equipment`, `electronics`, `headphones`, `pet`,
`golf`, `camping`, `toys`, `collectibles`, `vinyl` and `gaming`.

If your term has no exact match, the Actor suggests related niches that do have
results and continues with the closest match, so a run never comes back empty
just because of a spelling difference.

##### `searchTasks[].keyword` — string, default `""`

An optional extra term. Supplying a keyword switches that task to keyword
discovery, which is useful for brand terms or product-specific themes.

##### `searchTasks[].country` — string, default `""`

Restricts the task to a market. Accepts a 2-letter code (`US`, `GB`, `CA`, `AU`,
`DE`, `FR`, `IT`, `ES`, `NL`, `IN`, `JP`, `BR`, `MX`) or a full country name
(`United States`, `United Kingdom`, `Germany`, `Australia`, …).

When a country is set, only stores registered in that market are exported.

##### `searchTasks[].maxResults` — integer, default `10`

How many stores to aim for from this task. The run's total is simply the sum of
`maxResults` across every task you add, so you control volume task by task.

#### `enableKeywordSearch` — boolean, default `false`

Forces keyword-style discovery. Leave it off when you only search by niche.

#### `keywords` — array of strings, default `[]`

A shortcut for API and MCP callers. Each entry becomes its own search task.

```json
{ "keywords": ["vintage clothing", "minimalist jewelry", "organic skincare"] }
```

#### `enableStoreDetails` — boolean, default `false`

Turns on direct processing of `storeUrls`. Use it when you already know which
stores you want, and you do not need discovery.

#### `storeUrls` — array of strings, default `[]`

Store websites to process. Both `gymshark.com` and the full form are accepted.

#### `enableLeadDetails` — boolean, default `true`

Adds email addresses, phone numbers, social profile links and a store
description. **Every store is still exported even when no contact data is
available** — this option never filters results. Free accounts are limited to
2 results per run and must upgrade for full output.

#### `enrichSocials` / `enrichContact` — boolean

Fine-grained switches under `enableLeadDetails`. Both default to the value of
`enableLeadDetails`, so you normally never need to set them. Set either to
`false` to save time when you only need one kind of detail.

#### `enrichThemeApps` — boolean, default `false`

Adds the active Shopify theme name and the list of installed Shopify apps. Leave
this off for faster runs.

#### How the run total is calculated

There is no separate "max items" field. The total comes from your tasks:

- Each search task contributes its own `maxResults`.
- Each entry in `storeUrls` contributes one result.
- The run stops as soon as it has exported that many records.

So three tasks at 20 each export up to 60 records, and a single task at 50
exports up to 50. This keeps your cost control in one obvious place — the same
place you already choose your searches.

Free accounts are capped at 2 results regardless of what the tasks add up to.

#### `webhookUrl` — string, default `""`

Optional destination that receives each record as soon as it is exported. The
dataset is always written regardless — the webhook is an *additional* delivery.
If delivery fails, the run continues and the failure is counted in the summary.

#### `webhookFormat` — string, default `json`

- `json` — the complete record, exactly as stored in the dataset.
- `slack` — a compact single-field message for a Slack incoming webhook.

#### `proxyConfiguration` — object

Connection settings for the run. The default is the **Apify residential
connection with a US location**, which is what you want in almost every case —
no extra setup is required.

You only need to change this if you want a different region, or if you want to
supply your own connection URLs instead. You normally never need to touch it.

***

### Full output schema

One JSON object per store. The structure is identical for every feature, so a
single spreadsheet template or loader works for all of them.

| Field | Type | Description |
| --- | --- | --- |
| `featureType` | string | `directory`, `keyword_search` or `store_details`. |
| `name` | string | Store brand name. |
| `storeName` | string | Same value as `name`; kept for convenience. |
| `storeUrl` | string | Store website. |
| `description` | string | null | Short store description. |
| `category` | string | null | Primary category when known. |
| `tags` | array of strings | Niche and product tags. |
| `country` | string | null | 2-letter country code. |
| `productsCount` | number | null | Approximate catalog size. |
| `avgPrice` | string | null | Average price signal when available. |
| `email` | string | null | Primary business email. |
| `emails` | array of strings | Up to three public business emails. |
| `phone` | string | null | Primary business phone. |
| `phones` | array of strings | Up to two public business phones. |
| `instagram` | string | null | Instagram profile. |
| `facebook` | string | null | Facebook profile. |
| `tiktok` | string | null | TikTok profile. |
| `twitter` | string | null | X / Twitter profile. |
| `youtube` | string | null | YouTube profile. |
| `linkedin` | string | null | LinkedIn profile. |
| `pinterest` | string | null | Pinterest profile. |
| `theme` | string | null | Active Shopify theme. |
| `apps` | array of strings | Installed Shopify apps. |
| `leadDetailsComplete` | boolean | `true` when lead details ran for this record. |
| `searchTaskLabel` | string | null | The task that produced this record. |
| `scrapedAt` | string | ISO 8601 timestamp. |

#### Field-by-field notes

##### `featureType` — string

Tells you how the record was produced:

- `directory` — found through niche/category discovery.
- `keyword_search` — found through a keyword.
- `store_details` — processed from a URL you supplied.

##### `name` and `storeName` — string

The brand name. When only a bare domain is published, a readable title is used
instead, so you rarely get `store.example.com` as a name. `storeName` is an
alias of `name` for compatibility with spreadsheet templates.

##### `storeUrl` — string

The store's website. This is the most reliable key for deduplication — normalise
it (lowercase, strip the trailing slash) before joining across runs.

##### `description` — string or null

A short, human-readable description of the store. Useful for segmentation and
for AI-based categorisation. `null` when no description is available.

##### `category` — string or null

The primary category, when one is known.

##### `tags` — array of strings

Niche and product tags associated with the store, for example
`["fashion clothing", "men accessories", "women clothing"]`. Great for building
segments and for filtering in a spreadsheet.

##### `country` — string or null

Two-letter country code (`US`, `GB`, `PK`, `ES`, …). Filter on this when you set
a country in a search task.

##### `productsCount` — number or null

Approximate number of products in the catalog. A strong signal for store size —
useful for prioritising larger merchants in outbound sales.

##### `avgPrice` — string or null

Average price signal when available. Kept as a string because some sources
include a currency symbol.

##### `email` and `emails` — string / array of strings

Public business contact addresses. Addresses on the store's own domain are
preferred, with well-known personal mailbox providers used as a fallback.
Platform, analytics and infrastructure addresses are filtered out, so you get
addresses you can actually use.

##### `phone` and `phones` — string / array of strings

Public business phone numbers, normalised to a consistent format. Digits found
in scripts are ignored unless they are genuine contact links, which keeps
product IDs and random numbers out of your list.

##### Social fields — `instagram`, `facebook`, `tiktok`, `twitter`, `youtube`, `linkedin`, `pinterest`

Full profile URLs for the brand's own accounts. Only genuine profile links are
kept — posts, reels, share links and default platform accounts are removed. A
field is `null` when the store has no account on that network.

##### `theme` — string or null

The active Shopify theme. A strong buying signal for theme developers, agencies
and app developers. Requires `enrichThemeApps: true`.

##### `apps` — array of strings

Installed Shopify apps. Excellent for app developers and agencies looking for
stores that might benefit from a competing product. Requires
`enrichThemeApps: true`.

##### `leadDetailsComplete` — boolean

`true` when lead details were attempted for this record. A `true` value does
**not** guarantee an email address — some stores simply do not publish one. The
store is exported either way.

##### `searchTaskLabel` — string or null

The search task that produced this record, for example `clothing (US)`. Use it
to split one run's output into separate lists.

##### `scrapedAt` — string

ISO 8601 timestamp of when the record was produced.

#### Example record

```json
{
  "featureType": "directory",
  "name": "Eden's Echo Farmstead",
  "storeName": "Eden's Echo Farmstead",
  "storeUrl": "https://edensechofarmstead.com",
  "description": "Small-batch home goods made in the Pacific Northwest.",
  "category": "Home & Living",
  "tags": ["home decor", "handmade", "small batch"],
  "country": "US",
  "productsCount": 5374,
  "avgPrice": null,
  "email": "online@edensechofarmstead.com",
  "emails": ["online@edensechofarmstead.com"],
  "phone": "(520) 386-0556",
  "phones": ["(520) 386-0556"],
  "instagram": "https://www.instagram.com/edensechofarmstead",
  "facebook": null,
  "tiktok": null,
  "twitter": null,
  "youtube": null,
  "linkedin": null,
  "pinterest": null,
  "theme": "Dawn",
  "apps": ["Klaviyo", "Judge.me"],
  "leadDetailsComplete": true,
  "searchTaskLabel": "clothing (US)",
  "scrapedAt": "2026-09-25T00:15:42.113Z"
}
```

#### Dataset views

The dataset ships with three ready-made views:

| View | What it shows |
| --- | --- |
| `overview` | Every record with the most useful columns. |
| `leads` | Records grouped around contact data and social profiles. |
| `tech_stack` | Records grouped around theme and installed apps. |

#### Run summary (`OUTPUT`)

The run also stores a summary object under the `OUTPUT` key, visible in the run
overview and at the run's key-value store.

| Field | Type | Description |
| --- | --- | --- |
| `totalPushed` | number | Records exported. |
| `spendingLimitReached` | boolean | `true` if the run stopped on budget. |
| `stoppedReason` | string | `completed`, `max_items`, `spending_limit`, `free_tier_limit` or `free_tier_blocked`. |
| `paywall.detected` | boolean | Whether plan signals were present. |
| `paywall.isPaying` | boolean | Whether the account is on a paid plan. |
| `paywall.pricingTier` | string | null | Plan tier, when known. |
| `paywall.blocked` | boolean | Whether the run was stopped before any work. |
| `paywall.limited` | boolean | Whether a free-tier cap was applied. |
| `paywall.freeTierMaxItems` | number | null | The configured free cap. |
| `errors` | array of strings | Safe, fixed failure sentences (usually empty). |

***

### Webhook setup

A webhook lets you receive each record in real time instead of waiting for the
run to finish. The dataset is **always** written — the webhook is an additional
delivery, so adding one can never cause you to lose data.

#### How to set it up

1. Open the Actor in the Apify Console and go to **Input**.
2. Paste your destination into **Webhook URL** (for example your Zapier catch
   URL, a Make hook, a Slack incoming webhook address, or your own service).
3. Choose **Webhook format**:
   - `json` — the full record, identical to the dataset row. Best for CRMs,
     databases, Make and Zapier.
   - `slack` — a compact single-field message. Best for Slack channels.
4. Run the Actor. Each record is delivered as soon as it is exported.

#### Webhook payload — `json`

The body is exactly the dataset record, so it contains only your results:

```json
{
  "featureType": "directory",
  "name": "Eden's Echo Farmstead",
  "storeUrl": "https://edensechofarmstead.com",
  "description": "Small-batch home goods made in the Pacific Northwest.",
  "category": "Home & Living",
  "tags": ["home decor", "handmade"],
  "country": "US",
  "productsCount": 5374,
  "email": "online@edensechofarmstead.com",
  "emails": ["online@edensechofarmstead.com"],
  "phone": "(520) 386-0556",
  "phones": ["(520) 386-0556"],
  "instagram": "https://www.instagram.com/edensechofarmstead",
  "theme": "Dawn",
  "apps": ["Klaviyo", "Judge.me"],
  "leadDetailsComplete": true,
  "searchTaskLabel": "clothing (US)",
  "scrapedAt": "2026-09-25T00:15:42.113Z"
}
```

#### Webhook payload — `slack`

A single-field message designed for a Slack incoming webhook:

```json
{
  "text": ":shopping_bags: *Eden's Echo Farmstead*\n<https://edensechofarmstead.com|Visit store>\n*Category:* Home & Living\n*Country:* US\n*Products:* 5374\n*Email:* online@edensechofarmstead.com\n*Phone:* (520) 386-0556\n*Instagram:* https://www.instagram.com/edensechofarmstead\n*Theme:* Dawn"
}
```

#### Delivery guarantees and failure handling

- Delivery is **best effort**. A failed delivery never stops the run and never
  removes the record from the dataset.
- Failures are counted and reported in the run summary notes.
- Payloads contain **only your record's own output fields**. No internal
  structures, no connection details, no credentials.
- Only records that were actually exported are delivered.

#### Common webhook destinations

| Destination | Suggested format |
| --- | --- |
| Zapier | `json` |
| Make | `json` |
| n8n | `json` |
| Slack incoming webhook | `slack` |
| A custom CRM or database | `json` |

***

### MCP usage

This Actor works with the [Apify MCP Server](https://docs.apify.com/integrations/mcp),
so you can discover and run it directly from an AI assistant.

#### Setup

1. Open your AI tool (Claude Desktop, Cursor, Windsurf, VS Code Copilot, or any
   MCP-compatible client).
2. Add the Apify MCP Server to your configuration.
3. Connect it with your Apify API token.

#### Calling the Actor through MCP

Once connected, your assistant can start a run with natural language, for example:

- *"Run the Shopify Real-Time Data Scraper for the clothing niche in the US and give me 25 stores with contact details."*
- *"Find 50 Shopify stores in the skincare niche and include their installed apps."*
- *"Enrich these store websites for me: gymshark.com, allbirds.com, warburtons.com."*

The assistant will fill in the input fields for you, using the schema documented
above, and return the dataset results.

#### Tips for MCP users

- Always state the **niche** or **keywords** you care about.
- State the **country** if you only want one market.
- State the **result count** you want for each task — it maps to that task's `maxResults`.
- Mention whether you want **lead details** and **installed apps** so the
  assistant sets the right switches.
- Free Apify accounts receive up to 2 results per run.

***

### Free tier & monetization

This Actor follows a transparent **freemium** model.

#### Plan detection

At the start of every run the Actor reads the platform's account signals and
decides how to behave. The logic is deliberately defensive:

1. If the account is marked as paying, the run is **not** limited.
2. Otherwise, if a plan tier is present, any tier other than the free tier counts
   as paying.
3. If **neither** signal is present — for example during local development — the
   run is **not** limited, so you can test freely.
4. Gating is only applied when a platform signal is actually present.

#### Behaviour by account type

| Scenario | Logs | Dataset | Exit |
| --- | --- | --- | --- |
| Paying account | `Paying user — full output.` | Full output, no cap | Normal completion |
| Free account, `limit` mode | Warning to upgrade | Up to `FREE_TIER_MAX_ITEMS` (default 2) | Clean, with a clear notice |
| Free account, `block` mode | Warning to upgrade | Nothing | Clean, before any work |
| Local development | No gating messages | Full output | Normal completion |

In every case the run ends **gracefully** — it never crashes, never shows a
system-error state, and never charges you for work it will not export.

#### Transparency in the run summary

Every run publishes a `paywall` object so you can always see exactly what was
applied:

```json
{
  "totalPushed": 2,
  "spendingLimitReached": false,
  "stoppedReason": "free_tier_limit",
  "paywall": {
    "detected": true,
    "isPaying": false,
    "pricingTier": "FREE",
    "blocked": false,
    "limited": true,
    "freeTierMaxItems": 2
  },
  "errors": []
}
```

#### Owner configuration

These settings live in the Apify Console under **Settings → Environment** and can
be changed without a new build.

| Variable | Default | Meaning |
| --- | --- | --- |
| `FREE_TIER_MODE` | `limit` | `limit` caps results for free accounts; `block` exports nothing and stops immediately. |
| `FREE_TIER_MAX_ITEMS` | `2` | Maximum results exported for a free account in `limit` mode. |
| `PPE_EVENT_NAME` | `result` | Name of the pay-per-event used for billing. Must match the event configured in the Actor's monetization settings. |

> **Note on per-account caps:** there is no "max users per account" setting for
> this Actor. It processes stores, not users, so a per-account user cap does not
> map to anything meaningful here. `FREE_TIER_MAX_ITEMS` is the only cap applied.

#### Respecting your spending limit

If you set a maximum spend for a single run, the Actor honours it:

- The run stops as soon as the next result would exceed your limit.
- The run summary reports `spendingLimitReached: true` and
  `stoppedReason: "spending_limit"`.
- The run finishes cleanly with a clear notice.
- No further work is done after the budget is reached, so you never pay for
  results you did not receive.

***

### Performance

- **Concurrent search tasks** — several tasks run at once instead of one after another.
- **Concurrent lead details** — contact enrichment is bounded and parallelised.
- **Incremental output** — each record is saved as soon as it is ready, so memory
  stays light on long runs.
- **Fast paths** — when a store already shows enough contact data, the run does
  not look for more.
- **No duplicates** — results are deduplicated across tasks within a run.

#### How long should a run take?

Lead details add a little extra time per business. As a rough guide:

| `enableLeadDetails` | `enrichThemeApps` | Approximate time per business |
| --- | --- | --- |
| `false` | `false` | 1–3 seconds |
| `true` | `false` | 4–10 seconds |
| `true` | `true` | 6–14 seconds |

Actual timing depends on how responsive each store's site is. If you are
collecting tens of thousands of records, consider raising the run memory and
keeping lead details on only when you need contact data.

***

### Data quality

The Actor is tuned to keep output usable rather than merely complete.

- **Readable names** — a brand title is preferred over a bare domain.
- **Relevant email addresses** — addresses on the store's own domain are
  preferred, and platform, analytics and infrastructure addresses are removed.
- **Real phone numbers** — numbers are only accepted from genuine contact links
  or clearly formatted text, so product IDs and script values are excluded.
- **Real social profiles** — posts, reels, share links and default platform
  accounts are filtered out; only genuine profile links survive.
- **Consistent formatting** — phone numbers are normalised to a single format.
- **Nothing is filtered out** — if a field is missing, the record is still
  exported with `null` or an empty array. Completeness is never traded for
  enrichment quality.

### Error handling and privacy

- Every failure is converted into a **fixed, plain sentence** before it is shown.
- Technical details, hosts, addresses, connection settings and credentials are
  never written to logs, the dataset, webhook deliveries or the run summary.
- Logs never contain an unprocessed error object, even at the highest log level.
- Webhook deliveries contain only the record's own output fields.
- The run is configured with a generous time limit so long runs are not cut
  short unnecessarily.

### Responsible use

- Only public, business-facing information is collected.
- Please use the output in line with applicable data-protection and
  anti-spam laws, and with each platform's terms.
- Do not use the Actor for surveillance, harassment or unsolicited bulk messaging.

***

### FAQ

#### Does the Actor filter out stores without contact data?

No — and this is deliberate. **Enable lead details** adds contact and social data
as an *addition*. Every business found is exported, whether or not details were
available. Filtering would make the duration of a run unpredictable and would
make pricing impossible. Lead details simply add a little extra time per business.

#### Why do some records have no email address?

Not every store publishes a contact address. A missing email is normal and never
causes the store to be dropped — you still get its name, website, country, tags,
product count, social profiles and description.

#### Why is my free run limited to 2 results?

Free Apify accounts are intentionally limited to **2 results per run** so that the
Actor is usable as a trial. Upgrade to a paid Apify plan to get full, unlimited
data. The limit is applied before any work starts, and the run finishes cleanly
with a clear notice — it is a policy restriction, not a failure.

#### What is the difference between `limit` and `block` mode?

`limit` (the default) exports up to `FREE_TIER_MAX_ITEMS` records for a free
account and then finishes cleanly. `block` exports nothing and stops immediately,
still with a clear upgrade notice. Both are configured by the Actor owner.

#### What happens if I have a paid Apify plan?

You get full, unlimited output up to the total your tasks add up to. The run
logs `Paying user — full output.` and applies no cap.

#### Can I control my costs?

Yes, in two ways:

1. Set `maxResults` on each task to the exact number of records you want — the
   run total is the sum of those numbers.
2. Set a maximum spend for the run in Apify. The Actor stops as soon as that
   budget is reached and reports `spendingLimitReached: true`.

#### Why did my run stop before my task total?

The most common reasons are:

- Your **spending limit** was reached — check `spendingLimitReached` in the summary.
- You are on a **free account** — the run is capped at 2 results.
- The niche you chose returned **fewer matching stores** than you asked for.

#### My run returned 0 results. What should I check?

- Confirm the niche spelling — try a broader or more common term.
- Try removing the country filter.
- Try a nearby niche; the log will suggest related niches that do have results.
- Try again later if the warning mentions a temporary failure.

#### Can I combine a search with my own list of stores?

Yes. Leave `enableDirectoryBrowse` on and provide `storeUrls` with
`enableStoreDetails` enabled. Both sets of records are exported, tagged by
`featureType`.

#### Will I get duplicate stores?

Duplicates are removed within a run, matched on the store website. When you merge
results across several runs, deduplicate on `storeUrl` — lowercase it and strip
the trailing slash first.

#### Does the Actor respect a per-run spending limit?

Yes. Once your budget is exhausted the run stops immediately, no further work is
done, and the summary reports `stoppedReason: "spending_limit"`.

#### How do I get theme and installed-app data?

Set `enrichThemeApps` to `true`. The `theme` and `apps` fields will then be
populated. It adds a little extra time per business, so leave it off when you do
not need it.

#### Can I receive results in Slack?

Yes. Set `webhookUrl` to your Slack incoming webhook address and
`webhookFormat` to `slack`. For everything else, use `json`.

#### Can I run this from the API or MCP?

Yes. The input schema is fully documented above, and the Actor is ready for the
Apify MCP Server. See [MCP usage](#mcp-usage).

#### Does the Actor include source-system data?

No. The output is a curated, documented record structure. Internal source
structures are never included in the dataset or in webhook deliveries.

#### Is the collection method documented anywhere?

No. This documentation intentionally describes only **outcomes** — what you get
and how to use it.

#### Something failed and the message looks generic. Why?

By design. Failures are rewritten into fixed, safe sentences so that no host,
address, connection setting or credential can ever leak into a log, a dataset
row, a webhook delivery or the run summary. If a run fails repeatedly, try again
in a few minutes, lower `maxResults` on your tasks, or contact support.

***

### Development

```bash
npm install
npm run build
npm start:local
```

`npm start:local` reads `local.input.json` and writes newline-delimited JSON to
`output/local_results.jsonl`. Copy `local.input.example.json` to get started.
Local runs are never gated unless you explicitly set the plan variables.

#### Connections: local vs platform

These are deliberately kept separate:

- **On Apify**, runs always use the Apify residential connection from the
  `proxyConfiguration` input (US residential by default). No other connection
  settings are consulted.
- **Locally**, `npm start:local` uses the `PROXY_*` values from `.env` so you can
  exercise the Actor on a developer machine. Those values are never read by a
  platform run, so a testing connection can't be picked up by accident.

Useful environment variables for local testing:

| Variable | Purpose |
| --- | --- |
| `LOCAL_INPUT` | Path to an alternative input file. |
| `PROXY_HOST` / `PROXY_PORT` | Local-only connection settings. |
| `PROXY_USERNAME` / `PROXY_PASSWORD` | Local-only connection credentials. |
| `PROXY_COUNTRY` | Region for the local connection. |
| `FREE_TIER_MODE` | `limit` or `block`. |
| `FREE_TIER_MAX_ITEMS` | Free-tier result cap. |
| `APIFY_USER_IS_PAYING` | Set to `true` or `false` to simulate a platform account. |
| `APIFY_USER_PRICING_TIER` | e.g. `FREE`, `GOLD`. |

See `.env.example` for the full list. Never commit real credentials — `.env` is
git-ignored.

### License

ISC

# Actor input Schema

## `enableDirectoryBrowse` (type: `boolean`):

Discover stores by niche from the merchant directory. Enabled by default.

## `searchTasks` (type: `array`):

Define niches, keywords, or countries to discover. Multiple tasks run concurrently.

## `enableKeywordSearch` (type: `boolean`):

Enable full-text multi-keyword discovery across merchant niches.

## `keywords` (type: `array`):

Keywords to search across store directories.

## `enableStoreDetails` (type: `boolean`):

Enrich specific store domains or URLs directly without niche searching.

## `storeUrls` (type: `array`):

List of Shopify store website URLs to enrich (e.g. https://gymshark.com).

## `enableLeadDetails` (type: `boolean`):

Enrich stores with verified contact emails, phone numbers, and social media handles. Every merchant is still exported even if some contacts are missing. Adds a little extra time per business.

## `enrichThemeApps` (type: `boolean`):

Detect active Shopify theme name and list of installed Shopify apps for each store.

## `webhookUrl` (type: `string`):

Optional. Every record is always saved to the run's dataset — this webhook is an additional real-time push. When set, each new record is also POSTed to this URL (CRM, Slack incoming webhook, Zapier, Make, Google Sheets).

## `webhookFormat` (type: `string`):

json = full store record object; slack = Slack-formatted message block.

## `proxyConfiguration` (type: `object`):

Apify residential proxy (US) is enabled by default for reliable store and directory collection.

## Actor input object example

```json
{
  "enableDirectoryBrowse": true,
  "searchTasks": [
    {
      "niche": "clothing",
      "country": "US",
      "maxResults": 10
    }
  ],
  "enableKeywordSearch": false,
  "enableStoreDetails": false,
  "enableLeadDetails": true,
  "enrichThemeApps": false,
  "webhookUrl": "",
  "webhookFormat": "json",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `allResults` (type: `string`):

Complete dataset with every field from this run.

## `leads` (type: `string`):

Stores with contact emails, phone numbers, and social media.

## `techStack` (type: `string`):

Stores with detected Shopify theme and installed apps.

## `summary` (type: `string`):

Totals, stop reason, and the paywall state for this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "enableDirectoryBrowse": true,
    "searchTasks": [
        {
            "niche": "clothing",
            "country": "US",
            "maxResults": 10
        }
    ],
    "enableKeywordSearch": false,
    "enableStoreDetails": false,
    "enableLeadDetails": true,
    "enrichThemeApps": false,
    "webhookUrl": "",
    "webhookFormat": "json",
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("b2b_leads/shopify-real-time-data-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "enableDirectoryBrowse": True,
    "searchTasks": [{
            "niche": "clothing",
            "country": "US",
            "maxResults": 10,
        }],
    "enableKeywordSearch": False,
    "enableStoreDetails": False,
    "enableLeadDetails": True,
    "enrichThemeApps": False,
    "webhookUrl": "",
    "webhookFormat": "json",
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("b2b_leads/shopify-real-time-data-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "enableDirectoryBrowse": true,
  "searchTasks": [
    {
      "niche": "clothing",
      "country": "US",
      "maxResults": 10
    }
  ],
  "enableKeywordSearch": false,
  "enableStoreDetails": false,
  "enableLeadDetails": true,
  "enrichThemeApps": false,
  "webhookUrl": "",
  "webhookFormat": "json",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call b2b_leads/shopify-real-time-data-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,b2b_leads/shopify-real-time-data-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/irFLk8OCc2Nk9Wqwz/builds/LVcCbnydKqXrqBvcN/openapi.json
