Shopify Merchant Scraper
Pricing
from $4.99 / 1,000 results
Shopify Merchant Scraper
Pricing
from $4.99 / 1,000 results
Rating
0.0
(0)
Developer
Scraper Engine
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Shopify Merchant Scraper — Emails, Phones, Socials & Theme Data
Shopify Merchant Scraper turns a list of Shopify storefront URLs into one structured lead row per store: verified merchant identity, contact email and phone, seven-platform social links, product and collection counts, currency, theme identity, and a basic app/tech-stack inventory — delivered as clean JSON with no HTML parsing required. Paste your store URLs and start the run to see leads land in your dataset live.
What is Shopify Merchant Scraper?
Shopify Merchant Scraper is a lead-generation Actor that crawls a list of public Shopify storefronts and returns one merchant-lead row per store — contact details, social profiles, catalogue size, currency/market data, theme, and a basic app inventory, all as typed JSON. No Shopify account, login, or API key is required — it reads only what a normal site visitor and the store's own public JSON endpoints (/meta.json, /products.json) already expose. It's built for B2B prospecting teams, growth marketers, dropshipping researchers, and agencies who need a bulk, structured view of many Shopify merchants at once.
What Shopify merchant data is publicly available to scrape?
Shopify storefronts publish store identity, catalogue size, and market settings to any visitor and to unauthenticated JSON endpoints; only account-level backend data sits behind Shopify's own OAuth-gated Admin API.
| Data category | Publicly available | Requires Shopify Admin API (store owner's own OAuth grant) |
|---|---|---|
| Store name & merchant identity | ✅ page <title> + /meta.json | — |
| Contact email / phone | ✅ only if the merchant publishes it on the site | ❌ if undisclosed, it isn't available anywhere, including to the merchant's own Admin API consumers |
| Social profile links | ✅ if linked from the storefront | — |
| Product & collection counts | ✅ published counts via /meta.json | — |
| Currency, ships-to countries, accepted card brands | ✅ via /meta.json | — |
| Theme name & version | ✅ embedded in the storefront's own page HTML | — |
| Installed-app inventory (partial) | ✅ inferred from theme-extension and app-proxy URL patterns on the page | — |
| Orders, customers, revenue, inventory levels | ❌ | ✅ only for the store's own owner, after that merchant authorizes an app via OAuth |
Shopify Merchant Scraper only returns publicly visible data — what any visitor sees. Nothing behind a login wall.
What data can I extract with Shopify Merchant Scraper?
The Actor returns 35 fields per storefront, covering merchant identity, contact and social data, and commerce/theme/tech details.
| Field | Description |
|---|---|
storeName | Title of the storefront's homepage <title> tag |
merchantName | Registered merchant name, from /meta.json |
domain | Apex domain parsed from the input URL |
myshopifyDomain | The store's underlying *.myshopify.com domain |
storefrontStatus | open, password_protected, unreachable, or no_public_catalogue |
merchantCity / merchantProvince | Merchant's registered city/province, when disclosed via /meta.json |
merchantDescription | The merchant's own short store description |
url | Original input URL |
scrapedAt | UTC ISO-8601 timestamp (millisecond precision) of the scrape |
🧑💼 Contact & social fields
| Field | Description |
|---|---|
email | First validated contact email found (mailto: links + contact-page text) |
phone | First validated contact phone found (tel: hrefs first, then filtered free-text) |
facebook / instagram / twitter / tiktok / youtube / pinterest / linkedin | Linked social profile URL, or null if not published |
hasEmail / hasPhone | Booleans derived from email / phone |
socialPlatformCount | Count of the 7 social platforms found on the storefront |
📊 Commerce, theme & tech-stack fields
| Field | Description |
|---|---|
productCount | Authoritative published product count (from /meta.json, or a capped /products.json fallback — see the limitation below) |
publishedCollectionsCount | Authoritative published collection count |
currency | Default selling currency |
moneyFormat | The store's price-display format string |
shipsToCountries | Country codes the store ships to |
acceptedCardBrands | Card brands accepted via Shop Pay |
offersShopPayInstallments | Whether Shop Pay Installments is offered |
themeName / themeVersion | Active Shopify theme identity, read from the embedded Shopify.theme object |
themeAppExtensions | App handles detected from theme-app-extension asset paths in the page HTML |
appProxyHandles | App handles detected from /apps/<handle> app-proxy links on the storefront |
appCount | Distinct app handles detected (union of the two fields above) |
leadQuality | high / medium / low / none — derived only from confirmed fields already on the row (email present, phone present, ≥1 social found, storefront open) |
🤖 Add-on: Need additional Shopify data?
If your workflow needs full product catalogues rather than merchant leads, pair this Actor with Shopify Products Scraper (product- and variant-level data per store) or Shopify Scraper (title, brand, variants, pricing, images from homepages, collections, or single product pages). If you need a deeper store-profile report instead of contact leads — catalogue size, price range, and stock coverage — see Shopify Store Scraper.
How does Shopify Merchant Scraper differ from the official Shopify API?
Shopify's own Admin API only becomes usable after the target store's owner authorizes your app through Shopify's OAuth flow — per Shopify's published authentication docs (shopify.dev, checked 2026-08-15), an app only gains access to a store's data after that store's own owner approves the requested scopes. That makes it unusable for pulling lead data across a list of merchants you don't operate.
| Feature | Shopify Admin API | Shopify Merchant Scraper |
|---|---|---|
| Access requirement | OAuth approval from the store's own owner | No account, login, or API key |
| Scope | Only the one store that authorized your app | Any public Shopify storefront you list |
| Bulk cross-merchant leads | Not supported — one authorized store at a time | Built for this: a list of storefront URLs in one run |
| Contact/social extraction | Not part of the Admin API's product/order resources | Extracted directly (email, phone, 7 social platforms) |
| Setup per store | App review + OAuth flow, repeated per merchant | Paste URLs and start the run |
Use the Admin API when you operate the store yourself and need order, customer, or inventory management. Use Shopify Merchant Scraper when you need publicly available lead data across many merchants you have no relationship with.
How to use Shopify Merchant Scraper
Run it from the Apify Console — no separate signup or API key beyond your Apify account.
- Open Shopify Merchant Scraper on its Apify Store listing and click Try for free / Start.
- Provide the required input —
startUrls, one or more Shopify storefront URLs. - Optionally set
maxItems,concurrency,requestDelay, orproxyConfiguration. - Click Start and watch leads land in the Output tab as each store finishes.
- Download results as JSON or CSV, or stream them via the Apify API.
How to scale to bulk merchant lead extraction
startUrls is an array — bulk-paste a list, upload a file, or pipe it in from another Actor's dataset. There is no separate "run per store" step: one run processes the whole list, up to maxItems (max 10,000), with up to concurrency (max 50) stores scraped in parallel.
What can you do with Shopify merchant data?
- 🎯 B2B prospecting — outreach teams use
emailandphoneto build cold-outreach lists, filtering byleadQualityto work the best-fit leads first. - 🛒 Dropshipping & niche research — researchers use
productCount,themeName, andappProxyHandlesto see what a niche of stores sells and which apps power them. - 📊 Market sizing — analysts use
currencyandshipsToCountriesto segment a store list by target market. - 🤝 Influencer & agency outreach — agencies use
instagram,tiktok, andsocialPlatformCountalongsideemailto find brands active on specific social channels before pitching. - 🤖 AI agents & RAG pipelines — feed the typed JSON rows (
merchantDescription,themeName,productCount,leadQuality) into a lead-scoring agent or vector store to auto-prioritize an outreach queue.
How does Shopify Merchant Scraper handle rate limits and blocking?
The Actor starts every store with a direct connection (no proxy) so you don't burn proxy units on stores that don't need one. If a request comes back with a blocking status (403, 406, 407, 429, 451, 503) or a connection error, it escalates through an Apify Proxy connection ladder — direct → datacenter → residential — and stays on the escalated tier for the rest of that run rather than bouncing back and forth. Requests retry with exponential backoff plus jitter (up to 4 attempts for pages, 3 for JSON endpoints). If a store's homepage still can't be reached after retries, the Actor emits an honest row with storefrontStatus: "unreachable" instead of crashing the run or silently dropping the store.
⚠️ There is no headless browser here — pages are fetched as raw HTML via HTTP requests. Content injected purely by client-side JavaScript after page load (rather than present in the server-rendered HTML) will not be captured.
⬇️ Input
| Parameter | Required | Type | Description | Example value |
|---|---|---|---|---|
startUrls | ✅ Yes | array | One or more Shopify storefront URLs (e.g. https://kyliecosmetics.com/). Bulk paste, upload a list, or pipe from another actor. | ["https://kyliecosmetics.com/"] |
maxItems | ❌ No | integer | Hard cap on how many of the provided URLs will be processed. Use this to keep small test runs cheap. Default 100, min 1, max 10000. | 100 |
concurrency | ❌ No | integer | How many stores to scrape in parallel. Default 10, min 1, max 50. | 10 |
requestDelay | ❌ No | number | Delay (seconds) between requests within a single store, to be polite to the target site. Default 0.5, min 0, max 10. | 0.5 |
proxyConfiguration | ❌ No | object | Use Apify Proxy, custom proxy URLs, or no proxy. Default {"useApifyProxy": false}. | {"useApifyProxy": false} |
Example input
{"startUrls": ["https://kyliecosmetics.com/","https://www.allbirds.com/"],"maxItems": 100,"concurrency": 10,"requestDelay": 0.5,"proxyConfiguration": { "useApifyProxy": false }}
⬆️ Output
One typed JSON row per storefront, pushed to the dataset the moment each store finishes — the schema is identical across every run. Export from the Output tab to JSON, CSV, Excel, or any other format the Apify dataset export supports, or pull it via the API/apify_client.
This Actor uses Apify's Pay-Per-Event pricing: it bills the row_result event once per storefront processed. ⚠️ Every URL you submit charges a row_result row, even when the result is a blank lead — an unreachable, password_protected, or no_public_catalogue storefront still produces and charges one row, since the Actor writes a result for every input rather than silently skipping failures. Filter those out downstream with storefrontStatus == "open" if you only want to pay attention to usable leads. If your run hits its configured charge limit for row_result, the Actor stops dispatching new stores and exits cleanly rather than continuing to accrue charges.
Example output
{"storeName": "Kylie Cosmetics by Kylie Jenner | Kylie Jenner Fragrances | Kylie Skin","merchantName": "Kylie Cosmetics","domain": "kyliecosmetics.com","myshopifyDomain": "kylie-cosmetics.myshopify.com","storefrontStatus": "open","email": "customerservice@kyliecosmetics.com","phone": "1-877-916-6128","facebook": "https://www.facebook.com/KylieCosmetics/","instagram": "https://www.instagram.com/kyliecosmetics/","twitter": "https://twitter.com/kyliecosmetics","tiktok": null,"youtube": null,"pinterest": null,"linkedin": null,"productCount": 238,"publishedCollectionsCount": 226,"currency": "USD","moneyFormat": "${{amount}}","merchantCity": "Los Angeles","merchantProvince": "California","merchantDescription": "Shop Kylie Cosmetics by Kylie Jenner.","shipsToCountries": ["US", "CA", "GB", "AU"],"acceptedCardBrands": ["visa", "mastercard", "american_express"],"offersShopPayInstallments": true,"themeName": "KYLIE","themeVersion": "12.0.0","themeAppExtensions": ["klaviyo", "yotpo-reviews"],"appProxyHandles": ["loyalty"],"appCount": 3,"hasEmail": true,"hasPhone": true,"socialPlatformCount": 3,"leadQuality": "high","url": "https://kyliecosmetics.com/","scrapedAt": "2026-08-13T10:03:57.462Z"}
How does it work?
For each store, the Actor first requests the homepage directly (no proxy unless already escalated). It reads the page <title>, scans the raw HTML for social-platform URL patterns, pulls the theme name/version out of the storefront's own embedded Shopify.theme script object, and infers installed apps from theme-app-extension and /apps/<handle> proxy paths already present on the page. It then requests a small set of contact pages (/pages/contact, /contact, /pages/contact-us, /policies/contact-information) and extracts the first validated email and phone number found. Finally it fetches the store's own /meta.json endpoint for the authoritative merchant name, description, location, currency, shipping countries, card brands, and published product/collection counts — falling back to a capped /products.json?limit=250 count only when /meta.json doesn't supply one. Only data already public on the storefront is returned, and the output schema stays the same regardless of how a store's theme or layout looks.
Integrations
Shopify Merchant Scraper runs on the Apify platform, so it works with anything that can call the Apify API.
Calling Shopify Merchant Scraper programmatically
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("shopify-merchant-scraper").call(run_input={"startUrls": ["https://kyliecosmetics.com/"],"maxItems": 100,"concurrency": 10,})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
Works in Go, Ruby, Node.js, cURL — any language that can make an HTTP request.
No-code tools (n8n, Make)
In n8n, use the HTTP Request node pointed at the Actor's run-sync-get-dataset-items endpoint with your Apify token to get leads back in the same workflow run. In Make, an equivalent HTTP module call triggers a run and reads the resulting dataset items directly into your scenario — no custom code needed either way.
⚖️ Is it legal to scrape Shopify merchant leads?
Scraping publicly available business contact data from Shopify storefronts is generally legal, but the contact fields this Actor returns (email, phone) can constitute personal data under privacy law when they identify an individual, so GDPR/CCPA framing applies to how you store and use them. Shopify Merchant Scraper only returns data the merchant has already made publicly visible on their storefront — nothing behind a login. You are responsible for having a lawful basis to store and process that contact data (e.g. legitimate interest for B2B outreach, honoring opt-out/unsubscribe requests) and for complying with each store's own terms of service and applicable anti-spam law in your outreach. Consult legal counsel if your use case involves bulk storage of personal data.
❓ Frequently asked questions
What Shopify merchant lead fields does Shopify Merchant Scraper return?
The core fields are email, phone, the 7 social profile links (facebook, instagram, twitter, tiktok, youtube, pinterest, linkedin), leadQuality, and productCount. See What data can I extract for the full 35-field list.
Does Shopify Merchant Scraper require a Shopify account or login?
No. It sends unauthenticated HTTP requests to each storefront's public pages and public JSON endpoints (/meta.json, /products.json) — the same data any site visitor can see, with no Shopify credentials involved.
How many merchant leads can I extract in one run?
Up to maxItems storefronts per run — default 100, maximum 10,000 — with up to concurrency (max 50) stores processed in parallel.
What happens if a storefront is unreachable, password-protected, or has no public catalogue?
The Actor always returns a row, never a silent gap. storefrontStatus is set to unreachable if the homepage can't be fetched after retries, password_protected if the store gates its homepage behind Shopify's password screen, or no_public_catalogue if neither /meta.json nor the products feed returns usable data. In each case the other fields are returned as null or empty rather than guessed.
Can I scrape multiple Shopify stores at once?
Yes. startUrls accepts an array of storefront URLs, and concurrency (default 10, max 50) controls how many are scraped in parallel within one run.
Does Shopify Merchant Scraper work with Claude, ChatGPT, and other AI agent tools?
It's callable as a standard HTTP/REST endpoint via the Apify API or apify_client, so any agent framework capable of making an HTTP call can trigger a run and read back the resulting JSON.
How does Shopify Merchant Scraper compare to other Shopify scrapers in this account?
It's the only one of the account's Shopify Actors built specifically for merchant leads rather than product catalogues: Shopify Products Scraper and Shopify Scraper return product/variant/pricing data, and Shopify Store Scraper returns a store-profile report (theme, apps, catalogue size, price range, stock coverage). Shopify Merchant Scraper is the one that returns contact email, phone, social links, and a derived leadQuality score.
Does Shopify Merchant Scraper return data in a format LLMs can use directly?
Yes. Every row is typed, normalized JSON with consistent field names across runs — no HTML parsing or selectors needed. Pass it directly to an LLM, index it into a vector store, or feed it to an agent tool.
What happens when Shopify changes its storefront layout or endpoints?
The output schema is designed to stay stable regardless of a store's individual theme or layout, since extraction relies on Shopify's own structured endpoints (/meta.json, embedded Shopify.theme object) rather than page-specific CSS selectors. No specific update turnaround time is published.
Can I use Shopify Merchant Scraper without managing proxies or browser infrastructure?
Yes. The Actor runs direct by default and only escalates to Apify Proxy (datacenter, then residential) automatically when a store temporarily blocks requests — you don't configure or pay for proxy infrastructure unless you explicitly opt in via proxyConfiguration.
Which Shopify merchant fields work best for AI training data and RAG indexing?
For RAG, index merchantDescription, storeName, and themeName — the free-text fields with the most useful natural-language content. For structured training data, productCount, publishedCollectionsCount, socialPlatformCount, and leadQuality are the most consistently typed across records.
🔗 Related scrapers
| Scraper | What it extracts |
|---|---|
| Shopify Products Scraper | Product- and variant-level data (price, stock, options) per store |
| Shopify Scraper | Product title, brand, variants, pricing, and images from homepages, collections, or single product pages |
| Shopify Store Scraper | Store-profile report: theme, installed apps, catalogue size, collection count, price range, stock coverage |
💬 Your feedback
Found a bug or missing a field? Open an issue from the Actor's Issues tab on the Apify Console so it can be tracked and fixed.