Shopify Store Scraper - All Products & Live Prices
Pricing
from $0.50 / 1,000 product delivereds
Shopify Store Scraper - All Products & Live Prices
Scrape the full product catalogue of any Shopify store: titles, variants, SKUs, live prices, compare-at discounts, stock status, images, tags, product types and collections. Bulk multi-store, no API key, no login. Duplicate products from paginated boards are removed so you are never billed twice.
Pricing
from $0.50 / 1,000 product delivereds
Rating
0.0
(0)
Developer
DONGMIN KIM
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Shopify Store Scraper — All Products, Variants, Prices & Stock from Any Store

Point it at any Shopify store and get the entire product catalogue: every product, every variant, SKUs, live prices, compare-at discounts, stock status, images, tags, product types and optionally the store's collections.
Multiple stores per run. No API key, no store permission, no login.
A live run pulled 600 products in under a second from one store.
Why this is reliable
Every Shopify storefront publishes /products.json — it is part of the platform, not an oversight, and it has been stable for years. There is no HTML to parse here and no markup churn to chase, which is why this actor does not break the way theme-scraping tools do.
Two things it handles that a naive reader does not:
- Duplicate products. Shopify paginates by offset over a catalogue that shifts between requests, so the same product genuinely comes back on two pages — measured at 100 repeats in 600 rows on a live store. Those are dropped, so you are never billed twice for one product.
- Bot walls. A minority of stores sit behind a WAF that throttles datacenter IPs. The actor starts on cheap datacenter proxies and escalates to residential only for the stores that actually need it.
Input
{"storeUrls": ["gymshark.com", "https://allbirds.com/collections/mens"],"maxProductsPerStore": 1000,"onSaleOnly": true,"inStockOnly": false,"minPrice": 0,"productTypeContains": ["shoes"],"includeVariants": true,"includeCollections": false}
Bare domains or any URL on the store both work — it normalises to the origin.
Every option
The same wording you see in the Apify console, with the JSON key for API and MCP callers.
| Option | What it does | Default |
|---|---|---|
Shopify stores — storeUrls (required) | Bare domains or any URL on the store — gymshark.com, https://allbirds.com/collections/mens, shop.example.co.uk. Multiple stores per run. | — |
Max products per store — maxProductsPerStore | Products arrive 250 per request, so this is the main cost control. Leave high for a full catalogue. | 1000 |
On sale only — onSaleOnly | Keep only products whose compare-at price is above the live price. Filtered products are not billed. | false |
In stock only — inStockOnly | Keep only products with at least one available variant. | false |
Minimum price — minPrice | In the store's own currency. 0 disables. | 0 |
Maximum price — maxPrice | In the store's own currency. 0 disables. | 0 |
Product type or tag contains — productTypeContains | Keep only products whose product type or tags contain one of these. Case-insensitive. | — |
Include variant detail — includeVariants | Attach the full variant array (SKU, price, availability, options) to each product. Turn off for a slimmer dataset. | true |
Also scrape collections — includeCollections | Add one row per collection with its title, handle and product count — a map of how the store organises its catalogue. | false |
Skip non-Shopify domains — skipNonShopify | Check each domain is Shopify before scraping. Costs one cheap request per store and avoids wasted work. | true |
Concurrency — concurrency | Stores processed in parallel. | 3 |
Proxy — proxyConfiguration | Leave the default. A minority of stores sit behind a WAF that throttles datacenter IPs; the actor starts cheap and escalates to residential only for the stores that need it. | {"useApifyProxy":true} |
Output
One row per product. With Include collections on, each store also writes one row per
collection, marked type: "collection".
{"storeDomain": "allbirds.com","productId": 6806723428554,"title": "Men's Strider - Medium Grey","handle": "mens-strider","url": "https://allbirds.com/products/mens-strider","vendor": "Allbirds","productType": "Shoes","tags": "sale,mens,wool","createdAt": "2025-07-24T09:12:00-04:00","publishedAt": "2025-08-01T10:00:00-04:00","updatedAt": "2026-08-09T22:41:13-04:00","minPrice": 91,"maxPrice": 95,"compareAtPrice": 130,"onSale": true,"discountPercent": 30,"inStock": true,"variantCount": 13,"availableVariantCount": 2,"imageUrl": "https://cdn.shopify.com/…","imageCount": 7,"options": [{ "name": "Size", "values": ["9", "10", "11"] }],"description": "Soft & light. Wool upper.","variants": [{ "variantId": 1, "title": "9", "sku": "AB-9", "price": 91, "compareAtPrice": 130, "available": false, "option1": "9" }]}
Every field
You are billed per row delivered, so here is everything a row can contain. A field is absent when the store did not publish it.
Product rows
| Field | What it is |
|---|---|
storeDomain | Which store the row came from, so a multi-store run reads without a join. |
productId | Shopify's numeric product id. |
title | Product title. |
handle | The URL slug. It is the stable key across renames, and what you join on against a store's own exports. |
url | Canonical product URL. |
vendor | Brand as the store records it. |
productType | The store's own product type. |
tags | The store's tag string. |
createdAt | When the product was created in the store's admin — earlier than publishedAt, and the better signal for how long something has existed. |
publishedAt | When it went live on the storefront. |
updatedAt | Last edit of any kind, so a diff between runs tells you something changed. |
description | Body copy with HTML stripped. |
options | The option axes, each with name and values — e.g. Size and Colour. |
imageUrl | First image. |
imageCount | How many images the product has, without pulling them all. |
variantCount | Number of variants. |
availableVariantCount | How many of those are in stock — the difference is the size curve breaking. |
inStock | true when any variant is available. |
minPrice / maxPrice | Live price range across variants. |
compareAtPrice | Highest compare-at price found, which is how Shopify represents the pre-markdown price. |
onSale | true when a compare-at price sits above the live price. |
discountPercent | That markdown as a whole-number percent. |
variants | Per-variant rows: variantId, title, sku, price, compareAtPrice, available, requiresShipping, grams, option1–option3. Dropped when Include variants is off. |
onSale and discountPercent are derived from the compare-at price, so you get the
markdown without doing the arithmetic.
Collection rows (type: "collection")
| Field | What it is |
|---|---|
type | Always "collection" on these rows; product rows have no type. |
storeDomain | Which store the collection belongs to. |
collectionId | Shopify's numeric collection id. |
title | Collection title. |
handle | Its URL slug. |
url | Canonical collection URL. |
productsCount | How many products the store reports in it. |
updatedAt | Last time the collection changed. |
description | Collection body copy with HTML stripped. |
Who this is for
- Price analysts — schedule it and diff the price ladder over time, variant by variant.
- Buyers and merchandisers — what a competitor sells, in what types, at what price, and what is selling out.
- Sourcing and dropshipping teams — whole catalogues, filtered by price band, without asking anyone for access.
- Promotion trackers —
onSaleOnlyshows exactly what a store has marked down and by how much.
Common uses
- Competitor price monitoring — schedule it and diff prices over time.
- Assortment analysis — what a competitor sells, in what types, at what price ladder.
- Discount tracking —
onSaleOnly: trueshows exactly what a store has marked down and by how much. - Stock intelligence —
availableVariantCountreveals what is selling out. - Dropshipping and sourcing — pull catalogues and filter by price band.
- Feeding a comparison site or a price-tracking product.
Pricing
Pay per product delivered. Products removed by your filters, duplicate products, non-Shopify domains, and failed stores cost nothing.
Starting a run costs $0.00001 — the platform's $0.00001 minimum, charged once per GB of memory, and these Actors run on 512 MB.
Other Actors in this family
Same engines, same billing, no account or API key on any of them.
YouTube & video
- YouTube Scraper — No API Key, Any URL or Search — Any YouTube URL or search term in, videos out — with subtitles, comments and sponsor deals as add-ons.
- Download YouTube Subtitles in Bulk — SRT, VTT & Text — Bulk subtitles from videos, channels or playlists — text, SRT, VTT or RAG chunks.
- Export YouTube Comments to CSV — Replies and Likes — Every comment and reply thread, with likes, authors and creator flags.
- List Every Video on a YouTube Channel — Export to CSV — A channel's whole back catalogue plus a subscriber and RSS summary row.
- Find YouTube Sponsors — Brand Deals, Codes & Links — Which brands pay which creators, with the campaign link, the code and the timestamp.
- YouTube Search API — Bulk Results, No Quota — Many search terms at once, every result as a row, filtered before you are billed.
- Track Deleted YouTube Videos & Title Changes — What a channel quietly changed: deleted videos, rewritten titles, view velocity.
- YouTube Creator Email Finder & Sponsor Lookup — A channel list into leads: the published email, audience bands, and who already sponsors them.
- Export a YouTube Playlist to CSV — Every Video — Any playlist as a table, with each video position in it.
Search demand
- AnswerThePublic Alternative — Autocomplete Keyword API — One seed into hundreds of real keywords from Google, YouTube and Amazon autocomplete.
- Google Trends API — Today's Trending Searches, No Key — Today's trending searches by country, with traffic bands and the news behind them.
E-commerce
- New Shopify Product Alerts — Competitor Drop Tracker — Only what a store launched since the last run. Scanning is free.
- Shopify Store Email Finder — Qualified B2B Leads — A domain list into qualified leads: contact email, size, price band, and whether the shop still trades.
- Website Tech Stack & Email Finder for B2B Lists — Any domain list into leads: contact email, what the site runs on, and the marketing tags it carries.
Hiring
- Greenhouse, Lever & Ashby Job Scraper — No API Key — Paste a company domain, get its open roles from Greenhouse, Ashby, Lever or SmartRecruiters.
- Ghost Job Detector — Track Reposts, Closures & Edits — What changed on a careers page: opened, closed, quietly reposted, or a ghost job.
Run it from code
Nothing here needs a login to the source, only your Apify token.
HTTP — start a run and wait for the rows:
curl -X POST "https://api.apify.com/v2/acts/gganbukim~shopify-product-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "content-type: application/json" \-d @input.json
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('gganbukim/shopify-product-scraper').call(input);const { items } = await client.dataset(run.defaultDatasetId).listItems();
Python
from apify_client import ApifyClientclient = ApifyClient(os.environ["APIFY_TOKEN"])run = client.actor("gganbukim/shopify-product-scraper").call(run_input=input)items = client.dataset(run["defaultDatasetId"]).list_items().items
Scheduled or event-driven — attach a schedule to run it on a cron, or a webhook to push each finished run into your own endpoint. It also connects through Apify's Zapier, Make, n8n and LangChain integrations, and is reachable from an MCP server if you are driving it from an agent.
Standby / API mode — the run above is synchronous: one call in, rows out, no polling. That is the shape to use if you are calling this per request rather than in a batch.
Errors, limits and what you are charged for
- You pay for delivered rows only. A row your filters removed, a page that failed, a retry — none of it is billed. Starting a run costs $0.00001: the platform minimum, charged once per gigabyte, and this Actor runs on 512 MB.
- A run that delivers nothing still costs the start fee and nothing else. If the input resolved to zero items, the run fails loudly with the reason rather than finishing green on an empty dataset.
- Blocking is handled by changing address, not by waiting. The Actor starts on cheap datacenter proxies and moves up only after a tier has actually been refused several times in a row, then drops back down once the cheap tier answers cleanly again. You are not paying for residential bandwidth that was never needed.
- Rate limits belong to the source, not to this Actor. Very large inputs are worked through in batches; the run reports how many items succeeded, were filtered, and failed, so a partial result is never presented as a complete one.
- Dataset retention follows your Apify plan. Export what you need, or push it out with a webhook, if you want it past that window.
Is this legal?
This Actor reads pages and public endpoints that anyone can open in a browser without an account. It does not log in, does not defeat a paywall, and does not touch anything behind authentication.
Scraping public data is broadly lawful in the US and the EU, and courts have repeatedly said so — but "public" is not the same as "unrestricted", and what you may then do with the data is a separate question from whether you may collect it. Personal data pulls in the GDPR and similar regimes whatever the source, so if your rows contain people, you need a lawful basis for keeping them.
Apify publishes a fuller treatment in Is web scraping legal? and an ethical scraping guide. None of this is legal advice; if the use is commercial and the data is personal, ask someone qualified.
Something wrong, or missing?
Open an issue on the Actor's Issues tab — it goes straight to the developer and is the fastest route. Include the run ID; it carries the input and the log, which is usually enough to reproduce the problem without another round trip.
Sources change without warning, and a field that quietly goes null is worth reporting even if the run succeeded. A broken parser looks exactly like a quiet day in the data until someone says so.
FAQ
Will I get blocked or rate-limited? This reads /products.json, the endpoint Shopify itself serves on every shop for exactly this purpose. It is not defended, which is why the Actor runs on datacenter IPs and why the price can be what it is. Nothing here signs in to a store.
Does this need Shopify API access? No. It reads the public storefront endpoint that every Shopify store serves.
Does it work on every Shopify store? Almost all. A small number disable the endpoint (reported as a clear error) or sit behind a WAF (handled by proxy escalation).
Is the domain Shopify? The actor checks before scraping and skips non-Shopify domains rather than wasting your budget on them.
Can I get inventory quantities? No — Shopify exposes availability as a boolean publicly, not exact counts. availableVariantCount is the closest public signal.
Can I run it on a schedule? Yes, via Apify Schedules, webhooks, or the API. Also available over MCP for AI agents.
Is it legal to scrape Shopify stores? This reads the public /products.json endpoint
that Shopify itself serves on every storefront for exactly this purpose — no login,
nothing bypassed, no rate limit worked around. Prices, stock and titles are facts about
products on public sale. Each store's own terms are a separate contract question. Not
legal advice.
How much does 1,000 products cost? $0.50, plus $0.00002 for the run. Products your filters remove, duplicate rows from paginated boards, and stores that fail are never billed.
Can I export the results to Excel or Google Sheets? Yes. Every run's dataset downloads as CSV, Excel, JSON, XML or RSS from the Storage tab, or straight from the API if you want a live link a spreadsheet can pull.
Can I connect it to Zapier, Make or n8n? Yes — Apify publishes integrations for all three, plus webhooks that fire when a run finishes. A common setup is a schedule here and a webhook into your own database or Slack.
Do I need to write code? No. Fill the form in the console and press Start. If you do want code, the Apify client libraries for Python and JavaScript call this the same way, and it is available over MCP so an AI agent can call it directly.