Universal E-commerce Scraper avatar

Universal E-commerce Scraper

Pricing

from $2.50 / 1,000 results

Go to Apify Store
Universal E-commerce Scraper

Universal E-commerce Scraper

Turn any e-commerce category or product page into clean, structured data. Paste a category page or product URL.Some sites may block bots, but if reachable, the process is fully automated.

Pricing

from $2.50 / 1,000 results

Rating

0.0

(0)

Developer

Lofomachines

Lofomachines

Maintained by Community

Actor stats

0

Bookmarked

84

Total users

23

Monthly active users

1.5 days

Issues response

9 days ago

Last modified

Share

Universal E-commerce Scraper — Product Details, Prices, Specifications & Variants

Turn online store categories and product pages into a ready-to-use product dataset. Paste a URL to discover products, follow pagination, and collect product details in one run. Download JSON, CSV or Excel for competitor research, catalog enrichment, price monitoring and AI workflows.

No CSS selectors, separate URL extractor or coding required. Start with a small sample, check the results, then expand your run.

This Actor requires a paid Apify plan. A free-plan run returns an informational message instead of product data.

What can you collect?

  • Product names, descriptions, brands, SKU and barcode identifiers.
  • Prices, currencies, published original prices and discounts.
  • Main images and image galleries.
  • Availability, categories, ratings and review counts when published.
  • Technical specifications, preserving the store's original field names.
  • Product variants when exposed by the store's catalog or structured data.
  • Public payment promotions and installment plans, including available cards and extra discounts when visible.
  • Shipping conditions, published delivery charges and times, carriers, accepted payment methods, returns and warranty, when identifiable on the product page.
  • Product URLs, discovery URLs and extraction timestamps for traceability.

Fields depend on what the store publishes and what the Actor can access. Missing core values are omitted; payment plans and delivery/policy fields use "unavailable" when they cannot be identified. This is an adaptive scraper, not a guarantee of complete coverage on every website.

Start in three steps

  1. Paste a category, store homepage or product URL into Start URLs.
  2. Set Max items to your desired number. Keep Get product details enabled for richer results.
  3. Click Start, then open the dataset and download your preferred format.

A category URL can discover and enrich its products in a single run. You do not need to combine it with a separate URL extractor.

Example: Naldo category with 10 products

{
"startUrls": [{"url": "https://www.naldo.com.ar/climatizacion"}],
"maxItems": 10
}

Product details are enabled by default, including for API and MCP requests that only supply these two fields. For supported VTEX stores such as Naldo, a public catalog can provide descriptions, specifications, galleries and variants directly, without opening every product in a browser.

Example: multiple product pages

{
"startUrls": [
{"url": "https://example.com/products/blue-jacket"},
{"url": "https://example.com/products/green-jacket"}
],
"maxItems": 2,
"scrapeProductDetails": true
}

The example.com addresses illustrate the format; replace them with real product URLs.

Input options

OptionWhat it means
Start URLsOne or more public category or product URLs. Page type is detected automatically.
Max itemsMaximum saved products across all URLs. The Console suggests 100; 0 means no item limit. Run time and site limits still apply.
Get product detailsEnabled by default. Visit discovered products when needed; turn off for faster category-card results. Native catalog results can still contain details.
Proxy settingsOptional advanced override. Normally the Actor selects an available Apify proxy.

How automatic extraction works

The Actor first reads public catalogs, structured product information and recognizable page elements. For ambiguous fields, TypeSafe Jev chooses among values actually found on the page. The scraper copies the selected value and normalizes it in code. This keeps the AI from inventing prices or product attributes.

The adaptive layer runs automatically in the backend: no AI settings, prompts or API keys are required in the input. It learns selectors for a page structure and reuses them within the run, reading fresh values from each product. Changed structures trigger new decisions; uncertain selections are discarded. A Gemini fallback remains available for layouts that candidate extraction cannot resolve.

To keep large catalogs economical, valid learned selectors survive changes in optional page elements. Jev learning is limited to three attempts per domain and page type in a run, with a possible provider retry per attempt. Fields already present in reliable product data, including published warranty specifications, require no AI decision. Lightweight HTTP product retrieval runs concurrently while browser visits remain separately limited.

With automatic networking, the Actor first tries a direct HTTP connection and switches to the proxy if the store blocks it. This avoids residential traffic charges on accessible pages. Explicit proxy settings are respected; JavaScript browser visits still use the configured proxy. Actual savings depend on the store's access restrictions.

Categories whose products and page links are already present can be collected without opening a browser. Product collection starts immediately; interactive “load more” controls and incomplete pages still use browser support. A missing financing offer alone does not trigger an extra browser visit. The item limit is a maximum: a category containing 42 products returns 42, even when the limit is 200.

The same approach can select observed next-page and load-more controls. Every navigation must change the product set before it is accepted. Infinite scrolling is handled by observing newly loaded products. The Actor does not assume that an AI prediction proves a page has more products.

Supported native catalog paths include Shopify, WooCommerce and VTEX. Megatone category pagination and Frávega product details also have dedicated support. Other stores use page extraction, with platform profiles where available. Pagination includes next-page links, numbered controls, load-more buttons and infinite scrolling; unusual interfaces may require additional support.

A product's details are requested before its final record is saved. If a detail page cannot be read, the Actor can retain the discovered category-card data and label it listing-only, so incomplete records are visible.

Output example

Illustrative record — exact fields and values depend on the source:

{
"url": "https://example.com/products/blue-jacket",
"sourceUrl": "https://example.com/category/jackets",
"title": "Blue Waterproof Jacket",
"brand": "Example Outdoor",
"description": "Waterproof jacket with an adjustable hood.",
"price": 79.9,
"originalPrice": 99.9,
"discountPercent": 20.02,
"currency": "EUR",
"availability": "InStock",
"sku": "JKT-BLU-M",
"image": "https://example.com/images/jacket-front.jpg",
"images": ["https://example.com/images/jacket-front.jpg"],
"specifications": {"Material": "Recycled polyester", "Waterproof rating": "10000 mm"},
"variants": [{"sku": "JKT-BLU-M", "title": "Medium / Blue", "price": 79.9}],
"paymentPlans": [{"installments": 3, "installmentAmount": 26.63, "currency": "EUR", "interestFree": true, "paymentMethods": ["Visa"]}],
"shipping": {"rawText": "Delivery in Italy within 3–5 business days. Cost calculated by postal code.", "source": "dom", "scope": "product-page"},
"shippingCost": "unavailable",
"deliveryTime": {"rawText": "3–5 business days", "source": "dom", "scope": "product-page"},
"carriers": "unavailable",
"paymentMethods": "unavailable",
"returnsPolicy": "unavailable",
"warranty": "unavailable",
"detailStatus": "detail-page",
"extractionMethod": "http",
"scrapedAt": "2026-09-16T08:00:00+00:00"
}

Understand detail coverage

detailStatusMeaning
apiDetails came from a supported native catalog, including supported VTEX stores and Frávega.
detail-pageThe product page supplied descriptive and specification/variant data.
partialThe product page was read but the richer detail fields were not found.
listing-onlyOnly category-card information was available, or detail visits were disabled.

These labels describe the extraction path and observed fields; they do not certify that every attribute on the site was captured. Prices and availability can vary by location, seller or selected variant. Keep JSON for nested specifications and variants; CSV/Excel is convenient for flat comparisons.

Delivery and policy objects preserve the original conditions in rawText, or in conditions for structured data. A shipping charge requiring a postal code or checkout remains "unavailable". These fields describe public information, not a shipping quote for your address. fieldEvidence, when present, identifies adaptive selectors and their decision confidence; confidence is not a guarantee of accuracy. A cached selector reuses the original template confidence while reading the current page's value.

paymentPlans contains only promotions publicly visible during the run. Its value is "unavailable" when no plan can be identified. These offers can vary by card, bank, postal code, selected variant, stock and date; the product's base price is never replaced by a financing amount.

Some stores publish general provider terms instead of a repayment quote. Those records preserve the conditions in rawText, identify the provider, and use scope: "conditional-provider-offer" with eligibility: "unavailable". Purchase thresholds are never converted into installment amounts, and approval or eligibility is never assumed.

If an offer identifies a restricted issuing bank only by issuerId, that restriction is preserved and bank is "unavailable"; the offer must not be assumed valid for every bank.

Oscar Barbieri, Megatone and Cetrogar have dedicated support for their public financing data. Offers retain their product, card and bank associations. For Cetrogar, validity dates and eligible weekdays are checked; repeated category filters are preserved. These sources are read without opening a browser when available.

For these supported sources, financingStatus reports complete, partial or unavailable. partial means some source responses could not be recovered, so the returned plans may be incomplete. complete means the discovered source was checked, not that every customer qualifies. paymentPlans remains "unavailable" when no plan can be verified. Calculated amounts identify amountSource; cashback is not deducted automatically. Special plans such as Plan Z retain their label without interpreting a payment-system code as the installment count.

Offers from recommended products are excluded. Each plan includes paymentMethods when its cards can be identified from the offer; otherwise that field is "unavailable". A list of cards accepted by the store does not automatically apply to every installment plan. Some offers publish only a maximum installment count and card, without an amount; those records omit installmentAmount and totalAmount. Supported public calculators can also return the applicable bank and the retailer's displayed total. DOM-based totals are calculated from the displayed installment amount and may differ slightly because of rounding.

Competitor price monitoring: schedule category runs and compare price, availability and published discounts over time.

Product catalog enrichment: add descriptions, images and technical attributes to supplier catalogs or internal spreadsheets.

Market research: compare assortments, brands and product specifications across stores.

AI assistants and product search: feed structured product records into research assistants, recommendation workflows and retrieval systems.

Claude, MCP and automation

Use the Actor through Apify's MCP server, or call it via the Apify API. The input is the same: startUrls and an optional maxItems. Keep URLs as plain web addresses, without Markdown formatting.

The updated Python SDK accepts the MCP run origin. If you encounter a validation error mentioning meta.origin and MCP, check that your integration uses the updated latest build rather than an older pinned build.

Connect dataset results to Make, Zapier, n8n, Google Sheets or your own workflows using Apify integrations and APIs. Nested specification/variant fields may need mapping in your destination tool.

Frequently asked questions

Can it scrape a category and open each product in one run?

Yes. Get product details is on by default. Public catalog endpoints are used when they provide the data directly; otherwise the Actor attempts product pages, with browser fallback when needed.

Does it handle pagination automatically?

It attempts supported next-page, numbered, load-more and infinite-scroll patterns. It stops at the item limit, the end of accessible results, repeated pages or its safety limits. 0 removes the item cap, not those other limits. VTEX legacy search exposes at most 2,500 positions per query; larger catalogs should be split into categories.

Will every specification always be available?

No. Stores publish different fields, and some data is loaded through custom interactions or restricted endpoints. Check detailStatus and sample results before scheduling large runs. If you report an issue, include the URL, input and run link.

Does AI run on every product?

AI assistance is always enabled in the backend with Gemini; there is no input setting or model selector. It does not require an AI call for every product. Native catalogs and standard extraction are tried first. AI is used to help interpret unfamiliar layouts and can reuse selectors within the run. Model changes alone cannot fix inaccessible pages or missing source data.

Why are some records listing-only?

The product page could not be read, or detail visits were turned off. The category card is retained as a useful partial result rather than presented as a complete product record.

Can I scrape an entire store?

Use a supported store homepage or provide its category URLs with maxItems: 0. Category URLs are usually the best way to control scope. Filtered URLs, custom storefronts, access restrictions and run limits may reduce coverage.

How much does a run cost?

Cost depends on the Actor's current pricing, number of pages, browser use, proxies and any AI work. Check the Pricing tab and begin with 10 products to measure your store's actual results and cost.

Discover more tools for your workflow

Build your next data workflow with more Actors by Lofomachines — explore specialized scrapers and automation tools for your business.