Product Page Scraper - E-commerce Product Data to JSON avatar

Product Page Scraper - E-commerce Product Data to JSON

Pricing

$1.00 / 1,000 results

Go to Apify Store
Product Page Scraper - E-commerce Product Data to JSON

Product Page Scraper - E-commerce Product Data to JSON

Extract structured product data from a public product-page URL: title, price, currency, availability, brand, rating, review count, images, SKU and description via JSON-LD and microdata. Deterministic, no proxy and no headless browser needed.

Pricing

$1.00 / 1,000 results

Rating

0.0

(0)

Developer

Ahmed Moussa

Ahmed Moussa

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

14 days ago

Last modified

Share

E-commerce Product Extractor (single page)

Given a single public product-page URL (or a small bounded batch of URLs), this actor deterministically extracts structured product data:

url, status, title, price, currency, availability, brand, rating,
review_count, images[], sku, description, raw_prices[], method,
parse_confidence, extracted_at, error

What it is (honest scope)

This is single-page product extraction, not bulk store crawling. You give it the URL of one product page; it fetches that page once and parses the product data out of it. It does not spider an entire shop, follow category pages, or paginate through catalogues (that would require proxies and carries ToS/legal risk).

How extraction works (deterministic, code-only)

Extraction is pure code, tried most-reliable-first:

  1. JSON-LD / schema.org Productoffers (price, currency, availability), brand, aggregateRating (rating, review count), sku/mpn/gtin, image, description. This is the highest-confidence path.
  2. OpenGraph / product meta tagsog:title, og:image, product:price:amount, product:price:currency, product:availability, ...
  3. Plain <meta> / <title> / <h1> + visible-text heuristics — currency regex for price, keyword regex for availability.

The layer used is reported in method, and a code-owned parse_confidence (high/medium/low/none) is attached to every record.

Cost-safety ($0 idle, $0 uncovered per run)

  • No proxy — direct bounded HTTP GET.
  • No headless browser — static HTML fetch only.
  • No AI / LLM — pure deterministic parsing.
  • No paid third-party API.

The only cost is Apify platform compute for the run itself.

Always-on security (SSRF-guarded, fail-closed)

  • Private / loopback / link-local / reserved IPs are blocked (SSRF guard), re-validated on every redirect hop.
  • A domain blocklist rejects login-walled / litigation-magnet sites.
  • Hard caps: 5s connect / 10s read timeout, 2 MB body, 3 redirects.
  • The actor never raises — every URL yields a structured record (with status and error populated on failure: failed / blocked / empty).

Input

fieldtypenotes
urlstringsingle product-page URL
urlsarray of stringoptional bounded batch (capped at 50 per run)

Output

One dataset record per URL with the fields listed at the top.