Amazon Product Scraper | $0.50 per 1K Results avatar

Amazon Product Scraper | $0.50 per 1K Results

Pricing

from $0.50 / 1,000 amazon product results

Go to Apify Store
Amazon Product Scraper | $0.50 per 1K Results

Amazon Product Scraper | $0.50 per 1K Results

Scrape Amazon products by keyword, URL, ASIN, category, or seller across 20 marketplaces. Extract prices, ratings, reviews, sellers, images, variants, availability, and bestseller ranks. Run in Console or by API, schedule Tasks, and pay $0.50 per 1,000 results on paid plans.

Pricing

from $0.50 / 1,000 amazon product results

Rating

0.0

(0)

Developer

Tovuk

Tovuk

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Amazon Product Scraper is a structured Amazon data scraper for search, product, category-chart, bestseller, seller, review, offer, and question pages. Paid-plan prices start at $0.50 per 1,000 Dataset results, with no start fee, usage surcharge, or add-on event. Results are normalized products when Amazon exposes a usable record, or one transparent diagnostic item when a handled run produces no product.

Use the Amazon product scraper from the Apify Console, through the Amazon scraper API, from Python or JavaScript, or as a scheduled Task. You do not need to provide a proxy, cookie, captcha service, Amazon account, or third-party API key. Start with one query, URL, or Amazon Standard Identification Number (ASIN), then increase the bounded run limits after checking the first dataset.

Why use this Amazon data scraper?

  • Compare product prices, discounts, stock signals, shipping text, and visible sellers
  • Monitor search position, sponsored placement, bestseller rank, badges, and category visibility
  • Enrich product catalogs with ASINs, identifiers, specifications, variants, images, and descriptions
  • Analyze ratings, review counts, visible review samples, questions, and customer-photo signals
  • Build marketplace research datasets across 20 Amazon domains with consistent normalized fields
  • Feed spreadsheets, databases, dashboards, alerts, scheduled workflows, and AI agents

The Actor keeps search, product, review, offer, seller, bestseller, ASIN, and question workflows in one tested output contract. For a single workflow and a more specific Store listing, use one of the focused Amazon Actors linked in the related-Actors section after completing your first run.

How to scrape Amazon product data

  1. Open the Input tab.
  2. Enter one search query, supported Amazon URL, or ASIN.
  3. Keep maxItems at 10 and maxPages at 1 for the first run.
  4. Click Start, inspect the Dataset, and open the Actor outputs for the run summary or diagnostic.
  5. Confirm the fields and marketplace, then increase limits or save the input as a Task.

Amazon product scraper input example

Use this input for a low-cost first run:

{
"query": "usb c cable",
"marketplace": "amazon.com",
"maxItems": 10,
"maxPages": 1
}

The Actor infers a search operation, requests up to 10 products, and writes all available product fields to the default Dataset. If a handled run produces no product, it writes one billed recordType: "diagnostic" Dataset item explaining the cause, next action, retryability, and relevant Actor documentation. DIAGNOSTIC preserves the exact immutable diagnostic publication body, and a terminal RUN_SUMMARY confirms the settled Dataset outcome. If Dataset visibility or publication state is temporarily undecidable, the Actor writes only provisional, non-billed recovery evidence and deliberately withholds a terminal summary and any further Dataset write.

An actual run of the input above returned this Dataset item on July 24, 2026. Amazon listings change over time, so treat the values as a continuous input-to-output example rather than a guaranteed current offer:

{
"recordType": "product",
"marketplace": "amazon",
"operation": "search",
"title": "Anker USB C to USB C Cable, 60W Fast Charging Cable (2-Pack, 6 ft, Black) | For iPhone 17 Series, iPad mini 6, and More",
"asin": "B088NRLMPV",
"searchPage": 1,
"resultPosition": 1,
"priceText": "$9.99",
"rating": 4.7,
"reviewCount": 85571,
"extractedAt": "2026-07-24T11:45:21.067Z"
}

What you can collect

Available fields depend on the page, marketplace, delivery context, and content Amazon renders during the run.

  • Product identity: title, ASIN, product and parent IDs, brand, manufacturer, model and part numbers, SKU, GTIN, MPN, URL, and source URL
  • Search and category-chart data: search page, result position, sponsored position, category path, breadcrumbs, browse-node IDs, badges, and structured bestseller ranks
  • Pricing and fulfillment: price, list price, sale price, discount, currency, condition, availability, stock, quantity, shipping, Prime eligibility, and delivery context
  • Seller and offer data: seller name and ID, seller URL, ships-from, sold-by, fulfillment, offer count, visible price range, and structured offer details
  • Product content: description, feature bullets, specifications, package quantity, color, size, style, variation theme, variants, primary image, image gallery, videos, keywords, and tags
  • Customer signals: rating, rating count and distribution, review count, visible review samples, question count, visible question samples, and customer photos
  • Extraction metadata: language, region, detail-enrichment provenance, requested features, extraction quality, and per-feature coverage

The Actor preserves additional normalized fields when the output contract gains new data. It does not discard rich fields unless you set outputFields.

Choose a target

Provide at least one target field. Unknown fields are rejected before a data request starts.

FieldLimitUse
query1Run one Amazon search
searchQueries50Run several search phrases
startUrls200Provide supported Amazon page URLs for one compatible operation
productUrls200Collect product-page records
asins200Resolve ten-character ASINs in the selected marketplace
categoryUrls100Collect bestseller or category-chart pages; use startUrls for search-category URLs
categoryIds100Build bestseller targets from category IDs
sellerUrls100Collect seller or storefront pages

Each URL can contain up to 2,048 characters. The Actor accepts HTTPS URLs only from supported Amazon domains; private URLs are rejected.

Select an operation

Leave operation set to auto to infer one compatible operation for the whole run. Every supplied target must match that inferred operation. Split incompatible page types into separate runs. An explicit operation does not transform one page type into another, so every supplied target must match the selected operation.

OperationMatching target and output
searchSearch query or search URL; normalized product-result records
productProduct URL or ASIN; visible product details
reviewsReview page; product record with visible review samples and counts
offersOffer page; product record with visible offer, seller, price, and shipping fields
sellerSeller or storefront URL; visible storefront and product fields
bestsellersBestseller or category-chart URL, or category ID; ranked product records
asinsASIN list; normalized product records
questionsQuestion page; product record with visible question samples and counts

Review, offer, seller, and question data are embedded in normalized product-centric records. The Actor does not promise one complete dataset row for every individual review, offer, seller, or answer.

Set marketplace and delivery context

The Actor supports 20 Amazon marketplaces:

amazon.ae, amazon.ca, amazon.co.jp, amazon.co.uk, amazon.com, amazon.com.au, amazon.com.be, amazon.com.br, amazon.com.mx, amazon.com.tr, amazon.de, amazon.es, amazon.fr, amazon.in, amazon.it, amazon.nl, amazon.pl, amazon.sa, amazon.se, and amazon.sg.

Use one marketplace per run. The selected marketplace, every supplied Amazon URL, and every result identity URL must resolve to that same marketplace or one of its subdomains. URL-free runs default to amazon.com. You can also provide:

FieldPurpose
deliveryLocationCity, ZIP code, or postal code for delivery-localized content
proxyCountryCodeCountry hint for managed retrieval, without proxy credentials
proxyCityInternational city hint for managed retrieval
languageLanguage or locale hint such as en-US or de-DE

Use a city or postal code instead of a personal street address.

Filter Amazon products and review samples

The Actor applies filters after parsing Amazon records. This avoids returning a row merely because its source card was fetched. An active filter can return fewer than maxItems when Amazon exposes too few matching records.

FieldEffect
minRatingKeep products at or above a rating from 0 to 5
minReviewsKeep products at or above a parsed review count
titleIncludeKeywordsKeep titles containing at least one term
titleExcludeKeywordsRemove titles containing any term
excludeSponsoredRemove rows identified by a sponsored badge or ad position
maxReviewsLimit nested reviewSamples after filtering; 0 keeps every available sample
reviewRatingsKeep nested review samples in selected 1-to-5-star buckets
reviewKeywordsKeep nested review samples containing at least one term
reviewsStartDateKeep nested review samples on or after an inclusive YYYY-MM-DD date
reviewsEndDateKeep nested review samples on or before an inclusive YYYY-MM-DD date
reviewSortOrder nested review samples by relevance, date, or rating

Title and review-keyword matches are case-insensitive. Review controls filter the samples Amazon exposes on the fetched page; they do not create a separate complete row for every review.

Control run size

FieldRangeDefault
maxItems0 for no item ceiling; 1 to 1,000,000,000 for an explicit ceiling10
maxPages1 to 1,000,000,0001
continuationCursorSigned opaque cursor from a previous CONTINUATION recordEmpty
enablePaginationBooleantrue
maxConcurrency1 to 44
timeoutSecs30 to 3,600240

maxItems is the optional product-record ceiling. Set it to 0 when pages, targets, timeout, and the Apify spending limit should be the only run bounds. The final product count can be lower when a page has fewer records, Amazon limits pagination, a filter removes records, or the spending limit is reached. Large runs can take several minutes, so increase timeoutSecs and the Apify run timeout together. A direct target that produces no usable record completes with one billed actionable diagnostic item instead of disappearing or inventing a product.

If you override the Apify run timeout, keep it at least 60 seconds above timeoutSecs so the Actor can persist its final summary and any diagnostic.

For large search, seller, and bestseller jobs, the Actor splits source work into bounded internal requests of at most one source page and 1,000 service results. It follows the service's opaque source cursor across those chunks, including page 1,001 and later, without page arithmetic. It streams, validates, globally deduplicates, and repositions each completed request in the same Dataset instead of retaining the full run in memory. These internal safety boundaries are not user-facing result or page ceilings. The Actor also stops before publishing rows that exceed the run spending limit.

Before starting the request, the Actor normalizes and deduplicates exact target values so repeated ASINs, queries, identifiers, or URLs do not create duplicate service work.

CONTINUATION is written before request creation, after the returned request ID is known, before every result-page read, and after every durable result page. It records the request phase, first and last request IDs, request count, separate result and source cursors, aggregate fetched and published counts, target progress, and source-plan completion. continuationCursor is an authenticated opaque cursor for a fresh run; copy it exactly and keep the same targets, filters, marketplace, operation, and output fields. Never edit or decode an opaque cursor, and never use the raw internal request ID as an input.

When present, terminal RUN_SUMMARY is authoritative over provisional recovery records. It distinguishes fetched products, published products, published diagnostics, and total billed Dataset items. It identifies whether a diagnostic is present and repeats its safe code, message, next action, retryability, and documentation link. It also reports request identity, requested and effective product limits, result pages, elapsed time, and a continuationStatus of complete, item-limit, spending-limit, interrupted, or not-started. An unresolved Dataset-visibility or publication decision intentionally has no terminal RUN_SUMMARY: RECOVERY_STATUS carries visibility-only evidence, while PUBLICATION_RECOVERY preserves the exact unbilled diagnostic body awaiting fenced recovery.

The Actor automatically drains every stored-result cursor and launches the next bounded source request while the run has capacity. Request creation uses one deterministic, account-scoped idempotency key per logical source chunk, so retrying a lost POST response resolves to the original request. Deduplication by validated product URL spans all chunks, and result positions remain contiguous. If an API reader disconnects, the Apify run continues independently; retain the Apify run ID, reconnect later, and page through its stored default Dataset.

On an Apify migration or process restart, the Actor reads CONTINUATION, polls the exact in-flight request or replays the same idempotent create, scans the default Dataset, and reconciles product identities through the fenced backend ledger before publishing again. A Dataset row written just before interruption is therefore not published or billed a second time. The runtime also reserves 30 seconds before ACTOR_TIMEOUT_AT to persist its latest checkpoint, summary, or diagnostic. For a fresh run after an item, spending, or time boundary, copy the authenticated opaque continuationCursor; the Actor verifies its current version, cryptographic authenticity, and logical-input fingerprint before resuming.

Limit dataset columns

The Dataset contract declares 96 properties. Product items use up to 88: 86 projectable Amazon fields plus Actor-owned recordType and extractedAt. Diagnostic items use eight additional diagnostic-only fields plus the shared recordType, marketplace, and optional operation. Leave outputFields empty to return every available product field. To reduce product width, select up to 86 supported result fields:

{
"query": "mechanical keyboard",
"maxItems": 50,
"maxPages": 2,
"outputFields": [
"title",
"asin",
"priceText",
"rating",
"reviewCount",
"imageUrls"
]
}

Required identity, provenance, and Actor-owned fields stay present when applicable. The Actor always retains recordType, marketplace, operation, and extractedAt. Product rows also retain sourceUrl and url so every product remains identifiable.

Use dataset views

The Dataset includes these views:

  • Overview shows identity, rank, price, rating, stock, seller, image, URL, and extraction time
  • Diagnostics shows stable cause, explanation, recommended action, retryability, context, and documentation
  • Pricing and availability shows price, discount, stock, shipping, delivery, seller, fulfillment, and structured offer fields
  • Product details shows descriptions, identifiers, catalog attributes, categories, ranks, specifications, variants, and media
  • Reviews and questions shows ratings, rating distribution, counts, visible samples, and customer photos
  • Extraction quality shows source and detail provenance, parser classification, requested features, and coverage

Open the Dataset tab to export JSON, CSV, Excel, XML, RSS, or JSONL. You can also read the Dataset through the Apify API, save inputs as Tasks, schedule recurring runs, trigger webhooks, and connect supported Apify integrations.

Use the Amazon product data scraper API from Python

Install the official apify-client package, keep your API token in the APIFY_TOKEN environment variable, and call the same validated input used in the Console:

import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("tovuk/amazon-product-scraper").call(
run_input={
"query": "usb c cable",
"marketplace": "amazon.com",
"maxItems": 10,
"maxPages": 1,
}
)
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())

The call waits for completion and then reads the default Dataset. For asynchronous systems, start a run through the REST API, retain its run ID, poll until the run reaches a terminal status, and then read defaultDatasetId.

Use a focused Actor when every run has one intent. Each listing shares this Actor's tested runtime and output contract while exposing a smaller input surface:

The Amazon Reviews Scraper remains private until its complete review workflow passes the same live acceptance standard as the product Actors.

Amazon product scraper FAQ

Can I scrape more than 1,000 Amazon products in one run?

Yes. Set maxItems to 0 for no separate item ceiling, then set maxPages, timeoutSecs, and the Apify maximum cost per run for the work you want. Large listing jobs use bounded internal requests and one deduplicated Dataset. Amazon can expose fewer matching products, and active filters or spending limits can lower the final count.

Can an API client or agent continue from where it stopped?

Yes. When only the API client disconnects, keep the Apify run ID, poll that run after reconnecting, then page through its default Dataset. Apify migrations and same-run process restarts recover automatically from durable checkpoints. To continue in a fresh run after an item, spending, or time boundary, copy CONTINUATION.continuationCursor into the new input without changing the logical target or output fields. A complete state means the explicit source plan finished; item-limit or spending-limit explains why it stopped earlier.

Can I use the Amazon product scraper from Python?

Yes. Use the Python example above with apify-client, or call the Actor REST API from any language. The Console, API, scheduled Tasks, and integrations all use the same validated input and Dataset contract.

Can I scrape Amazon products by ASIN, seller, category, or keyword?

Yes. Use asins, sellerUrls, categoryUrls or categoryIds, and query or searchQueries. Keep one compatible operation and marketplace in each run.

Does the Actor return Amazon reviews and offers?

Product records can include visible review samples, review counts, offer counts, sellers, prices, and shipping fields when Amazon exposes them. The product scraper does not promise one complete row for every review or marketplace offer.

Pricing and spending limits

This Actor uses one pay-per-event charge. One transparent record price covers all available fields. The current rates below apply until the scheduled pricing change on August 8, 2026 at 02:23 UTC. The live Pricing panel remains the authoritative source for your plan's price.

Billing itemPriceWhen it applies
Actor start$0No start event is configured, so starting a run is free
Default Dataset itemPlan price belowOne visible product record or one transparent handled diagnostic
Other events$0No query, page, field, review, offer, seller, or add-on events
Platform usage surcharge$0The Dataset-item price includes platform usage
Apify planPer Dataset item10 items100 items1,000 items
Free$0.05000$0.50000$5.00$50.00
Bronze$0.00050$0.00500$0.050$0.50
Silver$0.00050$0.00500$0.050$0.50
Gold$0.00050$0.00500$0.050$0.50
Platinum$0.00050$0.00500$0.050$0.50
Diamond$0.00050$0.00500$0.050$0.50

The Free-plan Dataset-item price is exactly 100 times the paid-plan Dataset-item price. Bronze, Silver, Gold, Platinum, and Diamond all use the same $0.00050 rate, so upgrading to any paid plan unlocks the full 99% price reduction while the Free plan remains suitable for bounded evaluation runs.

The release configuration omits apify-actor-start. You pay no start fee, and the run does not receive the five-second compute coverage attached to Apify's synthetic start event.

One Dataset-item charge maps to one visible product record or one transparent diagnostic item. Apify automatically charges apify-default-dataset-item for each accepted row written to the default Dataset. The Actor does not issue a separate custom charge. Product items use recordType: "product". When a handled invalid, empty, or recoverable run publishes no product and Dataset state is decidable, it writes exactly one recordType: "diagnostic" item and preserves the same immutable body under DIAGNOSTIC for direct access by users and agents.

Set a positive maxItems value to bound product results, or use 0 to rely on page, target, timeout, and spending bounds. A diagnostic can add one Dataset item beyond the product count. Set the Apify maximum cost per run for an independent monetary cap. The Actor checks remaining capacity before service work and before every result page and stops when Apify reports that the spending limit was reached. If the platform reports zero remaining event capacity, it can prevent the diagnostic Dataset item too; KVS retains the safe outcome without claiming a billed row. Dataset-visibility uncertainty is a second honest zero-row exception: the Actor records provisional recovery evidence, issues no Dataset write, and waits for a later safe reconciliation before producing a terminal summary.

Data coverage and limitations

  • The Actor collects Amazon pages available without signing in
  • Amazon can vary content by marketplace, location, language, session, and time
  • A field can be null, empty, or absent when Amazon does not render it
  • Search and chart pagination can end before the configured page limit
  • Review and question modes return visible samples in product-centric records, not private account content or guaranteed complete histories
  • Offer data reflects visible page content and does not guarantee every marketplace seller or buy-box change
  • Numeric bestseller rank is returned only when the page exposes a supported rank signal
  • One run captures current values; schedule repeated runs to build a time series
  • The Actor does not provide historical prices from before your first run
  • Amazon page changes and transient blocking can reduce output; handled failures complete with a retryable diagnostic

Troubleshooting

Input is rejected

Provide at least one supported target and remove unknown fields. Check array limits, URL length, marketplace domain, ASIN shape, and operation-target compatibility.

Apify validates the declared input schema before starting the Actor. A structurally invalid API or Console input has no Actor run, Dataset, key-value store, or billable event, so follow the platform validation message and the generated Input/API documentation. An input that starts but fails semantic validation completes successfully with one billed diagnostic Dataset item, immutable DIAGNOSTIC, and terminal RUN_SUMMARY when event capacity remains and Dataset publication is decidable. An unresolved publication or visibility state stays provisional and never risks another billed write merely to produce a summary.

The run returns no products

Open the diagnostic Dataset item or the durable DIAGNOSTIC Actor output and read its message, nextAction, and documentationUrl. Confirm that the target opens without signing in for the selected marketplace, then retry with one query, maxItems set to 10, and maxPages set to 1.

Some fields are missing

Amazon did not expose every field on the source page. Check extractionQuality, featureCoverage, and rawTextExcerpt in the Extraction quality view.

Prices or delivery text use the wrong region

Align marketplace, deliveryLocation, proxyCountryCode, and language. Create separate Tasks for different marketplaces or delivery regions.

The run times out

Reduce the target count, maxItems, or maxPages. Increase timeoutSecs only after a smaller run succeeds.

A run reaches its spending limit

The Actor publishes only items that fit the remaining event budget and records the budget-limited state in DIAGNOSTIC and RUN_SUMMARY. When no product was published and capacity remains, the single diagnostic row is the billed result; it is explicitly typed and never pretends to be a product.

Privacy and lawful use

Use this Actor only for lawful collection and processing of Amazon data available without signing in. Review Amazon's terms and the laws that apply to your use case.

Do not submit credentials, cookies, access tokens, private URLs, personal street addresses, or other private data. Configure Apify storage retention and delete stored inputs or datasets when your policy requires it.

Unofficial Actor: This independent Actor is not affiliated with, endorsed by, or sponsored by Amazon. Amazon is a trademark of Amazon.com, Inc. or its affiliates.

Support

Open the Actor's Issues tab when you need help. Include the run ID, operation, marketplace, sanitized input, and a shareable Amazon target example.

Never post credentials, cookies, access tokens, private customer data, or full personal addresses in an issue.