Amazon Product Scraper | $0.50 per 1K Results
Pricing
from $0.50 / 1,000 amazon product results
Amazon Product Scraper | $0.50 per 1K Results
Scrape Amazon products by keyword, URL, ASIN, category, or seller across 20 marketplaces. Extract prices, ratings, reviews, sellers, images, variants, availability, and bestseller ranks. Run in Console or by API, schedule Tasks, and pay $0.50 per 1,000 results on paid plans.
Pricing
from $0.50 / 1,000 amazon product results
Rating
0.0
(0)
Developer
Tovuk
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Amazon Product Scraper is a structured Amazon data scraper for search, product, category-chart, bestseller, seller, review, offer, and question pages. Paid-plan prices start at $0.50 per 1,000 Dataset results, with no start fee, usage surcharge, or add-on event. Results are normalized products when Amazon exposes a usable record, or one transparent diagnostic item when a handled run produces no product.
Use the Amazon product scraper from the Apify Console, through the Amazon scraper API, from Python or JavaScript, or as a scheduled Task. You do not need to provide a proxy, cookie, captcha service, Amazon account, or third-party API key. Start with one query, URL, or Amazon Standard Identification Number (ASIN), then increase the bounded run limits after checking the first dataset.
Why use this Amazon data scraper?
- Compare product prices, discounts, stock signals, shipping text, and visible sellers
- Monitor search position, sponsored placement, bestseller rank, badges, and category visibility
- Enrich product catalogs with ASINs, identifiers, specifications, variants, images, and descriptions
- Analyze ratings, review counts, visible review samples, questions, and customer-photo signals
- Build marketplace research datasets across 20 Amazon domains with consistent normalized fields
- Feed spreadsheets, databases, dashboards, alerts, scheduled workflows, and AI agents
The Actor keeps search, product, review, offer, seller, bestseller, ASIN, and question workflows in one tested output contract. For a single workflow and a more specific Store listing, use one of the focused Amazon Actors linked in the related-Actors section after completing your first run.
How to scrape Amazon product data
- Open the Input tab.
- Enter one search query, supported Amazon URL, or ASIN.
- Keep
maxItemsat 10 andmaxPagesat 1 for the first run. - Click Start, inspect the Dataset, and open the Actor outputs for the run summary or diagnostic.
- Confirm the fields and marketplace, then increase limits or save the input as a Task.
Amazon product scraper input example
Use this input for a low-cost first run:
{"query": "usb c cable","marketplace": "amazon.com","maxItems": 10,"maxPages": 1}
The Actor infers a search operation, requests up to 10 products, and writes all available product fields to the default Dataset. If a handled run produces no product, it writes one billed recordType: "diagnostic" Dataset item explaining the cause, next action, retryability, and relevant Actor documentation. DIAGNOSTIC preserves the exact immutable diagnostic publication body, and a terminal RUN_SUMMARY confirms the settled Dataset outcome. If Dataset visibility or publication state is temporarily undecidable, the Actor writes only provisional, non-billed recovery evidence and deliberately withholds a terminal summary and any further Dataset write.
An actual run of the input above returned this Dataset item on July 24, 2026. Amazon listings change over time, so treat the values as a continuous input-to-output example rather than a guaranteed current offer:
{"recordType": "product","marketplace": "amazon","operation": "search","title": "Anker USB C to USB C Cable, 60W Fast Charging Cable (2-Pack, 6 ft, Black) | For iPhone 17 Series, iPad mini 6, and More","asin": "B088NRLMPV","searchPage": 1,"resultPosition": 1,"priceText": "$9.99","rating": 4.7,"reviewCount": 85571,"extractedAt": "2026-07-24T11:45:21.067Z"}
What you can collect
Available fields depend on the page, marketplace, delivery context, and content Amazon renders during the run.
- Product identity: title, ASIN, product and parent IDs, brand, manufacturer, model and part numbers, SKU, GTIN, MPN, URL, and source URL
- Search and category-chart data: search page, result position, sponsored position, category path, breadcrumbs, browse-node IDs, badges, and structured bestseller ranks
- Pricing and fulfillment: price, list price, sale price, discount, currency, condition, availability, stock, quantity, shipping, Prime eligibility, and delivery context
- Seller and offer data: seller name and ID, seller URL, ships-from, sold-by, fulfillment, offer count, visible price range, and structured offer details
- Product content: description, feature bullets, specifications, package quantity, color, size, style, variation theme, variants, primary image, image gallery, videos, keywords, and tags
- Customer signals: rating, rating count and distribution, review count, visible review samples, question count, visible question samples, and customer photos
- Extraction metadata: language, region, detail-enrichment provenance, requested features, extraction quality, and per-feature coverage
The Actor preserves additional normalized fields when the output contract gains new data. It does not discard rich fields unless you set outputFields.
Choose a target
Provide at least one target field. Unknown fields are rejected before a data request starts.
| Field | Limit | Use |
|---|---|---|
query | 1 | Run one Amazon search |
searchQueries | 50 | Run several search phrases |
startUrls | 200 | Provide supported Amazon page URLs for one compatible operation |
productUrls | 200 | Collect product-page records |
asins | 200 | Resolve ten-character ASINs in the selected marketplace |
categoryUrls | 100 | Collect bestseller or category-chart pages; use startUrls for search-category URLs |
categoryIds | 100 | Build bestseller targets from category IDs |
sellerUrls | 100 | Collect seller or storefront pages |
Each URL can contain up to 2,048 characters. The Actor accepts HTTPS URLs only from supported Amazon domains; private URLs are rejected.
Select an operation
Leave operation set to auto to infer one compatible operation for the whole run. Every supplied target must match that inferred operation. Split incompatible page types into separate runs. An explicit operation does not transform one page type into another, so every supplied target must match the selected operation.
| Operation | Matching target and output |
|---|---|
search | Search query or search URL; normalized product-result records |
product | Product URL or ASIN; visible product details |
reviews | Review page; product record with visible review samples and counts |
offers | Offer page; product record with visible offer, seller, price, and shipping fields |
seller | Seller or storefront URL; visible storefront and product fields |
bestsellers | Bestseller or category-chart URL, or category ID; ranked product records |
asins | ASIN list; normalized product records |
questions | Question page; product record with visible question samples and counts |
Review, offer, seller, and question data are embedded in normalized product-centric records. The Actor does not promise one complete dataset row for every individual review, offer, seller, or answer.
Set marketplace and delivery context
The Actor supports 20 Amazon marketplaces:
amazon.ae, amazon.ca, amazon.co.jp, amazon.co.uk, amazon.com, amazon.com.au, amazon.com.be, amazon.com.br, amazon.com.mx, amazon.com.tr, amazon.de, amazon.es, amazon.fr, amazon.in, amazon.it, amazon.nl, amazon.pl, amazon.sa, amazon.se, and amazon.sg.
Use one marketplace per run. The selected marketplace, every supplied Amazon URL, and every result identity URL must resolve to that same marketplace or one of its subdomains. URL-free runs default to amazon.com. You can also provide:
| Field | Purpose |
|---|---|
deliveryLocation | City, ZIP code, or postal code for delivery-localized content |
proxyCountryCode | Country hint for managed retrieval, without proxy credentials |
proxyCity | International city hint for managed retrieval |
language | Language or locale hint such as en-US or de-DE |
Use a city or postal code instead of a personal street address.
Filter Amazon products and review samples
The Actor applies filters after parsing Amazon records. This avoids returning a row merely because its source card was fetched. An active filter can return fewer than maxItems when Amazon exposes too few matching records.
| Field | Effect |
|---|---|
minRating | Keep products at or above a rating from 0 to 5 |
minReviews | Keep products at or above a parsed review count |
titleIncludeKeywords | Keep titles containing at least one term |
titleExcludeKeywords | Remove titles containing any term |
excludeSponsored | Remove rows identified by a sponsored badge or ad position |
maxReviews | Limit nested reviewSamples after filtering; 0 keeps every available sample |
reviewRatings | Keep nested review samples in selected 1-to-5-star buckets |
reviewKeywords | Keep nested review samples containing at least one term |
reviewsStartDate | Keep nested review samples on or after an inclusive YYYY-MM-DD date |
reviewsEndDate | Keep nested review samples on or before an inclusive YYYY-MM-DD date |
reviewSort | Order nested review samples by relevance, date, or rating |
Title and review-keyword matches are case-insensitive. Review controls filter the samples Amazon exposes on the fetched page; they do not create a separate complete row for every review.
Control run size
| Field | Range | Default |
|---|---|---|
maxItems | 0 for no item ceiling; 1 to 1,000,000,000 for an explicit ceiling | 10 |
maxPages | 1 to 1,000,000,000 | 1 |
continuationCursor | Signed opaque cursor from a previous CONTINUATION record | Empty |
enablePagination | Boolean | true |
maxConcurrency | 1 to 4 | 4 |
timeoutSecs | 30 to 3,600 | 240 |
maxItems is the optional product-record ceiling. Set it to 0 when pages, targets, timeout, and the Apify spending limit should be the only run bounds. The final product count can be lower when a page has fewer records, Amazon limits pagination, a filter removes records, or the spending limit is reached. Large runs can take several minutes, so increase timeoutSecs and the Apify run timeout together. A direct target that produces no usable record completes with one billed actionable diagnostic item instead of disappearing or inventing a product.
If you override the Apify run timeout, keep it at least 60 seconds above timeoutSecs so the Actor can persist its final summary and any diagnostic.
For large search, seller, and bestseller jobs, the Actor splits source work into bounded internal requests of at most one source page and 1,000 service results. It follows the service's opaque source cursor across those chunks, including page 1,001 and later, without page arithmetic. It streams, validates, globally deduplicates, and repositions each completed request in the same Dataset instead of retaining the full run in memory. These internal safety boundaries are not user-facing result or page ceilings. The Actor also stops before publishing rows that exceed the run spending limit.
Before starting the request, the Actor normalizes and deduplicates exact target values so repeated ASINs, queries, identifiers, or URLs do not create duplicate service work.
CONTINUATION is written before request creation, after the returned request ID is known, before every result-page read, and after every durable result page. It records the request phase, first and last request IDs, request count, separate result and source cursors, aggregate fetched and published counts, target progress, and source-plan completion. continuationCursor is an authenticated opaque cursor for a fresh run; copy it exactly and keep the same targets, filters, marketplace, operation, and output fields. Never edit or decode an opaque cursor, and never use the raw internal request ID as an input.
When present, terminal RUN_SUMMARY is authoritative over provisional recovery records. It distinguishes fetched products, published products, published diagnostics, and total billed Dataset items. It identifies whether a diagnostic is present and repeats its safe code, message, next action, retryability, and documentation link. It also reports request identity, requested and effective product limits, result pages, elapsed time, and a continuationStatus of complete, item-limit, spending-limit, interrupted, or not-started. An unresolved Dataset-visibility or publication decision intentionally has no terminal RUN_SUMMARY: RECOVERY_STATUS carries visibility-only evidence, while PUBLICATION_RECOVERY preserves the exact unbilled diagnostic body awaiting fenced recovery.
The Actor automatically drains every stored-result cursor and launches the next bounded source request while the run has capacity. Request creation uses one deterministic, account-scoped idempotency key per logical source chunk, so retrying a lost POST response resolves to the original request. Deduplication by validated product URL spans all chunks, and result positions remain contiguous. If an API reader disconnects, the Apify run continues independently; retain the Apify run ID, reconnect later, and page through its stored default Dataset.
On an Apify migration or process restart, the Actor reads CONTINUATION, polls the exact in-flight request or replays the same idempotent create, scans the default Dataset, and reconciles product identities through the fenced backend ledger before publishing again. A Dataset row written just before interruption is therefore not published or billed a second time. The runtime also reserves 30 seconds before ACTOR_TIMEOUT_AT to persist its latest checkpoint, summary, or diagnostic. For a fresh run after an item, spending, or time boundary, copy the authenticated opaque continuationCursor; the Actor verifies its current version, cryptographic authenticity, and logical-input fingerprint before resuming.
Limit dataset columns
The Dataset contract declares 96 properties. Product items use up to 88: 86 projectable Amazon fields plus Actor-owned recordType and extractedAt. Diagnostic items use eight additional diagnostic-only fields plus the shared recordType, marketplace, and optional operation. Leave outputFields empty to return every available product field. To reduce product width, select up to 86 supported result fields:
{"query": "mechanical keyboard","maxItems": 50,"maxPages": 2,"outputFields": ["title","asin","priceText","rating","reviewCount","imageUrls"]}
Required identity, provenance, and Actor-owned fields stay present when applicable. The Actor always retains recordType, marketplace, operation, and extractedAt. Product rows also retain sourceUrl and url so every product remains identifiable.
Use dataset views
The Dataset includes these views:
- Overview shows identity, rank, price, rating, stock, seller, image, URL, and extraction time
- Diagnostics shows stable cause, explanation, recommended action, retryability, context, and documentation
- Pricing and availability shows price, discount, stock, shipping, delivery, seller, fulfillment, and structured offer fields
- Product details shows descriptions, identifiers, catalog attributes, categories, ranks, specifications, variants, and media
- Reviews and questions shows ratings, rating distribution, counts, visible samples, and customer photos
- Extraction quality shows source and detail provenance, parser classification, requested features, and coverage
Open the Dataset tab to export JSON, CSV, Excel, XML, RSS, or JSONL. You can also read the Dataset through the Apify API, save inputs as Tasks, schedule recurring runs, trigger webhooks, and connect supported Apify integrations.
Use the Amazon product data scraper API from Python
Install the official apify-client package, keep your API token in the APIFY_TOKEN environment variable, and call the same validated input used in the Console:
import osfrom apify_client import ApifyClientclient = ApifyClient(os.environ["APIFY_TOKEN"])run = client.actor("tovuk/amazon-product-scraper").call(run_input={"query": "usb c cable","marketplace": "amazon.com","maxItems": 10,"maxPages": 1,})items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
The call waits for completion and then reads the default Dataset. For asynchronous systems, start a run through the REST API, retain its run ID, poll until the run reaches a terminal status, and then read defaultDatasetId.
Related Amazon scrapers
Use a focused Actor when every run has one intent. Each listing shares this Actor's tested runtime and output contract while exposing a smaller input surface:
- Amazon Search Scraper for keyword and search-result product records
- Amazon Bestsellers Scraper for ranked products and category charts
- Amazon ASIN Scraper for direct product lookup from ASIN lists
- Amazon Seller Scraper for storefront and seller-page product records
The Amazon Reviews Scraper remains private until its complete review workflow passes the same live acceptance standard as the product Actors.
Amazon product scraper FAQ
Can I scrape more than 1,000 Amazon products in one run?
Yes. Set maxItems to 0 for no separate item ceiling, then set maxPages, timeoutSecs, and the Apify maximum cost per run for the work you want. Large listing jobs use bounded internal requests and one deduplicated Dataset. Amazon can expose fewer matching products, and active filters or spending limits can lower the final count.
Can an API client or agent continue from where it stopped?
Yes. When only the API client disconnects, keep the Apify run ID, poll that run after reconnecting, then page through its default Dataset. Apify migrations and same-run process restarts recover automatically from durable checkpoints. To continue in a fresh run after an item, spending, or time boundary, copy CONTINUATION.continuationCursor into the new input without changing the logical target or output fields. A complete state means the explicit source plan finished; item-limit or spending-limit explains why it stopped earlier.
Can I use the Amazon product scraper from Python?
Yes. Use the Python example above with apify-client, or call the Actor REST API from any language. The Console, API, scheduled Tasks, and integrations all use the same validated input and Dataset contract.
Can I scrape Amazon products by ASIN, seller, category, or keyword?
Yes. Use asins, sellerUrls, categoryUrls or categoryIds, and query or searchQueries. Keep one compatible operation and marketplace in each run.
Does the Actor return Amazon reviews and offers?
Product records can include visible review samples, review counts, offer counts, sellers, prices, and shipping fields when Amazon exposes them. The product scraper does not promise one complete row for every review or marketplace offer.
Pricing and spending limits
This Actor uses one pay-per-event charge. One transparent record price covers all available fields. The current rates below apply until the scheduled pricing change on August 8, 2026 at 02:23 UTC. The live Pricing panel remains the authoritative source for your plan's price.
| Billing item | Price | When it applies |
|---|---|---|
| Actor start | $0 | No start event is configured, so starting a run is free |
| Default Dataset item | Plan price below | One visible product record or one transparent handled diagnostic |
| Other events | $0 | No query, page, field, review, offer, seller, or add-on events |
| Platform usage surcharge | $0 | The Dataset-item price includes platform usage |
| Apify plan | Per Dataset item | 10 items | 100 items | 1,000 items |
|---|---|---|---|---|
| Free | $0.05000 | $0.50000 | $5.00 | $50.00 |
| Bronze | $0.00050 | $0.00500 | $0.050 | $0.50 |
| Silver | $0.00050 | $0.00500 | $0.050 | $0.50 |
| Gold | $0.00050 | $0.00500 | $0.050 | $0.50 |
| Platinum | $0.00050 | $0.00500 | $0.050 | $0.50 |
| Diamond | $0.00050 | $0.00500 | $0.050 | $0.50 |
The Free-plan Dataset-item price is exactly 100 times the paid-plan Dataset-item price. Bronze, Silver, Gold, Platinum, and Diamond all use the same $0.00050 rate, so upgrading to any paid plan unlocks the full 99% price reduction while the Free plan remains suitable for bounded evaluation runs.
The release configuration omits apify-actor-start. You pay no start fee, and the run does not receive the five-second compute coverage attached to Apify's synthetic start event.
One Dataset-item charge maps to one visible product record or one transparent diagnostic item. Apify automatically charges apify-default-dataset-item for each accepted row written to the default Dataset. The Actor does not issue a separate custom charge. Product items use recordType: "product". When a handled invalid, empty, or recoverable run publishes no product and Dataset state is decidable, it writes exactly one recordType: "diagnostic" item and preserves the same immutable body under DIAGNOSTIC for direct access by users and agents.
Set a positive maxItems value to bound product results, or use 0 to rely on page, target, timeout, and spending bounds. A diagnostic can add one Dataset item beyond the product count. Set the Apify maximum cost per run for an independent monetary cap. The Actor checks remaining capacity before service work and before every result page and stops when Apify reports that the spending limit was reached. If the platform reports zero remaining event capacity, it can prevent the diagnostic Dataset item too; KVS retains the safe outcome without claiming a billed row. Dataset-visibility uncertainty is a second honest zero-row exception: the Actor records provisional recovery evidence, issues no Dataset write, and waits for a later safe reconciliation before producing a terminal summary.
Data coverage and limitations
- The Actor collects Amazon pages available without signing in
- Amazon can vary content by marketplace, location, language, session, and time
- A field can be
null, empty, or absent when Amazon does not render it - Search and chart pagination can end before the configured page limit
- Review and question modes return visible samples in product-centric records, not private account content or guaranteed complete histories
- Offer data reflects visible page content and does not guarantee every marketplace seller or buy-box change
- Numeric bestseller rank is returned only when the page exposes a supported rank signal
- One run captures current values; schedule repeated runs to build a time series
- The Actor does not provide historical prices from before your first run
- Amazon page changes and transient blocking can reduce output; handled failures complete with a retryable diagnostic
Troubleshooting
Input is rejected
Provide at least one supported target and remove unknown fields. Check array limits, URL length, marketplace domain, ASIN shape, and operation-target compatibility.
Apify validates the declared input schema before starting the Actor. A structurally invalid API or Console input has no Actor run, Dataset, key-value store, or billable event, so follow the platform validation message and the generated Input/API documentation. An input that starts but fails semantic validation completes successfully with one billed diagnostic Dataset item, immutable DIAGNOSTIC, and terminal RUN_SUMMARY when event capacity remains and Dataset publication is decidable. An unresolved publication or visibility state stays provisional and never risks another billed write merely to produce a summary.
The run returns no products
Open the diagnostic Dataset item or the durable DIAGNOSTIC Actor output and read its message, nextAction, and documentationUrl. Confirm that the target opens without signing in for the selected marketplace, then retry with one query, maxItems set to 10, and maxPages set to 1.
Some fields are missing
Amazon did not expose every field on the source page. Check extractionQuality, featureCoverage, and rawTextExcerpt in the Extraction quality view.
Prices or delivery text use the wrong region
Align marketplace, deliveryLocation, proxyCountryCode, and language. Create separate Tasks for different marketplaces or delivery regions.
The run times out
Reduce the target count, maxItems, or maxPages. Increase timeoutSecs only after a smaller run succeeds.
A run reaches its spending limit
The Actor publishes only items that fit the remaining event budget and records the budget-limited state in DIAGNOSTIC and RUN_SUMMARY. When no product was published and capacity remains, the single diagnostic row is the billed result; it is explicitly typed and never pretends to be a product.
Privacy and lawful use
Use this Actor only for lawful collection and processing of Amazon data available without signing in. Review Amazon's terms and the laws that apply to your use case.
Do not submit credentials, cookies, access tokens, private URLs, personal street addresses, or other private data. Configure Apify storage retention and delete stored inputs or datasets when your policy requires it.
Unofficial Actor: This independent Actor is not affiliated with, endorsed by, or sponsored by Amazon. Amazon is a trademark of Amazon.com, Inc. or its affiliates.
Support
Open the Actor's Issues tab when you need help. Include the run ID, operation, marketplace, sanitized input, and a shareable Amazon target example.
Never post credentials, cookies, access tokens, private customer data, or full personal addresses in an issue.