# Amazon Scraper (`mlg14/amazon-scraper`) Actor

Extract public Amazon product listings and product details from search, category, and product URLs. Export prices, ratings, images, availability, features, and search positions.

- **URL**: https://apify.com/mlg14/amazon-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Amazon Scraper

Scrape Amazon products from search pages, category URLs, product URLs, and ASINs. Export Amazon product data to CSV, JSON, or Excel for catalog research, price tracking, assortment analysis, and scheduled monitoring. This scraper provides a practical Amazon API alternative for public listing and product information without requiring product API credentials.

The actor accepts several start URLs in one run, follows search pages, removes duplicate products by ASIN, and stops at the requested result limit. Search results can be returned quickly from listing cards, or enriched with individual product pages for descriptions, features, seller text, stock signals, delivery text, and variants. It records the source page and collection time so changes can be compared across runs.

### What data can you extract from Amazon?

Each output row represents one product. Search cards provide the quickest set of fields; product pages can supply the additional detail fields. A blank value means the information was absent from the public page, could not be read reliably, or did not apply to that input. Prices and availability depend on the Amazon domain, delivery area, product variation, and time of collection.

| Field | Description | Example |
| --- | --- | --- |
| `asin` | Product identifier used for deduplication | `B004YAVF8I` |
| `url` | Canonical product URL | `https://www.amazon.com/dp/B004YAVF8I` |
| `title` | Displayed product title | `Logitech M185 Compact Ambidextrous 2.4 GHz Wireless Mouse - Swift Grey` |
| `brand` | Brand shown on a product page | `Logitech` |
| `price` | Current price as a number | `14.9` |
| `currency` | Displayed currency symbol or code | `$` |
| `listPrice` | Former or list price when displayed | `17.99` |
| `discountPercent` | Calculated reduction from the list price | `17.18` |
| `shippingPrice` | Separate shipping charge when available | `null` |
| `stars` | Displayed average rating | `4.4` |
| `reviewsCount` | Displayed count of ratings or reviews | `45106` |
| `answeredQuestions` | Answered question count when present | `null` |
| `thumbnailImage` | Main product image URL | `https://m.media-amazon.com/images/I/51WN5aXZWIL._AC_SY300_SX300_QL70_FMwebp_.jpg` |
| `images` | Image URLs found on the product page | array of URLs |
| `inStock` | Availability signal from product page | `true` |
| `inStockText` | Availability wording | `In Stock` |
| `description` | Product description text | `Logitech Wireless Mouse M185...` |
| `features` | Feature bullets | array of text |
| `breadCrumbs` | Category path shown on product page | `Electronics › Computers & Accessories › ...` |
| `variantAsins` | Other ASINs found in variation data | array of ASINs |
| `variantAttributes` | Selected variation labels | `1 Pack`, `Swift Grey` |
| `delivery` | Delivery message shown for the current location | `FREE delivery Thursday, October 1...` |
| `returnPolicy` | Return policy message when present | `null` |
| `sellerName` | Featured seller or merchant text | `Sold by Amazon Resale and Fulfilled by Amazon.` |
| `sellerId` | Featured seller identifier when exposed | `null` |
| `bestsellerRanks` | Displayed ranking text | `#1 in Computer Mice` |
| `searchTerm` | Keyword from the search URL | `wireless mouse` |
| `searchPage` | Search results page number | `1` |
| `searchPosition` | Position on that page | `1` |
| `isSponsored` | Whether the result displays a sponsored label | `false` |
| `sourceUrl` | Product input or results page that produced the row | `https://www.amazon.com/s?k=wireless+mouse` |
| `scrapedAt` | UTC time at which the row was collected | ISO timestamp |

`price`, `listPrice`, and `shippingPrice` are numeric fields so they can be sorted, charted, or compared. `currency` is separate because a symbol alone may not distinguish every market. A discount is computed only when both numeric prices are available and the list price exceeds the current price. The scraper does not infer stock from a search card: `inStock` is populated from a product page only. `searchPage` and `searchPosition` describe the page returned during that run, not a permanent rank.

### How to scrape Amazon products

1. Paste a public Amazon search, category, or product URL into **Category or product URLs**. A ten-character ASIN is also accepted. For a simple keyword search on amazon.com, enter a phrase under **Search terms**.
2. Set **Maximum products** and, if useful, a per-URL limit. Start with 30–100 items while checking whether the chosen search or category produces the expected products.
3. Leave **Visit product pages** enabled for detailed records. Turn it off when you need a faster listing-level sample containing names, displayed prices, ratings, review counts, images, and search positions.
4. Run the actor. Review a few output rows, especially price, currency, stock, and delivery fields, against the market and location you intended.
5. Download the dataset as JSON, CSV, or Excel, or use its dataset URL in a recurring data pipeline.

A category URL can carry Amazon's own filters for department, price, brand, or rating. Copy the filtered public URL after applying those filters in a browser. The actor preserves the query parameters, then adds its own page number when following results. A direct product URL yields one detailed row for that product, unless the page is missing or unavailable. Multiple search terms and URLs can be combined; the same ASIN appears only once within a run.

### Input

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `categoryOrProductUrls` | Array of URL objects or strings | Empty | Public search, category, or product URLs; bare ASIN strings also work. |
| `searchTerms` | Array of strings | Empty | Keyword searches on amazon.com. |
| `maxItems` | Integer | `100` | Total unique products to output. `0` removes the total cap. |
| `maxItemsPerStartUrl` | Integer | `0` | Maximum results from each search or category URL. `0` removes this cap. |
| `maxSearchPagesPerStartUrl` | Integer | `7` | Maximum pages visited per search URL, with a hard cap of seven. |
| `scrapeProductDetails` | Boolean | `true` | Visit each result's product page for additional fields. |
| `proxyCountry` | String | Empty | Optional two-letter country code for the request location. |
| `proxyConfiguration` | Object | Managed proxy | Optional proxy settings. |

At least one URL, ASIN, or search term is required. The URL host must be a supported Amazon retail domain. A product URL can be pasted alongside one or more search URLs. If the same ASIN appears in several searches, the first encountered row is retained; its `sourceUrl`, `searchPage`, and `searchPosition` reflect that first appearance. Put the most important search first when those fields matter. Search terms use amazon.com; use a full country-specific URL for a different marketplace.

For example, the following input collects up to 60 product rows from two searches and one product. It visits product pages so the detailed fields can be populated:

```json
{
  "categoryOrProductUrls": [
    { "url": "https://www.amazon.com/s?k=wireless+mouse" },
    { "url": "https://www.amazon.com/dp/B004YAVF8I" }
  ],
  "searchTerms": ["usb c hub"],
  "maxItems": 60,
  "maxItemsPerStartUrl": 40,
  "maxSearchPagesPerStartUrl": 3,
  "scrapeProductDetails": true
}
```

### Output example

The final golden run returned 35 products. One directly requested product produced the following values. The long description and feature list are shortened here for readability; the actual row contains the displayed text and the same set of documented fields.

```json
{
  "asin": "B004YAVF8I",
  "url": "https://www.amazon.com/dp/B004YAVF8I",
  "title": "Logitech M185 Compact Ambidextrous 2.4 GHz Wireless Mouse - Swift Grey",
  "brand": "Logitech",
  "price": 14.9,
  "currency": "$",
  "listPrice": 17.99,
  "discountPercent": 17.18,
  "stars": 4.4,
  "reviewsCount": 45106,
  "thumbnailImage": "https://m.media-amazon.com/images/I/51WN5aXZWIL._AC_SY300_SX300_QL70_FMwebp_.jpg",
  "inStock": true,
  "inStockText": "In Stock",
  "description": "Logitech Wireless Mouse M185. A simple, reliable mouse with plug-and-play wireless...",
  "features": ["Compact Mouse: With a comfortable and contoured shape..."],
  "breadCrumbs": "Electronics › Computers & Accessories › Computer Accessories & Peripherals › Keyboards, Mice & Accessories › Mice",
  "bestsellerRanks": "#1 in Computer Mice",
  "sellerName": "Amazon Resale",
  "variantAttributes": ["1 Pack", "Swift Grey", "USB Receiver"],
  "sourceUrl": "https://www.amazon.com/dp/B004YAVF8I"
}
```

The golden input included that direct product URL and the phrases `wireless mouse`, `usb c hub`, and `mechanical keyboard`. It capped the run at 35 rows, allowed two pages per search, and left detail visits off for search rows. This checks both direct product extraction and listing extraction. In that run, all 35 output rows had an ASIN, canonical URL, title, price, star rating, and thumbnail image. Individual items may show different prices or fields on a later date or from another location.

### Use cases

- **Price and promotion tracking:** Record product prices and displayed list prices on a schedule, then compare the same ASINs over time. Keep the marketplace and location consistent to avoid treating regional offers as price changes.
- **Category research:** Collect products from a filtered department or subcategory URL. Compare the visible assortment, ratings, and review counts for the specific slice you selected.
- **Search placement checks:** Save `searchTerm`, `searchPage`, and `searchPosition` for a selected set of phrases. Use repeated runs to see which products move into or out of the collected results. Search placement can vary by shopper and location.
- **Catalog enrichment:** Start with a list of ASINs or product URLs and gather titles, descriptions, features, category paths, images, and availability for internal product records.
- **Inventory watchlists:** Request known products regularly and flag a change in `inStock` or `inStockText`. Treat these as observed page signals, not guaranteed inventory counts.
- **Market comparisons:** Run equivalent category searches on different supported Amazon domains, then compare prices, ratings, and product mix after normalizing currencies and locations in your analysis.

### How much does it cost to scrape Amazon?

The configured result price is **$0.001 per output product**, or **$1.00 per 1,000 products**. This is a charge per delivered result. The number of pages or detail requests affects run duration and the chance of encountering a blocked page, but does not change the per-result price shown here. The platform usage for this pay-per-result configuration is included in the result charge. Check the current price shown on the actor page before scheduling a large run.

| Delivered products | Result charge |
| ---: | ---: |
| 100 | $0.10 |
| 1,000 | $1.00 |
| 5,000 | $5.00 |

A direct product URL that produces one row costs $0.001. A search yielding 620 distinct ASINs costs $0.62. If several input URLs repeat the same ASIN, that product is pushed once and charged once in that run. A request that cannot produce a product row is not counted as a delivered product. A `maxItems` setting is the simplest way to cap both output volume and the result charge. Turn detail visits off for exploratory searches when speed matters, then rerun the chosen ASINs with detail visits enabled.

### Tips for best results

Start with narrow search or category URLs. Amazon sometimes limits ordinary keyword searches to a handful of accessible pages. A filtered category URL can expose a different set of matching products and make collection more useful than simply requesting deeper pages of a broad keyword. For a large category, split the work by its available filters, such as brand or price range, and combine the datasets later by ASIN.

Keep product detail visits enabled when you need brand, features, stock, seller, category, or variant fields. A search card does not consistently display those fields, so a listing-only run should be treated as a product discovery pass. Direct product URLs always use the product page. If you only need title, current displayed price, rating, review count, image, and search position, listing-only mode can reduce request volume substantially.

Choose a suitable country domain and proxy location together. A product may have a price on one marketplace and no offer on another. The same domain can display different delivery text, availability, or seller offers to visitors from different places. Set `proxyCountry` only when location matters to your task; otherwise, record the domain and run time alongside the prices you analyze.

For repeated monitoring, use the same URL list and input order each time. Search positions and the first retained source URL depend on input order. Retain `asin` as the join key across snapshots, and include `scrapedAt` when calculating change over time. If Amazon substitutes a different ASIN for a selected variation, compare the variant identifiers before assuming two records describe the same purchasable option.

### Limits

Search pages are capped at seven pages per start URL. Amazon can expose fewer pages, especially for broad or changing keyword searches. Different filter URLs may reach additional products, but overlap must be expected. The actor removes duplicate ASINs within one run; it does not claim that a category is exhaustively covered. Product counts shown on the site can also differ from the number of accessible cards.

Amazon can return verification pages, temporary errors, or missing product pages. The actor rotates the proxy session and retries a few times, then moves on rather than producing a misleading product record from a verification page. A page can return HTTP 200 while containing only a verification prompt; that response is treated as unusable. If the first search page fails repeatedly, that input may yield no rows. A direct product URL can also be unavailable or return a 404.

Public pages vary by domain, layout, location, and product type. `shippingPrice`, `returnPolicy`, `sellerId`, `answeredQuestions`, variants, and stock text are conditional fields. Some products show a price range or variation-dependent price instead of a single stable offer. Use the captured URL and collection time when investigating a surprising value. The displayed review count is the count shown on the page; it is not a collection of individual reviews.

Search results can include sponsored placements. `isSponsored` records a displayed sponsorship label when present; its absence does not prove that placement is organic under every layout. The scraper does not log in, place orders, collect customer accounts, or access private product information. It does not bypass the need to validate scraped data for decisions that require exact price, stock, or shipping commitments.

### Automated workflow examples

A scheduled workflow can request “Collect the first 100 wireless mouse products from this filtered Amazon search URL, with price, rating, review count, and search position.” A second workflow can request “Refresh these ten product URLs every morning and flag ASINs whose displayed price or stock signal changed.” Both use the same input schema and dataset fields, which makes the resulting rows straightforward to compare.

### FAQ

#### Is it legal to scrape Amazon product pages?

This actor reads publicly visible product and listing pages. Your use of the data must follow applicable law and the site's terms. Avoid collecting personal information or using product data in a way that infringes others' rights. For any regulated or high-impact use, obtain appropriate advice for your jurisdiction and intended purpose.

#### Do I need to configure a proxy?

The default input uses a managed proxy. You can supply proxy settings or choose a country when a particular shopping location is relevant. Amazon may still return a verification page; the actor detects it, changes sessions, and retries within a limited budget. A proxy country is not a guarantee that every product will display an offer for that location.

#### How fast is a run?

Listing-only mode uses search pages and is generally much faster than visiting each product page. With product detail visits enabled, every distinct search result can require an extra page load. Duration also depends on page size, retries, and any verification pages encountered. The 35-row golden run completed successfully, but that single run is not a promise of fixed throughput for larger or different inputs.

#### Can I schedule and monitor repeated runs?

Yes. Save an input with a bounded `maxItems`, schedule it, and compare each completed dataset by `asin` and `scrapedAt`. Check the run status and row count as well as field fill. If a run produces fewer rows than usual, inspect whether a search page changed, a product became unavailable, or verification pages interrupted collection.

#### Can I export to a spreadsheet?

Yes. Download the default dataset as CSV or Excel, or send its rows to a spreadsheet integration. Numeric price and rating fields can be used for sorting and formulas. Preserve the `currency` column, because a numeric price alone does not identify the marketplace currency.

#### Why is a field empty?

The page may not show it for that product or location, or the field may require a product detail visit while the run used listing-only mode. Seller identifiers, shipping charges, return messages, and variants are especially conditional. Review the product URL, run settings, and location before treating an empty value as a defect.

#### Can I collect more than seven search pages?

The actor stops at seven pages for each search URL because deeper ordinary search pagination may not remain accessible. Divide a broad category using public filters, or provide several narrower category URLs. Deduplicate combined datasets by ASIN if you combine separate runs.

### Integrations

The output dataset can be read through the platform API, sent by webhook when a run finishes, connected to scheduling and automation services, or exported to a spreadsheet or data warehouse. CSV and JSON are available for one-off analysis. When connecting a recurring workflow, keep the input limits and location settings visible in its configuration so future readers can interpret changes in product count, price, and availability.

### Support

Open an issue on the Issues tab; we reply within 24h and add fields on request.

# Actor input Schema

## `categoryOrProductUrls` (type: `array`):

Public Amazon search, category, or product URLs. Product ASINs are also accepted as strings.

## `searchTerms` (type: `array`):

Search amazon.com for each term. Use category URLs for filters and other country domains.

## `maxItems` (type: `integer`):

Maximum number of unique products across all inputs. Set 0 for no total limit.

## `maxItemsPerStartUrl` (type: `integer`):

Maximum unique products from each search or category URL. Set 0 for no per-URL limit.

## `maxSearchPagesPerStartUrl` (type: `integer`):

Maximum search pages to visit for each URL. Capped at seven because deeper search pages can be unavailable.

## `scrapeProductDetails` (type: `boolean`):

Visit each result product page for brand, stock, features, seller, and more. This takes longer than listing-only mode.

## `proxyCountry` (type: `string`):

Optional two-letter country code for the network location. Results and offers can vary by location.

## `proxyConfiguration` (type: `object`):

Proxy configuration. A managed proxy is used unless overridden.

## Actor input object example

```json
{
  "categoryOrProductUrls": [
    {
      "url": "https://www.amazon.com/s?k=wireless+mouse"
    }
  ],
  "searchTerms": [],
  "maxItems": 100,
  "maxItemsPerStartUrl": 0,
  "maxSearchPagesPerStartUrl": 7,
  "scrapeProductDetails": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `products` (type: `string`):

All scraped products in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "categoryOrProductUrls": [
        {
            "url": "https://www.amazon.com/s?k=wireless+mouse"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/amazon-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "categoryOrProductUrls": [{ "url": "https://www.amazon.com/s?k=wireless+mouse" }] }

# Run the Actor and wait for it to finish
run = client.actor("mlg14/amazon-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "categoryOrProductUrls": [
    {
      "url": "https://www.amazon.com/s?k=wireless+mouse"
    }
  ]
}' |
apify call mlg14/amazon-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/amazon-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/btLxwni9FkjWRhs40/builds/kCyJH1oebvujq6gBS/openapi.json
