# Etsy Scraper — listings, prices, shops, ratings (`xhrdev/etsy-scraper`) Actor

Scrape Etsy listings past DataDome: price, shop, rating, review count, stock, materials and images. Search by keyword or scrape listing IDs directly. No browser. Powered by xhr.dev.

- **URL**: https://apify.com/xhrdev/etsy-scraper.md
- **Developed by:** [xhrdev](https://apify.com/xhrdev) (community)
- **Categories:** E-commerce
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

**Scrape [Etsy](https://www.etsy.com) listings past DataDome** — price, shop, rating, review count, stock on hand, shipping origin, materials, category and the full image set. Search by keyword, pass listing IDs you already hold, or hand it listing and category URLs.

No browser is launched. The DataDome challenge is solved as plain HTTP by [xhr.dev](https://xhr.dev), so a page costs a few hundred milliseconds of compute instead of the several seconds and gigabyte of RAM a stealth browser needs.

### Quickstart

Paste this into the input and press **Run**.

```json
{
  "queries": [
    "leather wallet"
  ],
  "scrapeDetails": true,
  "maxItems": 20
}
```

### Search is thin — this is Etsy, not the Actor

Etsy server-renders only about **a dozen results per page** and loads the rest in the browser as you scroll. There is no fix for that from an HTTP-only scraper, and any Actor claiming hundreds of results per search fetch is either driving a browser or counting differently.

So treat search as a way to *discover* listings, and raise `maxPagesPerQuery` to go wider rather than expecting depth from one page. `scrapeDetails` is on by default precisely because the listing page is where Etsy's good data lives.

### Input

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `queries` | string\[] | `["leather wallet"]` | Keywords to search. See the note on search depth below. |
| `itemIds` | string\[] | `[]` | Etsy listing IDs to fetch directly — the number in the URL, e.g. `1214703903`. |
| `startUrls` | array | `[]` | `/listing/` URLs are read as products; `/search`, `/c/` and `/market/` URLs are read as listings. |
| `scrapeDetails` | boolean | `true` | On, every listing found is opened for its full record. Off, you get only what search results show — cheaper, much thinner. See below. |
| `maxPagesPerQuery` | integer | `1` | How deep to page each search, 1–250. |
| `maxItems` | integer | `50` | Stop once this many listings have been scraped. |
| `proxyConfiguration` | object | Apify RESIDENTIAL | Which proxy to route through. Leave it on residential — see [Proxies](#proxies-read-this-before-changing-anything). |
| `maxRetries` | integer | `3` | Attempts per page before giving up on it, 1–5. Each attempt takes a fresh proxy session. |
| `timeoutSecs` | integer | `120` | How long one page — challenge, solve and all — may take before the attempt fails. 10–300. |
| `maxConcurrency` | integer | `4` | Pages worked on at once, 1–10. |

Nothing is required — the defaults run a real search out of the box.

### Output — with `scrapeDetails: true` (the default)

One dataset item per listing, from the listing page itself.

| Field | Type | Description |
| --- | --- | --- |
| `id` | string | Etsy listing ID. |
| `title` | string | Listing title. |
| `url` | string | Canonical listing URL, tracking parameters stripped. |
| `shop` | string | Shop name. |
| `price` | number | Current price. |
| `currency` | string | ISO currency code, e.g. `USD`. |
| `availability` | string | `InStock`, `OutOfStock`, etc. |
| `quantityAvailable` | number | Units the shop has on hand. |
| `shipsFrom` | string | Country code the item ships from. |
| `rating` | number | Average rating out of 5. |
| `reviewCount` | number | Number of reviews. |
| `category` | string | Full category path, e.g. `Bags & Purses < Wallets & Money Clips < Wallets`. |
| `material` | string | Primary material where the shop set one. |
| `description` | string | Full listing description. |
| `images` | string\[] | Full-size image URLs. |

### Output — with `scrapeDetails: false`

One item per search result card. No listing page is fetched, so this is billed at the cheaper listing rate.

| Field | Type | Description |
| --- | --- | --- |
| `id` | string | Etsy listing ID. |
| `title` | string | Listing title. |
| `url` | string | Canonical listing URL. |
| `price` | number | Price as shown on the card. |
| `currency` | string | Currency **symbol** as rendered, e.g. `$` — the card does not carry an ISO code. |
| `image` | string | Thumbnail URL (255px), not the full-size image. |
| `shopId` | string | Numeric shop ID. The shop *name* is only on the listing page. |
| `rating` | number | Average rating, when the card shows one. |
| `reviewCount` | number | Review count, when the card shows one. |

#### Example record

Taken verbatim from a live run. Long text and image lists are trimmed here for readability; the real record carries them in full.

```json
{
  "id": "4461240474",
  "title": "Personalized Leather Cash Wallet for Men, Full Grain Slim Front Pocket Card Holder, Handmade Minimalist Wallet, Gift for him",
  "url": "https://www.etsy.com/listing/4461240474/leather-cash-wallet-slim-front-pocket",
  "shop": "AmericanLeatherGift",
  "price": 37.79,
  "currency": "USD",
  "availability": "InStock",
  "quantityAvailable": 359,
  "shipsFrom": "US",
  "rating": 4.8,
  "reviewCount": 56,
  "category": "Bags & Purses < Wallets & Money Clips < Wallets",
  "material": "Leather",
  "description": "Personalized Leather Cash Wallet for Men Upgrade everyday carry with this personalized wallet crafted from premium full grain leather. Designed for modern simplicity, this leather cash wallet offers a refined alternative to bulky traditional wallets. Compact … (truncated here for readability)",
  "images": [
    "https://i.etsystatic.com/41443453/r/il/b11a3f/7994642756/il_fullxfull.7994642756_ld98.jpg",
    "https://i.etsystatic.com/41443453/r/il/deb604/7782209464/il_fullxfull.7782209464_1lsp.jpg",
    "https://i.etsystatic.com/41443453/r/il/3ccf3a/7830135935/il_fullxfull.7830135935_4qh6.jpg",
    "… 17 more"
  ]
}
```

#### A row from the same search with `scrapeDetails: false`

```json
{
  "id": "4544756667",
  "title": "Green Alligator Wallet - Handmade Luxury Bifold Wallet with Red Python Leather Interior - Exotic Leather Wallet for Men - Personalized Gift",
  "url": "https://www.etsy.com/listing/4544756667/green-alligator-wallet-handmade-luxury",
  "price": 570,
  "currency": "$",
  "image": "https://i.etsystatic.com/13350861/r/il/97feff/8302896792/il_255x319.8302896792_4bxf.jpg",
  "shopId": "13350861",
  "rating": 4.8,
  "reviewCount": 470
}
```

### Pricing

Two rates, and which one you pay depends on `scrapeDetails`:

| Event | Price | When |
| --- | --- | --- |
| `product-detail` | **$15.00 / 1,000** ($0.015 each) | A listing page was fetched and parsed — `scrapeDetails: true`. |
| `listing-item` | **$3.00 / 1,000** ($0.003 each) | A search-result row only — `scrapeDetails: false`. |

Charged on delivered rows only; failed pages cost nothing. The five-fold difference is real work: detail mode fetches one page per listing, listing mode gets ~12 rows from a single fetch.

**Use listing mode to survey, detail mode to collect.** A price-tracking job that already holds the IDs should pass `itemIds` and skip search entirely.

### Proxies: read this before changing anything

Set this to Apify **RESIDENTIAL** and leave it there. It is not a style preference, and it is the single most common reason a run comes back empty.

DataDome decides *which* challenge to serve based on the exit IP. We measured this from Apify against twelve DataDome-protected sites, same code, minutes apart:

| Exit IP | Challenge served | Result |
| --- | --- | --- |
| Datacenter (Apify default, or no proxy) | The hard captcha | **3 of 12 sites passed** |
| Residential | Interstitial, or no challenge at all | **12 of 12 sites passed**, 2–5s each |

A datacenter address does not make Etsy slower. It changes the problem into a different, much harder one. Every attempt already pins a fresh residential session automatically, so a burnt exit node gets a genuinely new IP rather than a retry down the same dead pipe.

If a run fails wholesale, check the proxy group before anything else.

### Run it from the API or CLI

Replace `<TOKEN>` with your Apify API token.

Run and get the results back in one call:

```bash
curl -X POST "https://api.apify.com/v2/acts/xhrdev~etsy-scraper/run-sync-get-dataset-items?token=<TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"queries":["leather wallet"],"scrapeDetails":true,"maxItems":20}'
```

Start a run without waiting:

```bash
curl -X POST "https://api.apify.com/v2/acts/xhrdev~etsy-scraper/runs?token=<TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"queries":["leather wallet"],"scrapeDetails":true,"maxItems":20}'
```

With the Apify CLI:

```bash
apify call xhrdev/etsy-scraper --input '{"queries":["leather wallet"],"scrapeDetails":true,"maxItems":20}'
```

#### Fetching results

Every run writes to a dataset. Change `format` for JSON, CSV, or Excel:

```bash
## JSON
curl "https://api.apify.com/v2/datasets/<DATASET_ID>/items?token=<TOKEN>&clean=true&format=json"

## CSV
curl "https://api.apify.com/v2/datasets/<DATASET_ID>/items?token=<TOKEN>&clean=true&format=csv"

## Page through a large dataset
curl "https://api.apify.com/v2/datasets/<DATASET_ID>/items?token=<TOKEN>&offset=1000&limit=1000"
```

`<DATASET_ID>` comes back as `defaultDatasetId` on the run object.

### Limits and failure behaviour

**100 pages per run, hard.** Not configurable. It is a guard against a typo in `maxPagesPerQuery` turning into a bill and a hammering of Etsy. For more than that, split the work across runs.

**Failures never enter the dataset.** A page that could not be fetched is not written as a half-empty row — that would corrupt the clean table you are paying for. Instead every run writes a `RUN_SUMMARY` record to the key-value store:

```json
{
  "site": "etsy",
  "itemsScraped": 12,
  "pagesFetched": 13,
  "pagesFailed": 0,
  "failures": [],
  "solverHost": "https://trial.xhr.dev"
}
```

Read it at `https://api.apify.com/v2/key-value-stores/<STORE_ID>/records/RUN_SUMMARY`. `failures` holds up to 50 entries, each with the URL and why it failed.

**A run that scrapes nothing and failed at least one page exits as failed**, rather than reporting success over an empty dataset. If everything failed, the message points at the proxy group first, because that is nearly always the cause.

**Pagination is never speculative.** Page 2 is queued only after page 1 comes back and the site confirms how many pages exist, so you are not billed for fetching past the end of a short result set.

**Spend cap.** Set `maxTotalChargeUsd` on the run. The Actor stops fetching once it is reached, mid-run, rather than overshooting.

### Questions

**A run came back empty. Why?**
Check the proxy group first — a datacenter exit is the cause the overwhelming majority of the time. Then read `RUN_SUMMARY` in the key-value store for the per-page reasons.

**Is this affiliated with Etsy?**
No. This is an independent tool with no affiliation with, endorsement by, or connection to Etsy or any bot-protection vendor. Names and trademarks belong to their owners.

**How does it get past the challenge without a browser?**
It is solved as HTTP, by [xhr.dev](https://xhr.dev). No Chrome is launched, which is why a page costs a few hundred milliseconds of compute instead of the several seconds and ~1 GB of RAM a stealth browser needs.

**Can I use the solver directly, on a site that isn't Etsy?**
Yes — that is [DataDome Unblocker](https://apify.com/xhrdev/datadome-unblocker), which takes any URL and hands back the HTML and the clearance cookies. There is also [Akamai Unblocker](https://apify.com/xhrdev/akamai-unblocker) for Akamai Bot Manager.

**Can I run this on my own infrastructure?**
Yes. The solver these Actors call is a self-hosted Docker container, sold on a flat fee with no per-request pricing. See [xhr.dev](https://xhr.dev).

**Why is `currency` a `$` symbol in listing mode but `USD` in detail mode?**
Because that is what each page provides. The search card renders a symbol; the listing page carries a proper ISO code in its structured data. Nothing is normalised or guessed.

### Related Actors

- **[DataDome Unblocker](https://apify.com/xhrdev/datadome-unblocker)** — any URL behind DataDome, returns HTML plus clearance cookies.
- **[Akamai Unblocker](https://apify.com/xhrdev/akamai-unblocker)** — the same, for Akamai Bot Manager.
- **[leboncoin Scraper](https://apify.com/xhrdev/leboncoin-scraper)** — French classified ads.
- **[Anthropologie Scraper](https://apify.com/xhrdev/anthropologie-scraper)** — apparel and homeware.
- **[Grainger Scraper](https://apify.com/xhrdev/grainger-scraper)** — industrial supply lookups.

***

Built by [xhr.dev](https://xhr.dev). Independent tool, not affiliated with Etsy or any bot-protection vendor. Scrape only what you are permitted to access, and check the site's terms before you run anything at scale.

# Actor input Schema

## `queries` (type: `array`):

What to search Etsy for. Etsy server-renders about a dozen results per page and loads the rest in the browser, so searching is best used to discover listings — raise "Pages per search" to go wider.

## `itemIds` (type: `array`):

Etsy listing IDs to scrape directly — the number in an Etsy URL, e.g. 1214703903 from etsy.com/listing/1214703903/...

## `startUrls` (type: `array`):

Etsy URLs to scrape. /listing/ URLs are read as products; /search, /c/ and /market/ URLs are read as listings.

## `scrapeDetails` (type: `boolean`):

On by default, because the listing page is where Etsy's good data is: rating, review count, stock, shipping origin, materials and the full image set. Turn it off to take only what search results show — title, price, shop ID, thumbnail — much cheaper, much thinner.

## `maxPagesPerQuery` (type: `integer`):

How deep to page through each search. Roughly a dozen listings per page.

## `maxItems` (type: `integer`):

Stop once this many listings have been scraped, across everything in this run.

## `proxyConfiguration` (type: `object`):

Leave this on RESIDENTIAL. It is not a preference: the exit address decides which challenge DataDome serves, and datacenter ranges draw its most aggressive one. This Actor is built and verified around a residential exit, and a fresh session is pinned for every attempt automatically.

## `maxRetries` (type: `integer`):

How many times to try a page before giving up on it. Each attempt takes a fresh proxy session, so a dead exit node or a burnt IP gets a genuinely new chance.

## `timeoutSecs` (type: `integer`):

How long one page — challenge, solve and all — may take before it counts as a failed attempt.

## `maxConcurrency` (type: `integer`):

How many pages to work on at once. The solve path is HTTP-only and cheap, so this can go higher than a browser scraper would allow, but be a considerate guest on someone else's site.

## Actor input object example

```json
{
  "queries": [
    "leather wallet"
  ],
  "itemIds": [],
  "startUrls": [],
  "scrapeDetails": true,
  "maxPagesPerQuery": 1,
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "maxRetries": 3,
  "timeoutSecs": 120,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "leather wallet"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("xhrdev/etsy-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": ["leather wallet"] }

# Run the Actor and wait for it to finish
run = client.actor("xhrdev/etsy-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "leather wallet"
  ]
}' |
apify call xhrdev/etsy-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,xhrdev/etsy-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LwIM9zzsh9VxoYttm/builds/zyuZqHqV8cJqZCwaN/openapi.json
