# Reverb Listings Scraper (`scrapyx/reverb-listings-scraper`) Actor

Music gear listings from Reverb: price, condition, make, model, year, shop and shipping. Every run proves its filters were applied -- Reverb accepts an unrecognised filter value, ignores it, and answers with the whole 794,000-listing catalogue instead of an error.

- **URL**: https://apify.com/scrapyx/reverb-listings-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** E-commerce
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.75 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reverb Listings Scraper

Music gear from Reverb — guitars, amps, pedals, synths, drums, pro audio — with
price and currency, condition, make, model, year, the shop selling it, and
shipping rates per region.

Built on Reverb's own public REST API. HTTP only, no browser.

### What it is for

- **Price research on used gear** — what a 1979 Stratocaster actually sells for,
  by condition, by year, by region.
- **Inventory monitoring** — new listings for a make or model, newest first.
- **Market sizing** — how many listings exist in a category, and at what prices.
- **Dealer tracking** — every row carries the shop, whether it is a preferred
  seller, and its shipping rates.

### Input

| field | what it does |
| --- | --- |
| `queries` | Search terms: `stratocaster`, `les paul`, `moog`, `shure sm57`. |
| `categorySlugs` | `electric-guitars`, `effects-and-pedals`, `amps`, `drums-and-percussion`… |
| `conditions` | `new`, `used`, `b-stock`, `mint`, `excellent`, `very-good`, `good`, `fair`, `poor`, `non-functioning`. |
| `makes` | `Fender`, `Gibson`, `Boss`, `Shure`. |
| `priceMin`/`priceMax`, `yearMin`/`yearMax` | Price and year of manufacture. |
| `shipsTo`, `currency` | Two-letter country, three-letter currency. |
| `sortBy` | Newest, oldest, price low→high, price high→low. |
| `maxItems`, `perPage`, `maxConcurrency` | Limits and pacing. |

Give both search terms and categories and every term is run inside every
category.

### Five things about this data worth knowing before you trust a run

#### 1. A filter Reverb does not recognise is ignored, not rejected

Against `guitar`, whose unfiltered total is ~794,000:

| filter | listings reported |
| --- | --- |
| `condition=used` | 257,022 — real |
| `condition=zzqq` | **794,220** — silently unfiltered |
| `condition=brand-new` | **794,217** — silently unfiltered |
| `make=Fender` | 82,934 — real |
| `make=ZzqqNotAMake` | **794,222** — silently unfiltered |
| `ships_to=ZZ` | **794,220** — silently unfiltered |

Only a bad `currency` is refused out loud. `brand-new` is the nastiest of
these: it is exactly what the listings themselves display, and the slug that
actually filters is `new`.

So **every run spends one extra request** on the same query *without* the
filters, and stops if the filters did not move the result count. The summary
reports `baselineTotal`, `filtersHeld` and `totalMovedShare` so you can see the
arithmetic.

#### 2. Three numbers describe the result set and none is the limit

```
total         794,218   the real match count — honest
total_pages        50   a CONSTANT. It is 50 at perPage 1, 5, 20 and 50.
actual limit   20,000   rows, as an OFFSET, not a page number
```

Confirmed four ways — the break is always at offset > 20,000, never at a page
number. A scraper that trusts `total_pages` stops after 50 pages; one that
computes `total / perPage` asks for 15,884 pages and hits an HTTP 400 partway.

**And in practice you get far less than 20,000**, because deep paging re-serves
listings you have already seen. A measured walk of 112 pages returned 3,837
unique listings and **1,763 duplicates**. Those duplicates are dropped and
counted in `duplicateRowsDropped`. To go wider, slice by price band, category or
make rather than paging.

#### 3. With some search terms, the count is filtered by category and the rows are not

This is the subtle one. Searching `guitar` inside `electric-guitars`:

- `total` says 218,938 — the filtered number
- Reverb's own echo says `"guitar" in Electric Guitars`
- and **18 of 20 returned listings are acoustic guitars**

Measured across ten search terms against the same category, only two misbehave,
and both are terms that name a neighbouring category:

| search term | rows actually in the category |
| --- | --- |
| `guitar` | **2 / 20** |
| `acoustic` | **6 / 20** |
| stratocaster, les paul, telecaster, amp, pedal, fuzz, ibanez | 20 / 20 |

The total cannot detect this — it looks perfect in exactly this case. Only the
rows can. Every run with a category therefore reports `categoryHeldShare`, and a
run where most rows are from elsewhere is an **error**, not a page of the wrong
department.

#### 4. Sorting only works in a form nobody would guess

`sort=price|desc` works. `sort=price_desc` — the spelling every other API uses —
is accepted, ignored, and returns the default order with HTTP 200. You pick from
an enum and this Actor sends the working form, then **checks the returned order
really is sorted** and reports `sortVerified`.

One wrinkle behind that check: Reverb sorts on the listing's **native** amount
and shows you a **converted** one. A `price|asc` page came back `0.99 USD, 1.00
USD, …, 1.50 (native CAD), 1.48 (native AUD)` — three inversions in twenty rows,
every one at a currency boundary. Verification is therefore scoped to rows whose
listing currency is the currency shown.

#### 5. `condition` filters on a tier the rows do not report

`condition=used` narrows 794,218 → 257,022 correctly, and then **not one
returned row says "Used"** — they say Excellent, Very Good, Good, Fair. `used`
is a meta-condition covering the grades. So per-row verification is applied only
to `make`, whose vocabulary does match the rows (and even there, matched as a
prefix: a `Fender` search returns rows labelled both `Fender` and
`Fender Guitars`).

### Output

Three record types in one dataset, told apart by `recordType`:

- **`LISTING`** — `title`, `make`, `model`, `year`, `finish`, `price` +
  `currency` + `buyerPrice` (tax-inclusive where it applies) +
  `listingCurrency`, `condition` + `conditionSlug`, `categories`, `shopName` +
  `preferredSeller`, `shipping` (cheapest rate and region), `photos`,
  `listingUrl`.
- **`QUERY_SUMMARY`** — `matchesReported`, `totalPagesReported` (Reverb's
  constant 50, shown so the difference is visible), `baselineTotal`,
  `filtersHeld`, `totalMovedShare`, `makeHeldShare`, `categoryHeldShare`,
  `sortVerified`, `ceilingHit`, `duplicateRowsDropped`.
- **`ERROR`** — one row naming what went wrong, instead of a silent empty.

### Technical notes

- **HTTP-only**, `curl_cffi` with `chrome124`. No WAF, no challenge.
- **`Accept-Version: 3.0` is mandatory** — without it every call is a 400.
  Version 2.0 is deprecated, and its error message came back **in Italian**
  from a non-US exit IP: Reverb localises error strings by geography, so no
  code here matches on error text, only on status and structure.
- **HTTP 400 is an answer, not a fault** past the offset ceiling, so it is
  never retried.
- **Proxy is optional and off by default.** Prices follow the `currency`
  parameter, not the exit IP, so leave the pool unpinned if you enable it.
- **robots.txt was read per origin**, comments included. No AI-bot group, no
  blanket disallow; the four `/api/` rules cover account data and ad surfaces,
  none of which this Actor reads.

### Known limits

- **No price guide.** `/api/priceguide` returns 403 without authentication, so
  sold-price history is out of reach.
- \~4,000 unique listings per query in practice (see above). Slice to go wider.
- `perPage` is clamped at 50 by Reverb; larger values silently return 50.
- 13 category slugs in Reverb's tree are not unique (`cables`, `delay`,
  `archtop`, `left-handed`…). An ambiguous bare slug is refused with the
  qualified `root/slug` alternatives rather than resolved to whichever the
  parser saw last.

# Actor input Schema

## `queries` (type: `array`):

What to search for on Reverb - 'stratocaster', 'les paul', 'moog', 'shure sm57'. Each term becomes its own target with its own summary row.

## `categorySlugs` (type: `array`):

Reverb category slugs such as electric-guitars, effects-and-pedals, amps, drums-and-percussion. Given both, every search term is run inside every category. Reverb's API only filters by category UUID - four other plausible parameter spellings are accepted and silently ignored - so this Actor takes the readable slug and resolves the UUID itself. An unknown slug stops the run and lists what exists.

## `conditions` (type: `array`):

These are Reverb's filter SLUGS, not the labels shown on a listing. A listing displayed as 'Brand New' is filtered with 'new'; passing 'brand-new' is accepted and silently ignored. Note 'used' is a meta-condition: it correctly narrows the results, but the rows still come back graded Excellent, Very Good and so on.

## `makes` (type: `array`):

Brand names as Reverb spells them - Fender, Gibson, Boss, Roland, Shure. Verified per row afterwards and reported as makeHeldShare.

## `priceMin` (type: `integer`):

Lowest price, in the selected currency.

## `priceMax` (type: `integer`):

Highest price, in the selected currency.

## `yearMin` (type: `integer`):

Earliest year of manufacture. Useful for vintage searches - Reverb carries instruments back to the 1920s.

## `yearMax` (type: `integer`):

Latest year of manufacture.

## `shipsTo` (type: `string`):

Two-letter country code - US, DE, GB, JP. Restricts to listings that ship there. An unrecognised code is ignored by Reverb rather than rejected, so it is checked here first.

## `currency` (type: `string`):

Three-letter code - USD, EUR, GBP. This is the one filter Reverb does reject out loud when it is wrong.

## `sortBy` (type: `string`):

Reverb's API only honours sorting in a pipe form ('price|desc'); the conventional 'price\_desc' spelling is accepted and silently ignored. This Actor sends the working form and then checks the returned order really is sorted, reporting sortVerified.

## `maxItems` (type: `integer`):

Overall cap across every target. Note Reverb serves at most 20,000 rows per query however deep you page - and its own total\_pages field says 50 regardless of page size, so neither number is the limit. A run that reaches the wall reports ceilingHit; slice by price band, category or make to go wider.

## `perPage` (type: `integer`):

Reverb clamps this at 50 and does not say so - 51, 100 and 500 all return exactly 50 rows.

## `maxConcurrency` (type: `integer`):

Parallel requests across targets.

## `minRequestInterval` (type: `integer`):

Politeness delay between request starts. It paces starts only and does not hold a concurrency slot, so raising it slows the run without idling workers.

## `proxyConfiguration` (type: `object`):

Optional. Reverb's API did not challenge any TLS profile from a plain IP, so this is off by default. Leave the pool unpinned if you enable it: prices follow the currency parameter, not the exit IP.

## Actor input object example

```json
{
  "queries": [
    "stratocaster"
  ],
  "sortBy": "newest",
  "maxItems": 200,
  "perPage": 50,
  "maxConcurrency": 2,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "stratocaster"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/reverb-listings-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": ["stratocaster"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/reverb-listings-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "stratocaster"
  ]
}' |
apify call scrapyx/reverb-listings-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/reverb-listings-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/EPLR9X4UIYuhEWN5r/builds/9oh8Lz1GfPeVce36d/openapi.json
