# Lempertz.com Auction Results & Artist Scraper (`artsiom_k/lempertz-scraper`) Actor

Lempertz.com auction scraper — real realized prices, upcoming estimates, rich provenance detail, and cross-run artist rollups, with delta mode.

- **URL**: https://apify.com/artsiom\_k/lempertz-scraper.md
- **Developed by:** [Artsiom Kunitsyn](https://apify.com/artsiom_k) (community)
- **Categories:** E-commerce, Other
- **Stats:** 2 total users, 1 monthly users, 80.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $7.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## lempertz-scraper

Scrapes auction lots and a real cross-run artist rollup from
[Lempertz.com](https://www.lempertz.com) — Germany's oldest auction house (Cologne, Berlin, Munich,
Brussels). Real, open realized prices for historical sales, upcoming-sale estimates, and rich per-lot
detail (provenance, exhibitions, literature, expert certificates) most comparable actors don't expose
at all — all from one Actor.

### Contents

- [Key features](#key-features)
- [Output](#output)
- [Input](#input)
- [Input examples](#input-examples)
- [Incremental (delta) mode](#incremental-delta-mode)
- [How to scrape Lempertz.com](#how-to-scrape-lempertzcom)
- [You might also like](#you-might-also-like)
- [FAQ](#faq)

### 🔑 Key features

- **A single all-in realized price, plus the estimate range.** `price` is inclusive of buyer's
  premium (confirmed against the site's own "incl. premium" labeling) — the real final figure a
  buyer paid, not a hammer-only number.
- **Three entity types, one Actor.** `entityType: currentAuctions` (whatever the site currently shows
  under "Auctions" — see the caveat below), `auctionResults` (historical, with real realized
  prices — the default), or `artists` (a real rollup — see below).
- **Richer per-lot detail than comparable actors.** `provenance`, `exhibitions`, `literature`, and
  expert `certificate` text, plus a `cites` flag (CITES-restricted material) and a `droit_de_suite`
  flag (EU resale royalty applies) — fields `dorotheum-scraper`/`artcurial-scraper` don't report.
- **Genuine multi-attribution.** A lot can be jointly credited to more than one artist/maker (e.g. a
  manufacturer plus its designer) — `artists` is a list, not a single name field, and both get
  credited in the artist rollup.
- **A bulk-fetch discovery design.** One catalogue's entire lot list comes from a single request, not
  one request per lot — the historical archive (439+ catalogues) is reachable without a per-lot crawl.
- **Delta mode, tuned per entity type.** `auctionResults` defaults to the usual auto-incremental
  behavior (full scan first run, changes only after). `currentAuctions` always fully refreshes every
  run instead.
- **A real artist rollup, not a capped page.** `entityType: artists` reports uncapped running stats
  (lot count, average price, latest lot) accumulated from every `auctionResults` lot actually
  crawled — see [Output](#output)'s known gaps for exactly what that means and doesn't mean.

### 📋 Output

One dataset item per lot or artist, depending on `entityType` — see
[`.actor/dataset_schema.json`](.actor/dataset_schema.json) for the full field list, or the Output
tab's per-entity-type views for a readable table.

**Example lot record** (historical, sold, jointly attributed):

```json
{
  "source": "lempertz",
  "entity_type": "auctionResults",
  "external_id": "lempertz_139853",
  "url": "https://www.lempertz.com/en/catalogues/lot/1096-2/1023-a-pair-of-augsburg-silver-candelabra.html",
  "title": "A pair of Augsburg silver candelabra",
  "auction_number": 1096,
  "state": "Z",
  "sold": true,
  "price": 6820.0,
  "currency": "EUR",
  "estimate_low": 5500.0,
  "estimate_high": 7000.0,
  "literature": "Cf. an almost identical work by Billers illus. in: Seling 1980, no. 714, 842.",
  "artists": [
    { "artist_id": "19245", "artist_key": "19245", "name": "Allgöwer, Jakob Samuel" },
    { "artist_id": "19547", "artist_key": "19547", "name": "Biller, Friedrich Jakob" }
  ],
  "change_type": "new"
}
```

**Example artist rollup record:**

```json
{
  "source": "lempertz",
  "entity_type": "artists",
  "external_id": "lempertz_artist_29533",
  "name": "Doat, Taxile",
  "tracked_lot_count": 21,
  "tracked_total_price": 90148.0,
  "tracked_avg_price": 4292.76,
  "latest_lot_title": "A Sèvres porcelain vase \"La musique et la danse\" by Taxile Doat",
  "latest_auction_number": 1096
}
```

Results can be downloaded as JSON, CSV, or Excel from the Console's Output tab, or pulled via the
Apify API/dataset endpoint.

**Known gaps:**

- `price`/`estimate_low`/`estimate_high` are only set once a lot is actually sold (`sold: true`) /
  when the site publishes an estimate.
- A rare, real data gap: ~0.7% of sampled lots carry a "sold" state code with no price recorded on
  the site's own end — those still report `sold: true` but `price: null`, not a fabricated zero.
- `entityType: currentAuctions` is **not confirmed to be a clean "upcoming only" list** — the site's
  own `/en/auctions.html` can include catalogues that already have real results. Check each lot's own
  `sold`/`session_end` fields rather than trusting the entityType label alone.
- `entityType: artists` **does not crawl anything itself** — it reads a rollup that `auctionResults`
  runs build up over time as they process sold lots. Running `artists` before `auctionResults` has
  ever run pushes nothing (with a clear log message saying so). The rollup's numbers are real and
  uncapped, but only as complete as what's actually been crawled so far.
- `catalogue_uid` on a lot record is a raw site-internal id — **not** the same id used to discover
  and fetch that catalogue (a real naming collision confirmed on the site's own end; see
  [`docs/actors/lempertz-scraper.md`](../../docs/actors/lempertz-scraper.md) for the full story).

### ⚙️ Input

See [`.actor/input_schema.json`](.actor/input_schema.json) for the full JSON schema.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `entityType` | String | `auctionResults` | `currentAuctions`, `auctionResults`, or `artists`. |
| `startUrls` | Array of strings | *(none)* | Specific Lempertz catalogue URLs, lot URLs, or bare catalogue ids (e.g. `"1296-1"`) to scrape directly. Ignored for `entityType=artists`. Scope is always `"custom"` — no delisting-detection, no persisted baseline. Leave empty for the full crawl instead. |
| `maxItems` | Integer | `50` | Stop after pushing this many dataset items (lots or artists). Defaults to a fast, cheap preview (also what keeps an unconfigured run within Apify's automated 5-minute QA check). A full `auctionResults` crawl covers the entire historical archive — raise this or clear it (set to `null`) for that. |
| `mode` | String | `auto` | `auto` (recommended): the usual full-then-incremental behavior for `auctionResults`; always `full` for `currentAuctions` regardless of an existing baseline. `full`/`incremental` override this per run. Has no effect for `entityType=artists`. |
| `impersonate` | String | `chrome` (internal) | curl\_cffi TLS-impersonation target. No bot-management signal was observed anywhere on lempertz.com while building this Actor, so this is set internally by default. |
| `proxyConfiguration` | Object | `{"useApifyProxy": false}` | Apify Proxy config. Off by default — this Actor clears the site fine without one. |

### 🧪 Input examples

**Quick preview of recent historical results** (the default):

```json
{ "entityType": "auctionResults" }
```

**Full historical archive** (exhaustive, slow — `maxItems: null` explicitly overrides the 50-item
default):

```json
{ "entityType": "auctionResults", "maxItems": null }
```

**Artist rollup** (run `auctionResults` — ideally a full, unbounded run — at least once first):

```json
{ "entityType": "artists", "maxItems": null }
```

**A specific catalogue's results:**

```json
{ "entityType": "auctionResults", "startUrls": ["1296-1"] }
```

**Scheduled tracking run** — full, uncapped run (`maxItems` cleared — required for the baseline to
save and delistings to be detected):

```json
{ "entityType": "auctionResults", "mode": "incremental", "maxItems": null }
```

### 🔄 Incremental (delta) mode

Every `currentAuctions`/`auctionResults` run classifies each lot as `new`, `changed` (state/price
moved), `unchanged`, or `delisted`, using a state baseline persisted in a named Apify Key-Value Store
scoped to `entityType`. `entityType: artists` doesn't use delta mode at all — every run is a full
snapshot of the current rollup (see Output's known gaps).

- `auctionResults`: `mode: auto` (default) — first run for a scope pushes everything (`full`); later
  runs push only `new`/`changed`/`delisted` (`incremental`).
- `currentAuctions`: `mode: auto` always behaves as `full`, every run.
- A `startUrls`-scoped run is always partial and never updates the baseline or reports delistings.

Full design: [`docs/incremental-mode.md`](../../docs/incremental-mode.md).

### 🚀 How to scrape Lempertz.com

1. Open the Lempertz Auction Results & Artist Scraper in Apify Console and go to the **Input** tab.
2. Pick `entityType` (`currentAuctions`, `auctionResults`, or `artists`).
3. `maxItems` defaults to 50 (a quick preview) — clear it (set to `null`) for a full, uncapped run.
4. Click **Start**.
5. When the run finishes, browse results in the **Output** tab, or download as JSON/CSV/Excel, or
   fetch them via the API.
6. To track over time instead of scraping once: create a **Schedule** with `mode: auto`.

### 🔗 You might also like

- **[Dorotheum Auction Results & Artist Scraper](https://apify.com/artsiom_k/dorotheum-scraper)** —
  the first entry in the Auction houses collection.
- **[Artcurial Auction Results & Artist Scraper](https://apify.com/artsiom_k/artcurial-scraper)** —
  the second entry, with two realized-price figures per lot (hammer and all-in).

### ❓ FAQ

**Is it legal to scrape Lempertz.com?** It's legal to collect publicly available auction-result data
such as lot descriptions, prices, and sale information. Scrape it only with a legitimate purpose
under GDPR.

**How do I get only new/changed items?** Use `mode: auto` (or `incremental`) on a schedule — see
[Incremental mode](#incremental-delta-mode).

**Why is `price` null for a lot that's clearly listed?** That lot hasn't sold yet (or wasn't sold) —
check `estimate_low`/`estimate_high` instead. In a small number of cases the site marks a lot sold
with no price recorded on its own end; check `sold` rather than `price` if you need that distinction.

**Why does `entityType: artists` push nothing?** It reads a rollup built up by
`entityType: auctionResults` runs — it doesn't crawl anything on its own. Run `auctionResults`
(ideally a full, unbounded run) at least once first.

### Search keywords

lempertz scraper, auction house scraper, auction results scraper, art price data, art market
analytics, realized price data, hammer price data, auction price index, art collector data feed

# Actor input Schema

## `entityType` (type: `string`):

"currentAuctions": whatever the site's own /en/auctions.html currently links to — not confirmed to be a clean "upcoming only" list, see README. "auctionResults": historical, resolved lots with real realized prices (inclusive of buyer's premium) — the full archive by default, the default entityType. "artists": a real cross-run artist rollup, accumulated from auctionResults lots as they're crawled (not a fresh crawl of its own — run auctionResults at least once first, ideally a full/unbounded run, or this will push nothing). Each produces a different output shape (see dataset\_schema.json).

## `startUrls` (type: `array`):

Optional list of specific Lempertz catalogue URLs (e.g. https://www.lempertz.com/en/catalogues/detail/1296-1-african-and-oceanic-art.html), lot URLs (any lot from the catalogue you want), or bare catalogue ids (e.g. "1296-1") to scrape directly, instead of the full sitemap-driven crawl. Ignored for entityType=artists. Scope is always "custom", with no delisting-detection or persisted incremental baseline (a hand-picked list is necessarily partial) — leave empty for the full crawl instead.

## `maxItems` (type: `integer`):

Stop after pushing this many dataset items (lots or artists). Defaults to 50 — a fast, cheap preview, and what keeps an unconfigured run within Apify's automated 5-minute QA check. A full auctionResults crawl covers the entire historical archive (439+ catalogues, each with many lots) — raise this or clear it (set to null) for that; note a capped run never updates the incremental baseline (see mode below) or feeds the artist rollup for lots it never reaches.

## `mode` (type: `string`):

"auto" (recommended): for auctionResults, full scan on the first run for a scope, incremental (new/changed only) afterwards. For currentAuctions specifically, "auto" always behaves as a full refresh every run regardless of an existing baseline — upcoming sales are a small, frequently-changing dataset where a full refresh is more useful than a delta. "full": always push every item and refresh the baseline. "incremental": always push only new/changed items. Only a plain, unscoped crawl (no startUrls) can detect delistings or update the baseline. Has no effect for entityType=artists (every run there just reads the current rollup state).

## `impersonate` (type: `string`):

curl\_cffi browser TLS-impersonation target. No Cloudflare or other bot-management signal was observed anywhere on lempertz.com while building this Actor, so this defaults to "chrome" internally — override only if that stops working.

## `proxyConfiguration` (type: `object`):

Apify Proxy configuration. Leave off — this Actor clears the site fine without one in testing so far (it doesn't appear to sit behind Cloudflare or any other bot-management layer at all).

## Actor input object example

```json
{
  "entityType": "auctionResults",
  "maxItems": 50,
  "mode": "auto",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("artsiom_k/lempertz-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("artsiom_k/lempertz-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call artsiom_k/lempertz-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,artsiom_k/lempertz-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CrHfCBNTkoY6QXooQ/builds/UFmVqJgvf5sofIHFg/openapi.json
