🍷Wine Searcher data scraper avatar

🍷Wine Searcher data scraper

Pricing

from $1.20 / 1,000 results

Go to Apify Store
🍷Wine Searcher data scraper

🍷Wine Searcher data scraper

price, ratings, critic score, and the full per-merchant offers table (merchant name, location, price, bottle size, verified status) for each wine.

Pricing

from $1.20 / 1,000 results

Rating

0.0

(0)

Developer

MrDoe

MrDoe

Maintained by Community

Actor stats

0

Bookmarked

7

Total users

2

Monthly active users

2 days ago

Last modified

Categories

Share

Wine-Searcher Scraper

Scrapes wine listings and detail pages from wine-searcher.com — price in USD, ratings, critic score, a sample merchant offer and critic review by default, and as many of each as you ask for.

What it does

  • Listing pages: give it a search term or a wine-searcher.com search URL, and it finds every wine/vintage link on the results page.
  • Detail pages: for each wine it visits, it reads the page's own schema.org structured data (more reliable than scraping visual markup) for wine name, image, region, vintage, style, grape, critic score, and offer count - plus one merchant offer and one critic review by default, so the output looks complete without any configuration.
  • Whole-site mode: optionally keep following every wine link discovered on every page (not just the ones you started with), up to a limit you set.

Input

FieldTypeDescription
startUrlsarray of stringswine-searcher.com URLs to scrape directly — listing pages or specific vintage detail pages.
searchQueriesarray of stringsPlain wine search terms (e.g. "opus one") — turned into search URLs automatically. Alternative to startUrls.
scrapeWholeSitebooleanIf enabled, keeps following every wine link found on every page instead of just the ones you gave it. Default false.
maxItemsintegerStop once this many wine detail pages have been scraped. Default 5.
maxConcurrencyintegerNumber of pages processed in parallel (1–5). Default 1.
maxOffersintegerHow many individual merchant offers to include per wine. Default 1, which costs nothing extra - it's already on the page fetched for the wine's other fields. Set higher to fetch more (following the site's own offer pagination, 50/page - extra page loads for each 50 beyond the first) or very high to effectively get every offer. 0 leaves offers out entirely.
maxReviewsintegerHow many individual critic reviews to include per wine - free, no extra page loads. Default 1. 0 leaves reviews out entirely.
proxyConfigurationobjectApify Proxy settings. Defaults to a US exit, which is what makes offer prices come back in USD instead of whatever currency the underlying network happens to be in - wine-searcher.com detects offer currency from the visiting session's location.

Output

One dataset row per wine. With the default maxOffers/maxReviews of 1:

{
"url": "https://www.wine-searcher.com/find/opus+one+napa+valley+county+north+coast+california+usa/2019",
"wineName": "Opus One",
"image": "https://www.wine-searcher.com/images/labels/55/91/11715591.jpg",
"region": "Napa Valley, USA",
"vintage": "2019",
"style": "Red - Bold and Structured",
"grape": "Bordeaux Blend Red",
"price": 441.0,
"currency": "$",
"userRating": 4.5,
"userRatingCount": 31,
"criticScore": "97 / 100",
"criticReviewCount": 7,
"offerCount": 50,
"offerPageCount": 4,
"offers": [
{
"merchantName": "Catawiki",
"merchantRegion": "Amsterdam",
"merchantCountry": "Netherlands",
"price": 48307.0,
"currency": "NPR",
"bottleSize": "Bottle (750ml)",
"availability": "InStock",
"priceValidUntil": "2026-08-24T17:44:55Z",
"offerUrl": "https://www.catawiki.com/en/c/443?anchor_lot_id=106232853",
"offerSku": "b61d1fc53a57f41129852770ba2eab10"
}
],
"reviews": [
{
"criticName": "Owen Bargreen",
"score": "98 / 100",
"vintage": "2019 Vintage",
"tastedDate": "Jan 2026",
"note": "The 2019 Opus One is really coming into its own right now...",
"criticProfileUrl": "https://www.wine-searcher.com/critics-113-owen+bargreen"
}
]
}
  • offerCount/offerPageCount reflect the true scale regardless of maxOffers - offerCount is at least 50 here even though only 1 offer is included, and offerPageCount: 4 means up to ~200 offers exist across the site's own pagination. Raise maxOffers to pull more of them in (each additional 50 costs one more page load); set it very high to effectively get every offer that exists.
  • currency (inside each offer) is the site's own ISO 4217 code (e.g. "USD", "NPR"), not a symbol - avoids the ambiguity a bare "Rs" (Indian or Nepalese rupee?) would carry.
  • reviews are critic reviews only - wine-searcher.com doesn't publish individual written user reviews on its public pages, just an aggregate rating (userRating/userRatingCount above). Some critics gate their score/note behind a subscription - that comes through faithfully as "score": "? / 100" rather than being guessed or dropped.
  • Setting maxOffers: 0 or maxReviews: 0 leaves that key out of the row entirely, not an empty array.

Notes

  • Most fields come from the page's own schema.org structured data and are reliable whenever a wine has that data at all. price/currency at the top level (the site's own global average, not a computed one) and userRating/userRatingCount are the exception - the site doesn't publish those in structured form, so they're read from the visual page instead and may occasionally come back null even on a successful fetch.
  • Each page load takes several seconds — this is expected and by design for reliability, not a performance bug. Keep maxItems/maxConcurrency modest for testing, and expect a high maxOffers to take noticeably longer on wines with many offers (each additional 50 is a separate page load).
  • Every page fetch uses its own fresh browser and its own fresh proxy exit IP - live-verified that reusing either one for a second page gets it (and everything after it) blocked. This makes each page slower (a real ~20-40s browser start) but is what actually holds up in production; a page that still gets blocked once is retried on a fresh session/IP rather than failing the whole run.
  • The first attempt at each page uses a cheaper proxy exit; only a retry (after that page got blocked) escalates to a residential exit, which blocks less often but costs more. Most pages succeed on the cheap first attempt. Set proxyConfiguration yourself to use exactly that for every attempt instead - only leaving it unset gets this escalation.

Running the tests

test/test_parsers.py covers the pure HTML/JSON-LD extraction logic in src/parsers.py against small fixture pages - no browser or network involved. Not part of the deployed image (pytest isn't in requirements.txt); run it locally with:

pip install pytest beautifulsoup4
pytest test/