Allegro.pl Scraper — Offers, Prices & Sellers | $4/1K
Pricing
from $4.75 / 1,000 listings
Allegro.pl Scraper — Offers, Prices & Sellers | $4/1K
Scrape Allegro.pl offers by keyword: title, price, auction/buy-now format, images, category, and offer URL. No login, OAuth2, or proxy setup — DataDome anti-bot is handled for you. Pay per result.
Pricing
from $4.75 / 1,000 listings
Rating
0.0
(0)
Developer
Vitalii Bondarev
Maintained by CommunityActor stats
0
Bookmarked
8
Total users
3
Monthly active users
6 days ago
Last modified
Categories
Share
Allegro.pl Scraper — Offers, Prices & Sellers (No Login)
Built for Polish e-commerce sellers monitoring competitor pricing on Allegro, importers researching the Polish market, and price-intelligence tools covering the CEE region.
Pricing: Pay per offer — $6.95 / 1,000 results. First 10 results free.
No login. No OAuth2. No proxy setup. Allegro is localized to Poland — this actor handles access for you (Polish residential routing built in). Just give it a search phrase.
Scrape Allegro.pl — Poland's #1 marketplace (~20M active offers) — by keyword and collect structured offer data: title, price, format (auction/buy-now), images, category, and a direct offer link.
What you get
| Field | Description |
|---|---|
id | Allegro offer ID (numeric) |
name | Offer title |
price | Price amount (float) |
currency | Currency code (PLN) |
format | buy_now or auction |
imageUrl | Primary image URL |
images | All image URLs (list, highest resolution available) |
categoryId | Allegro category ID |
offerUrl | Direct link to the offer |
scraped_at | ISO 8601 run timestamp |
parse_confidence | 0.0–1.0 per-record quality score |
warnings | List of field-level quality codes |
stockand sellerlocationare not exposed on the public listing view, so they are returned asnull. For per-offer stock/specs, a detail-page mode can be added on request.
Inputs
| Input | Type | Required | Description |
|---|---|---|---|
phrase | string | Yes | Search keyword(s), e.g. "laptop", "nike shoes" |
maxItems | integer | No | Max offers to return. Default 100. Set 0 for no limit (capped at the listing's last page) |
country | string | No | Two-letter proxy country. Default pl (recommended for Allegro) |
How it works
- Fetch: requests
https://allegro.pl/listing?string=<phrase>&p=<page>through a managed access layer that routes via Polish residential IPs, returning the rendered HTML. - Extract: parses the
__listing_StoreStateSSR JSON embedded in the page — far more stable than scraping CSS classes (Allegro rotates class names; the JSON schema does not). - Paginate: reads
searchMeta.lastAvailablePageand walks pages untilmaxItemsis reached or the listing ends. - Normalize: each offer is flattened to the schema above, with a per-record confidence score.
vs. competitors
| Feature | This actor | HTML-scraper incumbents |
|---|---|---|
| Setup | Just a search phrase | Often need proxy/cred config |
| Access protection | Handled for you | Breaks frequently |
| Parser resilience | SSR-JSON (schema-stable) | CSS-class scraping (brittle) |
format (auction/buy-now) | Yes | Rarely parsed cleanly |
| Per-record quality score | Yes | No |
Use with AI agents (MCP)
Compatible with Claude, GPT, and other MCP-aware agents:
https://mcp.apify.com/?tools=bovi/allegro-offers
Pricing example
| Volume | Cost |
|---|---|
| 100 offers | $0.70 |
| 1,000 offers | $6.95 |
| 10,000 offers | $69.50 |
First 10 results are free.
FAQ
Do I need Allegro credentials or a developer account? No. This actor scrapes the public listing pages — no login, OAuth2, or API keys required.
Do I need to configure a proxy? No. Polish residential routing is built in — no external proxy or credentials needed.
What output formats are available? JSON (default), CSV, and Excel — via the Apify dataset export or API.
Why are stock and location null?
Those fields aren't shown on the public listing/search view. They live on individual offer detail pages — available via a detail-page mode on request.
CEE cluster
Allegro is the #1 CEE marketplace. Bundle with the OLX Scraper (classifieds) as a "Poland/CEE ecom data" stack.
Integrations
The JSON/dataset output drops into the tools you already run, no glue code:
- n8n / Make / Zapier — pipe every new dataset item into 500+ apps (Google Sheets, Airtable, Slack, HubSpot, your database): n8n, Make, Zapier.
- Webhooks — fire your own endpoint the moment a run finishes (docs).
- MCP server — expose this actor as a tool to Claude, Cursor, or any MCP client.
- API & SDKs — fetch the dataset as JSON, CSV, or Excel through the Apify REST API or the Python / JS SDKs.
See all Apify integrations.
More scrapers from our toolkit
Building a data pipeline? These actors pair well with this one — each runs on your own Apify account with the same pay-per-result pricing, no subscription:
- Amazon Products Scraper
- Autoscout24 Cars Scraper
- Competitor Price Monitor
- Craigslist Scraper
- Doordash Scraper
- Ebay Scraper
Chain any of them together from the Integrations tab (the Run succeeded trigger) to build a multi-step workflow — one actor's output feeds the next.
Usage statistics
This Actor creates a small, content-free summary at the end of each run. It is used only to monitor reliability and improve this Actor. A copy is saved as USAGE_STATS in your own Apify key-value store, so you can see the exact record created for your run.
Set disableUsageStats to true in the input to opt out. Nothing is sent then; your USAGE_STATS record only says that statistics were disabled.
Only these fields are recorded:
- schema version, Actor name and build number;
- UTC start and finish hour (not a precise timestamp);
- run duration, number of results and time to the first result, each as a coarse range;
- whether the result was empty, the end status, and an error type from a fixed list;
- memory setting and counts of charged events;
- names of the input fields you set, never their values;
- the selected option for input fields that offer a fixed list of choices (for example a sort order).
We do not collect input text, search terms, URLs, domains, usernames, email addresses, names, proxy credentials, tokens, scraped records, output items, raw error messages, stack traces, or hashes of any of those values. Records are kept for no longer than 13 months, used only as aggregated operational statistics, and never sold or shared.
Additional fields (Phase 2)
This Actor also records your Apify user ID, whether Apify marks the account as paying, the size range of list inputs, the selected country when the input offers a fixed list of countries, and one category from a fixed Actor taxonomy. We use these fields only for aggregate reliability, repeat-use and cross-Actor analysis; reports suppress any cell with fewer than five distinct users.
The same disableUsageStats: true input flag turns these fields off too. The user ID is removed after 13 months; we do not export, sell, share, or attempt to re-identify this data.
Run-outcome signals (v2)
To learn whether a run did what it was asked to do, the record also holds a few more coarse ranges and yes/no flags. None of them contains content:
- the result limit you asked for (a range, when the input has one) and what share of it was delivered;
- results delivered per input item you listed (a range);
- output quality as ranges: how fully the result fields were filled, the share of rows that look like errors, the share of duplicate rows, and how many different fields appeared. These are counted in memory while results are saved; no result content is kept;
- how the run was started (console, API, schedule, webhook, another Actor);
- how it ended: stopped by you, timed out, reached the requested limit, stopped by the charge limit, and how many times the platform moved the run;
- if this Actor reports it: how many items to process worked or failed (ranges) and one failure reason from a fixed list;
- a short code made from the names of the input fields you set, never their values.
Repeat-run fingerprint (v2)
When your Apify user ID is recorded (see above), the record also holds an 8-character one-way code made from your input (proxy settings left out) and this Actor's name. It only lets us see that the same account ran the same input again soon after an unsatisfying run; we never see the input itself. It is stored only in the database, never published, and reports use it in aggregate with the same five-user minimum. It is the one exception to the statement above that no hashes are collected, and disableUsageStats: true turns it off.