Falabella Product Search Scraper
Pricing
from $2.00 / 1,000 results
Falabella Product Search Scraper
Scrape Falabella product search results: prices, discounts, brands, and availability across Latin America.
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Danilo Frias
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Apify actor that scrapes public, logged-out product search results from falabella.com (Colombia default; Chile, Peru, and Argentina storefronts also supported).
What it scrapes
For a given search query, the actor walks the search-results pages
(https://www.falabella.com.co/falabella-co/search?Ntt=<query>&page=N) and
extracts one record per product listing:
| Field | Type | Description |
|---|---|---|
title | string | Product display name |
brand | string or null | Brand name as shown on the listing |
price_cop | float or null | Current price shown (lowest non-crossed-out price; COP on the Colombia storefront) |
original_price_cop | float or null | Crossed-out reference price when a discount is shown |
product_url | string | Canonical product page URL |
image_url | string or null | First product image URL (Falabella CDN) |
rating | float or null | Average rating (e.g. 4.57) when reviews exist |
Falabella server-renders search pages as Next.js, so the actor parses the
embedded __NEXT_DATA__ JSON with httpx + beautifulsoup4/lxml — no
JavaScript engine, no headless browser, no login, no session cookies beyond
the anonymous ones the site sets itself.
Who would use it
Price-comparison builders, retail analysts, and marketplace researchers who need structured Falabella catalog data (titles, brands, prices, ratings) for a keyword without running a browser farm.
Input
| Field | Type | Default | Description |
|---|---|---|---|
query | string | "audifonos bluetooth" | Product search query (required). |
country | string | "co" | Storefront country: co (Colombia), cl (Chile), pe (Peru), ar (Argentina). |
max_items | integer | 60 | Maximum product records to scrape. The actor paginates (&page=2, &page=3, ...) until this many records are collected or results run out. |
Sample output record
{"title": "Audífonos bluetooth WH-CH520","brand": "SONY","price_cop": 119900.0,"original_price_cop": 299900.0,"product_url": "https://www.falabella.com.co/falabella-co/product/prod13360109/Audifonos-bluetooth-Sony-WH-CH520","image_url": "https://media.falabella.com.co/falabellaCO/73319593_1/public","rating": 4.6594}
(Real record from a 2026-09-26 test run against the Colombia storefront.)
Coverage and limits
- Search results are paginated server-side; the actor stops at
max_itemsor when the result set is exhausted (the site reported ~1,415 results / 30 pages for the test query). price_copreflects the lowest current listed price; event/promo prices (eventPrice) and card prices (cmrPrice) are treated as current prices.- Nulls are normal: unrated products have
rating: null, and listings without a crossed-out price haveoriginal_price_cop: null. - One homepage request seeds anonymous cookies, then ~1 request/second between result pages (politeness throttle).
- Only the
search?Ntt=surface is scraped. Falabella redirects the legacy?text=parameter to the homepage, so the actor usesNtt(the parameter in the site's own schema.orgSearchAction). - Field names say
price_cop; on non-Colombia storefronts the price is in that country's listing currency.
Data rules
- Public, logged-out pages only. No login flows, no stored session cookies, no CAPTCHA solving, no proxy services, no headless browsers (plain HTTP).
- Product and brand names on public listings are fair game; the actor never collects anything identifying private individuals.
- Respects a ~1 request/second pace between pages and stops cleanly after three consecutive page failures.
Local test
# write inputmkdir -p storage/key_value_stores/defaultprintf '{"country":"co","max_items":60,"query":"audifonos bluetooth"}' \> storage/key_value_stores/default/INPUT.json# run with the shared venv/path/to/.venv/bin/python main.py# results land in storage/datasets/default/