# Falabella Product Search Scraper (`stilled_stalagmite/falabella-product-search-scraper`) Actor

Scrape Falabella product search results: prices, discounts, brands, and availability across Latin America.

- **URL**: https://apify.com/stilled\_stalagmite/falabella-product-search-scraper.md
- **Developed by:** [Danilo Frias](https://apify.com/stilled_stalagmite) (community)
- **Categories:** E-commerce
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Falabella Product Search Scraper

Apify actor that scrapes **public, logged-out product search results** from
[falabella.com](https://www.falabella.com.co/) (Colombia default; Chile,
Peru, and Argentina storefronts also supported).

### What it scrapes

For a given search query, the actor walks the search-results pages
(`https://www.falabella.com.co/falabella-co/search?Ntt=<query>&page=N`) and
extracts one record per product listing:

| Field | Type | Description |
|---|---|---|
| `title` | string | Product display name |
| `brand` | string or null | Brand name as shown on the listing |
| `price_cop` | float or null | Current price shown (lowest non-crossed-out price; COP on the Colombia storefront) |
| `original_price_cop` | float or null | Crossed-out reference price when a discount is shown |
| `product_url` | string | Canonical product page URL |
| `image_url` | string or null | First product image URL (Falabella CDN) |
| `rating` | float or null | Average rating (e.g. 4.57) when reviews exist |

Falabella server-renders search pages as Next.js, so the actor parses the
embedded `__NEXT_DATA__` JSON with `httpx` + `beautifulsoup4`/`lxml` — **no
JavaScript engine, no headless browser, no login, no session cookies beyond
the anonymous ones the site sets itself.**

#### Who would use it

Price-comparison builders, retail analysts, and marketplace researchers who
need structured Falabella catalog data (titles, brands, prices, ratings) for
a keyword without running a browser farm.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `query` | string | `"audifonos bluetooth"` | Product search query (required). |
| `country` | string | `"co"` | Storefront country: `co` (Colombia), `cl` (Chile), `pe` (Peru), `ar` (Argentina). |
| `max_items` | integer | `60` | Maximum product records to scrape. The actor paginates (`&page=2`, `&page=3`, ...) until this many records are collected or results run out. |

### Sample output record

```json
{
  "title": "Audífonos bluetooth WH-CH520",
  "brand": "SONY",
  "price_cop": 119900.0,
  "original_price_cop": 299900.0,
  "product_url": "https://www.falabella.com.co/falabella-co/product/prod13360109/Audifonos-bluetooth-Sony-WH-CH520",
  "image_url": "https://media.falabella.com.co/falabellaCO/73319593_1/public",
  "rating": 4.6594
}
```

(Real record from a 2026-09-26 test run against the Colombia storefront.)

### Coverage and limits

- Search results are paginated server-side; the actor stops at `max_items`
  or when the result set is exhausted (the site reported ~1,415 results /
  30 pages for the test query).
- `price_cop` reflects the lowest current listed price; event/promo prices
  (`eventPrice`) and card prices (`cmrPrice`) are treated as current prices.
- Nulls are normal: unrated products have `rating: null`, and listings
  without a crossed-out price have `original_price_cop: null`.
- One homepage request seeds anonymous cookies, then ~1 request/second
  between result pages (politeness throttle).
- Only the `search?Ntt=` surface is scraped. Falabella redirects the legacy
  `?text=` parameter to the homepage, so the actor uses `Ntt` (the parameter
  in the site's own schema.org `SearchAction`).
- Field names say `price_cop`; on non-Colombia storefronts the price is in
  that country's listing currency.

### Data rules

- Public, logged-out pages only. No login flows, no stored session cookies,
  no CAPTCHA solving, no proxy services, no headless browsers (plain HTTP).
- Product and brand names on public listings are fair game; the actor
  **never collects anything identifying private individuals**.
- Respects a ~1 request/second pace between pages and stops cleanly after
  three consecutive page failures.

### Local test

```bash
## write input
mkdir -p storage/key_value_stores/default
printf '{"country":"co","max_items":60,"query":"audifonos bluetooth"}' \
  > storage/key_value_stores/default/INPUT.json
## run with the shared venv
/path/to/.venv/bin/python main.py
## results land in storage/datasets/default/
```

# Actor input Schema

## `country` (type: `string`):

Falabella storefront country code. One of: co (Colombia), cl (Chile), pe (Peru), ar (Argentina).

## `max_items` (type: `integer`):

Maximum number of product records to scrape. The actor paginates through search results until this many records are collected or results run out.

## `query` (type: `string`):

Product search query, e.g. 'audifonos bluetooth' or 'lavadora'.

## Actor input object example

```json
{
  "country": "co",
  "max_items": 60,
  "query": "audifonos bluetooth"
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("stilled_stalagmite/falabella-product-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("stilled_stalagmite/falabella-product-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call stilled_stalagmite/falabella-product-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,stilled_stalagmite/falabella-product-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WqJvtVggvoQMLLNja/builds/lxpU3YdEiBTkTcKyk/openapi.json
