Peru Grocery Basket Aggregator (Metro, Wong, PlazaVea, Vivanda)
Pricing
from $0.40 / 1,000 results
Peru Grocery Basket Aggregator (Metro, Wong, PlazaVea, Vivanda)
Compare a grocery basket across Peru's four largest chains (Metro, Wong, PlazaVea, Vivanda): cheapest chain per item, basket totals per chain, and where switching stores saves money. Prices in PEN. Use specific item names with size hints (e.g. 'aceite vegetal 1L') for best results.
Pricing
from $0.40 / 1,000 results
Rating
0.0
(0)
Developer
Philip Kirkbride
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
13 hours ago
Last modified
Categories
Share
Peru Grocery Basket Aggregator (Metro · Wong · PlazaVea · Vivanda)
One Apify Actor that prices a whole grocery basket across Peru's four largest chains — Metro, Wong, PlazaVea and Vivanda — in a single run and answers three questions: which chain is cheapest per item, what a basket costs at each chain, and where switching stores saves money. All prices in PEN (S/).
Plain-HTTP fetch of each chain's public VTEX catalog (keyless JSON endpoint, no browser, no proxies, no login). Every item triggers one first-page search per chain (top-sale ordering, 50 results), matched across chains, and the run emits normalized offer records plus a basket summary.
Input
basket(optional since 0.2.3): 1–50 free-text items, e.g.["arroz extra 5kg", "leche evaporada chica", "huevos", "pollo entero", "aceite vegetal 1L", "pan integral"]; when omitted (empty run input{}) the default starter basket["arroz extra 5kg", "leche evaporada chica"]is priced instead. An explicit empty array is still rejected (minItems: 1). Including a size/format hint (5kg,1L,30un) materially improves cross-chain matching. An item whose text names a category in a chain's own tree (e.g.huevos,pollo entero) is automatically scoped to it — see Category auto-scoping below.categoryHint(optional): a free-text category scoped against each chain's own category tree (e.g.abarrotes); chains where it doesn't uniquely resolve fall back to full-text search withcategoryResolved: false. An explicit hint overrides the per-item auto-scoping for the whole basket.requestDelaySecs(default 1.0): politeness pause between basket items; within an item the four chains are queried concurrently but staggered — each chain starts 1/4 of this delay after the previous one. Regardless of the delay, requests to any single host are serialized (at most one in flight at a time) so tree/search/retry hits never stack on one VTEX host.
Input guidance (read before first run)
Chain full-text search bleeds across categories: a generic aceite can
surface "Trozos de Atún en Aceite" (tuna-in-oil) and huevos can surface
quail eggs as cheaper substitutions. Use specific item names with size
hints — aceite vegetal 1L, huevos de gallina, arroz extra 5kg — or
pass categoryHint (e.g. abarrotes, huevos). Substitutions still slip
through are flagged matched: false in the dataset so you can tell a real
low price from a lookalike product.
Exactly these inputs are implemented — tests/test_actor_schemas.py guards
the schema against normalize_input() drift.
Output
Dataset — one record per (item × chain) offer (plus one unmatched
marker when no chain has a plausible offer). Fields: item, itemIndex,
store, productId/productKey, name, brand, category, price,
listPrice, currency (always PEN), unitMultiplier,
measurementUnit, availableQuantity, grams, pieceCount,
pricePerKg, pricePerPiece, inStock, url, imageUrl,
categoryResolved, matched / matchBasis (cross-chain linkage),
isCheapest / compareBasis (the cheapest decision), collectedAt.
availableQuantity is passed through but carries the VTEX platform default
(99999) for most items — treat it as "effectively unlimited", never as real
stock.
Key-value store OUTPUT — the basket summary: items (cheapest chain
per item with name/price/url), totalsByChain (total + items + coverage per
chain), wholeBasket (cheapest vs priciest chain and the savings on shared
items), mixAndMatch (cheapest-per-item total), a human summaryLine, and
warnings (any chain that failed for an item — one flaky chain never kills
a run).
Category auto-scoping
VTEX full-text search bleeds across categories: a generic aceite surfaces
tuna-in-oil and huevos surfaces quail eggs as cheaper substitutions. Each
chain's category tree (/api/catalog_system/pub/category/tree/3, fetched
once per run and cached) is therefore propagated into matching: when no
categoryHint is given, each basket item's own text is resolved per chain
against that chain's tree, and a unique name match scopes that chain's
search (map=ft,c,c…, categoryResolved: true). Items that resolve nowhere
or ambiguously (e.g. café soluble) keep the plain full-text search with
categoryResolved: false; surviving substitutions are still flagged
matched: false. An explicit categoryHint wins over auto-scoping.
Cross-chain matching (best-effort, documented bias)
- Metro, Wong and Vivanda are Cencosud storefronts sharing one
product-ID space: two chains' offers with the same
productIdare the same product (matchBasis: productId). - PlazaVea has a separate id space: it is matched against the Cencosud
reference by normalized
grams(or equally unknown masses), a compatiblepieceCount, and ≥ 50% shared name tokens, accent/case-insensitively (matchBasis: name-unit) — seematching.py. - Each chain's offer is its own best-ranked candidate (token overlap + weight consistency + price tiebreak), not a forced cross-chain id group — a chain's genuine low price (e.g. Metro's own-brand 5kg arroz at S/17.40 vs the shared Costeño at S/20.90) must never be dropped.
- The cheapest decision compares uniformly: per-kg when every offer
has a per-kg price, per-piece when every offer has a pack count, nominal
otherwise.
matched: falseoffers (e.g. quail eggs for a generic "huevos" query) still compete — the flags tell you when the winner is a substitution, not the same product.
Cost & reliability
~4 API calls per basket item (one per chain, staggered ~1/4 of
requestDelaySecs apart) plus one
category-tree fetch per chain per run; a 10-item basket runs in ~15–20 s.
Measured cloud-run cost: see docs/FINDINGS.md
(≈ USD 0.001–0.002/run on the free tier — the four VTEX endpoints are
keyless and ~100% up).
Local development
uv venv .venv && uv pip install --python .venv/bin/python -r requirements.txt jsonschema.venv/bin/python -m unittest discover -s tests -v # 47 tests.venv/bin/python scripts/live_basket_run.py # live run -> storage/live/
tests/test_metro_parity.py imports the live sibling
metro-peru-scraper module from the repo checkout and asserts the vendored
VTEX client (src/peru_grocery_aggregator/vtex_client.py) is behaviorally
identical — the build context of an Apify Actor is folder-scoped, so the
client is vendored verbatim rather than imported, and this test keeps the
two copies from drifting.
What this Actor is not
Not a crawler (comparison product, one first page per item), no banner scraping, no browsers, no other stores, no calls to sibling Actors via Apify's API (the aggregator hits the VTEX endpoints directly — no stacked per-run billing).