Competitor Product Comparison Builder avatar

Competitor Product Comparison Builder

Pricing

from $2.00 / 1,000 products

Go to Apify Store
Competitor Product Comparison Builder

Competitor Product Comparison Builder

Normalize supplied competitor product URLs into a comparison dataset.

Pricing

from $2.00 / 1,000 products

Rating

0.0

(0)

Developer

Danial Maqbool

Danial Maqbool

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

15 hours ago

Last modified

Categories

Share

Normalize supplied competitor product URLs into a comparison dataset.

Example.com inputs are placeholders, not verified compatible content. Replace them with your permitted source. Run npm run smoke for an account-free deterministic example. JSON Feed, RSS and table examples use the separately recorded small live-test sources; they do not imply compatibility with every site.

What it does

Produces product records for pricing analysts. Values are extracted deterministically from public pages; unavailable fields remain null.

Common use cases

Normalize supplied competitor product URLs into a comparison dataset. Use the resulting dataset in scheduled tasks, API pipelines, spreadsheets, or agent workflows.

Input

FieldTypeDescription
productUrlsarrayPublic HTTP(S) URLs. No credentials or local-network targets.
maxPagesintegerMaximum pages scheduled in this bounded run.
maxProductsintegerMaximum visible product records returned.
renderJavaScriptbooleanOpt into Chromium for JavaScript content. More expensive; no access-control bypass.
respectRobotsTxtbooleanHonor robots rules and supported crawl delays. Unavailable rules fail closed.
groupingKeystringYour comparison group label; this does not assert product equivalence.
priceSelectorstringPrice Selector. See README for extraction behavior.
concurrencyintegerMaximum concurrent page handlers. Per-origin delays still apply.
requestDelayMsintegerMinimum spacing between request starts to the same origin.
timeoutSecsintegerMaximum individual request duration in seconds.
retriesintegerRetries for page handlers. Access blocks are not retried.
proxyConfigurationobjectOptional Apify or public custom proxy. Direct connections are the default.

Output

FieldType
sourceDomainstring or null
productNamestring or null
brandstring or null
skustring or null
pricenumber or null
currencystring or null
availabilitystring or null
ratingnumber or null
reviewCountnumber or null
shippingTextstring or null
canonicalUrlstring or null
imageUrlstring or null
groupingKeystring or null
extractedAtstring or null

Results are in the default Dataset. RUN_SUMMARY, FAILURES and DIAGNOSTICS are available in the default key-value store. Diagnostics are not billed entity events. JSON, CSV and spreadsheet exports use Apify Dataset.

Example

Replace example.com with a public source relevant to this Actor. Plain example.com has no products, jobs or opportunities; an empty feed there is expected. Run the included local smoke test for a deterministic working example.

{
"productUrls": [
"https://example.com/"
],
"maxPages": 1,
"renderJavaScript": false,
"maxProducts": 10
}

Example output

The following records come from synthetic local fixtures, not a live website or claimed customer data.

[
{
"sourceDomain": "127.0.0.1",
"productName": "Fixture Camera",
"brand": "Example Optics",
"sku": "P1",
"price": 99.5,
"currency": "USD",
"availability": "InStock",
"rating": 4.7,
"reviewCount": 12,
"shippingText": "Free shipping on orders over USD 50.",
"canonicalUrl": "https://fixture.example/product",
"imageUrl": "https://fixture.example/image.png",
"groupingKey": null,
"extractedAt": "2026-09-29T23:32:58.502Z"
}
]

Pricing model

One product event per visible result record. Event prices are configured in Apify Console, never in extraction code. The SDK enforces the run's maxTotalChargeUsd. Failed pages and diagnostic-only messages are not billed. Do not configure an additional automatic default-dataset-item event.

How it works

Crawlee BasicCrawler manages bounded requests and retries. Cheerio parses HTTP HTML. A DNS-pinned transport validates every redirect and checks robots rules. Chromium is opt-in and its page requests pass through the same transport. No external paid API is required.

Limits

  • The Actor does not claim that similarly named products are identical.
  • Currencies are preserved, not converted.
  • Public GET/HEAD content only; no login, form submission, CAPTCHA solving or browser stealth.
  • Responses are capped at 2 MiB decoded. Browser mode caps requests per page and blocks media, fonts, downloads, service workers and WebSockets.
  • Missing/blocked robots rules fail closed. Robots crawl delays above 60 seconds are not supported.
  • Start at 512 MB for static runs; use 2 GB for browser runs and measure representative targets.

Responsible use

Use only public content you are authorized to access. Respect source terms, licenses and applicable law. No private-account extraction, personal-data enrichment, credential input or access-control bypass is provided. Cookie values are never returned.

Local development

Requires Node.js 24 LTS. All commands below run from this Actor directory.

npm ci --ignore-scripts
npm run input:example
# Edit storage/key_value_stores/default/INPUT.json with your public URLs.
npm run start:dev
npm test
npm run typecheck
npm run lint
npm run validate
npm run build
npm start
npm run smoke
npx apify run
docker build -t competitor-product-comparison .

Browser testing: set PLAYWRIGHT_BROWSERS_PATH to an Actor-local .cache/browsers directory, then run npx playwright install chromium. Docker installs its own matching browser.

Deployment

Deploy only after validating the Docker build and a representative target. Deployment is not performed by installation or tests.

npx apify login
npx apify push

Then configure the exact event in Console, set pricing and spending limits, test a private run, complete PUBLICATION_CHECKLIST.md, and deliberately enable Store visibility.

API usage

Replace YOUR_USERNAME with your Apify account name. Keep the token in an environment variable rather than input JSON.

$curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~competitor-product-comparison/runs?maxTotalChargeUsd=1" -H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" --data-binary @examples/basic-input.json

Monitor the returned run ID and read its default Dataset. Store/API names and price configuration are account-side settings.

Verified platform example

The saved Console input performs a bounded HTTP run with one page and at most two result records. The source is a public scraping demonstration site; these records are sample content, not customer or current hiring data.

Use the saved default input in Apify Console, or the repository examples/basic-input.json. Replace its source with your own permitted URL and keep limits small for the first run.

The following unmodified record was returned by a successful Apify platform run on 2026-09-30:

[
{
"sourceDomain": "books.toscrape.com",
"productName": "A Light in the Attic",
"brand": null,
"sku": null,
"price": 51.77,
"currency": null,
"availability": null,
"rating": null,
"reviewCount": null,
"shippingText": null,
"canonicalUrl": "https://books.toscrape.com/index.html",
"imageUrl": "https://books.toscrape.com/media/cache/fe/72/fe72f0532301ec28892ae79a629a293c.jpg",
"groupingKey": null,
"extractedAt": "2026-09-30T12:53:10.167Z"
}
]

Introductory pricing: $2.00 per 1,000 product results, plus Apify platform usage. The standard Actor start event costs $0.00005 per GB (minimum one event). Empty output and duplicate records incur no primary result event; the start fee and any platform usage still apply. The SDK enforces the event spending limit; platform usage is billed separately.