Supplier Catalog Cleanup avatar

Supplier Catalog Cleanup

Pricing

$0.25 / completed catalog

Go to Apify Store
Supplier Catalog Cleanup

Supplier Catalog Cleanup

Normalize catalog titles, match reference products, flag duplicates and review uncertain matches with evidence. CSV or JSON input; JSON, CSV and HTML output.

Pricing

$0.25 / completed catalog

Rating

0.0

(0)

Developer

Gilad Ronen

Gilad Ronen

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

15 hours ago

Last modified

Share

What does Supplier Catalog Cleanup do?

Clean supplier product catalogs, find likely duplicates, and match products against your master catalog. Paste CSV or JSON and receive normalized titles, suggested categories, field-level match evidence, and a separate review queue. Use it from Apify Console or through the API.

This rules-based beta is designed for wholesalers, ecommerce teams, and agencies preparing supplier data for human review. It uses no external AI service or API key. Every decision includes traceable evidence. Your source catalog is preserved; the Actor does not update your inventory.

How to clean a supplier catalog

  1. Open the Input tab and run the prefilled synthetic example to see the report.
  2. Replace the example with your products in Supplier catalog (JSON), or clear that field and paste Supplier catalog (CSV). Supply one format at a time.
  3. Optionally add your master catalog under Reference catalog. Clear the sample reference if you do not need matching.
  4. Map column names if needed and supply your own category keywords. Clear the sample category rules before using a different taxonomy.
  5. Start the Actor. In Output, inspect the summary and product decisions, then download the cleaned CSV, review queue, JSON, or interactive HTML report.
  6. Review flagged products before importing any suggestions into a business system.

What can you use it for?

  • Prepare inconsistent supplier exports for inventory onboarding.
  • Compare supplier products with an existing master catalog.
  • Find likely duplicate products while preserving pack, size, color, and model distinctions.
  • Apply your own category vocabulary through keyword rules.
  • Build a recurring catalog-review workflow through Apify schedules, API calls, or integrations.

Input fields and limits

Use catalog or catalogCsv. The optional master catalog uses reference or referenceCsv. CSV is pasted as text; remote URLs and uploaded spreadsheet files are not fetched. See the Input tab for all settings.

Product fieldHow it is used
titleProduct name; required for a usable row.
idYour product identifier; reference IDs must be unique.
brandExplicit brand used in matching.
gtinGTIN-8/12/13/14 as text, including leading zeros; checksum checked.
mpnManufacturer part number, scoped to the same brand.
sku, supplierSKU scoped to the same supplier.
size, packSize, colorVariant attributes; missing values are unknown.
categoryExisting category preserved as supplied.

Maximum: 2,000 supplier rows plus 2,000 reference rows, 4 MB input, 50 columns per row, 500 characters per title, and 200 characters per other recognized field. Rows must be flat. Unused columns are ignored. Reports over 8 MB or any export over 9 MB require a smaller batch.

For different headers, set catalogColumns, for example {"id":"Item Code","title":"Product Name","brand":"Manufacturer"}. referenceColumns works the same way. CSV supports comma, semicolon, and tab delimiters, quoted fields, and UTF-8 BOM.

A small input example:

{
"catalog": [{"id":"S-1","title":"Acme Water 0.5 L","brand":"Acme"}],
"reference": [{"id":"M-1","title":"Acme Water 500ml","brand":"Acme","category":"Beverages"}]
}

Category rules use taxonomy, for example [{"category":"Kitchen","keywords":["mixing bowl","saucepan"]}]. A supplied category takes precedence, followed by an accepted reference category, then keyword suggestions. Ties remain unresolved.

Output and match evidence

The default dataset contains one complete catalog report, with a summary and a rows array. The product view displays one product per table row. Use the dedicated CSV exports for a flat spreadsheet.

Output fieldMeaning
originalTitle, normalizedTitleSource text and normalized product title.
statusmatched, review, new, or invalid.
matchedReferenceIdAccepted reference suggestion, when unambiguous.
candidatesUp to three candidate matches with evidence and conflicts.
duplicates, duplicateCountUp to five duplicate examples and the full count.
categoryCategory value, source, and supporting keywords.
needsReview, reviewReasonsWhether and why a product needs review.

Download catalog-cleaned.csv, review-queue.csv, OUTPUT (full JSON), and report.html. The HTML report supports search, filtering, evidence inspection, and downloads without network requests. Download and open it locally if Console serves it as an attachment. Spreadsheet formula-like strings are escaped in CSV; JSON preserves the original values.

The prefilled 12-row synthetic example produces 3 reference matches, 1 potential duplicate pair, and 8 rows needing review. A matched product can still need review because it has a potential duplicate.

How does product matching work?

Normalization handles Unicode, spacing, selected spelling variants, and common units such as 0.5 L and 500 ml. Matching checks validated GTIN, brand plus MPN, supplier plus SKU, or identical title tokens with an explicit brand. Conflicting variant attributes, missing comparison attributes, and competing identifiers prevent automatic acceptance.

Fuzzy title overlap ranks candidates but never accepts a match on its own. Scores are ranking heuristics, not probabilities. Duplicate suggestions are direct pairs; products are never merged automatically. Category suggestions use your keyword rules.

How much does catalog cleanup cost?

The launch price is $0.25 per completed catalog, within the limits above. One catalog-processed event covers the whole report; there is no separate charge per product, candidate, category, or duplicate. The pricing panel is the authoritative source for current prices and any startup fee.

Invalid input is rejected before a catalog event is charged. A configured startup fee may still apply. A new run is a new billable catalog, even for identical input. Resuming the same run reuses its verified result and recorded charge. Platform costs are included in the event price; there is no external model bill.

Frequently asked questions

Can it handle different sizes, packs, or colors?

It compares explicit attributes and recognizes common unit, pack, and English color expressions. Provide structured attributes for best results. Product-specific variants beyond the supported fields may need manual review.

Does “new” mean a product is unique?

No. It means no plausible reference candidate was found, or no reference was supplied. It is not a global uniqueness check.

Does it use AI or translate product names?

This version uses deterministic rules. It preserves Unicode text but does not translate, interpret images, infer brands, or provide semantic matching across languages. Some spelling variations and near matches will be missed.

Is it ready for automatic inventory updates?

It is a human-review aid. Suggestions have not been benchmarked on customer catalogs. Validate a representative sample before relying on them. The report does not record approval decisions or perform inventory updates.

Where is my catalog stored?

Input and output are stored in your Apify run storage under its access and retention settings. The Actor uses only its run's storage and sends no catalog data to an external AI service. Only submit data you are authorized to process.

How do I get help or integrate it?

Use the Issues tab for bugs and feature requests; include a small anonymized example and expected result. Use the API tab for the generated request examples. Do not post confidential catalog data in public issues.