Product Catalog Quality Checker — CSV, SKU & Prices avatar

Product Catalog Quality Checker — CSV, SKU & Prices

Pricing

from $1.00 / 1,000 audited product records

Go to Apify Store
Product Catalog Quality Checker — CSV, SKU & Prices

Product Catalog Quality Checker — CSV, SKU & Prices

Audit CSV catalogs or Apify product datasets for missing names, invalid prices and duplicate SKUs. Preserve originals, get per-product issue codes and a spreadsheet-safe CSV report. Arabic numbers supported. No AI API key.

Pricing

from $1.00 / 1,000 audited product records

Rating

0.0

(0)

Developer

ABDULWAHAB NASER RASHED ALQARAWI

ABDULWAHAB NASER RASHED ALQARAWI

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

19 hours ago

Last modified

Categories

Share

Find missing product names, invalid prices and duplicate SKUs before using a catalog in another workflow. Each delivered audit preserves the original record and includes normalized fields, a PASS/WARN/FAIL status, stable issue codes and plain-English explanations. Arabic names and Arabic/Persian digits are supported.

Try it

Run the prefilled four-product example. It deliberately includes one valid product, a duplicate SKU, an empty name, a negative price and a free sample. A successful Actor run means the audit completed; individual products can still FAIL the catalog checks.

{
"csvText": "name,sku,price,currency\nCoffee,A-1,2.750,KWD\n,A-1,-1,KWD\nSample,,0,KWD\nTea,B-1,1.500,KWD\n",
"maxResults": 100
}

Supply exactly one of csvText, csvUrl, records, or datasetId. Remove the prefilled CSV example before using another source. CSV is UTF-8 (optional BOM), with an explicit comma, semicolon or tab delimiter. Quoted multiline cells are accepted. A URL must be publicly accessible HTTP(S); no login or proxy is required. Dataset access is read-only and must be granted to this Actor; finish the source run first.

For existing WooCommerce/Zid scraper outputs, use datasetId or records with their name, sku, price, currency and url fields. For other exports map exact columns:

{
"records": [{"Product": "Cup", "Item code": "001", "Retail": "1.250"}],
"fieldMap": {"name": "Product", "sku": "Item code", "price": "Retail"}
}

Automatic aliases include name/title/اسم المنتج/الاسم, sku/رمز المنتج/كود المنتج, price/Regular price/Sale price/السعر/سعر المنتج, currency/العملة, url/permalink/رابط المنتج. Alias matching ignores case and surrounding header spaces. If multiple columns match, select one explicitly in fieldMap; no price precedence is guessed.

Checks

CodeMeaning
NAME_REQUIREDA nonblank text name is required.
PRICE_REQUIRED / PRICE_INVALID / PRICE_NEGATIVEPrice must be a finite, nonnegative decimal. Zero and -0 are accepted.
SKU_MISSING / SKU_INVALIDMissing or non-text SKU is a warning. Codes must be text to preserve leading zeros.
SKU_DUPLICATEEvery member of a duplicate SKU group is flagged as an error.
CURRENCY_FORMATOptional supplied currency must have three ASCII letters; actual currency validity is not certified.
URL_FORMATOptional supplied product URL must have HTTP(S), a host and no spaces or credentials.
SPREADSHEET_FORMULAA mapped text value could be interpreted as a spreadsheet formula.

SKU comparison uses Unicode NFC normalization and surrounding whitespace removal, remains case-sensitive and does not strip leading zeros. Duplicate checks cover all supplied input records, even if the delivery cap stops the report early. A repeated SKU can be legitimate across stores or variants: scope the input to the catalog you intend to check.

Prices accept dot decimals, Arabic ٫, and comma decimals when decimalSeparator is ",". Comma mode rejects dots in text prices (for example "1.234") to avoid misreading grouping separators. JSON numbers and Arabic decimal marks work in either mode. Grouping separators, currency symbols, booleans, nested objects, scientific notation in text and prices longer than 40 characters are not accepted. JSON numeric precision is limited by JSON parsing; use strings for exact prices and SKUs. Uppercase currency and trimmed text appear only in normalized; original is retained unchanged.

This is a data-quality audit, not an import validator for a particular commerce platform. It does not infer variable-product parent prices, resolve missing variants, check price accuracy, convert currencies, fetch product/image URLs or repair records. Variable-product parent rows without prices will be flagged. No external LLM or other paid API is called.

Output and limits

  • One dataset row per audited product, including products with errors: rowNumber (one-based data record, excluding CSV header), status, issueCodes, issues, normalized, original, observedAt.
  • RUN_REPORT gives delivered PASS/WARN/FAIL counts, total input records, unreported records and auditComplete. A successful partial delivery is explicitly marked; never treat it as a complete catalog audit.
  • Download Spreadsheet-safe audit CSV (CATALOG_REPORT.csv) from Output. It includes normalized fields and issue codes for delivered rows only, and escapes spreadsheet formula prefixes. Generic dataset exports preserve original values: choose JSON to keep originals without spreadsheet interpretation.
  • 1–5,000 input records, 100 distinct columns, 200 characters per header, 10,000 characters per mapped audit field, and 5 MB both for source CSV and decoded JSON records. Unmapped fields such as long descriptions are preserved within the same 5 MB total limit. CSV additionally has a 100,000-character cell parser limit. A CSV close to 5 MB may exceed the decoded-record limit and must be split. Malformed/duplicate headers and ragged records fail before any result fee. Entirely blank CSV lines are malformed records.
  • maxResults defaults to 100 and is capped at 5,000. Set it to your input size for full delivery, and set Apify's maximum cost per run. Processing has a 240-second deadline; the default run timeout is 300 seconds. Supported memory: 512 MB–1 GB.
  • Inputs and outputs use your Apify account storage and retention. Nothing writes back to a source dataset or store. When a run resumes with the same input and output dataset, already delivered rows are skipped. Start a new run if inputs change. Concurrent edits to a source dataset are unsupported.

Price

$0.001 per delivered product audit ($1 per 1,000) plus $0.001 per run start at supported memory sizes. Products with validation errors are still completed audits and are billed. Malformed input or failed fetching generates no result fee; the start fee still applies. There is no additional automatic dataset-row fee. Your displayed Apify pricing and platform terms govern charges. Development/testing consumes the author's platform resources.

Open-source foundation

Uses Frictionless Framework 5.18.1, under its MIT license, for explicit in-memory name/price schema checks. Product rules, Arabic price handling, reports and Apify integration are implemented separately. The installed dependency retains its upstream license and notices. Not an official Frictionless, WooCommerce or Zid product.

عربي

افحص ملف المنتجات CSV أو نتائج أدواتك على Apify. الأداة تكشف الاسم الناقص والسعر غير الصالح وتكرار SKU، وتطلع سبب كل مشكلة وتحافظ على البيانات الأصلية. السعر صفر صحيح، وSKU يُحفظ كنص حتى لا تضيع الأصفار. السعر دولار لكل ألف منتج مفحوص، مع 0.001 دولار لبدء التشغيل؛ المنتج الذي تظهر فيه أخطاء يُحسب لأنه تم فحصه. حد النتائج الافتراضي 100؛ ارفعه لعدد منتجاتك إذا تريد التقرير كاملاً. لا توجد تعديلات تلقائية على متجرك ولا جدول تشغيل يُنشأ تلقائياً.

Batch delivery and limits

Audits are saved in batches of up to 100 using native per-event charging. RUN_REPORT and the CSV include only rows actually delivered, including a spending-limited batch. A 5,000-row input is a cap; parsing complexity and the 240-second processing deadline still apply. Plain negative numeric CSV cells are no longer prefixed with an apostrophe; formula-like strings remain escaped. Original JSON records are unchanged.