Marketplace Image QA avatar

Marketplace Image QA

Pricing

from $15.00 / 1,000 image inspecteds

Go to Apify Store
Marketplace Image QA

Marketplace Image QA

Check product listing images for dimensions, missing variants, background issues, duplicates, and visual readiness.

Pricing

from $15.00 / 1,000 image inspecteds

Rating

0.0

(0)

Developer

junipr

junipr

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Inspect marketplace listing images and rendered page evidence for dimensions, aspect ratios, alt text, duplicates, variants, background metadata, and visual readiness.

Use this actor to turn structured image metadata, listing HTML, or an explicitly rendered public listing page into a normalized Apify dataset. Each paid row represents one listing image audit; blocked and no-result diagnostics remain free so invalid evidence does not look like successful extraction.

What It Returns

Every dataset row includes stable source and record identifiers, a checked timestamp, current and previous hashes, a change classification, issue codes, a concise recommendation, the charged event name, and strict actor-specific data. The data object contains only fields relevant to listing image audits.

Change classifications are observed when no prior snapshot is supplied, added for a new key, updated when normalized metadata changed, and unchanged when the normalized record matches. Diagnostic rows use diagnostic, contain no paid event name, and explain how to supply usable evidence.

Input Sources

  • Inline currentText, text, or html for deterministic runs.
  • items for structured current records and previousItems for structured history.
  • previousText for snapshot comparison.
  • targets to process several bounded sources in one run.
  • urls or sourceUrl only when renderPages is explicitly enabled for a public listing page.

Rendered page retrieval accepts HTTP and HTTPS only and inspects browser-visible image dimensions and alt text. It rejects credentials, redirects, localhost, private and reserved addresses, and DNS results that resolve privately. Navigation has a timeout, and rendered HTML and screenshots have explicit byte limits. Defaults use bounded records from the stated public source and make no network request.

Quick Start

{
"sourceId": "open-food-facts-6111035000058",
"sourceUrl": "https://world.openfoodfacts.org/product/6111035000058",
"text": "",
"currentText": "",
"previousText": "",
"items": [
{
"imageId": "open-food-facts-6111035000058-front",
"productId": "6111035000058",
"variant": "1.5 L",
"imageUrl": "https://images.openfoodfacts.org/images/products/611/103/500/0058/front_en.134.full.jpg",
"altText": "Sidi Ali natural mineral water, 1.5 L",
"width": 1536,
"height": 2048,
"rendered": false
}
],
"maxItems": 1,
"maxRecordsPerTarget": 5,
"maxTextBytes": 250000,
"fetchTimeoutMs": 15000,
"includeDigest": false,
"includeReport": true,
"dryRun": false,
"debug": false,
"renderPages": false,
"captureScreenshot": false
}

Set maxItems to cap source targets, maxRecordsPerTarget to cap normalized records, maxTextBytes to cap each supplied or fetched body, and fetchTimeoutMs to cap public retrieval time. dryRun validates input without charges or dataset output.

Data Quality

The parser normalizes and deduplicates listing image audits by the actor's stable identity field. It never fabricates missing titles, dates, authors, links, guests, journals, or categories. Missing values stay null or empty and produce explicit issue codes.

Possible issue codes include: missing-image-url, missing-alt-text, missing-dimensions, low-resolution, unusual-aspect-ratio, large-file, duplicate-image, missing-variant, nonstandard-background, source-unavailable. Review warning rows before using them in downstream publishing, outreach, reporting, or alerting.

Reports And Digests

The default key-value store receives a results JSON file, summary JSON, and Markdown report after the corresponding report charge succeeds. Recurring monitor actors also produce JSON and Markdown change digests. Report artifacts are skipped when the run budget cannot cover the report event.

Pay Per Event

EventPricePurpose
actor-start$0.05000Marketplace Image QA Actor Start
image-inspected$0.01500Marketplace Image QA Image Inspected
issue-detected$0.01300Marketplace Image QA Issue Detected
report-generated$0.10000Marketplace Image QA Report Generated

Platform usage is included in these event prices. The actor charges before paid dataset rows or report artifacts are written. If Apify rejects a charge because the run limit is reached, the actor records billing status and stops before the affected paid output.

Operational Guidance

Use stable sourceId values across scheduled runs, retain the prior normalized source snapshot, and compare changes by recordId and changeType. Start with one source and inspect the dataset before increasing limits. Keep rendered retrieval disabled when inline HTML or structured image evidence is available.

How Parsing Works

This actor accepts structured image metadata, listing HTML, or an explicitly rendered public listing page. Structured records take priority when they are supplied. Otherwise, the parser reads the source format directly, normalizes whitespace and identifiers, resolves supported links, removes duplicate identities, and stops at maxRecordsPerTarget.

The actor-specific data fields are: imageId, productId, variant, imageUrl, altText, width, height, aspectRatio, fileSizeBytes, backgroundColor, contentHash, duplicateOf, rendered, screenshotKey. These fields are validated by the strict dataset schema. Unknown top-level output fields and unknown fields inside data are not part of the contract.

Record identity is based on the strongest stable public identifier available for each listing image audit. Links and official IDs are preferred; a deterministic content identity is used only when the source omits a stronger identifier. Reordering a feed or inserting a new record does not turn unchanged historical records into updates.

Comparison Workflow

  1. Save the prior source body or structured records after a successful run.
  2. Supply that evidence as previousText or previousItems on the next run.
  3. Supply the current evidence through currentText, items, or an explicitly enabled rendered public page.
  4. Read added and updated rows as detected events. unchanged rows preserve the current normalized inventory.
  5. Store the current source evidence for the following run.

Comparisons use normalized record hashes, not timestamps generated by the actor. This keeps reruns deterministic and prevents a new checkedAt value from creating false changes.

Scheduling And Delivery

For scheduled monitoring, use one stable target per public source and conservative caps. Send alerts from rows whose changeType is added or updated; keep all rows in the dataset when a complete current inventory is required. The JSON summary is suitable for automation, while the Markdown report and optional digest are suitable for review or delivery.

Do not treat an empty paid dataset as a successful zero-change result without checking diagnostics. A blocked diagnostic indicates missing evidence, disabled rendered retrieval, an invalid public URL, a private address, a timeout, an oversized response, or an unsupported source shape.

Limits And Failure Handling

  • maxItems limits targets before retrieval or analysis.
  • maxRecordsPerTarget limits normalized records per target.
  • maxTextBytes limits inline and fetched source bodies.
  • fetchTimeoutMs limits each explicitly enabled public request.
  • Redirects and credentialed URLs are rejected.
  • Charge-limit failures stop before the affected paid rows or report artifacts.
  • Parser failures produce explicit diagnostics rather than invented records.

Use debug only when investigating source parsing or billing decisions. Avoid logging source bodies that may contain information you do not intend to retain.

Only process public, authorized, or owned sources. Do not use this actor to bypass access controls, collect private data, or make unsupported legal or compliance claims.