Review Dataset Quality Gate avatar

Review Dataset Quality Gate

Pricing

from $20.00 / 1,000 review batch auditeds

Go to Apify Store
Review Dataset Quality Gate

Review Dataset Quality Gate

Audit exported product, place, or hospitality review datasets for duplicates, missing fields, rating-range errors, timestamp errors, and suspicious repeated text.

Pricing

from $20.00 / 1,000 review batch auditeds

Rating

0.0

(0)

Developer

Cedric Günther

Cedric Günther

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Audits a supplied review export for missing fields, duplicate IDs, rating-range errors, timestamp problems, and repeated normalized text using deterministic rules. It is built for review-data pipeline owners, marketplace and hospitality analysts, data-quality and ingestion teams. The main result is structured, deterministic evidence that can be consumed from the default dataset or an Apify automation.

When to use this Actor

  • Reject malformed review batches before warehouse or model ingestion.
  • Find duplicate review identifiers and repeated-text groups in exports.
  • Validate ratings, required fields, and publication timestamps against an explicit observation time.

How it works

  • Validate the bounded inline review array and the declared rating/field controls.
  • Apply exact field, ID, numeric-range, UTC timestamp, and normalized-text rules without heuristic fraud scoring.
  • Sort and cap displayed findings deterministically while retaining truthful totals in the complete report.

The Actor validates only the declared product contract. It does not infer facts outside the supplied data or claim outcomes that the source material cannot prove.

Quick start

  1. Open the Actor's Input tab or create a Task from one of the public examples.
  2. Paste or adapt this bounded example.
  3. Click Start and inspect the default dataset plus the output links shown on the run page.
{
"batchId": "hotel-reviews-2026-09",
"reviews": [
{
"id": "R-1",
"itemId": "hotel-1",
"author": "Ava",
"rating": 5,
"text": "Clean room and friendly staff",
"publishedAt": "2026-09-20T10:00:00Z"
},
{
"id": "R-2",
"itemId": "hotel-1",
"author": "Ava",
"rating": 5,
"text": "Clean room and friendly staff",
"publishedAt": "2026-09-20T10:00:00Z"
}
],
"ratingMin": 1,
"ratingMax": 5,
"observedAt": "2026-09-21T00:00:00Z"
}

Expected result: a REPEATED_CONTENT finding plus a quality-summary record for the two-review batch.

Input

The quick-start example is intentionally small. These are the material controls; the Input tab remains authoritative for the complete current schema.

FieldPurpose and formatDefaultImportant bounds or interaction
batchIdStable caller-defined identifier copied to all quality evidence for this review export.No implicit defaultminimum length 1; maximum length 120
reviewsBounded inline review records evaluated for duplicates, required fields, rating range, and timestamp validity.No implicit defaultminimum items 1; maximum items 20000
ratingMinInclusive lower bound for valid numeric ratings and must be lower than ratingMax.No implicit defaultSee the Input tab for the current schema bound.
ratingMaxInclusive upper bound for valid numeric ratings and must be higher than ratingMin.No implicit defaultSee the Input tab for the current schema bound.
observedAtOptional explicit timestamp used to identify future-dated reviews reproducibly.No implicit defaultSee the Input tab for the current schema bound.
requiredFieldsReview field names that must contain usable values; this augments rather than rewrites the fixed quality checks.No implicit defaultSee the Input tab for the current schema bound.
maxRecordsOptional lower record ceiling for the submitted batch without increasing the product maximum.No implicit defaultminimum 1; maximum 20000
maxFindingsCaps detailed finding records while summary counts retain the complete observed defect totals.No implicit defaultminimum 1; maximum 1000

Unknown top-level fields and invalid field combinations fail validation rather than being guessed.

Output

The default dataset contains typed records. The run's Output tab links the dataset and any key-value-store reports declared by the current output schema.

FieldMeaning
recordTypeDiscriminates quality-finding and quality-summary rows.
batchIdCurrent dataset field.
codeCurrent dataset field.
severityCurrent dataset field.
reviewIdsCurrent dataset field.
fieldCurrent dataset field.
messageCurrent dataset field.
fingerprintStable identity for a deterministic finding across retries.
statusBatch-level PASSED or FAILED result based on the declared checks.
reviewCountCurrent dataset field.
findingCountCurrent dataset field.
engineVersionCurrent dataset field.

Representative current-schema dataset item:

{
"recordType": "quality-finding",
"batchId": "hotel-reviews-2026-09",
"code": "REPEATED_CONTENT",
"severity": "WARNING",
"reviewIds": [
"R-1",
"R-2"
],
"fingerprint": "example-stable-fingerprint",
"engineVersion": "1.0.0"
}

A valid batch with no findings succeeds with a quality-summary record, status PASSED, and findingCount 0.

Pricing and billing

This Actor uses PAY_PER_EVENT; platform usage is included in event prices. A charge is eligible only after the billable unit described below is durably completed. Validation failures and the non-billable failure classes in the product contract do not emit the custom completion event. The current live policy uses the same event price at every Store tier; no tier discount is active. The Apify Pricing tab is authoritative if a later approved pricing change takes effect.

EventWhat triggers itFREEBRONZESILVERGOLDPLATINUMDIAMOND
review-batch-auditedOne bounded readable review batch converted into a durable quality report.$0.02000000$0.02000000$0.02000000$0.02000000$0.02000000$0.02000000
apify-actor-startPlatform-managed Actor start event.$0.00005000$0.00005000$0.00005000$0.00005000$0.00005000$0.00005000

The Actor does not have Task-specific prices: public Tasks use this same live Actor pricing. Third-party costs are not implied; see the data and security section for external services actually contacted.

Limits and bounds

  • The default and hard record/finding bounds come from the current Input schema; callers may lower but not expand them.
  • Reviews must be supplied inline as bounded JSON objects; the Actor does not crawl review sites.
  • Future timestamp checks run only when observedAt is a valid explicit UTC timestamp.

These are product-facing limits, not targets. Use smaller inputs when you need faster feedback or simpler evidence.

Failure and edge-case behavior

  • Invalid configuration or over-limit input fails closed before a review-batch-audited event.
  • Data-quality problems are successful audited outputs with finding rows, not Actor crashes.
  • Repeated text is reported as a deterministic signal and is never relabeled as fraud or fake-review proof.

Operationally:

  • Keep batchId stable for the same delivery when comparing repeated processing evidence.
  • Use maxFindings to bound dataset volume; findingCount remains the truthful total.

Use with Tasks and automation

Public Tasks provide reusable saved inputs for distinct supported workflows. Start with the closest Example Task, review its visible fields and scope caveat, then save your own Task for schedules or repeated runs. Do not treat an Example Task as evidence that unsupported behavior exists.

Integration and API usage

Every saved Task can be started manually, through the Apify API, or from an Apify schedule. Run-completion webhooks can notify a downstream system after output is durable. Actor-to-Actor calls should consume the typed dataset/output links instead of scraping the Store page.

  • Use a scheduled Task as a pre-ingestion gate for recurring exports.
  • Route finding code, severity, and fingerprint into CI, ETL quarantine, or alerting logic.

No third-party integration is claimed unless it is named above and supported by the current product contract.

Data, privacy, and security

  • Submitted reviews and findings remain in the run's Apify storages according to account retention settings.
  • No external service is contacted and no credentials are required; remove unnecessary personal data before submission.

Set Apify storage retention and access according to the sensitivity of your inputs and outputs. This documentation does not create legal, privacy, compliance, or security certification.

Support and known limitations

  • The Actor does not determine authenticity, sentiment, policy compliance, or reviewer identity.
  • Normalization can identify equal text after specified transformations but cannot infer copied meaning across different wording.

For support, use the Actor Issues page. Include the run ID, a minimal reproducible input with sensitive values removed, the failing record or error code, and what you expected. Do not post credentials, private source files, customer data, or full confidential payloads.