# Review Dataset Quality Gate (`ceddl/review-dataset-quality-gate`) Actor

Audit exported product, place, or hospitality review datasets for duplicates, missing fields, rating-range errors, timestamp errors, and suspicious repeated text.

- **URL**: https://apify.com/ceddl/review-dataset-quality-gate.md
- **Developed by:** [Cedric Günther](https://apify.com/ceddl) (community)
- **Categories:** E-commerce, Automation, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $20.00 / 1,000 review batch auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Review Dataset Quality Gate

Audits a supplied review export for missing fields, duplicate IDs, rating-range errors, timestamp problems, and repeated normalized text using deterministic rules. It is built for review-data pipeline owners, marketplace and hospitality analysts, data-quality and ingestion teams. The main result is structured, deterministic evidence that can be consumed from the default dataset or an Apify automation.

### When to use this Actor

- Reject malformed review batches before warehouse or model ingestion.
- Find duplicate review identifiers and repeated-text groups in exports.
- Validate ratings, required fields, and publication timestamps against an explicit observation time.

### How it works

- Validate the bounded inline review array and the declared rating/field controls.
- Apply exact field, ID, numeric-range, UTC timestamp, and normalized-text rules without heuristic fraud scoring.
- Sort and cap displayed findings deterministically while retaining truthful totals in the complete report.

The Actor validates only the declared product contract. It does not infer facts outside the supplied data or claim outcomes that the source material cannot prove.

### Quick start

1. Open the Actor's **Input** tab or create a Task from one of the public examples.
2. Paste or adapt this bounded example.
3. Click **Start** and inspect the default dataset plus the output links shown on the run page.

```json
{
  "batchId": "hotel-reviews-2026-09",
  "reviews": [
    {
      "id": "R-1",
      "itemId": "hotel-1",
      "author": "Ava",
      "rating": 5,
      "text": "Clean room and friendly staff",
      "publishedAt": "2026-09-20T10:00:00Z"
    },
    {
      "id": "R-2",
      "itemId": "hotel-1",
      "author": "Ava",
      "rating": 5,
      "text": "Clean room and friendly staff",
      "publishedAt": "2026-09-20T10:00:00Z"
    }
  ],
  "ratingMin": 1,
  "ratingMax": 5,
  "observedAt": "2026-09-21T00:00:00Z"
}
```

Expected result: a REPEATED\_CONTENT finding plus a quality-summary record for the two-review batch.

### Input

The quick-start example is intentionally small. These are the material controls; the Input tab remains authoritative for the complete current schema.

| Field | Purpose and format | Default | Important bounds or interaction |
|---|---|---|---|
| `batchId` | Stable caller-defined identifier copied to all quality evidence for this review export. | No implicit default | minimum length 1; maximum length 120 |
| `reviews` | Bounded inline review records evaluated for duplicates, required fields, rating range, and timestamp validity. | No implicit default | minimum items 1; maximum items 20000 |
| `ratingMin` | Inclusive lower bound for valid numeric ratings and must be lower than ratingMax. | No implicit default | See the Input tab for the current schema bound. |
| `ratingMax` | Inclusive upper bound for valid numeric ratings and must be higher than ratingMin. | No implicit default | See the Input tab for the current schema bound. |
| `observedAt` | Optional explicit timestamp used to identify future-dated reviews reproducibly. | No implicit default | See the Input tab for the current schema bound. |
| `requiredFields` | Review field names that must contain usable values; this augments rather than rewrites the fixed quality checks. | No implicit default | See the Input tab for the current schema bound. |
| `maxRecords` | Optional lower record ceiling for the submitted batch without increasing the product maximum. | No implicit default | minimum 1; maximum 20000 |
| `maxFindings` | Caps detailed finding records while summary counts retain the complete observed defect totals. | No implicit default | minimum 1; maximum 1000 |

Unknown top-level fields and invalid field combinations fail validation rather than being guessed.

### Output

The default dataset contains typed records. The run's Output tab links the dataset and any key-value-store reports declared by the current output schema.

| Field | Meaning |
|---|---|
| `recordType` | Discriminates quality-finding and quality-summary rows. |
| `batchId` | Current dataset field. |
| `code` | Current dataset field. |
| `severity` | Current dataset field. |
| `reviewIds` | Current dataset field. |
| `field` | Current dataset field. |
| `message` | Current dataset field. |
| `fingerprint` | Stable identity for a deterministic finding across retries. |
| `status` | Batch-level PASSED or FAILED result based on the declared checks. |
| `reviewCount` | Current dataset field. |
| `findingCount` | Current dataset field. |
| `engineVersion` | Current dataset field. |

Representative current-schema dataset item:

```json
{
  "recordType": "quality-finding",
  "batchId": "hotel-reviews-2026-09",
  "code": "REPEATED_CONTENT",
  "severity": "WARNING",
  "reviewIds": [
    "R-1",
    "R-2"
  ],
  "fingerprint": "example-stable-fingerprint",
  "engineVersion": "1.0.0"
}
```

A valid batch with no findings succeeds with a quality-summary record, status PASSED, and findingCount 0.

### Pricing and billing

This Actor uses `PAY_PER_EVENT`; platform usage is included in event prices. A charge is eligible only after the billable unit described below is durably completed. Validation failures and the non-billable failure classes in the product contract do not emit the custom completion event. The current live policy uses the same event price at every Store tier; no tier discount is active. The Apify **Pricing** tab is authoritative if a later approved pricing change takes effect.

| Event | What triggers it | FREE | BRONZE | SILVER | GOLD | PLATINUM | DIAMOND |
|---|---|---:|---:|---:|---:|---:|---:|
| `review-batch-audited` | One bounded readable review batch converted into a durable quality report. | $0.02000000 | $0.02000000 | $0.02000000 | $0.02000000 | $0.02000000 | $0.02000000 |
| `apify-actor-start` | Platform-managed Actor start event. | $0.00005000 | $0.00005000 | $0.00005000 | $0.00005000 | $0.00005000 | $0.00005000 |

The Actor does not have Task-specific prices: public Tasks use this same live Actor pricing. Third-party costs are not implied; see the data and security section for external services actually contacted.

### Limits and bounds

- The default and hard record/finding bounds come from the current Input schema; callers may lower but not expand them.
- Reviews must be supplied inline as bounded JSON objects; the Actor does not crawl review sites.
- Future timestamp checks run only when observedAt is a valid explicit UTC timestamp.

These are product-facing limits, not targets. Use smaller inputs when you need faster feedback or simpler evidence.

### Failure and edge-case behavior

- Invalid configuration or over-limit input fails closed before a review-batch-audited event.
- Data-quality problems are successful audited outputs with finding rows, not Actor crashes.
- Repeated text is reported as a deterministic signal and is never relabeled as fraud or fake-review proof.

Operationally:

- Keep batchId stable for the same delivery when comparing repeated processing evidence.
- Use maxFindings to bound dataset volume; findingCount remains the truthful total.

### Use with Tasks and automation

Public Tasks provide reusable saved inputs for distinct supported workflows. Start with the closest Example Task, review its visible fields and scope caveat, then save your own Task for schedules or repeated runs. Do not treat an Example Task as evidence that unsupported behavior exists.

- [Audit a review export for quality](https://apify.com/ceddl/review-dataset-quality-gate/examples/audit-review-export): Check a safe inline review batch for duplicates, missing fields, rating ranges, and timestamp validity.
- [Find duplicate reviews](https://apify.com/ceddl/review-dataset-quality-gate/examples/find-duplicate-reviews): Find exact duplicate review rows in an exported review dataset.
- [Validate review ratings and timestamps](https://apify.com/ceddl/review-dataset-quality-gate/examples/validate-review-ratings-and-timestamps): Validate review ratings, timestamps, and required fields before analysis or import.

### Integration and API usage

Every saved Task can be started manually, through the Apify API, or from an Apify schedule. Run-completion webhooks can notify a downstream system after output is durable. Actor-to-Actor calls should consume the typed dataset/output links instead of scraping the Store page.

- Use a scheduled Task as a pre-ingestion gate for recurring exports.
- Route finding code, severity, and fingerprint into CI, ETL quarantine, or alerting logic.

No third-party integration is claimed unless it is named above and supported by the current product contract.

### Data, privacy, and security

- Submitted reviews and findings remain in the run's Apify storages according to account retention settings.
- No external service is contacted and no credentials are required; remove unnecessary personal data before submission.

Set Apify storage retention and access according to the sensitivity of your inputs and outputs. This documentation does not create legal, privacy, compliance, or security certification.

### Support and known limitations

- The Actor does not determine authenticity, sentiment, policy compliance, or reviewer identity.
- Normalization can identify equal text after specified transformations but cannot infer copied meaning across different wording.

For support, use the [Actor Issues page](https://apify.com/ceddl/review-dataset-quality-gate/issues). Include the run ID, a minimal reproducible input with sensitive values removed, the failing record or error code, and what you expected. Do not post credentials, private source files, customer data, or full confidential payloads.

# Actor input Schema

## `batchId` (type: `string`):

Input field batchId.

## `reviews` (type: `array`):

Input field reviews.

## `ratingMin` (type: `number`):

Input field ratingMin.

## `ratingMax` (type: `number`):

Input field ratingMax.

## `observedAt` (type: `string`):

Optional UTC timestamp used only for future-date checks.

## `requiredFields` (type: `array`):

Input field requiredFields.

## `maxRecords` (type: `integer`):

Input field maxRecords.

## `maxFindings` (type: `integer`):

Input field maxFindings.

## Actor input object example

```json
{
  "batchId": "sample-reviews",
  "reviews": [
    {
      "id": "R-1",
      "itemId": "hotel-1",
      "author": "Ava",
      "rating": 5,
      "text": "Clean room and friendly staff",
      "publishedAt": "2026-09-20T10:00:00Z"
    },
    {
      "id": "R-2",
      "itemId": "hotel-1",
      "author": "Ava",
      "rating": 5,
      "text": "Clean room and friendly staff",
      "publishedAt": "2026-09-20T10:00:00Z"
    }
  ],
  "ratingMin": 1,
  "ratingMax": 5
}
```

# Actor output Schema

## `evidence` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `reports` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "batchId": "sample-reviews",
    "reviews": [
        {
            "id": "R-1",
            "itemId": "hotel-1",
            "author": "Ava",
            "rating": 5,
            "text": "Clean room and friendly staff",
            "publishedAt": "2026-09-20T10:00:00Z"
        },
        {
            "id": "R-2",
            "itemId": "hotel-1",
            "author": "Ava",
            "rating": 5,
            "text": "Clean room and friendly staff",
            "publishedAt": "2026-09-20T10:00:00Z"
        }
    ],
    "ratingMin": 1,
    "ratingMax": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("ceddl/review-dataset-quality-gate").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "batchId": "sample-reviews",
    "reviews": [
        {
            "id": "R-1",
            "itemId": "hotel-1",
            "author": "Ava",
            "rating": 5,
            "text": "Clean room and friendly staff",
            "publishedAt": "2026-09-20T10:00:00Z",
        },
        {
            "id": "R-2",
            "itemId": "hotel-1",
            "author": "Ava",
            "rating": 5,
            "text": "Clean room and friendly staff",
            "publishedAt": "2026-09-20T10:00:00Z",
        },
    ],
    "ratingMin": 1,
    "ratingMax": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("ceddl/review-dataset-quality-gate").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "batchId": "sample-reviews",
  "reviews": [
    {
      "id": "R-1",
      "itemId": "hotel-1",
      "author": "Ava",
      "rating": 5,
      "text": "Clean room and friendly staff",
      "publishedAt": "2026-09-20T10:00:00Z"
    },
    {
      "id": "R-2",
      "itemId": "hotel-1",
      "author": "Ava",
      "rating": 5,
      "text": "Clean room and friendly staff",
      "publishedAt": "2026-09-20T10:00:00Z"
    }
  ],
  "ratingMin": 1,
  "ratingMax": 5
}' |
apify call ceddl/review-dataset-quality-gate --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ceddl/review-dataset-quality-gate"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZVSgMdRnmGFg2NPm1/builds/r9UrufuCX5AgzzUyb/openapi.json
