AI Dataset Cleaner & Schema Normalizer avatar

AI Dataset Cleaner & Schema Normalizer

Pricing

from $0.99 / 1,000 cleaned records

Go to Apify Store
AI Dataset Cleaner & Schema Normalizer

AI Dataset Cleaner & Schema Normalizer

Clean, normalize, deduplicate and validate Apify datasets or JSON. Infer schemas, fix inconsistent fields, flatten nested data and prepare reliable datasets for AI agents, RAG, CRM, analytics, APIs and databases.

Pricing

from $0.99 / 1,000 cleaned records

Rating

0.0

(0)

Developer

Flavian COMBES

Flavian COMBES

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Turn messy scraper output into reliable, schema-consistent data.

AI Dataset Cleaner & Schema Normalizer cleans whitespace and nulls, normalizes field names and types, removes duplicates, flattens JSON and infers a schema — directly from an Apify dataset or JSON input.

The normalization layer between scraped data and downstream AI/automation systems. Deterministic, with no LLM. Your source dataset stays read-only; results go to a new clean dataset in this run.

The problem it solves

Scraper output mixes field names, numeric strings, null placeholders, nesting and repeated rows. This dataset cleaner applies reproducible rules to make AI agent data, RAG datasets, CRM imports and analytics easier to use. Normalization consistency is not proof of semantic correctness.

Before → after

Before:

{"Company Name":" Acme Corp ","Email":" SALES@ACME.COM ","Employees":"25"}

After Smart clean:

{"company_name":"Acme Corp","email":"SALES@acme.com","employees":25}

Email local-part case is preserved by default. Normalized dedupe compares strings without changing the first output row's values.

Try it in 30 seconds

  1. Keep Source: Inline and Preset: Smart clean.
  2. Run the provided three-row example.
  3. Open Cleaned records: two rows remain, employees become numbers and N/A becomes null.
  4. Inspect Cleaning report, Inferred schema and Diagnostics.

No private dataset is needed for the default run. For an existing Apify dataset, switch Source to Dataset and select it using the READ-only picker. Published Tasks may instead expose the plain-text source_dataset field; it accepts the same dataset identifier without changing source permissions.

What it does

  • Recursive whitespace and configurable null-placeholder cleaning.
  • Preserve, snake_case or camelCase names; rename, include and exclude rules.
  • Collision-safe nested-object flattening without row expansion.
  • Conservative scalar types and email/URL/phone normalization.
  • Exact, normalized, composite-key and bounded optional fuzzy deduplication.
  • Population JSON Schema inference, supplied-schema validation and reports.
  • Streaming source pagination and publication with structural limits.

It does not scrape, crawl, call LLMs, fetch record URLs, infer missing facts or execute user code.

Common use cases

Use this JSON cleaner for scraped data cleaning before chaining Actors, JSON normalization before APIs/databases, CRM data cleaning for leads and contacts, structured data for analytics, and AI-ready data for ChatGPT, AI agents or a RAG pipeline. It prepares records, not embeddings or a vector index.

Example input

{
"source_mode": "inline",
"preset": "smart_clean",
"inline_records": [
{"Company Name":" Acme Corp ","Email":" SALES@ACME.COM ","Employees":"25","Website":"https://acme.com/"},
{"Company Name":"Acme Corp","Email":"sales@acme.com","Employees":"25","Website":"https://acme.com"},
{"Company Name":" Beta SAS ","Email":"N/A","Employees":"12","Website":" https://beta.example.com "}
]
}

Dataset input:

{"source_mode":"dataset","source_dataset_id":"YOUR_DATASET_ID","preset":"smart_clean","max_records":10000}

Published Task equivalent:

{"source_mode":"dataset","source_dataset":"YOUR_DATASET_ID","preset":"smart_clean","max_records":10000}

Example output

The default example produces two direct records:

[
{"company_name":"Acme Corp","email":"SALES@acme.com","employees":25,"website":"https://acme.com"},
{"company_name":"Beta SAS","email":null,"employees":12,"website":"https://beta.example.com"}
]

Inputs

The first section contains Source, Source dataset, Inline records, Maximum records (default 1,000) and Cleaning preset. Named presets fully control the advanced cleaning settings so Apify Console defaults cannot silently change their behavior. Choose Custom when you want the Advanced options to be authoritative. Unknown input options fail validation.

OptionsBehavior
field_name_style, rename_mapRecursive naming and original-key renames; collision suffixes preserve values.
flatten, flatten_separator, flatten_depthFlatten objects, default _ and depth 8; retain arrays and empty objects.
trim_strings, normalize_nulls, null_placeholders, remove_null_fieldsSmart placeholders: empty, N/A, NA, null, none, -, unknown. Missing fields are not invented. Zero/false remain values.
infer_typesConservative integers/decimals/booleans and ISO dates. Leading zeros, precision-risk IDs and ambiguous dates remain strings.
normalize_email, lowercase_email_local, normalize_url, normalize_phoneConservative formatting. Full email lowercase is opt in; invalid/ambiguous values remain with diagnostics.
include_fields, exclude_fieldsFinal top-level names after flattening; exclusion wins.
dedupe_mode, dedupe_keysnone, exact, normalized, key or fuzzy; keys can be composite.
fuzzy_field, fuzzy_threshold, fuzzy_max_candidatesOne string field, default similarity .92, ≤10,000 candidates; fuzzy off by default.
infer_schema, required_threshold, provided_schema, schema_invalid_behaviorPopulation inference (presence default .95); supplied validation defaults to retain + report.
include_cleaning_metadataDefault false; append collision-protected source index and consistency score.

See .actor/input_schema.json for field descriptions and ARCHITECTURE.md for exact transformation order and conversion rules. Binary decimal floats are approximate; disable inference when exact decimal precision matters.

Output

The default dataset contains direct cleaned objects with dynamic user fields, without a cleaned_record wrapper. Use Apify JSON/CSV/Excel exports, Dataset API, downstream Actors or agent workflows. Arrays remain arrays; no separate exporter is needed.

Default key-value store keyContents
CLEANING_REPORTCounts, configuration summary, billing confirmations, field statistics, warnings and status.
INFERRED_SCHEMADraft 2020-12 JSON Schema, or an explicit unconstrained artifact if inference is disabled/incomplete.
DIAGNOSTICSReason totals and up to 100 source-index samples; no record bodies.

Schema inference

This schema normalizer observes all processed structurally valid retained cleaned rows, including duplicates before dedupe and excluding metadata. It tracks missing versus null, mixed types, arrays and nested objects. Integer and float observations converge to number. Properties are sorted.

Default required threshold .95 means presence in at least 95% of observed object occurrences. Some observed rows can therefore fail that requirement. Set 1 for population-wide presence. Inference describes processed records, not future rows or semantic truth.

An optional supplied JSON Schema validates each cleaned row. Retain + report minimizes data loss; schema_invalid_behavior: "reject" excludes invalid rows from output and billing. V1 supports a bounded draft 2020-12 subset: common types/properties/items/required/enum/const/range/length constraints. References, regex, combinators and conditional/unevaluated constraints are rejected. Formats are annotations. No second pass validates against the inferred schema.

Deduplication

Exact deduplicate-dataset mode hashes canonical cleaned JSON independently of key order. Normalized additionally casefolds/NFKC-normalizes strings and collapses whitespace for fingerprints only. Key compares selected final fields; missing/null/empty components keep rows distinct. First occurrence wins; last retention is not supported.

Fuzzy requires one selected string field. It uses two-character prefix blocks, ≤100 representatives/block, values ≤256 characters and ≤10,000 candidates. Known larger sources skip fuzzy with a warning and continue exact dedupe. It misses cross-prefix matches and can match incorrectly; review it for your use case. Removed duplicates are neither published nor charged as cleaned records.

Cleaning presets

PresetUse
Smart cleanTrim/nulls/snake_case/flatten/conservative types/contact formatting/normalized dedupe/schema.
MinimalTrim and schema inference; preserve names/scalar strings/placeholders/nesting and retain duplicates.
Strict normalizationSmart clean plus null object-field removal and lowercase email local parts.
CustomAdvanced controls are authoritative; the form starts from Smart clean defaults and you can change them.

Strict does not mean semantically correct. Full email lowercase may not suit every mailbox.

Global report

Counts distinguish seen, cleaned, output, skipped, schema-rejected and duplicate rows. Reports include collision/rename/trim/null/contact/type counts, presence/null/type observations, transformations, fuzzy comparisons, runtime and confirmed custom events. Normal reconciliation is seen = output + skipped + rejected + duplicates; failures can add unpublished/uncertain rows.

Optional record_consistency_score metadata measures consistency, not truthfulness: 100 minus up to 10 points for top-level null fraction, 30 for supplied-schema invalidity and up to 10 for warnings. Metadata is off by default and excluded from inference.

Pricing / billing contract

One custom record-cleaned event is requested after each confirmed cleaned-record publication. No events for removed duplicates, structural skips, failed writes or schema rejections. Budget checks happen before more work; the Actor stops when exhausted. The platform configures pricing; code contains no commercial price constant.

Candidate launch price: $0.99 / 1,000 successfully cleaned records, pending Cloud benchmarks — not a configured Store price. Do not enable nonzero synthetic dataset-item pricing alongside the custom event: the code refuses it. Optional platform-managed apify-actor-start pricing may be configured later; source code never manually emits it.

Publication and billing are separate operations. Network/process failures can leave a published row with an unconfirmed charge; execution stops for output/ledger review. billable_result_count counts confirmed published eligible rows; platform_charged_event_count counts confirmed custom charges. They differ in unmonetized runs or failures. No real customer billing was used locally.

API / MCP / automation usage

Pass the same JSON to the Actor API, Console tasks or available Apify MCP integrations. Chain an upstream run's defaultDatasetId into source_dataset_id; use this run's defaultDatasetId downstream. Read reports/schema through defaultKeyValueStoreId.

Example after deployment, using the SDK's API client:

import asyncio
import os
from apify_client import ApifyClientAsync
async def run():
client = ApifyClientAsync(os.environ["APIFY_TOKEN"])
result = await client.actor("YOUR_ACCOUNT/ai-dataset-cleaner-schema-normalizer").call(
run_input={"source_mode": "dataset", "source_dataset_id": "SOURCE_ID", "max_records": 10000}
)
page = await client.dataset(result.default_dataset_id).list_items(limit=100, clean=False)
return page.items
asyncio.run(run())

This repository task does not publish the Actor. Replace placeholders after creating it. Ten future task templates: docs/PUBLISHED_TASKS.md.

Limits

LimitV1 ceiling
Rows1,000 inline; 100,000 dataset; default max_records 1,000
Source page500 rows
Nesting / flatten depth10 / 8
Object fields / key length500 / 256 characters
String / array / total nodes100,000 characters / 1,000 elements / 10,000 value nodes per record
Serialized record1,000,000 UTF-8 bytes before and after cleaning
Rule lists/maps100 entries per option
Inference nodes / transformation paths20,000 each
Diagnostic samples100; aggregate counters continue
Fuzzy10,000 candidates, 256-character field, 100 representatives/block

Unsafe records are skipped, never truncated or billed as successful. Inference exhaustion continues cleaning with an explicit warning and no partial constraints. Extreme pages can need substantial memory; benchmark your payloads. Start a fresh run after interruption. V1 does not resume and refuses preserved nonempty output storage.

Privacy and security

Source READ-only permissions are declared for LIMITED_PERMISSIONS; Cloud verification remains required. No external record-value API, URL fetching, browser, proxy, telemetry, user code, eval/exec or input-driven shell commands. Logs/diagnostics omit record bodies. Reports include field names/source IDs. Retention follows Apify; see SECURITY.md.

Technical notes

Python 3.12: pip install -e '.[dev]'. Checks: pytest -q, ruff check ., ruff format --check ., python -m compileall -q my_actor scripts tests, python scripts/validate_schemas.py --official, python scripts/smoke.py, python scripts/sdk_smoke.py.

Local execution: use fresh storage and python -m my_actor. Missing input runs the deterministic example. SDK smoke uses isolated temporary storage and simulated PPE with zero synthetic pricing. Docker uses Python 3.12 slim without browser/model downloads. Cloud and Docker evidence: VALIDATION.md, BENCHMARK_PLAN.md, REVIEW.md.

FAQ

AI-powered cleaning? No. “AI” describes downstream use; the engine uses deterministic rules without an LLM.

Does it change source data? No. Source is read-only and cleaned records go to run output.

Guaranteed phone/email/URL validity? No. Formatting is conservative, ambiguities remain and no lookups occur.

Does 00123 become 123? No. Leading zeros and precision-risk integer strings remain strings.

Can it fill missing fields or verify facts? No. It measures structural consistency, not truthfulness or semantic data quality.

CSV/Excel? Use Apify exports; arrays remain arrays.

Resume interrupted runs? No. Review output/billing, then start fresh; do not retry into existing output.

Established Cloud performance? No. Local benchmarks do not establish Cloud runtime, memory or cost. Complete validation before release/pricing.