AI Dataset Cleaner & Schema Normalizer
Pricing
from $0.99 / 1,000 cleaned records
AI Dataset Cleaner & Schema Normalizer
Clean, normalize, deduplicate and validate Apify datasets or JSON. Infer schemas, fix inconsistent fields, flatten nested data and prepare reliable datasets for AI agents, RAG, CRM, analytics, APIs and databases.
Pricing
from $0.99 / 1,000 cleaned records
Rating
0.0
(0)
Developer
Flavian COMBES
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Turn messy scraper output into reliable, schema-consistent data.
AI Dataset Cleaner & Schema Normalizer cleans whitespace and nulls, normalizes field names and types, removes duplicates, flattens JSON and infers a schema — directly from an Apify dataset or JSON input.
The normalization layer between scraped data and downstream AI/automation systems. Deterministic, with no LLM. Your source dataset stays read-only; results go to a new clean dataset in this run.
The problem it solves
Scraper output mixes field names, numeric strings, null placeholders, nesting and repeated rows. This dataset cleaner applies reproducible rules to make AI agent data, RAG datasets, CRM imports and analytics easier to use. Normalization consistency is not proof of semantic correctness.
Before → after
Before:
{"Company Name":" Acme Corp ","Email":" SALES@ACME.COM ","Employees":"25"}
After Smart clean:
{"company_name":"Acme Corp","email":"SALES@acme.com","employees":25}
Email local-part case is preserved by default. Normalized dedupe compares strings without changing the first output row's values.
Try it in 30 seconds
- Keep Source: Inline and Preset: Smart clean.
- Run the provided three-row example.
- Open Cleaned records: two rows remain, employees become numbers and
N/Abecomes null. - Inspect Cleaning report, Inferred schema and Diagnostics.
No private dataset is needed for the default run. For an existing Apify dataset, switch Source to Dataset and select it using the READ-only picker. Published Tasks may instead expose the plain-text source_dataset field; it accepts the same dataset identifier without changing source permissions.
What it does
- Recursive whitespace and configurable null-placeholder cleaning.
- Preserve, snake_case or camelCase names; rename, include and exclude rules.
- Collision-safe nested-object flattening without row expansion.
- Conservative scalar types and email/URL/phone normalization.
- Exact, normalized, composite-key and bounded optional fuzzy deduplication.
- Population JSON Schema inference, supplied-schema validation and reports.
- Streaming source pagination and publication with structural limits.
It does not scrape, crawl, call LLMs, fetch record URLs, infer missing facts or execute user code.
Common use cases
Use this JSON cleaner for scraped data cleaning before chaining Actors, JSON normalization before APIs/databases, CRM data cleaning for leads and contacts, structured data for analytics, and AI-ready data for ChatGPT, AI agents or a RAG pipeline. It prepares records, not embeddings or a vector index.
Example input
{"source_mode": "inline","preset": "smart_clean","inline_records": [{"Company Name":" Acme Corp ","Email":" SALES@ACME.COM ","Employees":"25","Website":"https://acme.com/"},{"Company Name":"Acme Corp","Email":"sales@acme.com","Employees":"25","Website":"https://acme.com"},{"Company Name":" Beta SAS ","Email":"N/A","Employees":"12","Website":" https://beta.example.com "}]}
Dataset input:
{"source_mode":"dataset","source_dataset_id":"YOUR_DATASET_ID","preset":"smart_clean","max_records":10000}
Published Task equivalent:
{"source_mode":"dataset","source_dataset":"YOUR_DATASET_ID","preset":"smart_clean","max_records":10000}
Example output
The default example produces two direct records:
[{"company_name":"Acme Corp","email":"SALES@acme.com","employees":25,"website":"https://acme.com"},{"company_name":"Beta SAS","email":null,"employees":12,"website":"https://beta.example.com"}]
Inputs
The first section contains Source, Source dataset, Inline records, Maximum records (default 1,000) and Cleaning preset. Named presets fully control the advanced cleaning settings so Apify Console defaults cannot silently change their behavior. Choose Custom when you want the Advanced options to be authoritative. Unknown input options fail validation.
| Options | Behavior |
|---|---|
field_name_style, rename_map | Recursive naming and original-key renames; collision suffixes preserve values. |
flatten, flatten_separator, flatten_depth | Flatten objects, default _ and depth 8; retain arrays and empty objects. |
trim_strings, normalize_nulls, null_placeholders, remove_null_fields | Smart placeholders: empty, N/A, NA, null, none, -, unknown. Missing fields are not invented. Zero/false remain values. |
infer_types | Conservative integers/decimals/booleans and ISO dates. Leading zeros, precision-risk IDs and ambiguous dates remain strings. |
normalize_email, lowercase_email_local, normalize_url, normalize_phone | Conservative formatting. Full email lowercase is opt in; invalid/ambiguous values remain with diagnostics. |
include_fields, exclude_fields | Final top-level names after flattening; exclusion wins. |
dedupe_mode, dedupe_keys | none, exact, normalized, key or fuzzy; keys can be composite. |
fuzzy_field, fuzzy_threshold, fuzzy_max_candidates | One string field, default similarity .92, ≤10,000 candidates; fuzzy off by default. |
infer_schema, required_threshold, provided_schema, schema_invalid_behavior | Population inference (presence default .95); supplied validation defaults to retain + report. |
include_cleaning_metadata | Default false; append collision-protected source index and consistency score. |
See .actor/input_schema.json for field descriptions and ARCHITECTURE.md for exact transformation order and conversion rules. Binary decimal floats are approximate; disable inference when exact decimal precision matters.
Output
The default dataset contains direct cleaned objects with dynamic user fields, without a cleaned_record wrapper. Use Apify JSON/CSV/Excel exports, Dataset API, downstream Actors or agent workflows. Arrays remain arrays; no separate exporter is needed.
| Default key-value store key | Contents |
|---|---|
CLEANING_REPORT | Counts, configuration summary, billing confirmations, field statistics, warnings and status. |
INFERRED_SCHEMA | Draft 2020-12 JSON Schema, or an explicit unconstrained artifact if inference is disabled/incomplete. |
DIAGNOSTICS | Reason totals and up to 100 source-index samples; no record bodies. |
Schema inference
This schema normalizer observes all processed structurally valid retained cleaned rows, including duplicates before dedupe and excluding metadata. It tracks missing versus null, mixed types, arrays and nested objects. Integer and float observations converge to number. Properties are sorted.
Default required threshold .95 means presence in at least 95% of observed object occurrences. Some observed rows can therefore fail that requirement. Set 1 for population-wide presence. Inference describes processed records, not future rows or semantic truth.
An optional supplied JSON Schema validates each cleaned row. Retain + report minimizes data loss; schema_invalid_behavior: "reject" excludes invalid rows from output and billing. V1 supports a bounded draft 2020-12 subset: common types/properties/items/required/enum/const/range/length constraints. References, regex, combinators and conditional/unevaluated constraints are rejected. Formats are annotations. No second pass validates against the inferred schema.
Deduplication
Exact deduplicate-dataset mode hashes canonical cleaned JSON independently of key order. Normalized additionally casefolds/NFKC-normalizes strings and collapses whitespace for fingerprints only. Key compares selected final fields; missing/null/empty components keep rows distinct. First occurrence wins; last retention is not supported.
Fuzzy requires one selected string field. It uses two-character prefix blocks, ≤100 representatives/block, values ≤256 characters and ≤10,000 candidates. Known larger sources skip fuzzy with a warning and continue exact dedupe. It misses cross-prefix matches and can match incorrectly; review it for your use case. Removed duplicates are neither published nor charged as cleaned records.
Cleaning presets
| Preset | Use |
|---|---|
| Smart clean | Trim/nulls/snake_case/flatten/conservative types/contact formatting/normalized dedupe/schema. |
| Minimal | Trim and schema inference; preserve names/scalar strings/placeholders/nesting and retain duplicates. |
| Strict normalization | Smart clean plus null object-field removal and lowercase email local parts. |
| Custom | Advanced controls are authoritative; the form starts from Smart clean defaults and you can change them. |
Strict does not mean semantically correct. Full email lowercase may not suit every mailbox.
Global report
Counts distinguish seen, cleaned, output, skipped, schema-rejected and duplicate rows. Reports include collision/rename/trim/null/contact/type counts, presence/null/type observations, transformations, fuzzy comparisons, runtime and confirmed custom events. Normal reconciliation is seen = output + skipped + rejected + duplicates; failures can add unpublished/uncertain rows.
Optional record_consistency_score metadata measures consistency, not truthfulness: 100 minus up to 10 points for top-level null fraction, 30 for supplied-schema invalidity and up to 10 for warnings. Metadata is off by default and excluded from inference.
Pricing / billing contract
One custom record-cleaned event is requested after each confirmed cleaned-record publication. No events for removed duplicates, structural skips, failed writes or schema rejections. Budget checks happen before more work; the Actor stops when exhausted. The platform configures pricing; code contains no commercial price constant.
Candidate launch price: $0.99 / 1,000 successfully cleaned records, pending Cloud benchmarks — not a configured Store price. Do not enable nonzero synthetic dataset-item pricing alongside the custom event: the code refuses it. Optional platform-managed apify-actor-start pricing may be configured later; source code never manually emits it.
Publication and billing are separate operations. Network/process failures can leave a published row with an unconfirmed charge; execution stops for output/ledger review. billable_result_count counts confirmed published eligible rows; platform_charged_event_count counts confirmed custom charges. They differ in unmonetized runs or failures. No real customer billing was used locally.
API / MCP / automation usage
Pass the same JSON to the Actor API, Console tasks or available Apify MCP integrations. Chain an upstream run's defaultDatasetId into source_dataset_id; use this run's defaultDatasetId downstream. Read reports/schema through defaultKeyValueStoreId.
Example after deployment, using the SDK's API client:
import asyncioimport osfrom apify_client import ApifyClientAsyncasync def run():client = ApifyClientAsync(os.environ["APIFY_TOKEN"])result = await client.actor("YOUR_ACCOUNT/ai-dataset-cleaner-schema-normalizer").call(run_input={"source_mode": "dataset", "source_dataset_id": "SOURCE_ID", "max_records": 10000})page = await client.dataset(result.default_dataset_id).list_items(limit=100, clean=False)return page.itemsasyncio.run(run())
This repository task does not publish the Actor. Replace placeholders after creating it. Ten future task templates: docs/PUBLISHED_TASKS.md.
Limits
| Limit | V1 ceiling |
|---|---|
| Rows | 1,000 inline; 100,000 dataset; default max_records 1,000 |
| Source page | 500 rows |
| Nesting / flatten depth | 10 / 8 |
| Object fields / key length | 500 / 256 characters |
| String / array / total nodes | 100,000 characters / 1,000 elements / 10,000 value nodes per record |
| Serialized record | 1,000,000 UTF-8 bytes before and after cleaning |
| Rule lists/maps | 100 entries per option |
| Inference nodes / transformation paths | 20,000 each |
| Diagnostic samples | 100; aggregate counters continue |
| Fuzzy | 10,000 candidates, 256-character field, 100 representatives/block |
Unsafe records are skipped, never truncated or billed as successful. Inference exhaustion continues cleaning with an explicit warning and no partial constraints. Extreme pages can need substantial memory; benchmark your payloads. Start a fresh run after interruption. V1 does not resume and refuses preserved nonempty output storage.
Privacy and security
Source READ-only permissions are declared for LIMITED_PERMISSIONS; Cloud verification remains required. No external record-value API, URL fetching, browser, proxy, telemetry, user code, eval/exec or input-driven shell commands. Logs/diagnostics omit record bodies. Reports include field names/source IDs. Retention follows Apify; see SECURITY.md.
Technical notes
Python 3.12: pip install -e '.[dev]'. Checks: pytest -q, ruff check ., ruff format --check ., python -m compileall -q my_actor scripts tests, python scripts/validate_schemas.py --official, python scripts/smoke.py, python scripts/sdk_smoke.py.
Local execution: use fresh storage and python -m my_actor. Missing input runs the deterministic example. SDK smoke uses isolated temporary storage and simulated PPE with zero synthetic pricing. Docker uses Python 3.12 slim without browser/model downloads. Cloud and Docker evidence: VALIDATION.md, BENCHMARK_PLAN.md, REVIEW.md.
FAQ
AI-powered cleaning? No. “AI” describes downstream use; the engine uses deterministic rules without an LLM.
Does it change source data? No. Source is read-only and cleaned records go to run output.
Guaranteed phone/email/URL validity? No. Formatting is conservative, ambiguities remain and no lookups occur.
Does 00123 become 123? No. Leading zeros and precision-risk integer strings remain strings.
Can it fill missing fields or verify facts? No. It measures structural consistency, not truthfulness or semantic data quality.
CSV/Excel? Use Apify exports; arrays remain arrays.
Resume interrupted runs? No. Review output/billing, then start fresh; do not retry into existing output.
Established Cloud performance? No. Local benchmarks do not establish Cloud runtime, memory or cost. Complete validation before release/pricing.