Data Cleaner & Normalizer API — JSON/CSV Dedupe & Clean
Pricing
from $1.00 / 1,000 item cleaneds
Data Cleaner & Normalizer API — JSON/CSV Dedupe & Clean
Data Cleaner API for JSON/CSV normalization, deduplication and CRM cleanup. Trim or collapse whitespace, lowercase emails, normalize phones and ISO dates, remove empty fields or rows, deduplicate by key and process nested objects with bounded batch limits.
Pricing
from $1.00 / 1,000 item cleaneds
Rating
0.0
(0)
Developer
Rosario Vitale
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Data Cleaner API — JSON/CSV Normalize, Dedupe & Clean
Why use this Actor?
Clean and normalize JSON/CSV records for CRM, analytics and data pipelines. Trim or collapse whitespace, lowercase emails, normalize phones and ISO dates, remove empty fields or rows, deduplicate by key and process nested objects with bounded batch limits.
Features
- Items (JSON array) — The records to clean — a JSON array of objects (maximum 10,000 per run).
- Trim whitespace — Trim leading/trailing spaces from all string values.
- Collapse inner spaces — Collapse runs of inner whitespace into a single space.
- Lowercase emails — Detect email-looking values and lowercase them.
- Clean phone numbers — Strip spaces, dashes and brackets from phone-looking values (keeps a leading +).
- Normalize dates to ISO — Best-effort convert date-looking values to ISO format (YYYY-MM-DD).
- Remove empty fields — Drop keys whose value is null or an empty string.
- Drop empty rows — Skip records that have no remaining fields after cleaning.
- Deduplicate by field — Optional field name. If set, keeps only the first record for each non-empty unique value of this field (e.g. "email").
- Clean nested values — Apply the same string cleaning rules recursively inside nested objects and arrays.
Use cases
- Crm data cleanup.
- Lead normalization.
- Deduplication before import.
- Json and csv pipeline preparation.
Example input
{"items": [{"name": " Alice ","email": "ALICE@example.com "}],"trimWhitespace": true,"collapseSpaces": true,"lowercaseEmails": true,"cleanPhones": true,"normalizeDates": false}
Pricing & cost control
Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.
FAQ
What is this Actor for?
It is designed for CRM data cleanup, lead normalization, deduplication before import.
Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.
How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.
Search keywords
data cleaner json csv
Clean and normalize messy datasets in one step. Feed a JSON array of records and get back tidy, consistent data — perfect as the cleaning step before deduplication, import, or analysis.
Features
- ✂️ Trim & collapse whitespace in every string field.
- 📧 Lowercase emails so
Mario.Rossi@EXAMPLE.combecomesmario.rossi@example.com. - ☎️ Normalize phone numbers — strips spaces, dashes and brackets, keeps a leading
+. - 📅 Normalize dates to ISO
YYYY-MM-DD(optional, best-effort). - 🧹 Drop empty fields and empty rows.
- 🔁 Deduplicate by any field (e.g. keep one record per
email). - 🧬 Recursive cleaning — nested objects and arrays are normalized too.
- 🛡️ Safe batch processing — validates input and caps runs at 10,000 records.
- 💳 Fair PPE billing — only records actually returned are charged as cleaned results.
Input
| Field | Type | Description |
|---|---|---|
items | array | JSON array of objects to clean (required). |
trimWhitespace | boolean | Trim spaces. Default true. |
collapseSpaces | boolean | Collapse inner whitespace. Default true. |
lowercaseEmails | boolean | Lowercase email values. Default true. |
cleanPhones | boolean | Strip phone formatting. Default true. |
normalizeDates | boolean | Convert dates to ISO. Default false. |
removeEmptyValues | boolean | Drop null/empty fields. Default false. |
dropEmptyRows | boolean | Skip empty records. Default true. |
dedupeKey | string | Field to deduplicate by (optional). Empty/missing keys are not collapsed together. |
recursive | boolean | Clean nested objects and arrays. Default true. |
Example input
{"items": [{ "name": " Mario Rossi ", "email": "Mario.Rossi@EXAMPLE.com ", "phone": "+39 (333) 123-4567" },{ "name": "Mario Rossi", "email": "mario.rossi@example.com", "phone": "+393331234567" }],"dedupeKey": "email"}
Output
One cleaned record per kept item:
{ "name": "Mario Rossi", "email": "mario.rossi@example.com", "phone": "+393331234567" }
Export as JSON, CSV, or Excel, or pull via the Apify API.
Common use cases
- Clean scraped leads before importing to a CRM.
- Normalize emails and phone numbers for matching and dedup.
- Tidy any dataset before deduplication, analysis, or upload.
- A reliable cleaning step in an automated data pipeline.
Notes
- Pure in-memory processing — no external services, nothing to break over time.
- Date normalization is best-effort; ambiguous formats may not convert.
How to use
Add your JSON records to Items, choose the normalization rules, optionally set a deduplication field, and run the Actor. The output dataset preserves the record structure while applying only the cleaning rules you enable.
Pricing and cost estimation
This Actor uses pay-per-event pricing. Billing is applied to records successfully written to the output dataset, plus the small Actor start event shown in the Store pricing panel. Invalid or removed rows are not charged as cleaned output records.
FAQ and support
The Actor is deterministic and does not use AI to guess values. Date normalization intentionally accepts only unambiguous year-month-day style dates. For reproducible issues, include a minimal input record and the run ID in the Actor Issues tab.