CRM Contact Cleanup & Duplicate Review
Pricing
from $0.08 / 1,000 processed rows
CRM Contact Cleanup & Duplicate Review
Review your CRM export before importing. Compare website, email and address formatting suggestions, spot possible duplicates, and retain original information. No contact verification or automatic merging. Base price: $0.10 per 1,000 accepted review rows.
Pricing
from $0.08 / 1,000 processed rows
Rating
0.0
(0)
Developer
Critical Distinction
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Know what to change before your next CRM import.
Turn an existing contact export into a clear review list. Compare suggested formatting changes for website URLs, email addresses and U.S.-leaning addresses, see which contacts may be duplicates, and keep the original information available while you decide what to import. Other columns are preserved, not cleaned.
Use this tool to prepare and review a file you already have. It does not find new contacts, verify contact details, merge records or connect to your CRM.
Start · Prepare a file · Pricing and limits · Help · Technical reference
What you receive
| Result | How it helps |
|---|---|
| A read-only review page | Compare possible duplicates, see differences and find rows needing attention. |
| A complete CSV or JSON review file | Keep original information and review fields together while preparing your own approved import copy. |
| Supporting review files | Follow the matching evidence and inspect rows that need attention or carry a decision. |
You remain in control. A formatting suggestion is not verification, and a possible duplicate is not a confirmed match. The review page does not save approvals or change your CRM. The accepted-row Dataset is a separate review table, not the complete file to use for an import.
Explanatory example: Two supplied contacts share a website and address but have differently capitalized email text. The result flags a possible match for you to review; it does not merge them. A third contact's invalid email stays visible as an unresolved issue.
Start here
Open the fictional contact example to try the workflow without supplying your contact list. Sign in to Apify, review the displayed price and maximum cost, then start the example. When it finishes, open Start here: Contact review in the run's Output section.
Built-in form example:
Select Input source → Try the fictional example, check the displayed maximum cost, and click Start. The built-in example processes 12 fictional rows through the normal priced workflow: $0.0012 at the base rate. Plan discounts may apply.
When the run finishes, open Start here: Contact review. Compare the suggestions and possible matches, then use Complete review file to obtain the file for your review. The example lets you see the result without supplying customer data.
Try a small file of your own
- Keep a backup of your export. Choose My stored data / file, then Source Route → File upload. Upload a CSV, JSON Lines or JSON array file and use Match your columns to identify your ID, website, email and address columns. Check the tested file sizes below first.
- Set Source Row Cap explicitly and enable Preview Only (check source,
no files). A successful check reports
Previewedand the accepted count in the run status message; it does not create download files. Turn Preview Only off to run cleanup. - After completion, open the review page and compare the complete review file with your source. Record decisions separately, prepare a copy containing only approved changes, and resolve remaining issues before a test import into your CRM. Keep leading-zero IDs as text and disable formula evaluation when importing untrusted CSV data into a spreadsheet.
For a worked example, use the fictional three-row CSV and the copyable file and column mapping below. It shows a possible duplicate pair and a contact with an invalid email that still needs a decision. The worked handoff shows exactly which cells a reviewer chooses to change and why the resulting file still needs an import check.
Use My confidential records for an inline JSON array, or select an existing Apify Dataset or stored file. Fill only one input route. Remove the previous route's data when switching; an empty or null data field still counts as supplied. Input-route details are below.
Pricing and limits
This Actor uses Pay Per Event pricing.
Base price: $0.10 per 1,000 accepted review rows—equivalent to $0.0001 per accepted Dataset row. Three accepted rows cost $0.0003; 12 cost $0.0012. A row is billable when accepted into the review Dataset, even when unchanged or still needing a human decision. Review files and this Actor's platform usage are included. Preview produces no billable result rows.
Check the current pricing page and maximum cost before starting. Plan discounts may change the effective run price. Account subscriptions, taxes and unrelated services are separate.
Check your file size before uploading. Current testing covers these two separate profiles:
| Accepted rows | Source columns per row |
|---|---|
| 1,000 | 6 |
| 12 | 261 |
These tests do not establish other combinations or arbitrary field contents; 1,000 rows with 261 columns is not a validated combination. The technical schema retains a 10,000-row maximum for older integrations, not as supported customer capacity. Set Source Row Cap explicitly. Exceeding the cap stops the run rather than silently dropping rows. Physical blank CSV/JSONL lines do not count; accepted rows without a usable mapped value do.
A successful preview checks the source; it does not establish broader capacity, CRM import readiness or the success of a later run's output and billing steps.
Contact 4.0 input-name migration
Starting with Contact 4.0, rename dedupeKeyMode to legacyMatchMode in any saved input that includes it, keeping the same value. Inputs without that property need no change. The values off, keys_only and keys_and_candidates, the default and the Console prefill are unchanged. The setting remains inactive: it does not change cleanup or matching. To keep using the old property name, select the retained 3.2.1 build. Result formats, contact-input encryption and charging are unchanged.
Privacy and Contact 3.0 migration
Your review files and results contain the contact data you provide. Set suitable access and retention controls before using customer data. The Actor does not delete your source or result storage.
Both inline fields, confidentialRecords and records, use Apify secret input
encryption in stored run input. This setting does not extend to uploaded
files, source storage, Dataset rows, review pages or downloaded results.
For existing API and saved-input users: continue submitting the same JSON
array. Direct reads of stored INPUT.records return ciphertext rather than
that array. Keep your original input securely or use the intended result files;
never copy ciphertext into a new run as a retry payload. Older saved inputs may
still contain plaintext—the schema change does not retroactively encrypt or
delete them. The contact-array format and result schemas are unchanged. Older saved inputs may still need the property rename described in Contact 4.0 input-name migration above.
Recovery and help
A stopped run may already have saved results or incurred charges. Check the
existing run before starting another one. Its Dataset, stored files, OUTPUT
and any RECOVERY_STATE.json help determine what happened. Follow the
recovery guide; a new run can
create duplicate results and charges.
Questions before you start? Ask about file fit or the review workflow in the Issues tab. Sign in to Apify to post. Describe your goal, file format and approximate row and column counts; do not upload your contact list. For a problem with an existing run, include its run ID and what you expected. Issue discussions are public: use fictional examples and never post private files, credentials or access links.
Integration recipes
- CRM export: upload and map columns → review the complete file → prepare an approved import copy → test it using your CRM's own import tools.
- Apify workflow: provide one selected Dataset as
source, grant READ access and setsourceMaxRowsexplicitly. Retrieve the complete review file through the finished run's Output links, not by exporting the review Dataset. - API: submit with a run cap. Wait for completion and inspect{"inputSource":"records","records":[{"recordId":"demo-1", "email":"team@example.invalid"}]}
OUTPUT; do not blindly retry an uncertain run-creation request.
Advanced reference
The detailed file, mapping, import and recovery guides below describe the existing workflow. Use them when preparing your own import or connecting an API-based process; the example above is the starting point for a first run.
Your read-only review page and full files
The default Key-Value Store contains the buyer workflow package:
| File | Use it for | Important boundary |
|---|---|---|
CONTACT_REVIEW.html | Start with readable possible matches, differences, changed and unresolved rows | Self-contained, no scripts or external assets; bounded display with full-file links |
FOCUSED_REVIEW.json | Consume the prioritized unique candidate pairs in another tool | Versioned view; up to 100 pairs, with total and omitted counts |
CLEANED_CONTACT_REVIEW.csv or CLEANED_CONTACT_REVIEW.json | Compare the whole file before preparing an approved import copy | Includes review fields; it is not itself an approved import copy |
RECORDS_TO_REVIEW.json | Work only the rows that need attention or carry a human decision | Includes blank/rejected identities that are intentionally absent from the Dataset |
POSSIBLE_MATCH_GROUPS.json | Inspect all raw URL, email and address groups, including weak context | Group membership does not establish the same person |
RUN_SUMMARY.json | Confirm route, format, counts, and exact artifact keys | Written at artifacts_ready, before Dataset publication |
The Dataset is a secondary accepted-row review surface. It preserves each accepted original row as typed JSON beside safe cleaned values, workflow facts, and possible-match references. It is not the reimport file and excludes structurally blank or rejected source rows.
The Complete review file (CSV/JSON record listing) Output link opens a record listing; select the file there. The HTML page provides a direct link to the selected CSV or JSON file. Use a CSV-aware tool with formula evaluation disabled when working with unknown source data; preserved formula-like values are not spreadsheet-sanitized.
| What you see | What it means | Your decision |
|---|---|---|
| Changed / formatting suggestion | A supported value has a proposed formatting change | Compare it with the original; this is not verification or your approval |
No cleanup issue flagged (no_review) | The cleanup rules did not flag a review issue | Check your CRM's rules; this does not mean approved, unique or import-ready |
| Possible duplicate | Supplied values have a comparison signal in common | Check identity; choose no survivor or merge solely from this result |
| Review required | The original or a suggested value needs a human decision | Correct it from a reliable source or leave it explicitly unresolved |
The four original workflow roles remain unchanged. The two additional review
views add no billing event. They are prepared under byte limits and acknowledged
before the first billable Dataset write. The final structured OUTPUT is written
only after those files and all Dataset batches are acknowledged. It records the final
operation state, not a storage readback or exactly-once
transaction proof.
Try the copyable CSV example below for a complete first-file walkthrough. It requires no source-repository access.
Try a small CSV file
Save this fictional example as first-file.csv:
crm_id,website,email_address,mailing_address,segment,note00017, alpha.example/?utm_source=crm ,team@alpha.example,100 Example Road Chicago IL 60601,trial,"Keep comma, and leading zero ID"00018,https://alpha.example,TEAM@ALPHA.EXAMPLE,100 Example Rd Chicago IL 60601,customer,Review the possible match; do not merge automatically00019,https://beta.example,not-an-email,warehouse behind blue door,trial,Keep original values when review is required
Choose Input source → My stored data / file. Set the source type to
file, upload the CSV with the Source file picker, and choose format csv.
Under Match your columns, enter the four source names below. The JSON/API shape remains:
{"recordId":"crm_id","url":"website","email":"email_address","address":"mailing_address"}
Set sourceMaxRows: 3 and previewOnly: true to check the source and estimate.
Preview consumes platform resources but emits no billable Dataset rows. Then
set previewOnly: false, workflow.exportMode: "side_by_side", and
workflow.exportFormat: "csv" to produce the workflow files.
Open CONTACT_REVIEW.html, then CLEANED_CONTACT_REVIEW.csv. The original IDs must remain 00017, 00018
and 00019, with segment and note unchanged. Review IDs 00017 and 00018 as possible
matches because their normalized URL and address signals agree; inspect the
mailbox-case difference rather than assuming email ownership. Do not merge them
automatically. ID 00019's email and address require a decision. Use
RECORDS_TO_REVIEW.json and POSSIBLE_MATCH_GROUPS.json and the handoff below.
A Dataset export is not the whole-file reimport artifact.
Worked approval-to-import handoff
Keep the original CSV immutable. Record its file hash, the run ID, source column,
original value, proposed value and your accept, keep or defer decision.
Identify cells by the file identity and source row number, counting CSV data
records from one after the header; embedded newlines are not extra rows. A CRM
ID can be blank or duplicated and is not a safe sole key. Reject stale values or
duplicate/contradictory decisions before applying anything.
This fictional operator chooses the following decisions; they are examples of human judgment, not automatic approvals from the Actor:
| Source row / CRM ID | Column | Decision | Value in approved copy |
|---|---|---|---|
1 / 00017 | website | Accept | https://alpha.example/ |
1 / 00017 | mailing_address | Keep | Original address |
2 / 00018 | website, mailing_address | Keep | Original values |
2 / 00018 | email_address | Accept | TEAM@alpha.example |
3 / 00019 | website | Keep | Original value |
3 / 00019 | email_address, mailing_address | Defer | Original unresolved values |
Make a new import copy with only the original six columns, applying only the
two accepted cell changes. Do not import helper columns or change the immutable
run's humanDecision fields. Keep the possible-match rows separate. The result is:
crm_id,website,email_address,mailing_address,segment,note00017,https://alpha.example/,team@alpha.example,100 Example Road Chicago IL 60601,trial,"Keep comma, and leading zero ID"00018,https://alpha.example,TEAM@alpha.example,100 Example Rd Chicago IL 60601,customer,Review the possible match; do not merge automatically00019,https://beta.example,not-an-email,warehouse behind blue door,trial,Keep original values when review is required
This file is not ready for the example destination: its nonempty email rule
rejects row 3's not-an-email. Keep that blocker visible until a human supplies a
verified correction or chooses a destination-supported treatment. Do not invent
an email, silently blank it, drop the row or claim the address is verified.
Other CRMs have their own required fields and validation rules.
Parse both files with a CSV-aware tool and verify identical headers, row count, row order and every unchanged cell; exactly those two approved cells may differ. Preserve leading zeros, Unicode, quoted commas/newlines and formula-like strings as literal data. Use text/CSV import settings that avoid spreadsheet evaluation. Recreate the copy from the original and the same decisions when needed; do not apply decisions blindly to a previously changed file. Resolve blockers before a test destination import, then check its results and rollback options before production. No CRM access, automatic merge or measured time-saving is implied.
Review and reimport
- Read
RUN_SUMMARY.jsonand confirm the source route, source format, export mode, export format, source count, accepted count, review count, possible match count, and four artifact keys. - Open
RECORDS_TO_REVIEW.json. Useidentity.sourceRowNumberand an optional business ID to locate each source row. Resolve everyreview_requireditem and preserve the resolution outside this immutable run artifact. - Start with
FOCUSED_REVIEW.jsonor the HTML possible-match section. Multiple matching field families rank first, then normalized mailbox matches, then single location/website matches. Ties use business ID and source identity. Each pair appears once with combined evidence and differences; A-B and B-C overlap does not invent A-C identity. Domain-only, host-only, building-only and provider-alias groups stay inPOSSIBLE_MATCH_GROUPS.jsonas context. - Inspect the corresponding rows in
CLEANED_CONTACT_REVIEW.*. Inside_by_side, compare the source and generated fields. Inreplace_supported_fields, compare replaced cells with the displaced original fields. - Prepare a separate import copy containing the destination's original columns and only approved cell changes, as above. Resolve its destination blockers, keep the decision record and source backup, and use your CRM's own test-import, validation and rollback controls.
humanDecision begins as unreviewed; the Actor does not invent an approval
or rejection. A possible-match group never deletes a row or chooses which row
should survive.
The focused view displays at most 100 pairs. The HTML displays at most 20 changed or unresolved rows, 12 original fields per displayed row and 512 characters per displayed value. It labels shortened values and omitted work. Full values remain in the import and machine files. Shared mailboxes can belong to roles or families; addresses can be shared premises, and websites can be hosting platforms. Even multiple matching fields require a human identity decision. These counts do not measure review minutes. Exact CSV originals may contain spreadsheet formulas; the HTML displays text inertly, while the import CSV intentionally preserves it.
Recovery: inspect first, never retry blindly
Key-Value Store publication is ordered before Dataset publication, and final
OUTPUT is last. There is no cross-storage transaction, rollback, or
exactly-once guarantee. If execution returns a caught failure,
RECOVERY_STATE.json records the last acknowledged durable milestone and the
first read-only action.
| Observed state | What is known | First safe action |
|---|---|---|
publish_review_views failure | Original workflow files may be staged; a review-view PUT failed or its acknowledgement is uncertain. No Dataset write was attempted | GET the two review keys and RUN_SUMMARY.json; do not infer final completion without OUTPUT |
| Deterministic Dataset rejection | Four artifacts and RUN_SUMMARY.json were acknowledged; the attempted Dataset batch committed zero items | GET the four KVS records and inspect the recorded Dataset attempt before any new Dataset POST |
| Dataset transport/5xx failure | Four artifacts were acknowledged; the attempted Dataset batch may or may not have committed | GET the Dataset and four KVS records; reconcile source-row identities and counts before deciding whether any write is safe |
Final OUTPUT write failure | Buyer artifacts and exact Dataset batches were acknowledged; final status commit is unknown | GET OUTPUT, the four KVS records, and Dataset items before considering an overwrite |
| Hard process death | Only the last separately acknowledged durable milestone is known | Inspect KVS, Dataset, and OUTPUT; do not infer later steps from local preparation |
GET-only inspection is the common first step. For deterministic rejection,
recorded Dataset cardinality is zero_rejected.
For transport or server ambiguity, it is unknown_commit and
committedDatasetItemCount is null. Both states set
automaticDatasetRetryForbidden: true. A successful GET inspection can inform
an operator decision, but it cannot retroactively prove a transaction or erase
an unknown commit.
Support diagnostics: share facts, not contacts
Before opening the public Issues route,
collect only the run ID, UTC time, input route type, workflow mode/format,
source and accepted counts, presence of the four artifact keys, final
OUTPUT or RECOVERY_STATE state, and the exact error type/status. If a
reproduction is needed, use the smallest fictional input that still fails.
Do not paste credentials, authorization headers, signed storage URLs, raw contact values, source files, or complete Dataset/KVS bodies. Redact customer and account identifiers before sharing logs. Use the Issues tab for product questions and problems with a run. We do not offer a guaranteed response time or service-level agreement.
Migrating from Contact 1.0
Contact 2.0 package: Cargo
2.0.0, Actor version2.0, andcontact-cleanup-output-v3describe one whole-file workflow contract. Deployment availability, Store visibility, and pricing are point-in-time Apify properties; confirm them before starting a run.
The Contact 2.0 changes below preserved the earlier contact-row and source-input formats. When moving to Contact 4.0, also apply the input-name migration above if your saved input contains dedupeKeyMode. In 3.0, both inline fields are encrypted in stored INPUT; see the migration note above. The new
confidentialRecords lane is additive. Omitted workflow uses deterministic
side-by-side CSV defaults, and the accepted compatibility controls are not
runtime-version selectors.
Contact 2.0 changes the output package:
- the Dataset is an accepted-row review projection, not the whole-file
reimport file; blank and rejected source identities remain in
RECORDS_TO_REVIEW.json; - the default Key-Value Store leads with one cleaned reimport file,
RECORDS_TO_REVIEW.json,POSSIBLE_MATCH_GROUPS.json, andRUN_SUMMARY.json; - final
OUTPUTusescontact-cleanup-output-v3and is written only after those KVS artifacts and accepted Dataset batches are acknowledged; and - caught cross-storage failures may leave a committed or commit-unknown
prefix. Inspect KVS, Dataset, and
OUTPUTwith GET requests before deciding whether any retry or overwrite is safe.
Consumers moving from Contact 1.0 should validate their Dataset and OUTPUT
readers against these shapes, use the Actor Output links for file discovery,
and preserve source backups until review and reimport are complete.
Minimal inline example
{"records": [{"recordId": "crm-1001","url": " example.invalid/about/?utm_source=crm ","email": "Person@Example.Invalid","address": "900 cedar street chicago il 60601","accountOwner": "Synthetic Team"}],"workflow": {"exportMode": "side_by_side","exportFormat": "csv","generatedFieldNamespace": "contactCleanup"}}
Inline records use the fixed recordId, url, email, and address names.
The committed Console prefill and smoke input contain 12 fictional rows that
exercise safe changes, review-required values, and possible-match groups.
API, task, CLI, and scheduler callers must submit their own input because
Console prefill is a UI convenience, not a runtime default.
Choose one input route
Routes are mutually exclusive. Do not combine inline records with source.
Do not combine confidentialRecords with either route, and do not submit an
empty selected inline array. Source-only controls sourceMaxRows, rowLimitBehavior, and a true previewOnly are invalid with either inline lane. Console may include neutral previewOnly: false with an explicit example or inline inputSource; that is accepted. Legacy calls without inputSource must omit source-only flags.
For file and Key-Value Store routes, declare csv, jsonl, or json_array,
or use auto when the record name, extension, or content type is sufficient.
Dataset rows are already JSON objects and do not take a format selector.
| Route | Required input | Mapping | Route behavior / limit |
|---|---|---|---|
| Sensitive inline | confidentialRecords | Fixed supported names | Schema accepts 1-1,000 rows; secret input storage protection; see the separately tested profiles above |
| File | source.type = "file" and source.file | Optional source.fieldMap | Preview first; use the separately tested profiles under Pricing and limits, not a general 1,000-row promise |
| Dataset | source.type = "dataset" and source.datasetId | Optional source.fieldMap | Reads complete object rows from an optional offset |
| Key-Value Store | source.type = "keyValueStore", storeId, and recordKey | Optional source.fieldMap | Reads one selected text record |
| Existing inline API field | records | Fixed supported names | Schema accepts 1-1,000 rows; same tested-profile limits; encrypted in stored INPUT from 3.0, with the same literal-row submission shape |
Sensitive inline records
confidentialRecords and records are top-level secret JSON arrays with the same row shape and 1-1,000 limit. It intentionally has no default, prefill, or
committed body example. When this Actor is invoked through Apify API or
Console, Apify encrypts the field before saving the run input; the Actor's
runtime input reader decrypts it only inside that run. A direct read of the
run's INPUT KVS record therefore returns the encrypted field rather than the
plaintext submitted rows.
This boundary protects the stored input field, not every later artifact. The
cleaned KVS files, Dataset review rows, final OUTPUT, logs you add outside
this Actor, and exported downloads keep their existing visibility and Apify
retention controls; the secret-input setting does not encrypt or erase them.
The review page and focused JSON contain supplied contact data just like the import files. They are not protected by the inline secret-input setting. This Actor does not delete source or result storage. Review the destination
storage and retention settings before processing customer data, and never send
raw contacts or storage bodies to public support channels.
For a retry, rollback, or replacement run, inspect prior storage first under the recovery guidance above, then resubmit the original input through the Actor's API or Console so Apify can encrypt it for that run. Do not treat the stored ciphertext as a portable plaintext payload or as proof that prior outputs were deleted.
auto is a selector for a supported resolved format; it is not an additional
format. The maintained equivalence suite closes these eight resolved cells:
inline JSON object, file CSV/JSONL/JSON array, Dataset JSON object, and
Key-Value Store CSV/JSONL/JSON array.
File preview
{"source": {"type": "file","file": "apify://key-value-stores/synthetic-source/records/contacts.csv","format": "csv","fieldMap": {"recordId": "crm_id","url": "website","email": "email_address","address": "mailing_address"}},"previewOnly": true,"sourceMaxRows": 1000,"rowLimitBehavior": "fail_over_limit","workflow": {"exportMode": "side_by_side","exportFormat": "csv"}}
Preview performs bounded reads and local workflow preparation but writes no
buyer artifacts, Dataset items, final OUTPUT, or custom charge. A successful
preview reports Previewed and the accepted count in the run status message;
it does not produce Output downloads. Set Preview Only to false to run cleanup
after checking the count and capacity limits. A successful preview does not
prove import readiness, provider storage, charge settlement, or a later execution.
Dataset execution
{"source": {"type": "dataset","datasetId": "synthetic-dataset-name","offset": 0,"fieldMap": {"recordId": "crm_id","url": "website","email": "email_address","address": "mailing_address"}},"previewOnly": false,"sourceMaxRows": 1000,"workflow": {"exportMode": "replace_supported_fields","exportFormat": "json"}}
The input schema retains a parser/runtime guard as high as 10,000 for backward
compatibility. That guard is not a supported buyer ceiling: the current
tested profiles are the two combinations listed above. Set sourceMaxRows
explicitly within those limits; do not infer support from
the schema maximum or the omitted-field runtime fallback.
The row-limit policy is always fail_over_limit. The first excess accepted
row fails before output rather than silently truncating the source.
Choose the cleaned-file shape
side_by_side (recommended for review)
Every original source field appears first and remains unchanged. Generated
fields are appended under a collision-safe namespace such as
contactCleanupCleanedUrl, contactCleanupReviewNeed, and
contactCleanupHumanDecision.
This mode makes the old and proposed values easy to compare. It does not replace a supported source value automatically.
replace_supported_fields
Only a mapped url, email, or address cell with an exact safe,
no_review replacement changes in place. The generated namespace retains the
displaced original supported values and the same decision fields. Original
fields outside those three mapped targets remain exact.
A missing safe replacement, a review-required result, blank input, or rejected row is never turned into an automatic replacement. This mode is deterministic projection, not verification, merge, or survivor selection.
csv
The CSV export is UTF-8 with a header. CSV-origin columns and order are preserved, followed by generated fields. JSON-origin source values are encoded as compact JSON text inside CSV cells so numbers, booleans, nulls, strings, arrays, and objects remain distinguishable. For example, a JSON string source value is represented by JSON text including its quote characters; it is not silently flattened to an untyped value.
json
The JSON export is a compact UTF-8 array. It preserves native JSON source types and source-defined field order under the workflow's deterministic projection contract.
CSV and JSON are semantically equivalent workflow choices, but different serializations are not byte-identical to each other. Replaying the same input and the same selected format produces the same maintained artifact bytes.
Cost and qualification boundary
The qualified local support profiles are 1,000 accepted rows at 6 source fields and 261 source fields at 12 rows. They are separate profiles.
The tests did not establish 1,000 rows with 261 fields or other combinations. The same
production-facade window also passed all eight resolved input route/format
cells, both export modes, both export formats, and exact same-format replay
for all four buyer artifacts. The 1,000-row profile has a local configured
estimate of $0.100000 for 1,000 synthetic default Dataset-item event
units.
That number is a deterministic local estimate derived from the configured fixture price. It is not an Actor price, a bill, or a settled charged-event count; a platform-owned event amount is not this package's price. The estimate agrees arithmetically with the point-in-time Store rate above, but the live pricing route remains authoritative immediately before a run. No custom charge call occurs in the locally qualified workflow; the synthetic default Dataset-item event remains platform-owned and must be proved on the live surface that owns it.
The 1,000-row run completed locally in about 0.178 seconds with a measured peak working set of 47,427,584 bytes. Those values describe one qualification host and fixture, not an SLA, production latency, concurrency, or memory guarantee.
The platform owns the synthetic default Dataset-item event charge. Check the current price and your run cap in Apify before execution.
Supported cleanup semantics
- URL: safe structural normalization such as scheme/host normalization and removal of known tracking fragments. The Actor does not crawl or verify the destination.
- Email: conservative syntax normalization and deterministic review signals. The Actor does not send mail, verify deliverability, prove inbox existence, or prove ownership.
- Address: conservative U.S.-leaning normalization with unit-aware comparison signals. The Actor does not geocode or certify postal deliverability.
- Possible matches: deterministic same-run candidate groups with stable IDs and reason codes. These do not prove identity and never merge records.
Unknown source fields are preserved as data; they are not interpreted as additional cleanup targets. The workflow is deterministic and does not call external enrichment or verification providers.
Integration recipes
API or task input
Submit the same JSON shapes shown above. Do not depend on Console prefill.
Submit either confidentialRecords or records as a JSON array through Apify
API or Console; both inline fields are encrypted in stored run input. For
already-stored data, use a fixed file/KVS reference or Dataset ID, explicit
mappings, an explicit sourceMaxRows <= 1000, and your chosen workflow object.
This input setting does not encrypt uploaded files, source storage or output
artifacts, and does not retroactively encrypt or delete older saved inputs.
Manage their access and retention separately.
Reading results
Open Contact review from Actor Output first; use the linked full import and machine files for complete data. The four original artifact roles and their hashes remain the compatibility contract; the two review views are additional files. Use the Dataset review view for accepted-row exploration. Dataset views change presentation only; they do not create a new export or replace the whole-file artifact.
Scheduled runs
Treat each run as a separately reviewed package. Store the run ID, source identity, selected workflow controls, exact build, and artifact keys. Do not automatically replay failed Dataset writes or infer version, price, or output behavior from a schedule, task, or moving build selector.
Explicit nonclaims
Contact Cleanup does not claim:
- deployment, Store availability, or price without a current Apify check;
- future pricing or charged-event settlement from a saved price observation;
- unlimited volume, production throughput, or an SLA;
- website reachability, email or postal deliverability, ownership, or identity;
- enrichment, geocoding, confidence scoring, or external verification;
- automatic merge, deletion, identity confidence scoring, or survivor selection;
- cross-storage transactions, rollback, exactly-once delivery, or storage readback; or
- that a generated schema, fixture, report, or successful local command proves live consumer behavior.
Permissions
The Actor runs with limited permissions. It reads only the input resources you select, writes its run-scoped Dataset and Key-Value Store outputs, and does not contact enrichment, verification, geocoding, or scraping providers.
Maintained evidence
The examples in this package are projected from
.actor/smoke_input.json through the same Contact 2.0 production workflow and
frozen in
tests/fixtures/workflow_v3_buyer_workflow_package.json. A causal regression
test fails if the runtime projection, fixture, buyer documentation markers,
four schema discovery surfaces, supported ceiling, pricing nonclaim, or
recovery guidance drifts.
The source package includes the maintained changelog and detailed fixture examples.