Job Feed Normalizer & Deduplicator
Pricing
from $0.50 / 1,000 processed records
Job Feed Normalizer & Deduplicator
Normalize supplied job records and identify duplicates by source ID or canonical URL, without fuzzy merging.
Pricing
from $0.50 / 1,000 processed records
Rating
0.0
(0)
Developer
Pietro Filippo
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Turn supplied job records into a consistent schema and identify exact duplicates. No source websites or AI services are called.
{"records": [{"feed":"acme","id":"123","name":"Engineer"},{"feed":"acme","id":"123","name":"Engineer"}],"mapping": {"source":"feed","jobId":"id","title":"name"}}
Mapping keys: recordId, source, jobId, url, title, company, location. Unmapped fields use the same input name. Dot paths such as job.title are supported. A source namespace must identify the company/feed whose IDs are unique; two unrelated sources may reuse the same numeric ID.
Each input row produces a row with inputIndex, the normalized fields and disposition: unique, duplicate or invalid. Duplicate rows include duplicateOfInputIndex; invalid rows include a reason. Keep only unique rows for your cleaned feed. The duplicate report remains available in the same dataset. SUMMARY contains counts.
Identity precedence: recognized Greenhouse/Lever URL identity; otherwise the exact supplied source plus job ID; otherwise canonical URL. Known tracking parameters and URL fragments are removed, but unknown query parameters are preserved. Rows are not merged by similar titles or company names. When an explicit generic source/ID exists it takes precedence over its URL; mixing ID-based and URL-only rows can leave duplicates, deliberately favoring separate records over false merges.
Maximum 10,000 input rows. A valid row requires a URL or source plus job ID; title alone is insufficient. Only mapped scalar fields are exported; input objects are not copied wholesale. Configured price: $0.50 per 1,000 valid processed input rows via record-processed, including duplicates, plus Apify's $0.00005 start event (one event at the default 256 MB memory). Invalid rows have no processing charge; the start fee still applies. Check the current Pricing tab.