Get started
Product
Back
Start here!
Ready-to-run tools for your AI agents and apps. Just pick one and go.
Browse 62,902 Actors
Apify platform
Apify Store
Actors for any job on the web
Actors
Build and run serverless programs
Integrations
Connect with apps and services
MCP
Give your AI access to Actors
Anti-blocking
Scrape without getting blocked
Proxy
Rotate scraper IP addresses
Open source
Crawlee
Web scraping and crawling library
Solutions
MCP server configuration
Configure your Apify MCP server with Actors and tools for seamless integration with MCP clients.
Start building
Apify for
Enterprise
Startups
Universities
Nonprofits
Use cases
Data for generative AI
Data for AI agents
Lead generation
Market research
View more →
Consulting
Apify Professional Services
Apify Partners
Developers
Documentation
Full reference for the Apify platform
Actor templates
Python, JavaScript, and TypeScript
Web scraping academy
Courses for beginners and experts
Monetize your code
Publish your Actors and get paid
Learn
API reference
CLI
SDK
Earn from your code
$1.5M paid out last month. Many developers earn over $3k.
Start earning now
Resources
Help and support
Advice and answers about Apify
Actor ideas
Get inspired to build Actors
Changelog
See what’s new on Apify
Customer stories
Find out how others use Apify
Company
About Apify
Contact us
Blog
Live events
Doing Good
Jobs
We're hiring!
Join our Discord
Talk to other builders
Pricing
Contact sales
Fuzzy Search Dataset Actor
from $0.001 / actor start
dtrungtin/fuzzy-search-dataset-actor
Search any Apify dataset using typo-tolerant fuzzy matching.
Rating
0.0
(0)
Developer
Tin
Actor stats
0
Bookmarked
3
Total users
Monthly active users
3 months ago
Last modified
Categories
Other
Developer tools
Share
broomwagon/dataset-deduper
Collapse duplicates in any dataset, in two passes you control. Exact matching ignores case, punctuation, and word order; an optional fuzzy pass catches the near-duplicates normalization cannot. Every decision is explained in an audit trail, and a cross-run ledger remembers what you already got.
Brandon Mensing
2
luminous_i/dataset-cleaner-deduplicator
Clean, normalize and deduplicate any dataset: field names, whitespace, empty values, numbers, plus exact and fuzzy duplicate removal.
Steve
searchapi/google-dataset-search-scraper
Extracts dataset titles, repositories, publishers, descriptions, formats, licenses, update dates, and dataset links from Google Dataset Search.
Search API
enosgb/crm-deduplication-tool
Detects and merges duplicate contacts in CRM databases using advanced fuzzy matching algorithms
Enos Melo
alizarin_refrigerator-owner/hubspot-company-enrichment-fuzzy-matcher-for-clay
Fuzzy match and enrich companies against your HubSpot CRM using multi-signal matching (domain, company name, phone, location). Returns HubSpot ID, lifecycle stage, deal status & confidence scores. Perfect for Clay workflows, lead deduplication, and outbound enrichment.
The Howlers
idiatech/apify-Dataset-Download
Download any dataset from the Apify platform automatically and in any format you want. Use this actor along with a Dataset toolbox automation tool.
idIA Tech
6
nibble/list-fuzzy-reconciler
Fuzzy-join two lists (CSV/JSON) on a key column with a similarity threshold. Returns matched pairs with scores, plus unmatched and near-match rows.
Simon Fletcher
fiery_dream/content-similarity-finder
Find duplicate and similar content with advanced fuzzy matching algorithms. Perfect for data cleaning and deduplication.
Cody Churchwell
scrapestorm/data-gov-uk-scraper---cheap
🔎 Easily collect dataset listings from data.gov.uk Provide one or multiple search URLs and extract dataset information such as 📄 Dataset Title 🏢 Published By 🕒 Last Updated 📝 Description 🔗 Dataset URL & more Perfect for open data research, government data monitoring & dataset discovery 📊🚀
Storm_Scraper
1
5.0
eszetael_lab/dataset-deduplicator-cleaner
Returns your records with duplicates removed and values normalized — same fields as the input, nothing added or renamed. Exact, normalized or fuzzy matching on keys you choose. A data quality report lands under QUALITY_REPORT. Takes a dataset ID or inline JSON, so it chains after any scraper.
Radosław Szal