Local Business Leads Scraper 🚀- MX-Verified Emails & WordPress
Pricing
from $3.00 / 1,000 business leads
Local Business Leads Scraper 🚀- MX-Verified Emails & WordPress
Finds local businesses of any category in any city worldwide (dentist, roofer, gym). Crawls each site for MX-verified emails, 5 socials + website platform (WordPress, Wix, Squarespace, GoDaddy, Shopify) - ideal for web design agency leads. Flags agency-owned mailboxes. Scores 0-100 A-F. $3/1k.
Pricing
from $3.00 / 1,000 business leads
Rating
5.0
(3)
Developer
Flash Scrape
Maintained by CommunityActor stats
4
Bookmarked
13
Total users
4
Monthly active users
an hour ago
Last modified
Categories
Share
Local Business Leads Scraper — any category, any city, with MX-verified emails and lead scoring
Any local business category, in any city on earth — and every row comes back with an MX-verified email, whose mailbox it is, the website platform the business runs, and a 0-100 lead score. Priced per delivered lead (current rate on this page's Pricing tab), with email verification included rather than sold as an add-on.
Dentists, gyms, lawyers, plumbers, roofers, salons, real estate agencies, restaurants and 50+ more curated categories — plus any term you type. No API key, no proxy, no separate scraper per niche.
What you get
- Any category, any city — ~50 categories mapped to exact OpenStreetMap tags, any other term attempted directly and then rescued by a business-name search, and any city on earth geocoded (
Austin, Texas,casablanca morocco,Dubai UAE). - Emails that are verified, not guessed — every address is MX-checked over DNS-over-HTTPS and graded
deliverable/risky/undeliverable. Verification is part of the price, not an upsell. - Whose mailbox it is —
own_domain, a free inbox, or the business's marketing agency. Emailing an agency mailbox never reaches the business, so this column decides whether a lead is worth a send. - Redesign pitch signals —
website_platform(WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more),mobile_viewport,copyright_year_stale,has_meta_pixel. A builder-tier site with no tracking is the highest-intent pitch there is. - A score you can audit — 0-100, an A-F grade, a hot/warm/cold tier, and a
score_breakdownobject showing the arithmetic that produced it. - Five ready-made Output views — Overview, Email-ready, Web-agency targets, Map & source, All columns. Pick one above the results table instead of scrolling a wall of 61 raw columns.
- Filters that cut your bill, not just the table — every filter drops the row before it is pushed and before it is charged, and the run log names each filter and how many rows it removed.
- One-click provenance —
google_maps_urlandosm_urlon every row, plus the ODbL attribution the licence requires you to keep.
Measured, not promised
Every number below is from a real run of this Actor on dentist / Austin, Texas. Inputs and dates are given so you can reproduce them.
| Run | Result |
|---|---|
| Console form untouched (cap 25, Website required on), pressed Save & start — 2026-08-08 | 25 of 25 asked for — 25 with a website, 25 with a phone, 10 with an email. The same input delivered 14 before the pool-depth fix described under How to use it |
| Console form as it opens, Max businesses raised to 100, pressed Save & start — 2026-08-08 | 55 businesses, 100% contactable — 96% with a phone, 55% with an MX-verified email. Under the 100 asked for, and the run says why: OpenStreetMap holds 171 dentist records in Austin and 55 of them are crawlable with a website inside the search area. That is the whole city, not a truncation |
Bare {} from the API (no filters at all) — 2026-08-08 | 100 businesses, 50% contactable — the honest filter-off case, see the fill-rate table below |
| Cold-email list preset — 2026-08-08 | 31 rows (every emailable dentist Austin has — the exact count varies by city and over time), 100% with an MX-verified email, 94% with a phone |
| Call list preset, cap 100 — 2026-08-08 | 56 rows, 100% with a phone — the same number on two independent runs, and the run reports it as the whole of Austin rather than a truncation |
| Web-design prospects preset, cap 100 — 2026-08-08 | 7 rows, every one scored ≤45 and reachable on at least one channel. The score ceiling alone matches 68 businesses in Austin, but only 7 of those publish a phone, email or social profile — the preset drops the other 61 rather than bill you for map pins you cannot contact |
| Full enrichment preset — 2026-08-08 | 100 rows plus 23 email_guess addresses, kept out of email and never billed as verified |
| Column stability across every run above | one identical 61-column tuple on every row — CSV headers never shift mid-export |
The gap between the Console-form rows and the bare {} row is the whole story of this Actor's honesty: the Console form pre-fills onlyWithWebsite, which is why it delivers 100% contactable businesses, while a bare API call keeps every mapped location including the ones with nothing to contact. Both numbers are published; neither is hidden. The same goes for the 55-of-100 line: a run that cannot reach your cap says so in its status message, with the counts that explain it.
The Console form opens at 25 businesses so your very first run finishes in a couple of minutes and costs cents — raise Max businesses once you like what comes back. That number is a Console pre-fill only: the API default is still 100, so existing API callers, tasks and schedules that send no maxItems are completely unaffected.
Quick start
{"category": "dentist","location": "Austin, Texas"}
That is a complete run — and it is also exactly what the form already contains. Open the Actor, press Save & start without touching anything, and you get businesses. Everything else has a working default: the Console form starts at 25 businesses (the API default stays 100), website crawling on, email verification on, 3 pages per website.
The input form at a glance
The sections in bold are the ones worth a look. The other two — radius search and crawler tuning — are technical, already tuned, and safe to ignore on a first run.
| Section | What it is for |
|---|---|
| Quick start | One dropdown that configures everything else — see the presets below |
| What to search | The business category and the city (one of each, or several at once) |
| Search a radius around a point | Technical. A circle around a coordinate instead of a city |
| Bring your own list | You supply the websites; OpenStreetMap discovery is skipped |
| Limits | How many businesses this run may deliver — your cost ceiling — and expandNearby, which widens the search around the city when the city itself runs out (each widened row labelled in query_location) |
| Email + platform enrichment | What gets read from each business website |
| Crawler tuning | Technical. Speed-versus-depth knobs for the crawl |
| Filters | Drop businesses before you are charged for them |
| Output | Which columns you get, and in what order |
Presets
The Use-case preset dropdown at the top configures the rest of the form for one specific outreach job. It is optional — leave it on Custom and nothing at all changes.
| Preset | What it sets | Measured on the default search (Austin dentists) |
|---|---|---|
| Cold-email list | onlyWithEmail, onlyVerifiedEmail | 31 leads (the whole of Austin; the count varies by city and over time), 100% with an MX-verified email, 94% with a phone |
| Web-design prospects | maxScore: 45, requireAnyContact | 7 leads, all in the thin/neglected half — no published email, no marketing tech, often no mobile viewport — and all reachable. maxScore: 45 on its own matches 68 businesses; the contact floor is what removes the 61 you could not pitch to |
| Call list | requirePhone | 56 leads, 100% with a phone number |
| Full enrichment | emailPatternGuess, maxPagesPerSite: 6 (every crawl toggle is already on by default) | 100 leads plus 23 guessed email_guess addresses, kept out of email |
Anything you set yourself wins, no preset ever widens your bill, and the run log names exactly what each preset applied and what it stood down on. The three guarantees are spelled out — and tested — under Presets: the three guarantees below.
What the Output tab looks like
Five curated views ship with the Actor. Pick one above the results table; the CSV / JSON / Excel export is unaffected and always carries every column you asked for.
| View | Columns | Use it for |
|---|---|---|
| Overview (default tab) | name, category, city, phone, email, website, lead_grade, lead_score | The eight columns that answer "is this a lead?" |
| Email-ready | name, email, email_status, email_type, phone, website, contact_page_url, lead_score | Loading a cold-email sequence — deliverability grade and mailbox owner side by side |
| Web-agency targets | name, website, website_platform, mobile_viewport, copyright_year_stale, has_meta_pixel, phone, email | Building a redesign pitch list from the neglect signals |
| Map & source | name, address, google_maps_url, osm_url, latitude, longitude, attribution | Verifying a row in one click, territory mapping, keeping the ODbL attribution with the data |
| All columns | all 61, in CSV order | Everything, when you want the full table |
Two deliberate choices in those views. website, contact_page_url, google_maps_url and osm_url render as clickable links. And the tri-state flags — mobile_viewport, copyright_year_stale, has_meta_pixel — render as text, not as a checkbox, because they are null when the site was never crawled and a checkbox cannot tell "no mobile viewport" apart from "we never looked". has_email / has_phone / has_website are never null, so those do get the real boolean widget.
What does it do?
This actor takes a plain-English business category (e.g. "dentist," "hair salon," "roofing contractor") and a location, maps the category to the right OpenStreetMap tags, and pulls every matching business in the area. It then crawls each business's own website — following the site's own contact/about/team links (including Shopify /pages/contact and non-English slugs) within the maxPagesPerSite budget — and extracts, from pages it has already downloaded:
- a contact email, MX-verified over DNS-over-HTTPS and graded
deliverable/risky/undeliverable - whose mailbox it is — the business's own domain, a free inbox, or its marketing agency (emailing that one never reaches the business)
- social profiles — Facebook, Instagram, LinkedIn, X/Twitter, YouTube
- which website platform it runs on — WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Joomla, Drupal, HubSpot CMS, and a "Custom / Next.js" bucket for hand-built sites
- marketing & booking tech — Meta Pixel, Google Analytics/GTM, Calendly, NexHealth, Klaviyo, WooCommerce and ~25 more
- the homepage's title and meta description — drop-in mail-merge personalization fields (and a data tripwire: a title naming a different business exposes a wrong OSM
websitetag) - web-agency pitch signals — a missing mobile viewport tag, a stale footer copyright year
- one-click source links —
google_maps_urlto the live Google listing,osm_urlto the OpenStreetMap element - a lead score 0-100, an A-F grade, and a hot/warm/cold tier
Why use it / who's it for
- Web design & marketing agencies — filter for a builder-tier platform (Wix, GoDaddy, Weebly) and
has_meta_pixel: falseto find businesses with a dated site and no tracking: the highest-intent redesign pitch there is. TheCustom / Next.jsvalue is the inverse signal — don't pitch a DIY-site rebuild to someone who already paid a developer. Build 0.1.21 adds two more neglect signals —mobile_viewport(missing = the site predates responsive design) andcopyright_year_stale— plus anonlyWithoutWebsitefilter that returns only businesses with no site at all: the first-website pitch list. - Freelancers on Fiverr/Upwork — generate a "100 dentists in Austin with verified emails" list on demand for any client vertical without building a new scraper per niche.
- SaaS sales teams — any B2B tool sold to local businesses (booking software, payment processors, review management) can target by category and city, and
has_booking_widgettells you who already has a competitor installed. - B2B lead-gen resellers — one actor covers any category, replacing dozens of niche scrapers.
- Franchise & market researchers — count and map competitor density for a category in a target city (set
onlyWithWebsite: falseto keep every mapped location, including ones with no contact details).
How to use it
The form is ready to run as it opens: hit Save & start without touching anything and you get Austin dentists that all have a website — a first run that takes a couple of minutes and costs cents. Measured on 2026-08-08, the untouched form (cap 25, Website required pre-ticked) delivers 25 businesses: 25 with a website, 25 with a phone, 10 with an email.
That number used to be 14. The cause was on our side, not the map's: only the two email filters deepened the candidate pool, so any run using Website required, Phone required or Social profile required filtered a pool sized for an unfiltered run and quietly came up short. Every filter that can drop a row now deepens the pool, and the filters that can be decided from the map data alone (website present, name keywords, permanently closed) are applied before any site is crawled, so no crawl budget is spent on a row that was never going to ship. Same-day measurements at cap 25 in Austin: onlyWithWebsite 14 → 25, requirePhone 10 → 25, onlyWithEmail 19 → 25.
Max businesses is still a ceiling, not a promise — a thin area really can run out. When that happens the run now says so in its status message with the counts behind it, instead of returning fewer rows without comment. A live example, dentist in Laramie, Wyoming at cap 25: "you asked for up to 25 and 1 could be delivered. That is the whole of Laramie, Albany County, Wyoming, United States, not a silent truncation: OpenStreetMap returned 4 'dentist' record(s) there, 1 became crawlable candidate(s), and 1 survived your filter(s) (onlyWithWebsite)." The run only claims an area is exhausted when no search came back at its row ceiling and every candidate was actually checked; otherwise it says which limit it hit and what to change.
- Pick a Business category from the dropdown, or leave it on Custom and type one in plain English (e.g.
medspa,HVAC contractor,funeral home). - Pick a City from the dropdown, or leave it on Custom and type any city on earth — city and region/country (e.g.
Austin, Texas). Capitalisation and the comma are optional:RABAT MAROCandcasablanca moroccoresolve too. - Leave Website required on unless you are doing density research — see the honest note below. (Web-design agencies can flip Businesses with no website instead to get the no-site prospect list; the two filters are mutually exclusive, and setting both fails the run with a clear error before anything is charged.)
- Run the actor. It geocodes the location, pulls matching places from OpenStreetMap, then crawls each business website for contact details, tech and reviews.
- Export as CSV, JSON, or Excel — or connect the dataset to your CRM/outreach tool via API/webhook.
Presets: the three guarantees
The preset table is at the top of this page. Behind it sit three guarantees, each of them tested:
- Anything you set yourself wins. A preset only fills in a field you left at its default. Set
maxScore: 100alongside the web-design preset and you keep 100 — the run log says so explicitly (left as you set them: maxScore (you chose 100)). - A preset never widens your bill. Every preset either adds a filter (fewer rows delivered, so fewer rows charged) or turns on enrichment that adds columns to rows you were already getting. None of them clears a filter you set or raises
maxItems. - It tells you what it did. One log line names the preset and lists exactly which settings it applied and which it left alone, and the run summary names the preset too.
Presets are Console and API: send "preset": "cold_email" from the API and you get the same behaviour. Existing API callers, tasks and schedules that send no preset are completely unaffected.
Which field wins: the precedence chain
Category and location each accept three inputs. Precedence is the same for both, highest first:
| Rank | Category | Location | Why |
|---|---|---|---|
| 1 | categories (list) | locations (list) | The plural list is an explicit multi-search request — it beats everything when non-empty. |
| 2 | categorySelect (dropdown) | citySelect (dropdown) | 20 categories with hand-verified OpenStreetMap tags; 25 metros verified against the live geocoder. |
| 3 | category (free text) | location (free text) | Anything else — several hundred more category terms are mapped, and any city on earth geocodes. |
Both dropdowns default to "" ("Custom"), which is why every input that worked before this option existed still resolves to exactly the same search.
Three search modes
| Mode | Set | What happens |
|---|---|---|
| City (default) | location (or locations) + category (or categories) | The location is geocoded and every matching business inside its bounding box is returned. Several categories x several cities run as a matrix in one run. |
| Radius | searchRadiusKm + centerLat + centerLon | Searches a circle around a coordinate instead of a city box — sales territories, franchise catchment areas, "everything within 5 km of this address". Replaces the bounding box entirely; no geocoding happens, so country stays empty. |
| Bring your own list | websiteList (or startUrls) | OpenStreetMap discovery is skipped entirely. The actor runs only the crawl + email verification + scoring pipeline over the websites you supply. Enrich a CRM export, a conference exhibitor list, or a list you bought elsewhere. |
When the city runs out: expandNearby
A single city often holds fewer businesses than you asked for — measured live: "dentist" in Austin, Texas tops out near 119 rows, and Round Rock, Texas holds 21. Until now the run could only tell you "try a larger area or a nearby city"; with expandNearby: true it does that for you:
- The same search is widened in growing rings around the city's centre (sized to the city, up to ~150 km) until your
maxItemsis met or the whole region is genuinely exhausted. - Every widened row is labelled: its
query_locationreadswithin ~16 km of Round Rock, Williamson County, Texas, United Statesinstead of the city name, so you can always tell expansion rows apart — or filter them out afterwards. - The same dedup and every pre-charge filter apply to ring rows, and the status message reports exactly how many delivered rows came from outside the city.
- Measured (2026-08-15, discovery-only):
dentist/Round Rock, Texas/maxItems: 100delivered 21 rows without the flag, 100 with it — 79 labelled ring rows, 0 duplicates.
It is off by default for API callers (existing inputs keep a byte-identical pull and bill) and pre-ticked in the Console form. It never fires when you set your own radius (searchRadiusKm) or bring your own list, when the city pull was truncated (deepening, not widening, is the fix there — the status message tells you), or during a mirror outage.
Searching several categories and cities at once
categories and locations are the plural versions of category and location, and they cross into a matrix — ["dentist","orthodontist"] x ["Austin, Texas","Dallas, Texas"] is four searches in one run. The single fields keep working exactly as before; the plural ones take precedence when non-empty.
Three things make the matrix safe rather than a footgun:
maxItemsis the whole run's budget, not a per-search one. The budget is split between the searches and results are interleaved, so the first city cannot eat the entire quota.- The same business found by two categories is delivered — and billed — once. Deduplication is on the OpenStreetMap object identity (
osm_type+osm_id), so a clinic tagged bothdentistandorthodontistappears one time. - 25 category x location combinations is the ceiling for one run. Above it the run fails immediately with a message naming the numbers, before any network request and before any charge — OpenStreetMap's public mirrors are a free shared resource.
Every row carries query_category and query_location, so you always know which search produced it.
Bring your own list (skip discovery)
Already have the businesses and only need the emails, socials, platform and score? Put the domains in websiteList (bare domains and full URLs both work; startUrls is accepted as an alias):
{ "websiteList": ["aloha-dental.com", "https://www.averyranchdental.com", "typotes.com"] }
No map data is fetched at all. Those rows differ from discovered rows in exactly three honest ways:
sourceisuser_supplied, notOpenStreetMap;attributionis null — the rows are not OSM-derived, so stamping the ODbL notice on them would be a false licence claim. The column is still present, so a mixed export keeps one stable header row;namecomes from the site's own<title>, falling back to the bare domain when the site does not answer. Nothing is invented;latitude,longitude,osm_idandosm_urlstay empty.
Everything else — the contact crawl, MX verification, mailbox-owner classification, tech fingerprinting, scoring and every filter — behaves identically.
How complete is the data? (measured, not estimated)
Most scrapers show you a perfect sample row and let you discover the gaps after you have paid. Here are the real fill rates from two reference runs of dentist in Austin, Texas on the current build — one with the website filter on (the whole city, n=55), one with every filter off (n=100):
onlyWithWebsite: true (n=55) | schema defaults, filter off (n=100) | |
|---|---|---|
phone | 96% | 48% |
website | 100% | 48% |
email | 55% | 26% |
website_platform | 71% | 34% |
opening_hours | 80% | 44% |
facebook | 73% | 33% |
| rows graded F | 0% | 51% |
On that same filter-off n=100 run, the crawl-derived fields measured: contact_page_url 18%, website_title 40%, website_description 35%, mobile_viewport 40%, copyright_year_stale flagged on 3 rows; google_maps_url, osm_url and attribution sat at 100%.
With onlyWithWebsite: true the email fill runs far higher than the filter-off numbers — the reference run above measured 55% for dentists, and a roofing contractor / Denver run delivered emails on 6 of 8 rows (75%): roughly three-quarters on website-verified trade categories.
Read that second column before you run. OpenStreetMap has no website for a large share of businesses, and — measured on the filter-off run — 50 of the 52 rows with no website had no phone, no email and no social profile either (the other 2 carried only an OSM phone): name, coordinates and usually an address, nothing contactable. With the filter off you pay for those rows. onlyWithWebsite and onlyWithEmail drop non-matching rows before you are charged, so they cut your bill rather than just tidying the output. Every run's status message now reports the contactable ratio it actually delivered.
Every filter now goes further: the actor over-fetches, deepening the candidate pool and — for filters that need the site crawled — crawling extra candidates in batches (up to 6× maxItems, the ceiling that makes a thin area terminate) until it has maxItems surviving rows or the pool is spent. This used to apply to onlyWithEmail / onlyVerifiedEmail only, which is why onlyWithWebsite, requirePhone and requireSocial quietly returned short. Measured 2026-08-08 on dentist / Austin, Texas at cap 25, one filter at a time: onlyWithWebsite 14 → 25, requirePhone 10 → 25, onlyWithEmail 19 → 25. Where the pool genuinely runs out first, the status message reports the counts instead of leaving you to guess — e.g. dentist in Laramie, Wyoming returns 1 row and says the map holds 4 dentist records there, 1 of them with a website.
The default is false (not true) so that existing API callers, scheduled tasks and density-research use cases keep getting every mapped location. The Apify console pre-fills it to true.
There is also the mirror filter, onlyWithoutWebsite: keep only businesses that list no website — the prospect list for web-design agencies pitching a first site (verified: a 15-row run delivered 15/15 rows without websites). It is mutually exclusive with onlyWithWebsite; setting both fails the run with a clear error before anything is charged. Expect these rows to be name + address + coordinates (and occasionally a phone) — with no site to crawl, no email/platform enrichment is possible.
Which categories does OpenStreetMap actually cover well?
This is the honest limit of an OSM-based source, and it matters more than any field:
- Dense: businesses with premises a mapper walks past — restaurants, cafés, dentists, pharmacies, hairdressers, gyms, hotels, shops, banks, clinics. (171 dentist records in the Austin bounding box.)
- Sparse: trades run from a van or a home office — plumbers (9 in the same box), electricians (5), chiropractors (0), roofers (~12). A metro of a million people can return single digits. That is what is mapped, not a bug.
~50 categories are curated and mapped to exact OSM tags — build 0.1.21 added pest control, photographer, moving company, self storage, funeral home and dry cleaner. Any other term is attempted as an OSM tag directly, and common phrasings are handled (auto repair shop → shop=car_repair, landscaping company → craft=gardener, insurance agency → office=insurance). Terms with no OSM tag now fall back to a business-name keyword search of OSM, and a mapped category that returns zero tagged places in the area is rescued by the same name-keyword query. Rows found only by name are labelled category_match: "name_keyword" (tag-matched rows say category) so you can filter them out if you only trust tag-confirmed rows. Measured: pest control in Denver returned 1 labelled row on 0.1.21 where the previous build returned 0 — sparse trades stay sparse, but no longer invisible. A term matching nothing at all still returns 0 rows with a suggestion and no charge — you are never billed for a run that found nothing.
Filters — every one of them runs before you are charged
Filters here are not a tidying step applied to an invoice you have already run up: a filtered row is dropped before it is pushed to the dataset and before the charge, so filtering cuts your bill.
Filtering does not shrink your delivery. Every filter in this table deepens the candidate pool to compensate, so maxItems means "this many rows I can use", not "this many businesses considered". Filters decidable from the map data alone — onlyWithWebsite, onlyWithoutWebsite, excludeKeywords, skipClosed — are applied before any website is crawled, so no crawl budget is spent on a row that was never going to ship. The rest are applied batch by batch as sites are crawled, stopping the moment enough rows survive. The pool is bounded at 6x maxItems candidates so a thin area terminates instead of crawling forever; when that bound or the map itself is what stopped the run, the status message says so and gives you the counts.
| Filter | Keeps only |
|---|---|
onlyWithWebsite | businesses that list a website (recommended, see the fill-rate table) |
onlyWithoutWebsite | businesses with no site — the first-website pitch list. Mutually exclusive with the above |
onlyWithEmail | rows where an email was found |
onlyVerifiedEmail | rows whose email passed MX verification (deliverable / risky) |
requirePhone | New. rows with a phone number (from OSM, a tel: link, or the site's schema.org markup) |
requireSocial | New. rows with at least one Facebook / Instagram / LinkedIn / X / YouTube profile |
requireAnyContact | New. rows reachable on at least one channel — phone or email or a social profile. The loosest contactability floor there is, and the one to reach for when you do not care which channel. It matters more than it sounds: on a bare {} run of the default search (100 rows, measured 2026-08-08) exactly 50 rows carried no phone, no email and no social profile at all, and without this switch you are billed for them. Off by default, so no existing run changes |
excludeKeywords | drops businesses whose name contains any of your keywords (case-insensitive). Only the name is matched, so clinic cannot knock out a business on Clinic Street. Use it to strip chains, franchises or your existing customers |
skipClosed | New. drops permanently-closed premises (see below) |
minScore / maxScore | maxScore is new. A score ceiling is the web-design agency filter: a low score means a thin online presence, which is exactly the redesign pitch list |
minRating / minReviewCount | New — read the warning below before using these |
skipClosed: what it actually removes
OpenStreetMap mappers retire a business without deleting it: the primary tag moves from amenity=restaurant to disused:amenity=restaurant, so the object keeps its name and address but no longer describes an operating business. skipClosed drops elements carrying a lifecycle-prefixed primary tag (disused:, abandoned:, was:, removed:, demolished:, razed:), a disused=yes / abandoned=yes flag, opening_hours=closed, or shop=vacant. A disused: tag alongside a live tag of the same kind (a former bank that is now a café) is not treated as closed.
Measured live on 2026-08-08 in the Austin bounding box: 199 such elements, 43 of them still carrying a business name — e.g. Chago's with disused:amenity=restaurant, Corner Store with abandoned:shop=fuel. These cannot reach a normal tagged search (["amenity"="restaurant"] cannot match disused:amenity), but they do reach the business-name fallback used for unmapped categories, which is where dead businesses were being delivered as fresh leads. Verified end to end: with skipClosed: false that closed restaurant is delivered and billed; with it true it is dropped before billing.
It is off by default so existing runs, tasks and API callers are unchanged. The Apify console pre-fills it to true.
Honest warning about
minRating/minReviewCount: OpenStreetMap carries no review data at all. A rating only exists when the business publishes schema.orgaggregateRatingon its own website — measured at roughly 3-8% of rows. So when you set a rating or review floor, rows with no rating are DROPPED, not kept: an unknown rating is not a passing rating, and nothing is ever invented to save a row. Setting either filter will cut your result count to a small fraction — an Austin restaurant run asking for 15 rows withminRating: 4.0crawled 86 candidates looking for them and delivered 0 rows, charging nothing (re-measured 2026-08-08 after the over-fetch was generalised, so this is the deep-search result, not a shallow one). If you need ratings on every business this is the wrong source; pair it with our Google Maps Places Scraper, or use thegoogle_maps_urlon every row.
Choosing your columns and sort order
outputFields— pick the columns you want (e.g.["name","email","phone","website","lead_grade"]) and the dataset carries only those. Every row still shares one identical column tuple, so CSV headers never shift mid-export, and the columns keep the documented ROW order regardless of the order you listed them in.nameandattributionare always included whatever you choose —attributionbecause the OpenStreetMap licence (ODbL) has to travel with the data into your CSV. Column names are matched case-insensitively (and-/space count as_, soLead Gradeworks). A single unknown name is logged and ignored; if none of the names you list exists, the run stops before billing rather than charging you full price for a name-only export.sortBy—score_desc(default, what every previous build did),name_asc,review_count_descorrating_desc. Sorting never removes a row; all filtering already happened. Rows missing the sort value (no rating, no review count) are placed last, because a missing value is unknown rather than zero.
Output fields
61 columns on every row (57 before this build; the four new ones are query_category, query_location, email_guess and email_guess_confidence). Re-measured on a 100-row dentist / Austin, Texas run with schema defaults: all 100 rows carried the same stable 61-column tuple, so CSV headers don't shift mid-export — and if you narrow the export with outputFields, every row still shares one identical (smaller) tuple. Fill rates are from the onlyWithWebsite: true reference run above; anything conditional says so. Fields marked New (0.1.21) show fill rates from the 100-row filter-off benchmark instead — read them accordingly, since roughly half of those rows had no website to crawl.
In the Console, the Output tab opens on the Overview view; the All columns tab shows every field listed below, in this order. Views only change what the Console table renders — a CSV, JSON or Excel export always carries every column the run produced (or exactly the ones you named in outputFields).
Identity & location — from OpenStreetMap
| Field | Fill | Description |
|---|---|---|
name | 100% | Business name |
category | 100% | OSM category tag (e.g. dentist, hairdresser, lawyer) |
category_match | 100% | New (0.1.21). How the row matched your category: category = matched the mapped OpenStreetMap tag; name_keyword = found by the business-name fallback search (used for unmapped terms, and as a rescue when a mapped tag returns zero places in the area). Filter on category if you only want tag-confirmed rows |
address | 96% | Street address assembled from OSM address tags |
city / state / postal_code | 93% / 93% / 93% | Address components |
country | 100% | From the geocode. Empty in radius mode and on user-supplied rows (no geocode happens there) |
query_category | 100% | New. Which of your input categories produced this row — the column that makes a multi-category run readable. Null on user-supplied rows |
query_location | 100% | New. Which of your input locations produced this row (the geocoder's resolved name, or the radius description). Null on user-supplied rows |
latitude / longitude | 100% | Coordinates |
osm_type / osm_id | 100% | OpenStreetMap source identifiers |
google_maps_url | 100% | New (0.1.21). A Google Maps link for the business — one click to the live listing with its current rating and review count (data this actor deliberately does not scrape; see the rating warning below). Need those fields at scale? Pair with our Google Maps Places Scraper |
osm_url | 100% | New (0.1.21). Link to the row's source OpenStreetMap element — instant provenance, and the place to fix bad map data |
Contact
| Field | Fill | Description |
|---|---|---|
phone | 96% | Primary phone. OSM tag first; falls back to a tel: link or JSON-LD telephone only when OSM has none |
phones | 73% | New. All tel: numbers found on the site. May include a call-tracking number — phone stays the trusted value |
website | 100% | Business website URL. A social page in OSM's website tag is routed to that social column instead, so this is always a real site |
domain | 100% | New. Bare registered domain — the field CRMs dedup on |
email | 55% | Primary contact email, chosen by mailbox ownership then deliverability |
emails | 55% | All emails for the row, primary first (previously excluded a primary that came from OSM) |
contact_page_url | 42% | New. The exact page the primary email was found on, so you can spot-check it. Null when the email came from OpenStreetMap rather than from a crawled page |
email_guess | 0% unless opted in | New, opt-in, and deliberately not an email. With emailPatternGuess: true, a business that has a website but publishes no address anywhere we crawled gets a pattern-guessed info@<domain> here. It is never promoted into email, never counts as has_email, never earns a lead-score point and never satisfies onlyWithEmail / onlyVerifiedEmail. Nobody checked that this mailbox exists — treat it as a lead, not an address |
email_guess_confidence | same | New. A statement about the domain, never the mailbox: mx_ok (the domain does run mail servers), no_mx (it does not — the guess is almost certainly dead), unchecked (email verification is off, so no lookup was made) |
Email quality — verification is included, not an add-on
| Field | Fill | Description |
|---|---|---|
email_status | 100% | deliverable / risky / undeliverable when verification is on and a candidate exists; missing when no email was found; found / missing when verifyEmails is off. Turning verification off changes the vocabulary |
email_provider | 55% | Mailbox host where identifiable (Google Workspace, Microsoft 365...). Only when verifyEmails is on and an email was found |
email_domain_match | 55% | Whether the email's domain matches the website's |
email_type | 55% | own_domain / free_mail / third_party / unknown. third_party means the address belongs to the business's marketing agency or web designer — it is deliverable but does not reach the business, and it is scored at half weight. Measured on the reference run: 12% of all harvested addresses were third-party, but only 10% of primary emails, because candidates are ranked by mailbox ownership before one is promoted |
email_types | 55% | The same classification for every entry in emails |
Socials
| Field | Fill | Description |
|---|---|---|
facebook | 73% | Facebook profile URL |
instagram | 55% | Instagram profile URL |
youtube | 25% | New. YouTube channel URL |
linkedin | 22% | New. LinkedIn company/profile URL |
twitter | 18% | New. X/Twitter profile URL |
Share buttons, tracking pixels and embedded posts are filtered out, so these are profile URLs rather than facebook.com/tr or instagram.com/p/....
Website & marketing tech
| Field | Fill | Description |
|---|---|---|
website_platform | 71% | CMS / site builder: WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, Custom / Next.js, Custom / React. WordPress plugins (Elementor, WP Rocket, Divi) report as WordPress, not as their own platform |
website_platform_status | 100% | New. Why website_platform is what it is: detected, unknown (site loaded, no fingerprint), site_unreachable, no_website, not_crawled. A null platform is explained rather than unexplained |
platform_version | 27% | New. Only when the site's own meta generator names the platform and a version — never a plugin's version passed off as the platform's |
site_generator | 42% | New. The raw meta name="generator" string |
tech | 84% | New. Marketing/booking/commerce tech found on the page (Meta Pixel, Google Analytics, GTM, Calendly, NexHealth, Klaviyo, WooCommerce, Yelp Reviews, live chat...) |
has_meta_pixel / has_google_analytics / has_booking_widget | see note | New. Booleans derived from tech. Tri-state: true/false when the site was read, null when it was unreachable or not crawled — false never means "we could not check" |
website_title | 40% | New (0.1.21). The crawled homepage's <title> — a drop-in mail-merge personalization field. Also a data tripwire: a title that clearly names a different business means OSM's website tag is wrong, and this column lets you catch that before you hit send |
website_description | 35% | New (0.1.21). The homepage's meta description — the other personalization field, and the same tripwire |
mobile_viewport | 40% | New (0.1.21). Whether the homepage declares a mobile viewport tag. A site without one predates responsive design — a concrete web-agency pitch signal |
copyright_year_stale | 3% | New (0.1.21). Flags a visibly out-of-date copyright year in the footer — a small but unambiguous neglected-site signal (flagged on 3 of the 100 benchmark rows) |
Business detail & reviews
| Field | Fill | Description |
|---|---|---|
opening_hours | 80% | From OSM; gap-filled from the site's JSON-LD only when OSM has none |
rating | 7% | New, and sparse — see the warning below. Star rating (1-5) the business publishes in its own schema.org markup |
review_count | 7% | New, sparse. Review count from the same markup |
price_range | 24% | New. schema.org priceRange (e.g. $$) |
Honest warning about
rating/review_count: these are not Google Maps ratings. They are only present when a business publishesaggregateRatingin its own JSON-LD, and most do not. Measured on-platform: 7% of website-bearing rows for dentists in Austin (4/55) and 13% for roofing contractors in Denver (1/8); the filter-off benchmark (n=100) measured 4%. An earlier 12-row roofer sample hit 33%, so the rate swings wildly with category, city and sample size — assume under 10% and treat anything higher as luck. Do not build a workflow that needs a rating on every row. If you need star ratings and review counts for every business, this is the wrong source — pair it with our Google Maps Places Scraper; every row'sgoogle_maps_urlalso takes you to the live listing in one click. OpenStreetMap carries no review data at all, and this actor never invents a substitute:lead_scoreis a data-completeness score, not a customer rating.
Lead scoring
| Field | Fill | Description |
|---|---|---|
lead_score | 100% | 0-100 completeness/reachability score |
lead_grade | 100% | A ≥80, B ≥65, C ≥50, D ≥35, else F |
lead_tier | 100% | hot ≥75, warm ≥50, else cold |
score_breakdown | 100% | Per-component points, plus signals_from — the list of extra signals that actually scored — so the score is auditable |
has_email / has_phone / has_website | 100% | Booleans for quick filtering |
Scoring rubric (sums to a true 100, so grade A is reachable — the filter-off benchmark's top row scored the full 100): email deliverable 40 / risky 25 / present-but-unverified 15 — halved when email_type is third_party; phone 20; website 15; socials 5 each capped at 10; extra signals 5 each capped at 15, drawn from website platform detected, opening hours, marketing tech found, and a published star rating — score_breakdown.signals_from names the ones that counted.
Provenance
| Field | Fill | Description |
|---|---|---|
source | 100% | OpenStreetMap for discovered rows, user_supplied for rows that came from your own websiteList |
attribution | 100% on OSM rows | The ODbL attribution string, so the licence travels with an exported CSV. Null on user_supplied rows — they are not OpenStreetMap data, so attaching the notice would be a false licence claim. The column is always present either way |
enriched_from_website | 33% | New. Lists only the fields where an OSM-overlapping value was taken from the site instead — rating, review_count, price_range (OSM carries none of these) plus phone / opening_hours / email where OSM was empty and the site's schema.org data filled the gap. It is not a full provenance map: emails, the five socials, website_platform, platform_version, site_generator and tech are always crawled from the site and are deliberately not repeated here |
Example output
Real rows from a run with onlyWithWebsite: true. This is a best-case slice, not a typical one — see the measured fill rates above; on the website-filtered reference run roughly half of rows carry an email.
| Name | Phone | Status | Type | Platform | Tech | Score | |
|---|---|---|---|---|---|---|---|
| Aloha Dental | +1-512-707-7300 | riverside@aloha-dental.com | deliverable | own_domain | WordPress | Meta Pixel, GTM, Yelp | 95 (A) |
| Avery Ranch Dental | +1-512-246-7645 | smile@averyranchdental.com | deliverable | own_domain | WordPress | GTM, reCAPTCHA | 95 (A) |
| Aviva Dental Care | +1 512 852 8528 | dr.apurva@avivadentalcare.com | deliverable | own_domain | WordPress | Google Analytics | 90 (A) |
Input examples
A preset plus the two dropdowns — the shortest useful input there is:
{"preset": "cold_email","categorySelect": "hair salon","citySelect": "Casablanca, Morocco","maxItems": 100}
The preset sets onlyWithEmail and onlyVerifiedEmail; everything else stays at its default. Add any field you want and yours wins — {"preset": "web_design", "maxScore": 100} keeps maxScore: 100 and the log says which setting the preset stood down on.
One category, one city — the classic run, unchanged:
{"category": "hair salon","location": "Miami, Florida","maxItems": 200,"crawlEmails": true,"onlyWithWebsite": true,"onlyWithEmail": false,"verifyEmails": true,"maxPagesPerSite": 3,"concurrency": 8}
Several categories across several cities, contactable rows only, narrow columns:
{"categories": ["dentist", "orthodontist"],"locations": ["Austin, Texas", "Dallas, Texas"],"maxItems": 200,"onlyWithWebsite": true,"requirePhone": true,"excludeKeywords": ["Aspen Dental"],"skipClosed": true,"outputFields": ["name", "email", "phone", "website", "lead_grade", "query_location"],"sortBy": "name_asc"}
A 5 km sales territory around one address:
{"category": "restaurant","searchRadiusKm": 5,"centerLat": 30.2672,"centerLon": -97.7431,"maxItems": 150,"onlyWithEmail": true}
Web-design prospect list — businesses with a site, but a weak one:
{"category": "hair salon","location": "Lyon, France","countryCode": "fr","onlyWithWebsite": true,"maxScore": 45,"skipClosed": true}
Enrich your own list (no map data fetched at all):
{"websiteList": ["aloha-dental.com", "averyranchdental.com", "typotes.com"],"verifyEmails": true,"emailPatternGuess": true}
How much does it cost?
Pay-per-result — the current per-lead rate is on this page's Pricing tab, and you are charged for delivered businesses only, not for API calls or compute. No subscription. New Apify users get platform free credits to test with.
A run's total is simply rows delivered x the per-lead rate. Be aware what the rows contain: on the reference run with onlyWithWebsite: true, ~55% carried an email, so 500 rows ≈ 275 emailable leads (roughly 1.8x the per-lead rate per emailable lead). With the website filter off, 26% carried an email on the filter-off benchmark (roughly three-quarters on website-verified trade categories). Filters run before billing, so onlyWithEmail: true is the cheapest way to buy emails specifically — and the adaptive over-fetch delivers as close to maxItems email rows as the city allows (measured 2026-08-08: asked 25, delivered 25 in Austin; the same input delivered 19 before the pool-depth fix, and the whole city tops out at 31, which the run tells you when you ask for more).
For comparison, the largest competing Maps-based scraper charges $0.004/place plus $0.002 for contacts plus $0.004 per record for email verification — about $0.010 at feature parity. Here, MX verification, mailbox-ownership classification and lead scoring are all included in the single per-lead rate.
Data source & licence
Business listings come from OpenStreetMap. © OpenStreetMap contributors, available under the Open Database License (ODbL) v1.0 — https://www.openstreetmap.org/copyright. Every row carries source and attribution fields; keep them if you redistribute or publish the data, as ODbL requires attribution.
The website-crawled fields — email, emails, socials, website_platform, tech, rating, review_count, price_range, phones — are not OSM-derived. They come from each business's own public website and are not covered by ODbL. Those fields are site-derived on every row. The enriched_from_website column is narrower than that: it flags only the fields that OSM could have supplied but didn't, so treat the list above — not that column — as the ODbL boundary.
Frequently asked questions
Is it legal to scrape local business data?
This actor reads publicly available data from OpenStreetMap (an open-data project, ODbL-licensed) and each business's own public website — the same information anyone could find by visiting the site. No login, no private data, no anti-bot circumvention, no Google Maps terms-of-service exposure.
Does it work outside the US? How should I type the location?
Yes — anywhere OpenStreetMap covers. Type the location however you like: the raw string is tried first, and only if that finds nothing is it automatically re-spelled (Title Case, and a comma inserted before the trailing country word) before the run is given up on. RABAT MAROC, Rabat Maroc, casablanca morocco and Rabat, Morocco all resolve to the same place. If the run still cannot geocode, the status message lists every spelling it tried instead of a generic failure.
If your city name exists in more than one country — Rabat is a city in both Morocco and Malta, Cambridge in both the UK and the US — set countryCode to an ISO 3166-1 alpha-2 code (ma, mt, us, fr, gb) to pin the geocoder. Verified: location: "Rabat" with countryCode: "mt" resolves to Rabat, Western Region, Malta; without it, Morocco wins. The pin is enforced: the fallback geocoder (Photon) has no country filter, so it is skipped whenever countryCode is set. A city that does not exist inside that country ends the run with a clear message and no charge, instead of quietly returning the same-named city elsewhere (location: "Austin" + countryCode: "ma" used to deliver Austin, Texas).
One non-US bug is fixed as of this build: the geocoder is now asked for English place names. It previously answered in the local language(s), and since the row's country column is derived from the geocoder's answer, every Moroccan row used to ship country: "Maroc ⵍⵎⵖⵔⵉⴱ المغرب" — a three-script run-on where Morocco belonged. Measured and fixed on 2026-08-08; Austin, Texas and Lyon, France are unaffected. Note that city and address still come straight from the OpenStreetMap tags a local mapper wrote, so a Moroccan row can legitimately read Témara تمارة — that is the map's own data and it is not rewritten.
How fresh is the data?
Business listings come from OpenStreetMap, which is community-maintained and updated continuously; most established businesses have accurate name/address/phone data. Website content (email, socials, platform, tech, ratings) is crawled live on every run, so that part is current at run time.
Why is email empty on some rows?
Three distinct reasons, and email_status plus website_platform_status tell you which: the business has no website in OSM (no_website), its site did not respond (site_unreachable), or the site simply does not publish an address anywhere on the pages crawled (missing). Cloudflare-obfuscated addresses are decoded, and since 0.1.21 the crawler follows the site's own contact/about/team links — including Shopify /pages/contact and non-English slugs — within the maxPagesPerSite budget. Raising maxPagesPerSite finds a few more.
Does it get star ratings and review counts?
Only when a business publishes them in its own website markup — 7-33% of website-bearing rows depending on the category (4% on the latest filter-off benchmark). See the warning in the output section. OpenStreetMap has no review data, so if ratings are essential for every row, pair this with our Google Maps Places Scraper — every row's google_maps_url opens the live Google listing in one click.
Why does website platform detection matter?
Knowing a business runs on Wix, GoDaddy or Weebly (often a dated, self-built site) versus WordPress or a custom build lets agencies pre-qualify leads before picking up the phone. Combined with has_meta_pixel: false and has_booking_widget: false, you get businesses that are visibly under-invested in their web presence — a far better redesign pitch than a cold list.
Do I need an API key or proxy?
No. Discovery runs on OpenStreetMap (Nominatim + Overpass) and enrichment crawls public websites directly — no API keys, no proxy configuration, and it runs from datacenter IPs. Since 0.1.21, Overpass mirrors are raced in pairs (at most two in flight) instead of tried strictly in sequence, so one slow mirror no longer stalls discovery. A small share of sites (measured: 8 of the 55 website-bearing rows in the reference run) block or fail to answer datacenter requests; those rows are marked site_unreachable rather than silently blamed on missing data.
Other Flash Scrape lead tools
- Google Maps Places Scraper — live Google Maps rating, review count, phone and website per place: the honest companion for the review data this actor deliberately does not invent. Every row's
google_maps_urllinks the two. - Restaurant Leads Scraper — restaurant leads enriched with POS, reservation and delivery tech stack (Toast, OpenTable, DoorDash, and more).
- Google Maps Leads Opener — pull business leads directly from Google Maps search results.
- Email Verifier — bulk-validate and clean an existing email list before you launch a campaign.
- Google Ads Transparency Scraper — see every Google ad a local competitor runs, with first/last shown dates: the other half of a local-market picture.
- Company Domain Enricher — turn a company name into its website domain and firmographic details.