Local Business Leads Scraper 🚀- MX-Verified Emails & WordPress avatar

Local Business Leads Scraper 🚀- MX-Verified Emails & WordPress

Pricing

from $3.00 / 1,000 business leads

Go to Apify Store
Local Business Leads Scraper 🚀- MX-Verified Emails & WordPress

Local Business Leads Scraper 🚀- MX-Verified Emails & WordPress

Finds local businesses of any category in any city worldwide (dentist, roofer, gym). Crawls each site for MX-verified emails, 5 socials + website platform (WordPress, Wix, Squarespace, GoDaddy, Shopify) - ideal for web design agency leads. Flags agency-owned mailboxes. Scores 0-100 A-F. $3/1k.

Pricing

from $3.00 / 1,000 business leads

Rating

5.0

(3)

Developer

Flash Scrape

Flash Scrape

Maintained by Community

Actor stats

4

Bookmarked

13

Total users

4

Monthly active users

an hour ago

Last modified

Share

Local Business Leads Scraper — any category, any city, with MX-verified emails and lead scoring

Any local business category, in any city on earth — and every row comes back with an MX-verified email, whose mailbox it is, the website platform the business runs, and a 0-100 lead score. Priced per delivered lead (current rate on this page's Pricing tab), with email verification included rather than sold as an add-on.

Dentists, gyms, lawyers, plumbers, roofers, salons, real estate agencies, restaurants and 50+ more curated categories — plus any term you type. No API key, no proxy, no separate scraper per niche.

What you get

  • Any category, any city — ~50 categories mapped to exact OpenStreetMap tags, any other term attempted directly and then rescued by a business-name search, and any city on earth geocoded (Austin, Texas, casablanca morocco, Dubai UAE).
  • Emails that are verified, not guessed — every address is MX-checked over DNS-over-HTTPS and graded deliverable / risky / undeliverable. Verification is part of the price, not an upsell.
  • Whose mailbox it isown_domain, a free inbox, or the business's marketing agency. Emailing an agency mailbox never reaches the business, so this column decides whether a lead is worth a send.
  • Redesign pitch signalswebsite_platform (WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more), mobile_viewport, copyright_year_stale, has_meta_pixel. A builder-tier site with no tracking is the highest-intent pitch there is.
  • A score you can audit — 0-100, an A-F grade, a hot/warm/cold tier, and a score_breakdown object showing the arithmetic that produced it.
  • Five ready-made Output views — Overview, Email-ready, Web-agency targets, Map & source, All columns. Pick one above the results table instead of scrolling a wall of 61 raw columns.
  • Filters that cut your bill, not just the table — every filter drops the row before it is pushed and before it is charged, and the run log names each filter and how many rows it removed.
  • One-click provenancegoogle_maps_url and osm_url on every row, plus the ODbL attribution the licence requires you to keep.

Measured, not promised

Every number below is from a real run of this Actor on dentist / Austin, Texas. Inputs and dates are given so you can reproduce them.

RunResult
Console form untouched (cap 25, Website required on), pressed Save & start — 2026-08-0825 of 25 asked for — 25 with a website, 25 with a phone, 10 with an email. The same input delivered 14 before the pool-depth fix described under How to use it
Console form as it opens, Max businesses raised to 100, pressed Save & start — 2026-08-0855 businesses, 100% contactable — 96% with a phone, 55% with an MX-verified email. Under the 100 asked for, and the run says why: OpenStreetMap holds 171 dentist records in Austin and 55 of them are crawlable with a website inside the search area. That is the whole city, not a truncation
Bare {} from the API (no filters at all) — 2026-08-08100 businesses, 50% contactable — the honest filter-off case, see the fill-rate table below
Cold-email list preset — 2026-08-0831 rows (every emailable dentist Austin has — the exact count varies by city and over time), 100% with an MX-verified email, 94% with a phone
Call list preset, cap 100 — 2026-08-0856 rows, 100% with a phone — the same number on two independent runs, and the run reports it as the whole of Austin rather than a truncation
Web-design prospects preset, cap 100 — 2026-08-087 rows, every one scored ≤45 and reachable on at least one channel. The score ceiling alone matches 68 businesses in Austin, but only 7 of those publish a phone, email or social profile — the preset drops the other 61 rather than bill you for map pins you cannot contact
Full enrichment preset — 2026-08-08100 rows plus 23 email_guess addresses, kept out of email and never billed as verified
Column stability across every run aboveone identical 61-column tuple on every row — CSV headers never shift mid-export

The gap between the Console-form rows and the bare {} row is the whole story of this Actor's honesty: the Console form pre-fills onlyWithWebsite, which is why it delivers 100% contactable businesses, while a bare API call keeps every mapped location including the ones with nothing to contact. Both numbers are published; neither is hidden. The same goes for the 55-of-100 line: a run that cannot reach your cap says so in its status message, with the counts that explain it.

The Console form opens at 25 businesses so your very first run finishes in a couple of minutes and costs cents — raise Max businesses once you like what comes back. That number is a Console pre-fill only: the API default is still 100, so existing API callers, tasks and schedules that send no maxItems are completely unaffected.

Quick start

{
"category": "dentist",
"location": "Austin, Texas"
}

That is a complete run — and it is also exactly what the form already contains. Open the Actor, press Save & start without touching anything, and you get businesses. Everything else has a working default: the Console form starts at 25 businesses (the API default stays 100), website crawling on, email verification on, 3 pages per website.

The input form at a glance

The sections in bold are the ones worth a look. The other two — radius search and crawler tuning — are technical, already tuned, and safe to ignore on a first run.

SectionWhat it is for
Quick startOne dropdown that configures everything else — see the presets below
What to searchThe business category and the city (one of each, or several at once)
Search a radius around a pointTechnical. A circle around a coordinate instead of a city
Bring your own listYou supply the websites; OpenStreetMap discovery is skipped
LimitsHow many businesses this run may deliver — your cost ceiling — and expandNearby, which widens the search around the city when the city itself runs out (each widened row labelled in query_location)
Email + platform enrichmentWhat gets read from each business website
Crawler tuningTechnical. Speed-versus-depth knobs for the crawl
FiltersDrop businesses before you are charged for them
OutputWhich columns you get, and in what order

Presets

The Use-case preset dropdown at the top configures the rest of the form for one specific outreach job. It is optional — leave it on Custom and nothing at all changes.

PresetWhat it setsMeasured on the default search (Austin dentists)
Cold-email listonlyWithEmail, onlyVerifiedEmail31 leads (the whole of Austin; the count varies by city and over time), 100% with an MX-verified email, 94% with a phone
Web-design prospectsmaxScore: 45, requireAnyContact7 leads, all in the thin/neglected half — no published email, no marketing tech, often no mobile viewport — and all reachable. maxScore: 45 on its own matches 68 businesses; the contact floor is what removes the 61 you could not pitch to
Call listrequirePhone56 leads, 100% with a phone number
Full enrichmentemailPatternGuess, maxPagesPerSite: 6 (every crawl toggle is already on by default)100 leads plus 23 guessed email_guess addresses, kept out of email

Anything you set yourself wins, no preset ever widens your bill, and the run log names exactly what each preset applied and what it stood down on. The three guarantees are spelled out — and tested — under Presets: the three guarantees below.

What the Output tab looks like

Five curated views ship with the Actor. Pick one above the results table; the CSV / JSON / Excel export is unaffected and always carries every column you asked for.

ViewColumnsUse it for
Overview (default tab)name, category, city, phone, email, website, lead_grade, lead_scoreThe eight columns that answer "is this a lead?"
Email-readyname, email, email_status, email_type, phone, website, contact_page_url, lead_scoreLoading a cold-email sequence — deliverability grade and mailbox owner side by side
Web-agency targetsname, website, website_platform, mobile_viewport, copyright_year_stale, has_meta_pixel, phone, emailBuilding a redesign pitch list from the neglect signals
Map & sourcename, address, google_maps_url, osm_url, latitude, longitude, attributionVerifying a row in one click, territory mapping, keeping the ODbL attribution with the data
All columnsall 61, in CSV orderEverything, when you want the full table

Two deliberate choices in those views. website, contact_page_url, google_maps_url and osm_url render as clickable links. And the tri-state flags — mobile_viewport, copyright_year_stale, has_meta_pixel — render as text, not as a checkbox, because they are null when the site was never crawled and a checkbox cannot tell "no mobile viewport" apart from "we never looked". has_email / has_phone / has_website are never null, so those do get the real boolean widget.

What does it do?

This actor takes a plain-English business category (e.g. "dentist," "hair salon," "roofing contractor") and a location, maps the category to the right OpenStreetMap tags, and pulls every matching business in the area. It then crawls each business's own website — following the site's own contact/about/team links (including Shopify /pages/contact and non-English slugs) within the maxPagesPerSite budget — and extracts, from pages it has already downloaded:

  • a contact email, MX-verified over DNS-over-HTTPS and graded deliverable / risky / undeliverable
  • whose mailbox it is — the business's own domain, a free inbox, or its marketing agency (emailing that one never reaches the business)
  • social profiles — Facebook, Instagram, LinkedIn, X/Twitter, YouTube
  • which website platform it runs on — WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Joomla, Drupal, HubSpot CMS, and a "Custom / Next.js" bucket for hand-built sites
  • marketing & booking tech — Meta Pixel, Google Analytics/GTM, Calendly, NexHealth, Klaviyo, WooCommerce and ~25 more
  • the homepage's title and meta description — drop-in mail-merge personalization fields (and a data tripwire: a title naming a different business exposes a wrong OSM website tag)
  • web-agency pitch signals — a missing mobile viewport tag, a stale footer copyright year
  • one-click source linksgoogle_maps_url to the live Google listing, osm_url to the OpenStreetMap element
  • a lead score 0-100, an A-F grade, and a hot/warm/cold tier

Why use it / who's it for

  • Web design & marketing agencies — filter for a builder-tier platform (Wix, GoDaddy, Weebly) and has_meta_pixel: false to find businesses with a dated site and no tracking: the highest-intent redesign pitch there is. The Custom / Next.js value is the inverse signal — don't pitch a DIY-site rebuild to someone who already paid a developer. Build 0.1.21 adds two more neglect signals — mobile_viewport (missing = the site predates responsive design) and copyright_year_stale — plus an onlyWithoutWebsite filter that returns only businesses with no site at all: the first-website pitch list.
  • Freelancers on Fiverr/Upwork — generate a "100 dentists in Austin with verified emails" list on demand for any client vertical without building a new scraper per niche.
  • SaaS sales teams — any B2B tool sold to local businesses (booking software, payment processors, review management) can target by category and city, and has_booking_widget tells you who already has a competitor installed.
  • B2B lead-gen resellers — one actor covers any category, replacing dozens of niche scrapers.
  • Franchise & market researchers — count and map competitor density for a category in a target city (set onlyWithWebsite: false to keep every mapped location, including ones with no contact details).

How to use it

The form is ready to run as it opens: hit Save & start without touching anything and you get Austin dentists that all have a website — a first run that takes a couple of minutes and costs cents. Measured on 2026-08-08, the untouched form (cap 25, Website required pre-ticked) delivers 25 businesses: 25 with a website, 25 with a phone, 10 with an email.

That number used to be 14. The cause was on our side, not the map's: only the two email filters deepened the candidate pool, so any run using Website required, Phone required or Social profile required filtered a pool sized for an unfiltered run and quietly came up short. Every filter that can drop a row now deepens the pool, and the filters that can be decided from the map data alone (website present, name keywords, permanently closed) are applied before any site is crawled, so no crawl budget is spent on a row that was never going to ship. Same-day measurements at cap 25 in Austin: onlyWithWebsite 14 → 25, requirePhone 10 → 25, onlyWithEmail 19 → 25.

Max businesses is still a ceiling, not a promise — a thin area really can run out. When that happens the run now says so in its status message with the counts behind it, instead of returning fewer rows without comment. A live example, dentist in Laramie, Wyoming at cap 25: "you asked for up to 25 and 1 could be delivered. That is the whole of Laramie, Albany County, Wyoming, United States, not a silent truncation: OpenStreetMap returned 4 'dentist' record(s) there, 1 became crawlable candidate(s), and 1 survived your filter(s) (onlyWithWebsite)." The run only claims an area is exhausted when no search came back at its row ceiling and every candidate was actually checked; otherwise it says which limit it hit and what to change.

  1. Pick a Business category from the dropdown, or leave it on Custom and type one in plain English (e.g. medspa, HVAC contractor, funeral home).
  2. Pick a City from the dropdown, or leave it on Custom and type any city on earth — city and region/country (e.g. Austin, Texas). Capitalisation and the comma are optional: RABAT MAROC and casablanca morocco resolve too.
  3. Leave Website required on unless you are doing density research — see the honest note below. (Web-design agencies can flip Businesses with no website instead to get the no-site prospect list; the two filters are mutually exclusive, and setting both fails the run with a clear error before anything is charged.)
  4. Run the actor. It geocodes the location, pulls matching places from OpenStreetMap, then crawls each business website for contact details, tech and reviews.
  5. Export as CSV, JSON, or Excel — or connect the dataset to your CRM/outreach tool via API/webhook.

Presets: the three guarantees

The preset table is at the top of this page. Behind it sit three guarantees, each of them tested:

  • Anything you set yourself wins. A preset only fills in a field you left at its default. Set maxScore: 100 alongside the web-design preset and you keep 100 — the run log says so explicitly (left as you set them: maxScore (you chose 100)).
  • A preset never widens your bill. Every preset either adds a filter (fewer rows delivered, so fewer rows charged) or turns on enrichment that adds columns to rows you were already getting. None of them clears a filter you set or raises maxItems.
  • It tells you what it did. One log line names the preset and lists exactly which settings it applied and which it left alone, and the run summary names the preset too.

Presets are Console and API: send "preset": "cold_email" from the API and you get the same behaviour. Existing API callers, tasks and schedules that send no preset are completely unaffected.

Which field wins: the precedence chain

Category and location each accept three inputs. Precedence is the same for both, highest first:

RankCategoryLocationWhy
1categories (list)locations (list)The plural list is an explicit multi-search request — it beats everything when non-empty.
2categorySelect (dropdown)citySelect (dropdown)20 categories with hand-verified OpenStreetMap tags; 25 metros verified against the live geocoder.
3category (free text)location (free text)Anything else — several hundred more category terms are mapped, and any city on earth geocodes.

Both dropdowns default to "" ("Custom"), which is why every input that worked before this option existed still resolves to exactly the same search.

Three search modes

ModeSetWhat happens
City (default)location (or locations) + category (or categories)The location is geocoded and every matching business inside its bounding box is returned. Several categories x several cities run as a matrix in one run.
RadiussearchRadiusKm + centerLat + centerLonSearches a circle around a coordinate instead of a city box — sales territories, franchise catchment areas, "everything within 5 km of this address". Replaces the bounding box entirely; no geocoding happens, so country stays empty.
Bring your own listwebsiteList (or startUrls)OpenStreetMap discovery is skipped entirely. The actor runs only the crawl + email verification + scoring pipeline over the websites you supply. Enrich a CRM export, a conference exhibitor list, or a list you bought elsewhere.

When the city runs out: expandNearby

A single city often holds fewer businesses than you asked for — measured live: "dentist" in Austin, Texas tops out near 119 rows, and Round Rock, Texas holds 21. Until now the run could only tell you "try a larger area or a nearby city"; with expandNearby: true it does that for you:

  • The same search is widened in growing rings around the city's centre (sized to the city, up to ~150 km) until your maxItems is met or the whole region is genuinely exhausted.
  • Every widened row is labelled: its query_location reads within ~16 km of Round Rock, Williamson County, Texas, United States instead of the city name, so you can always tell expansion rows apart — or filter them out afterwards.
  • The same dedup and every pre-charge filter apply to ring rows, and the status message reports exactly how many delivered rows came from outside the city.
  • Measured (2026-08-15, discovery-only): dentist / Round Rock, Texas / maxItems: 100 delivered 21 rows without the flag, 100 with it — 79 labelled ring rows, 0 duplicates.

It is off by default for API callers (existing inputs keep a byte-identical pull and bill) and pre-ticked in the Console form. It never fires when you set your own radius (searchRadiusKm) or bring your own list, when the city pull was truncated (deepening, not widening, is the fix there — the status message tells you), or during a mirror outage.

Searching several categories and cities at once

categories and locations are the plural versions of category and location, and they cross into a matrix — ["dentist","orthodontist"] x ["Austin, Texas","Dallas, Texas"] is four searches in one run. The single fields keep working exactly as before; the plural ones take precedence when non-empty.

Three things make the matrix safe rather than a footgun:

  • maxItems is the whole run's budget, not a per-search one. The budget is split between the searches and results are interleaved, so the first city cannot eat the entire quota.
  • The same business found by two categories is delivered — and billed — once. Deduplication is on the OpenStreetMap object identity (osm_type + osm_id), so a clinic tagged both dentist and orthodontist appears one time.
  • 25 category x location combinations is the ceiling for one run. Above it the run fails immediately with a message naming the numbers, before any network request and before any charge — OpenStreetMap's public mirrors are a free shared resource.

Every row carries query_category and query_location, so you always know which search produced it.

Bring your own list (skip discovery)

Already have the businesses and only need the emails, socials, platform and score? Put the domains in websiteList (bare domains and full URLs both work; startUrls is accepted as an alias):

{ "websiteList": ["aloha-dental.com", "https://www.averyranchdental.com", "typotes.com"] }

No map data is fetched at all. Those rows differ from discovered rows in exactly three honest ways:

  • source is user_supplied, not OpenStreetMap;
  • attribution is null — the rows are not OSM-derived, so stamping the ODbL notice on them would be a false licence claim. The column is still present, so a mixed export keeps one stable header row;
  • name comes from the site's own <title>, falling back to the bare domain when the site does not answer. Nothing is invented; latitude, longitude, osm_id and osm_url stay empty.

Everything else — the contact crawl, MX verification, mailbox-owner classification, tech fingerprinting, scoring and every filter — behaves identically.

How complete is the data? (measured, not estimated)

Most scrapers show you a perfect sample row and let you discover the gaps after you have paid. Here are the real fill rates from two reference runs of dentist in Austin, Texas on the current build — one with the website filter on (the whole city, n=55), one with every filter off (n=100):

onlyWithWebsite: true (n=55)schema defaults, filter off (n=100)
phone96%48%
website100%48%
email55%26%
website_platform71%34%
opening_hours80%44%
facebook73%33%
rows graded F0%51%

On that same filter-off n=100 run, the crawl-derived fields measured: contact_page_url 18%, website_title 40%, website_description 35%, mobile_viewport 40%, copyright_year_stale flagged on 3 rows; google_maps_url, osm_url and attribution sat at 100%.

With onlyWithWebsite: true the email fill runs far higher than the filter-off numbers — the reference run above measured 55% for dentists, and a roofing contractor / Denver run delivered emails on 6 of 8 rows (75%): roughly three-quarters on website-verified trade categories.

Read that second column before you run. OpenStreetMap has no website for a large share of businesses, and — measured on the filter-off run — 50 of the 52 rows with no website had no phone, no email and no social profile either (the other 2 carried only an OSM phone): name, coordinates and usually an address, nothing contactable. With the filter off you pay for those rows. onlyWithWebsite and onlyWithEmail drop non-matching rows before you are charged, so they cut your bill rather than just tidying the output. Every run's status message now reports the contactable ratio it actually delivered.

Every filter now goes further: the actor over-fetches, deepening the candidate pool and — for filters that need the site crawled — crawling extra candidates in batches (up to 6× maxItems, the ceiling that makes a thin area terminate) until it has maxItems surviving rows or the pool is spent. This used to apply to onlyWithEmail / onlyVerifiedEmail only, which is why onlyWithWebsite, requirePhone and requireSocial quietly returned short. Measured 2026-08-08 on dentist / Austin, Texas at cap 25, one filter at a time: onlyWithWebsite 14 → 25, requirePhone 10 → 25, onlyWithEmail 19 → 25. Where the pool genuinely runs out first, the status message reports the counts instead of leaving you to guess — e.g. dentist in Laramie, Wyoming returns 1 row and says the map holds 4 dentist records there, 1 of them with a website.

The default is false (not true) so that existing API callers, scheduled tasks and density-research use cases keep getting every mapped location. The Apify console pre-fills it to true.

There is also the mirror filter, onlyWithoutWebsite: keep only businesses that list no website — the prospect list for web-design agencies pitching a first site (verified: a 15-row run delivered 15/15 rows without websites). It is mutually exclusive with onlyWithWebsite; setting both fails the run with a clear error before anything is charged. Expect these rows to be name + address + coordinates (and occasionally a phone) — with no site to crawl, no email/platform enrichment is possible.

Which categories does OpenStreetMap actually cover well?

This is the honest limit of an OSM-based source, and it matters more than any field:

  • Dense: businesses with premises a mapper walks past — restaurants, cafés, dentists, pharmacies, hairdressers, gyms, hotels, shops, banks, clinics. (171 dentist records in the Austin bounding box.)
  • Sparse: trades run from a van or a home office — plumbers (9 in the same box), electricians (5), chiropractors (0), roofers (~12). A metro of a million people can return single digits. That is what is mapped, not a bug.

~50 categories are curated and mapped to exact OSM tags — build 0.1.21 added pest control, photographer, moving company, self storage, funeral home and dry cleaner. Any other term is attempted as an OSM tag directly, and common phrasings are handled (auto repair shopshop=car_repair, landscaping companycraft=gardener, insurance agencyoffice=insurance). Terms with no OSM tag now fall back to a business-name keyword search of OSM, and a mapped category that returns zero tagged places in the area is rescued by the same name-keyword query. Rows found only by name are labelled category_match: "name_keyword" (tag-matched rows say category) so you can filter them out if you only trust tag-confirmed rows. Measured: pest control in Denver returned 1 labelled row on 0.1.21 where the previous build returned 0 — sparse trades stay sparse, but no longer invisible. A term matching nothing at all still returns 0 rows with a suggestion and no charge — you are never billed for a run that found nothing.

Filters — every one of them runs before you are charged

Filters here are not a tidying step applied to an invoice you have already run up: a filtered row is dropped before it is pushed to the dataset and before the charge, so filtering cuts your bill.

Filtering does not shrink your delivery. Every filter in this table deepens the candidate pool to compensate, so maxItems means "this many rows I can use", not "this many businesses considered". Filters decidable from the map data alone — onlyWithWebsite, onlyWithoutWebsite, excludeKeywords, skipClosed — are applied before any website is crawled, so no crawl budget is spent on a row that was never going to ship. The rest are applied batch by batch as sites are crawled, stopping the moment enough rows survive. The pool is bounded at 6x maxItems candidates so a thin area terminates instead of crawling forever; when that bound or the map itself is what stopped the run, the status message says so and gives you the counts.

FilterKeeps only
onlyWithWebsitebusinesses that list a website (recommended, see the fill-rate table)
onlyWithoutWebsitebusinesses with no site — the first-website pitch list. Mutually exclusive with the above
onlyWithEmailrows where an email was found
onlyVerifiedEmailrows whose email passed MX verification (deliverable / risky)
requirePhoneNew. rows with a phone number (from OSM, a tel: link, or the site's schema.org markup)
requireSocialNew. rows with at least one Facebook / Instagram / LinkedIn / X / YouTube profile
requireAnyContactNew. rows reachable on at least one channel — phone or email or a social profile. The loosest contactability floor there is, and the one to reach for when you do not care which channel. It matters more than it sounds: on a bare {} run of the default search (100 rows, measured 2026-08-08) exactly 50 rows carried no phone, no email and no social profile at all, and without this switch you are billed for them. Off by default, so no existing run changes
excludeKeywordsdrops businesses whose name contains any of your keywords (case-insensitive). Only the name is matched, so clinic cannot knock out a business on Clinic Street. Use it to strip chains, franchises or your existing customers
skipClosedNew. drops permanently-closed premises (see below)
minScore / maxScoremaxScore is new. A score ceiling is the web-design agency filter: a low score means a thin online presence, which is exactly the redesign pitch list
minRating / minReviewCountNew — read the warning below before using these

skipClosed: what it actually removes

OpenStreetMap mappers retire a business without deleting it: the primary tag moves from amenity=restaurant to disused:amenity=restaurant, so the object keeps its name and address but no longer describes an operating business. skipClosed drops elements carrying a lifecycle-prefixed primary tag (disused:, abandoned:, was:, removed:, demolished:, razed:), a disused=yes / abandoned=yes flag, opening_hours=closed, or shop=vacant. A disused: tag alongside a live tag of the same kind (a former bank that is now a café) is not treated as closed.

Measured live on 2026-08-08 in the Austin bounding box: 199 such elements, 43 of them still carrying a business name — e.g. Chago's with disused:amenity=restaurant, Corner Store with abandoned:shop=fuel. These cannot reach a normal tagged search (["amenity"="restaurant"] cannot match disused:amenity), but they do reach the business-name fallback used for unmapped categories, which is where dead businesses were being delivered as fresh leads. Verified end to end: with skipClosed: false that closed restaurant is delivered and billed; with it true it is dropped before billing.

It is off by default so existing runs, tasks and API callers are unchanged. The Apify console pre-fills it to true.

Honest warning about minRating / minReviewCount: OpenStreetMap carries no review data at all. A rating only exists when the business publishes schema.org aggregateRating on its own website — measured at roughly 3-8% of rows. So when you set a rating or review floor, rows with no rating are DROPPED, not kept: an unknown rating is not a passing rating, and nothing is ever invented to save a row. Setting either filter will cut your result count to a small fraction — an Austin restaurant run asking for 15 rows with minRating: 4.0 crawled 86 candidates looking for them and delivered 0 rows, charging nothing (re-measured 2026-08-08 after the over-fetch was generalised, so this is the deep-search result, not a shallow one). If you need ratings on every business this is the wrong source; pair it with our Google Maps Places Scraper, or use the google_maps_url on every row.

Choosing your columns and sort order

  • outputFields — pick the columns you want (e.g. ["name","email","phone","website","lead_grade"]) and the dataset carries only those. Every row still shares one identical column tuple, so CSV headers never shift mid-export, and the columns keep the documented ROW order regardless of the order you listed them in. name and attribution are always included whatever you choose — attribution because the OpenStreetMap licence (ODbL) has to travel with the data into your CSV. Column names are matched case-insensitively (and -/space count as _, so Lead Grade works). A single unknown name is logged and ignored; if none of the names you list exists, the run stops before billing rather than charging you full price for a name-only export.
  • sortByscore_desc (default, what every previous build did), name_asc, review_count_desc or rating_desc. Sorting never removes a row; all filtering already happened. Rows missing the sort value (no rating, no review count) are placed last, because a missing value is unknown rather than zero.

Output fields

61 columns on every row (57 before this build; the four new ones are query_category, query_location, email_guess and email_guess_confidence). Re-measured on a 100-row dentist / Austin, Texas run with schema defaults: all 100 rows carried the same stable 61-column tuple, so CSV headers don't shift mid-export — and if you narrow the export with outputFields, every row still shares one identical (smaller) tuple. Fill rates are from the onlyWithWebsite: true reference run above; anything conditional says so. Fields marked New (0.1.21) show fill rates from the 100-row filter-off benchmark instead — read them accordingly, since roughly half of those rows had no website to crawl.

In the Console, the Output tab opens on the Overview view; the All columns tab shows every field listed below, in this order. Views only change what the Console table renders — a CSV, JSON or Excel export always carries every column the run produced (or exactly the ones you named in outputFields).

Identity & location — from OpenStreetMap

FieldFillDescription
name100%Business name
category100%OSM category tag (e.g. dentist, hairdresser, lawyer)
category_match100%New (0.1.21). How the row matched your category: category = matched the mapped OpenStreetMap tag; name_keyword = found by the business-name fallback search (used for unmapped terms, and as a rescue when a mapped tag returns zero places in the area). Filter on category if you only want tag-confirmed rows
address96%Street address assembled from OSM address tags
city / state / postal_code93% / 93% / 93%Address components
country100%From the geocode. Empty in radius mode and on user-supplied rows (no geocode happens there)
query_category100%New. Which of your input categories produced this row — the column that makes a multi-category run readable. Null on user-supplied rows
query_location100%New. Which of your input locations produced this row (the geocoder's resolved name, or the radius description). Null on user-supplied rows
latitude / longitude100%Coordinates
osm_type / osm_id100%OpenStreetMap source identifiers
google_maps_url100%New (0.1.21). A Google Maps link for the business — one click to the live listing with its current rating and review count (data this actor deliberately does not scrape; see the rating warning below). Need those fields at scale? Pair with our Google Maps Places Scraper
osm_url100%New (0.1.21). Link to the row's source OpenStreetMap element — instant provenance, and the place to fix bad map data

Contact

FieldFillDescription
phone96%Primary phone. OSM tag first; falls back to a tel: link or JSON-LD telephone only when OSM has none
phones73%New. All tel: numbers found on the site. May include a call-tracking number — phone stays the trusted value
website100%Business website URL. A social page in OSM's website tag is routed to that social column instead, so this is always a real site
domain100%New. Bare registered domain — the field CRMs dedup on
email55%Primary contact email, chosen by mailbox ownership then deliverability
emails55%All emails for the row, primary first (previously excluded a primary that came from OSM)
contact_page_url42%New. The exact page the primary email was found on, so you can spot-check it. Null when the email came from OpenStreetMap rather than from a crawled page
email_guess0% unless opted inNew, opt-in, and deliberately not an email. With emailPatternGuess: true, a business that has a website but publishes no address anywhere we crawled gets a pattern-guessed info@<domain> here. It is never promoted into email, never counts as has_email, never earns a lead-score point and never satisfies onlyWithEmail / onlyVerifiedEmail. Nobody checked that this mailbox exists — treat it as a lead, not an address
email_guess_confidencesameNew. A statement about the domain, never the mailbox: mx_ok (the domain does run mail servers), no_mx (it does not — the guess is almost certainly dead), unchecked (email verification is off, so no lookup was made)

Email quality — verification is included, not an add-on

FieldFillDescription
email_status100%deliverable / risky / undeliverable when verification is on and a candidate exists; missing when no email was found; found / missing when verifyEmails is off. Turning verification off changes the vocabulary
email_provider55%Mailbox host where identifiable (Google Workspace, Microsoft 365...). Only when verifyEmails is on and an email was found
email_domain_match55%Whether the email's domain matches the website's
email_type55%own_domain / free_mail / third_party / unknown. third_party means the address belongs to the business's marketing agency or web designer — it is deliverable but does not reach the business, and it is scored at half weight. Measured on the reference run: 12% of all harvested addresses were third-party, but only 10% of primary emails, because candidates are ranked by mailbox ownership before one is promoted
email_types55%The same classification for every entry in emails

Socials

FieldFillDescription
facebook73%Facebook profile URL
instagram55%Instagram profile URL
youtube25%New. YouTube channel URL
linkedin22%New. LinkedIn company/profile URL
twitter18%New. X/Twitter profile URL

Share buttons, tracking pixels and embedded posts are filtered out, so these are profile URLs rather than facebook.com/tr or instagram.com/p/....

Website & marketing tech

FieldFillDescription
website_platform71%CMS / site builder: WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, Custom / Next.js, Custom / React. WordPress plugins (Elementor, WP Rocket, Divi) report as WordPress, not as their own platform
website_platform_status100%New. Why website_platform is what it is: detected, unknown (site loaded, no fingerprint), site_unreachable, no_website, not_crawled. A null platform is explained rather than unexplained
platform_version27%New. Only when the site's own meta generator names the platform and a version — never a plugin's version passed off as the platform's
site_generator42%New. The raw meta name="generator" string
tech84%New. Marketing/booking/commerce tech found on the page (Meta Pixel, Google Analytics, GTM, Calendly, NexHealth, Klaviyo, WooCommerce, Yelp Reviews, live chat...)
has_meta_pixel / has_google_analytics / has_booking_widgetsee noteNew. Booleans derived from tech. Tri-state: true/false when the site was read, null when it was unreachable or not crawled — false never means "we could not check"
website_title40%New (0.1.21). The crawled homepage's <title> — a drop-in mail-merge personalization field. Also a data tripwire: a title that clearly names a different business means OSM's website tag is wrong, and this column lets you catch that before you hit send
website_description35%New (0.1.21). The homepage's meta description — the other personalization field, and the same tripwire
mobile_viewport40%New (0.1.21). Whether the homepage declares a mobile viewport tag. A site without one predates responsive design — a concrete web-agency pitch signal
copyright_year_stale3%New (0.1.21). Flags a visibly out-of-date copyright year in the footer — a small but unambiguous neglected-site signal (flagged on 3 of the 100 benchmark rows)

Business detail & reviews

FieldFillDescription
opening_hours80%From OSM; gap-filled from the site's JSON-LD only when OSM has none
rating7%New, and sparse — see the warning below. Star rating (1-5) the business publishes in its own schema.org markup
review_count7%New, sparse. Review count from the same markup
price_range24%New. schema.org priceRange (e.g. $$)

Honest warning about rating / review_count: these are not Google Maps ratings. They are only present when a business publishes aggregateRating in its own JSON-LD, and most do not. Measured on-platform: 7% of website-bearing rows for dentists in Austin (4/55) and 13% for roofing contractors in Denver (1/8); the filter-off benchmark (n=100) measured 4%. An earlier 12-row roofer sample hit 33%, so the rate swings wildly with category, city and sample size — assume under 10% and treat anything higher as luck. Do not build a workflow that needs a rating on every row. If you need star ratings and review counts for every business, this is the wrong source — pair it with our Google Maps Places Scraper; every row's google_maps_url also takes you to the live listing in one click. OpenStreetMap carries no review data at all, and this actor never invents a substitute: lead_score is a data-completeness score, not a customer rating.

Lead scoring

FieldFillDescription
lead_score100%0-100 completeness/reachability score
lead_grade100%A ≥80, B ≥65, C ≥50, D ≥35, else F
lead_tier100%hot ≥75, warm ≥50, else cold
score_breakdown100%Per-component points, plus signals_from — the list of extra signals that actually scored — so the score is auditable
has_email / has_phone / has_website100%Booleans for quick filtering

Scoring rubric (sums to a true 100, so grade A is reachable — the filter-off benchmark's top row scored the full 100): email deliverable 40 / risky 25 / present-but-unverified 15 — halved when email_type is third_party; phone 20; website 15; socials 5 each capped at 10; extra signals 5 each capped at 15, drawn from website platform detected, opening hours, marketing tech found, and a published star ratingscore_breakdown.signals_from names the ones that counted.

Provenance

FieldFillDescription
source100%OpenStreetMap for discovered rows, user_supplied for rows that came from your own websiteList
attribution100% on OSM rowsThe ODbL attribution string, so the licence travels with an exported CSV. Null on user_supplied rows — they are not OpenStreetMap data, so attaching the notice would be a false licence claim. The column is always present either way
enriched_from_website33%New. Lists only the fields where an OSM-overlapping value was taken from the site instead — rating, review_count, price_range (OSM carries none of these) plus phone / opening_hours / email where OSM was empty and the site's schema.org data filled the gap. It is not a full provenance map: emails, the five socials, website_platform, platform_version, site_generator and tech are always crawled from the site and are deliberately not repeated here

Example output

Real rows from a run with onlyWithWebsite: true. This is a best-case slice, not a typical one — see the measured fill rates above; on the website-filtered reference run roughly half of rows carry an email.

NamePhoneEmailStatusTypePlatformTechScore
Aloha Dental+1-512-707-7300riverside@aloha-dental.comdeliverableown_domainWordPressMeta Pixel, GTM, Yelp95 (A)
Avery Ranch Dental+1-512-246-7645smile@averyranchdental.comdeliverableown_domainWordPressGTM, reCAPTCHA95 (A)
Aviva Dental Care+1 512 852 8528dr.apurva@avivadentalcare.comdeliverableown_domainWordPressGoogle Analytics90 (A)

Input examples

A preset plus the two dropdowns — the shortest useful input there is:

{
"preset": "cold_email",
"categorySelect": "hair salon",
"citySelect": "Casablanca, Morocco",
"maxItems": 100
}

The preset sets onlyWithEmail and onlyVerifiedEmail; everything else stays at its default. Add any field you want and yours wins — {"preset": "web_design", "maxScore": 100} keeps maxScore: 100 and the log says which setting the preset stood down on.

One category, one city — the classic run, unchanged:

{
"category": "hair salon",
"location": "Miami, Florida",
"maxItems": 200,
"crawlEmails": true,
"onlyWithWebsite": true,
"onlyWithEmail": false,
"verifyEmails": true,
"maxPagesPerSite": 3,
"concurrency": 8
}

Several categories across several cities, contactable rows only, narrow columns:

{
"categories": ["dentist", "orthodontist"],
"locations": ["Austin, Texas", "Dallas, Texas"],
"maxItems": 200,
"onlyWithWebsite": true,
"requirePhone": true,
"excludeKeywords": ["Aspen Dental"],
"skipClosed": true,
"outputFields": ["name", "email", "phone", "website", "lead_grade", "query_location"],
"sortBy": "name_asc"
}

A 5 km sales territory around one address:

{
"category": "restaurant",
"searchRadiusKm": 5,
"centerLat": 30.2672,
"centerLon": -97.7431,
"maxItems": 150,
"onlyWithEmail": true
}

Web-design prospect list — businesses with a site, but a weak one:

{
"category": "hair salon",
"location": "Lyon, France",
"countryCode": "fr",
"onlyWithWebsite": true,
"maxScore": 45,
"skipClosed": true
}

Enrich your own list (no map data fetched at all):

{
"websiteList": ["aloha-dental.com", "averyranchdental.com", "typotes.com"],
"verifyEmails": true,
"emailPatternGuess": true
}

How much does it cost?

Pay-per-result — the current per-lead rate is on this page's Pricing tab, and you are charged for delivered businesses only, not for API calls or compute. No subscription. New Apify users get platform free credits to test with.

A run's total is simply rows delivered x the per-lead rate. Be aware what the rows contain: on the reference run with onlyWithWebsite: true, ~55% carried an email, so 500 rows ≈ 275 emailable leads (roughly 1.8x the per-lead rate per emailable lead). With the website filter off, 26% carried an email on the filter-off benchmark (roughly three-quarters on website-verified trade categories). Filters run before billing, so onlyWithEmail: true is the cheapest way to buy emails specifically — and the adaptive over-fetch delivers as close to maxItems email rows as the city allows (measured 2026-08-08: asked 25, delivered 25 in Austin; the same input delivered 19 before the pool-depth fix, and the whole city tops out at 31, which the run tells you when you ask for more).

For comparison, the largest competing Maps-based scraper charges $0.004/place plus $0.002 for contacts plus $0.004 per record for email verification — about $0.010 at feature parity. Here, MX verification, mailbox-ownership classification and lead scoring are all included in the single per-lead rate.

Data source & licence

Business listings come from OpenStreetMap. © OpenStreetMap contributors, available under the Open Database License (ODbL) v1.0https://www.openstreetmap.org/copyright. Every row carries source and attribution fields; keep them if you redistribute or publish the data, as ODbL requires attribution.

The website-crawled fields — email, emails, socials, website_platform, tech, rating, review_count, price_range, phones — are not OSM-derived. They come from each business's own public website and are not covered by ODbL. Those fields are site-derived on every row. The enriched_from_website column is narrower than that: it flags only the fields that OSM could have supplied but didn't, so treat the list above — not that column — as the ODbL boundary.

Frequently asked questions

This actor reads publicly available data from OpenStreetMap (an open-data project, ODbL-licensed) and each business's own public website — the same information anyone could find by visiting the site. No login, no private data, no anti-bot circumvention, no Google Maps terms-of-service exposure.

Does it work outside the US? How should I type the location?

Yes — anywhere OpenStreetMap covers. Type the location however you like: the raw string is tried first, and only if that finds nothing is it automatically re-spelled (Title Case, and a comma inserted before the trailing country word) before the run is given up on. RABAT MAROC, Rabat Maroc, casablanca morocco and Rabat, Morocco all resolve to the same place. If the run still cannot geocode, the status message lists every spelling it tried instead of a generic failure.

If your city name exists in more than one country — Rabat is a city in both Morocco and Malta, Cambridge in both the UK and the US — set countryCode to an ISO 3166-1 alpha-2 code (ma, mt, us, fr, gb) to pin the geocoder. Verified: location: "Rabat" with countryCode: "mt" resolves to Rabat, Western Region, Malta; without it, Morocco wins. The pin is enforced: the fallback geocoder (Photon) has no country filter, so it is skipped whenever countryCode is set. A city that does not exist inside that country ends the run with a clear message and no charge, instead of quietly returning the same-named city elsewhere (location: "Austin" + countryCode: "ma" used to deliver Austin, Texas).

One non-US bug is fixed as of this build: the geocoder is now asked for English place names. It previously answered in the local language(s), and since the row's country column is derived from the geocoder's answer, every Moroccan row used to ship country: "Maroc ⵍⵎⵖⵔⵉⴱ المغرب" — a three-script run-on where Morocco belonged. Measured and fixed on 2026-08-08; Austin, Texas and Lyon, France are unaffected. Note that city and address still come straight from the OpenStreetMap tags a local mapper wrote, so a Moroccan row can legitimately read Témara تمارة — that is the map's own data and it is not rewritten.

How fresh is the data?

Business listings come from OpenStreetMap, which is community-maintained and updated continuously; most established businesses have accurate name/address/phone data. Website content (email, socials, platform, tech, ratings) is crawled live on every run, so that part is current at run time.

Why is email empty on some rows?

Three distinct reasons, and email_status plus website_platform_status tell you which: the business has no website in OSM (no_website), its site did not respond (site_unreachable), or the site simply does not publish an address anywhere on the pages crawled (missing). Cloudflare-obfuscated addresses are decoded, and since 0.1.21 the crawler follows the site's own contact/about/team links — including Shopify /pages/contact and non-English slugs — within the maxPagesPerSite budget. Raising maxPagesPerSite finds a few more.

Does it get star ratings and review counts?

Only when a business publishes them in its own website markup — 7-33% of website-bearing rows depending on the category (4% on the latest filter-off benchmark). See the warning in the output section. OpenStreetMap has no review data, so if ratings are essential for every row, pair this with our Google Maps Places Scraper — every row's google_maps_url opens the live Google listing in one click.

Why does website platform detection matter?

Knowing a business runs on Wix, GoDaddy or Weebly (often a dated, self-built site) versus WordPress or a custom build lets agencies pre-qualify leads before picking up the phone. Combined with has_meta_pixel: false and has_booking_widget: false, you get businesses that are visibly under-invested in their web presence — a far better redesign pitch than a cold list.

Do I need an API key or proxy?

No. Discovery runs on OpenStreetMap (Nominatim + Overpass) and enrichment crawls public websites directly — no API keys, no proxy configuration, and it runs from datacenter IPs. Since 0.1.21, Overpass mirrors are raced in pairs (at most two in flight) instead of tried strictly in sequence, so one slow mirror no longer stalls discovery. A small share of sites (measured: 8 of the 55 website-bearing rows in the reference run) block or fail to answer datacenter requests; those rows are marked site_unreachable rather than silently blamed on missing data.

Other Flash Scrape lead tools

  • Google Maps Places Scraper — live Google Maps rating, review count, phone and website per place: the honest companion for the review data this actor deliberately does not invent. Every row's google_maps_url links the two.
  • Restaurant Leads Scraper — restaurant leads enriched with POS, reservation and delivery tech stack (Toast, OpenTable, DoorDash, and more).
  • Google Maps Leads Opener — pull business leads directly from Google Maps search results.
  • Email Verifier — bulk-validate and clean an existing email list before you launch a campaign.
  • Google Ads Transparency Scraper — see every Google ad a local competitor runs, with first/last shown dates: the other half of a local-market picture.
  • Company Domain Enricher — turn a company name into its website domain and firmographic details.