Local Business Leads Scraper: Verified Business Emails & Finder avatar

Local Business Leads Scraper: Verified Business Emails & Finder

Pricing

from $2.40 / 1,000 business leads

Go to Apify Store
Local Business Leads Scraper: Verified Business Emails & Finder

Local Business Leads Scraper: Verified Business Emails & Finder

Local business leads scraper, local business email finder and local business email scraper: any category, any city. Verified business emails (MX-checked), phones, socials on every row - local leads with emails, no API key, no proxy, $3 per 1,000. Business email finder for agencies.

Pricing

from $2.40 / 1,000 business leads

Rating

5.0

(3)

Developer

Flash Scrape

Flash Scrape

Maintained by Community

Actor stats

6

Bookmarked

33

Total users

22

Monthly active users

4 hours ago

Last modified

Share

Local Business Leads Scraper is a pay-per-result lead-generation actor that finds local businesses in any category and any city on earth and delivers MX-verified emails, phones, social profiles, website platform and a 0-100 lead score — built on OpenStreetMap, so it needs no API key, no proxy and no login.

Try it: Find local business leads with emails by category

Local business leads scraper and business email finder — verified business emails, phones and socials for any category in any city; a local business email scraper built for agencies.

What you get per lead

One row per business, 73 stable columns, with a Google Maps link and OpenStreetMap provenance on every row. Fill rates are measured on the reference run — dentist / Austin, Texas, onlyWithWebsite: true, n=55, 2026-08-08 (google_maps_url and osm_url also measured 100% on the same day's filter-off n=100 run); the full story is under Measured, not promised below.

FieldFilledWhat it is
MX-verified email + whose mailbox55%email graded deliverable / risky / undeliverable; email_type says own domain, free inbox or the business's marketing agency
Phone96%OSM tag first, then the site's own tel: link
Social profilesFacebook 73% · Instagram 55% · YouTube 25% · LinkedIn 22%profile URLs the business's own site links
Website platform + tech71%WordPress, Wix, Shopify, Squarespace… plus Meta-pixel / Google-Ads-tag flags
Lead score + grade100%0-100 completeness/reachability score and an A-F grade — not a customer rating
Google Maps link100%google_maps_url opens a Google Maps search for the business in one click; nothing is scraped from Google
OpenStreetMap provenance100%osm_url, coordinates and the ODbL attribution string

$3 per 1,000 delivered leads ($0.003 per lead) on the free plan; paid plans pay less (Pricing tab). Every filter — onlyWithEmail, onlyVerifiedEmail, onlyWithWebsite, requirePhone, requireAnyContact and the rest — runs before you are charged, MX verification is included in that price, and a run that finds nothing charges nothing. A scheduled pricing record raises this to $0.005 per delivered lead on 2026-09-14 (notified 2026-08-30); the Pricing tab on this page is authoritative.

Run report: leads with phone, verified email and website from a real run

This Actor takes a business category and a city and returns one row per local business, 73 stable columns wide and every row scored 0-100 — with an MX-verified email, whose mailbox that email is, the phone, the social profiles and the website platform wherever the business's own site exposes them (measured fill rates below). It costs $3 per 1,000 delivered leads ($0.003 per lead on the free plan; the Pricing tab always carries the current rate, and a change to $0.005 per lead is scheduled for 2026-09-14), and email verification is part of that price rather than an add-on.

Dentists, gyms, lawyers, plumbers, roofers, salons, real estate agencies, restaurants and 50+ more curated categories, plus any term you type. No API key, no proxy, no separate scraper per niche.

Key facts:

  • $3 per 1,000 delivered leads ($0.003 per lead; rising to $0.005 per lead on 2026-09-14 under a scheduled pricing record) — MX email verification, mailbox-ownership classification and lead scoring are included in the single per-lead rate; filtered rows are dropped before billing, and a run that finds nothing charges nothing.
  • No API key, no proxy, no login — discovery runs on OpenStreetMap and enrichment crawls each business's own public website, from datacenter IPs.
  • 95 curated categories (220+ terms), any city on earth — plus a radius search around a coordinate, or bring your own website list and skip discovery entirely.
  • 73 stable columns on every row — export as CSV, JSON or Excel, or connect the dataset to your CRM via API/webhook; column headers never shift mid-export.
  • Built-in monitoring and alertsonlyNewBusinesses turns a schedule into a new-business alert, and webhookUrl posts a Slack / Discord / JSON digest whenever a run delivers rows.

What it does not do, stated up front so nothing here is a surprise after you have paid:

  • It does not scrape Google Maps. Listings come from OpenStreetMap, so there are no Google star ratings and no Google review counts. Pair it with our Google Maps Places Scraper when you need those.
  • The rating and review_count columns are sparse. They exist only when a business publishes a rating in its own website markup, measured at about 7% of rows (4 of 55 on the 2026-08-08 reference run), and they are never a Google rating.
  • Van-based trades are thinly mapped. OpenStreetMap held 171 dentists in the Austin bounding box but 9 plumbers and 5 electricians. The coverage section below names which categories are dense.
  • A guessed email is never sold as a verified one. A pattern-guessed address stays in email_guess and is never promoted into email.

Use from an AI agent

  • MCP: point Claude, ChatGPT, Cursor or any MCP client at https://mcp.apify.com?tools=flash_scraper/local-business-leads; the tool is named after the Store slug and takes this actor's input unchanged. Keywords for the server's search-actors tool: business leads, local business, lead generation, business emails, verified business emails. Tool-name spellings, payment without an Apify token and measured timings: Use it from an AI agent (MCP).
  • Smallest useful call (Python apify-client; the same JSON works in the Console, the REST API and n8n/Make/Zapier):
from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("flash_scraper/local-business-leads").call(run_input={"category": "dentist", "location": "Austin, Texas", "maxItems": 10, "crawlEmails": False})
rows = client.dataset(run["defaultDatasetId"]).list_items().items
  • Output contract: the same 73 columns on every row, headers never shift mid-export; email_guess is never promoted into email. The full field list with measured fill rates is under What data you get; every run also writes a machine-readable RUN_SUMMARY record to its key-value store.

What you get

  • Any category, any city — 220+ category terms (95 curated categories plus their aliases) mapped to exact OpenStreetMap tags, any other term attempted directly and then rescued by a business-name search, and any city on earth geocoded (Austin, Texas, casablanca morocco, Dubai UAE).
  • Emails that are verified, not guessed — every address is MX-checked over DNS-over-HTTPS and graded deliverable / risky / undeliverable. Verification is part of the price, not an upsell.
  • Whose mailbox it isown_domain, a free inbox, or the business's marketing agency. Emailing an agency mailbox never reaches the business, so this column decides whether a lead is worth a send.
  • Redesign pitch signalswebsite_platform (WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more), mobile_viewport, copyright_year_stale, has_meta_pixel, has_google_ads_tag. A builder-tier site with no tracking is the highest-intent pitch there is.
  • A score you can audit — 0-100, an A-F grade, a hot/warm/cold tier, and a score_breakdown object showing the arithmetic that produced it.
  • Five ready-made Output views — Overview, Email-ready, Web-agency targets, Map & source, All columns. Pick one above the results table instead of scrolling a wall of 73 raw columns.
  • Filters that cut your bill, not just the table — every filter drops the row before it is pushed and before it is charged, and the run log names each filter and how many rows it removed.
  • Schedulable — only what is newonlyNewBusinesses turns a weekly schedule into a new-business alert: the first run is the baseline, every later run delivers only businesses that were not delivered before, and a run with nothing new delivers 0 rows and bills nothing. How it works.
  • One-click provenancegoogle_maps_url (a Google Maps search for the business's name and address, so it opens the listing search rather than a bare map pin) and osm_url on every discovered row, plus the ODbL attribution the licence requires you to keep.

Is there a Google Maps scraper alternative that does not scrape Google Maps?

Yes: this Actor discovers businesses on OpenStreetMap and then crawls each business's own public website, so no listing, rating or review is ever taken from Google Maps and no Google page is ever scraped. That is the whole design, not a setting you have to switch on.

On the 2026-08-08 reference run (dentist / Austin, Texas, onlyWithWebsite: true, n=55) that design delivered a phone on 96% of rows, an MX-verified email on 55% and a detected website platform on 71%, with no proxy input and no proxy cost, at $3 per 1,000 delivered leads (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).

What that buys you:

  • No proxy bill and no proxy setup. There is no proxy input on this Actor and no proxy cost in a run: OpenStreetMap's public endpoints and most business websites answer datacenter IPs, so runs go direct. A minority refuse them — measured at 8 of the 55 website-bearing rows in the reference run — and those rows say site_blocked rather than pretending the business is gone.
  • No Google Maps terms-of-service exposure. What you receive is open-licensed map data plus pages the businesses publish themselves.
  • Fields a Maps listing does not carry. MX-verified emails, mailbox ownership (own_domain / free inbox / the business's marketing agency), website_platform, has_meta_pixel and mobile_viewport all come from the business's own site, and the 0-100 lead score is computed from those signals together with the phone and website the map supplies.
  • A Google Maps link on every discovered row anyway. google_maps_url is a constructed search link (?api=1&query=<name>, <address>), so you can open the real listing in one click. Nothing is read from Google to build it.

What you give up is equally concrete. There are no Google star ratings or review counts, and OpenStreetMap maps some trades thinly, which the coverage section below quantifies. If you need Google's own review data, run our Google Maps Places Scraper alongside this one.

Is this a small business leads scraper with MX verified emails?

Yes. It is a small business leads scraper for any category and any city: every row is a real local business found on OpenStreetMap, and every email is MX verified against the domain's mail servers before delivery, so what you get is verified email leads rather than guessed addresses. Local leads with emails, phones and social profiles arrive in one table, and rows removed by a filter are never billed.

Measured, not promised

Every number below is from a real run of this Actor on dentist / Austin, Texas. Inputs and dates are given so you can reproduce them.

RunResult
Console form untouched, pressed Save & startToday's untouched form is dentist / Austin, Texas at a cap of 25 (the form's starting value since 2026-08-29; it was 100 before) with Website required, Skip closed and Widen to nearby areas pre-ticked and no keyword exclusions — every delivered row lists a website. No count is quoted here on purpose: the 2026-08-08 measurement of this row was taken at cap 25 with a two-chain exclusion list (Aspen Dental, Walmart) the form pre-filled at the time and no longer does, so its figure cannot be reproduced from the form as it opens now. The next row is the same search with the same website filter at cap 100 (the API default)
Console form as it opened on 2026-08-08 (it then pre-filled excludeKeywords: ["Aspen Dental", "Walmart"], which today's form does not), Max businesses raised to 100, pressed Save & start55 businesses, 100% contactable — 96% with a phone, 55% with an MX-verified email. Under the 100 asked for, and the run says why: OpenStreetMap holds 171 dentist records in Austin and 55 of them are crawlable with a website inside the search area. That is the whole city, not a truncation
Bare {} from the API (no filters at all) — 2026-08-08100 businesses, 50% contactable — the honest filter-off case, see the fill-rate table below
Cold-email list preset — 2026-08-0831 rows (every emailable dentist Austin has — the exact count varies by city and over time), 100% with an MX-verified email, 94% with a phone
Call list preset, cap 100 — 2026-08-0856 rows, 100% with a phone — the same number on two independent runs, and the run reports it as the whole of Austin rather than a truncation
Web-design prospects preset, cap 100 — 2026-08-087 rows, every one scored ≤45 and reachable on at least one channel. The score ceiling alone matches 68 businesses in Austin, but only 7 of those publish a phone, email or social profile — the preset drops the other 61 rather than bill you for map pins you cannot contact
Full enrichment preset — 2026-08-08100 rows plus 23 email_guess addresses, kept out of email and never billed as verified
Column stability across every run aboveone identical column tuple on every row (61 columns on those runs; 73 on the current build — the 12 additions of 2026-08-29 are appended at the END, so existing CSV column positions are unchanged) — CSV headers never shift mid-export

Every 2026-08-08 row that was started from the Console form (the preset rows included) ran with the two-chain excludeKeywords list the form pre-filled at the time and no longer does; the bare {} row did not. The counts stay as historical figures.

The gap between the Console-form rows and the bare {} row is the whole story of this Actor's honesty: the Console form pre-fills onlyWithWebsite, which is why it delivers 100% contactable businesses, while a bare API call keeps every mapped location including the ones with nothing to contact. Both numbers are published; neither is hidden. The same goes for the 55-of-100 line: a run that cannot reach your cap says so in its status message, with the counts that explain it.

Max businesses starts at 25 in the Console form — raise it once the first run looks right. The form also pre-ticks requireAnyContact (since 2026-08-29), so a first run never bills you for a business with no phone, no email and no social profile — untick it if you want every listing. At $3 per 1,000 that first run costs at most $0.075 in leads (at most $0.125 once the scheduled 2026-09-14 change to $5 per 1,000 lands) and finishes in a couple of minutes. The API default is unchanged at 100: API calls, tasks and schedules that send no maxItems keep the ceiling they always had.

Quick start

The shortest complete input is a business category and a location. Everything else has a working default:

{
"category": "dentist",
"location": "Austin, Texas"
}

That is a complete run, and it is also exactly what the form already contains. Open the Actor, press Save & start without touching anything, and you get businesses. The defaults behind it: Max businesses starts at 25 in the form (raise it once the first run looks right; an API call that sends no maxItems gets 100), website crawling on, email verification on, 3 pages per website.

The input form at a glance

Eight sections, in the order the form shows them. A first run only ever touches the first three; the rest are already tuned and safe to ignore.

SectionWhat it is for
Quick startOne dropdown, preset, that fills in the rest of the form for one outreach job — see the presets below. Custom changes nothing
What to searchThe business category and the city (dropdown or free text, or the categories / locations lists for several at once), an optional countryCode and keyword (matched against the name, and since 2026-08-29 the cuisine / healthcare:speciality tags too), maxItems — how many businesses this run may deliver, i.e. your cost ceiling; the form starts at 25, raise it once the first run looks right (API default 100) — and expandNearby, pre-ticked, which widens the search around the city when the city itself runs out (each widened row labelled in query_location)
Filters — businesses are dropped BEFORE you are chargedWebsite / email / phone / social floors, excludeKeywords (arrives empty — nothing is excluded until you type a keyword), skipClosed (pre-ticked), and the score and rating floors
🔔 MonitoringonlyNewBusinesses — put the run on a schedule and get only the businesses that were not delivered before
EnrichmentWhat gets read from each business website: emails, social profiles, phones, optional pattern-guessed addresses, plus the maxPagesPerSite and concurrency crawl knobs
Other ways to searchA radius around a coordinate (searchRadiusKm + centerLat / centerLon) instead of a city, or bring your own list (websiteList / startUrls) and skip OpenStreetMap discovery entirely
OutputoutputFields — which columns you get, and in what order — and sortBy
🔔 AlertswebhookUrl — a Slack / Discord / JSON digest whenever a run delivers rows

Presets

The Use-case preset dropdown at the top configures the rest of the form for one specific outreach job. It is optional — leave it on Custom and nothing at all changes.

PresetWhat it setsMeasured on the default search (Austin dentists)
Cold-email listonlyWithEmail, onlyVerifiedEmail31 leads (the whole of Austin; the count varies by city and over time), 100% with an MX-verified email, 94% with a phone
Web-design prospectsmaxScore: 45, requireAnyContact7 leads, all in the thin/neglected half — no published email, no marketing tech, often no mobile viewport — and all reachable. maxScore: 45 on its own matches 68 businesses; the contact floor is what removes the 61 you could not pitch to
Call listrequirePhone56 leads, 100% with a phone number
Full enrichmentemailPatternGuess, maxPagesPerSite: 6 (every crawl toggle is already on by default)100 leads plus 23 guessed email_guess addresses, kept out of email

Anything you set yourself wins, no preset ever widens your bill, and the run log names exactly what each preset applied and what it stood down on. The three guarantees are spelled out — and tested — under Presets: the three guarantees below.

What the Output tab looks like

Six curated views ship with the Actor. Pick one above the results table; the CSV / JSON / Excel export is unaffected and always carries every column you asked for.

ViewColumnsUse it for
Overview (default tab)name, category, city, state, phone, email, email_status, website, rating, lead_grade, lead_score, google_maps_urlThe twelve columns that answer "is this a lead?" — plus one click to the live Google Maps listing
Email-readyname, email, email_status, email_type, phone, website, contact_page_url, lead_scoreLoading a cold-email sequence — deliverability grade and mailbox owner side by side
Business detailsname, category, speciality, cuisine, brand, is_chain, opening_hours, price_range, rating, review_count, wheelchair, operator, cityWhat the business actually is: speciality / cuisine / brand tags, chain or independent, hours, price range, and the sparse rating / review count
Web-agency targetsname, website, website_platform, mobile_viewport, copyright_year_stale, has_meta_pixel, has_google_ads_tag, phone, email, website_platform_statusBuilding a redesign pitch list from the neglect signals
Map & sourcename, address, google_maps_url, osm_url, latitude, longitude, osm_last_edited, osm_check_date, attributionVerifying a row in one click, territory mapping, row freshness, keeping the ODbL attribution with the data
All columnsall 73, in CSV orderEverything, when you want the full table

Three deliberate choices in those views. website, contact_page_url, google_maps_url (labelled Map link) and osm_url render as clickable links; rating, review_count, lead_score, latitude and longitude render as numbers (sortable). The tri-state flags — mobile_viewport, copyright_year_stale, has_meta_pixel, has_google_ads_tag — render as text, not as a checkbox, because they are null when the site was never crawled and a checkbox cannot tell "no mobile viewport" apart from "we never looked". has_email / has_phone / has_website / is_chain are never null, so those do get the real boolean widget.

How do I find local businesses whose website is neglected enough to pitch a redesign?

Set maxScore: 45 with requireAnyContact (the Web-design prospects preset), then read the neglect columns on the rows that come back. The score cap keeps both kinds of prospect — a business with a neglected site and a business with no site at all — so add onlyWithWebsite: true if you only want the ones that already have a site to replace. Four columns carry the pitch:

Measured 2026-08-08 on Austin dentists at cap 100: the Web-design prospects preset delivered 7 rows, every one scored 45 or below and reachable on at least one channel. The score ceiling alone matched 68 businesses; the contact floor dropped the other 61 — unbilled — rather than sell you map pins you cannot contact.

  • website_platform — WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more. A builder-tier platform usually means a self-built site.
  • mobile_viewportfalse means the homepage declares no mobile viewport tag, so the site predates responsive design.
  • copyright_year_stale — a visibly out-of-date footer copyright year.
  • has_meta_pixelfalse means nobody is measuring anything on the site.

A builder-tier site with no mobile viewport, a stale copyright year and no tracking pixel is the strongest redesign signal this Actor can give you. The Web-agency targets output view shows exactly those columns and nothing else.

Read email_type before you send. It says whether the mailbox belongs to the business, to a free inbox, or to the marketing agency that already holds the account, so you can drop the leads where you would only be pitching a competitor.

How do I find local businesses that have no website, for a web-design pitch list?

Type the trade and the city and tick Businesses with no website (onlyWithoutWebsite: true), and every delivered row is a business that lists no site at all. That is the first-website pitch list for web designers and local SEO agencies, verified at 15 of 15 rows without websites on a 15-row run.

Size the list honestly: on the filter-off benchmark of 2026-08-08 (n=100), 50 of the 52 businesses with no website had no phone, no email and no social profile either, so these rows are name, address and coordinates — which is also why nobody else can cold-email them.

Expect name, address, coordinates and occasionally a phone. With no site to crawl, no email, platform or tech enrichment is possible on these rows, which is also what keeps them uncrowded: nobody else can cold-email them either. Every row still carries google_maps_url, one click to a Google Maps search for the business's name and address.

The filter is mutually exclusive with onlyWithWebsite. Setting both stops the run with an explanatory status message before anything is charged, and nothing is billed. Like every other filter it runs before billing, and a thin city can be widened with expandNearby.

What does it do?

It turns a business category and a city into one row per local business — 73 stable columns with an MX-verified email, whose mailbox it is, phone, socials, website platform and a 0-100 lead score — at $3 per 1,000 delivered leads. Measured 2026-08-08 on dentist / Austin, Texas with the website filter on: 55 businesses, 96% with a phone, 55% with an MX-verified email (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).

This actor takes a plain-English business category (e.g. "dentist," "hair salon," "roofing contractor") and a location, maps the category to the right OpenStreetMap tags, and pulls every matching business in the area. It then crawls each business's own website — following the site's own contact/about/team links (including Shopify /pages/contact and non-English slugs) within the maxPagesPerSite budget — and extracts, from pages it has already downloaded:

  • a contact email, MX-verified over DNS-over-HTTPS and graded deliverable / risky / undeliverable
  • whose mailbox it is — the business's own domain, a free inbox, or its marketing agency (emailing that one never reaches the business)
  • social profiles — Facebook, Instagram, LinkedIn, X/Twitter, YouTube
  • which website platform it runs on — WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Joomla, Drupal, HubSpot CMS, and a "Custom / Next.js" bucket for hand-built sites
  • marketing & booking tech — Meta Pixel, Google Analytics/GTM, Calendly, NexHealth, Klaviyo, WooCommerce and ~25 more
  • the homepage's title and meta description — drop-in mail-merge personalization fields (and a data tripwire: a title naming a different business exposes a wrong OSM website tag)
  • web-agency pitch signals — a missing mobile viewport tag, a stale footer copyright year
  • one-click source linksgoogle_maps_url, a Google Maps search for the business's name and address, and osm_url to the OpenStreetMap element
  • a lead score 0-100, an A-F grade, and a hot/warm/cold tier

Why use it / who's it for

  • Web design & marketing agencies — filter for a builder-tier platform (Wix, GoDaddy, Weebly) and has_meta_pixel: false to find businesses with a dated site and no tracking: the highest-intent redesign pitch there is. The Custom / Next.js value is the inverse signal — don't pitch a DIY-site rebuild to someone who already paid a developer. Build 0.1.21 adds two more neglect signals — mobile_viewport (missing = the site predates responsive design) and copyright_year_stale — plus an onlyWithoutWebsite filter that returns only businesses with no site at all: the first-website pitch list.
  • Freelancers on Fiverr/Upwork — generate a "100 dentists in Austin with verified emails" list on demand for any client vertical without building a new scraper per niche.
  • SaaS sales teams — any B2B tool sold to local businesses (booking software, payment processors, review management) can target by category and city, and has_booking_widget tells you who already has a competitor installed.
  • B2B lead-gen resellers — one actor covers any category, replacing dozens of niche scrapers.
  • Franchise & market researchers — count and map competitor density for a category in a target city (set onlyWithWebsite: false to keep every mapped location, including ones with no contact details).

How to use it

The form is ready to run as it opens: hit Save & start without touching anything and you get Austin dentists that all list a website, at a cap of 25 (raise it once the first run looks right) with Website required, Skip closed and Widen to nearby areas pre-ticked and no keyword exclusions — at $3 per 1,000 that is at most $0.075 of leads (at most $0.125 after the scheduled 2026-09-14 change to $5 per 1,000). (The count that used to be quoted here was measured on 2026-08-08 at cap 25 with a chain-exclusion list the form pre-filled at the time and no longer does, so it is not repeated.)

Filtered runs used to come up short of the cap. The cause was on our side, not the map's: only the two email filters deepened the candidate pool, so any run using Website required, Phone required or Social profile required filtered a pool sized for an unfiltered run and quietly came up short. Every filter that can drop a row now deepens the pool, and the filters that can be decided from the map data alone (website present, name keywords, permanently closed) are applied before any site is crawled, so no crawl budget is spent on a row that was never going to ship. Same-day comparisons at cap 25 in Austin (Console form, 2026-08-08, which then pre-filled a two-chain exclusion list the form no longer carries) showed Website required, Phone required and Email required each filling the cap after the fix where they had come up short before; those counts are not repeated here because the form no longer reproduces them.

Max businesses is still a ceiling, not a promise — a thin area really can run out. When that happens the run now says so in its status message with the counts behind it, instead of returning fewer rows without comment. A live example, dentist in Laramie, Wyoming at cap 25: "you asked for up to 25 and 1 could be delivered. That is the whole of Laramie, Albany County, Wyoming, United States, not a silent truncation: OpenStreetMap returned 4 'dentist' record(s) there, 1 became crawlable candidate(s), and 1 survived your filter(s) (onlyWithWebsite)." The run only claims an area is exhausted when no search came back at its row ceiling and every candidate was actually checked; otherwise it says which limit it hit and what to change.

  1. Pick a Business category from the dropdown, or leave it on Custom and type one in plain English (e.g. medspa, HVAC contractor, funeral home).
  2. Pick a City from the dropdown, or leave it on Custom and type any city on earth — city and region/country (e.g. Austin, Texas). Capitalisation and the comma are optional: RABAT MAROC and casablanca morocco resolve too.
  3. Leave Website required on unless you are doing density research — see the honest note below. (Web-design agencies can flip Businesses with no website instead to get the no-site prospect list; the two filters are mutually exclusive, and setting both stops the run with an explanatory status message (nothing is billed) with a clear error before anything is charged.)
  4. Run the actor. It geocodes the location, pulls matching places from OpenStreetMap, then crawls each business website for contact details, tech and reviews.
  5. Export as CSV, JSON, or Excel — or connect the dataset to your CRM/outreach tool via API/webhook.

Presets: the three guarantees

The preset table is at the top of this page. Behind it sit three guarantees, each of them tested:

  • Anything you set yourself wins. A preset only fills in a field you left at its default. Set maxScore: 100 alongside the web-design preset and you keep 100 — the run log says so explicitly (left as you set them: maxScore (you chose 100)).
  • A preset never widens your bill. Every preset either adds a filter (fewer rows delivered, so fewer rows charged) or turns on enrichment that adds columns to rows you were already getting. None of them clears a filter you set or raises maxItems.
  • It tells you what it did. One log line names the preset and lists exactly which settings it applied and which it left alone, and the run summary names the preset too.

Presets are Console and API: send "preset": "cold_email" from the API and you get the same behaviour. Existing API callers, tasks and schedules that send no preset are completely unaffected.

Which field wins: the precedence chain

Category and location each accept three inputs. Precedence is the same for both, highest first:

RankCategoryLocationWhy
1categories (list)locations (list)The plural list is an explicit multi-search request — it beats everything when non-empty.
2categorySelect (dropdown)citySelect (dropdown)20 categories with hand-verified OpenStreetMap tags; 25 metros verified against the live geocoder.
3category (free text)location (free text)Anything else — over a hundred more category aliases are mapped, and any city on earth geocodes.

Both dropdowns default to "" ("Custom"), which is why every input that worked before this option existed still resolves to exactly the same search.

Three search modes

ModeSetWhat happens
City (default)location (or locations) + category (or categories)The location is geocoded and every matching business inside its bounding box is returned. Several categories x several cities run as a matrix in one run.
RadiussearchRadiusKm + centerLat + centerLonSearches a circle around a coordinate instead of a city box — sales territories, franchise catchment areas, "everything within 5 km of this address". Replaces the bounding box entirely; no geocoding happens, so country stays empty.
Bring your own listwebsiteList (or startUrls)OpenStreetMap discovery is skipped entirely. The actor runs only the crawl + email verification + scoring pipeline over the websites you supply. Enrich a CRM export, a conference exhibitor list, or a list you bought elsewhere.

What happens when the city runs out of businesses?

Set expandNearby: true and the same search is widened in growing rings around the city centre until your maxItems is met or the region is genuinely exhausted. It is needed because a single city often holds fewer businesses than you asked for: measured live, "dentist" in Austin, Texas tops out near 119 rows, and Round Rock, Texas holds 21.

  • The rings are sized to the city and reach up to about 150 km.
  • Every widened row is labelled: its query_location reads within ~16 km of Round Rock, Williamson County, Texas, United States instead of the city name, so you can always tell expansion rows apart — or filter them out afterwards.
  • The same dedup and every pre-charge filter apply to ring rows, and the status message reports exactly how many delivered rows came from outside the city.
  • Measured (2026-08-15, discovery-only): dentist / Round Rock, Texas / maxItems: 100 delivered 21 rows without the flag, 100 with it — 79 labelled ring rows, 0 duplicates.

It is off by default for API callers (existing inputs keep a byte-identical pull and bill) and pre-ticked in the Console form. It never fires when you set your own radius (searchRadiusKm) or bring your own list, when the city pull was truncated (deepening, not widening, is the fix there — the status message tells you), or during a mirror outage.

Searching several categories and cities at once

categories and locations are the plural versions of category and location, and they cross into a matrix — ["dentist","orthodontist"] x ["Austin, Texas","Dallas, Texas"] is four searches in one run. The single fields keep working exactly as before; the plural ones take precedence when non-empty.

Three things make the matrix safe rather than a footgun:

  • maxItems is the whole run's budget, not a per-search one. The budget is split between the searches and results are interleaved, so the first city cannot eat the entire quota.
  • The same business found by two categories is delivered — and billed — once. Deduplication is on the OpenStreetMap object identity (osm_type + osm_id), so a clinic tagged both dentist and orthodontist appears one time.
  • 25 category x location combinations is the ceiling for one run. Above it the run fails immediately with a message naming the numbers, before any network request and before any charge — OpenStreetMap's public mirrors are a free shared resource.

Every row carries query_category and query_location, so you always know which search produced it.

Bring your own list (skip discovery)

Already have the businesses and only need the emails, socials, platform and score? Put the domains in websiteList (bare domains and full URLs both work; startUrls is accepted as an alias):

{ "websiteList": ["aloha-dental.com", "https://www.averyranchdental.com", "typotes.com"] }

No map data is fetched at all. Those rows differ from discovered rows in exactly three honest ways:

  • source is user_supplied, not OpenStreetMap;
  • attribution is null — the rows are not OSM-derived, so stamping the ODbL notice on them would be a false licence claim. The column is still present, so a mixed export keeps one stable header row;
  • name comes from the site's own <title>, falling back to the bare domain when the site does not answer. Nothing is invented; latitude, longitude, osm_id and osm_url stay empty.

Everything else — the contact crawl, MX verification, mailbox-owner classification, tech fingerprinting, scoring and every filter — behaves identically.

How complete is the data? (measured, not estimated)

On the reference run of dentist in Austin, Texas with onlyWithWebsite: true (n=55), 96% of rows carried a phone, 55% an MX-verified email, 71% a detected website platform and 100% a website. With every filter off (n=100) the same city measured 48% phone, 26% email and 34% platform, because roughly half of mapped businesses list no website to crawl.

Both reference runs were taken on 2026-08-08 (the website-filtered run and the bare API run in the Measured, not promised table above), and a roofing contractor / Denver run delivered emails on 6 of 8 rows (75%) — roughly three-quarters on website-verified trade categories.

Both reference runs are in the table below, so you can see the gaps before you pay rather than after:

onlyWithWebsite: true (n=55)schema defaults, filter off (n=100)
phone96%48%
website100%48%
email55%26%
website_platform71%34%
opening_hours80%44%
facebook73%33%
rows graded F0%51%

On that same filter-off n=100 run, the crawl-derived fields measured: contact_page_url 18%, website_title 40%, website_description 35%, mobile_viewport 40%, copyright_year_stale flagged on 3 rows; google_maps_url, osm_url and attribution sat at 100%.

With onlyWithWebsite: true the email fill runs far higher than the filter-off numbers — the reference run above measured 55% for dentists, and a roofing contractor / Denver run delivered emails on 6 of 8 rows (75%): roughly three-quarters on website-verified trade categories.

Read that second column before you run. OpenStreetMap has no website for a large share of businesses, and — measured on the filter-off run — 50 of the 52 rows with no website had no phone, no email and no social profile either (the other 2 carried only an OSM phone): name, coordinates and usually an address, nothing contactable. With the filter off you pay for those rows. onlyWithWebsite and onlyWithEmail drop non-matching rows before you are charged, so they cut your bill rather than just tidying the output. Every run's status message now reports the contactable ratio it actually delivered.

Every filter now goes further: the actor over-fetches, deepening the candidate pool and — for filters that need the site crawled — crawling extra candidates in batches (up to 24× maxItems, hard-capped at 30,000 candidates — the ceiling that makes a thin area terminate) until it has maxItems surviving rows or the pool is spent. This used to apply to onlyWithEmail / onlyVerifiedEmail only, which is why onlyWithWebsite, requirePhone and requireSocial quietly returned short. Compared 2026-08-08 on dentist / Austin, Texas at cap 25, one filter at a time, through a Console form that then pre-filled a two-chain exclusion list it no longer carries: onlyWithWebsite, requirePhone and onlyWithEmail each filled the cap after the change where each had come up short before (the exact counts are not repeated because today's form cannot reproduce them). Where the pool genuinely runs out first, the status message reports the counts instead of leaving you to guess — e.g. dentist in Laramie, Wyoming returns 1 row and says the map holds 4 dentist records there, 1 of them with a website.

The default is false (not true) so that existing API callers, scheduled tasks and density-research use cases keep getting every mapped location. The Apify console pre-fills it to true.

There is also the mirror filter, onlyWithoutWebsite: keep only businesses that list no website — the prospect list for web-design agencies pitching a first site (verified: a 15-row run delivered 15/15 rows without websites). It is mutually exclusive with onlyWithWebsite; setting both stops the run with an explanatory status message (nothing is billed) with a clear error before anything is charged. Expect these rows to be name + address + coordinates (and occasionally a phone) — with no site to crawl, no email/platform enrichment is possible.

Which business categories does OpenStreetMap cover well, and which are sparse?

OpenStreetMap covers businesses with premises a mapper walks past, and is thin on trades run from a van or a home office: the Austin bounding box held 171 dentist records but 9 plumbers, 5 electricians, 0 chiropractors and about 12 roofers. This is the honest limit of an OSM-based source, and it matters more than any field:

Two dated counts from the same city: 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged chain outlets (measured 2026-08-29), and on a 163-element sample of named Austin dentists taken the same day, 12% had not been edited since before 2020 and 13% carried a mapper's on-the-ground check_date.

  • Dense: businesses with premises a mapper walks past — restaurants, cafés, dentists, pharmacies, hairdressers, gyms, hotels, shops, banks, clinics. (171 dentist records in the Austin bounding box.)
  • Sparse: trades run from a van or a home office — plumbers (9 in the same box), electricians (5), chiropractors (0), roofers (~12). A metro of a million people can return single digits. That is what is mapped, not a bug.

95 categories (220+ terms with aliases) are curated and mapped to exact OSM tags — build 0.1.21 added pest control, photographer, moving company, self storage, funeral home and dry cleaner. Any other term is attempted as an OSM tag directly, and common phrasings are handled (auto repair shopshop=car_repair, landscaping companycraft=gardener, insurance agencyoffice=insurance). Terms with no OSM tag now fall back to a business-name keyword search of OSM, and a mapped category that returns zero tagged places in the area is rescued by the same name-keyword query. Rows found only by name are labelled category_match: "name_keyword" (tag-matched rows say category) so you can filter them out if you only trust tag-confirmed rows. Measured: pest control in Denver returned 1 labelled row on 0.1.21 where the previous build returned 0 — sparse trades stay sparse, but no longer invisible. A term matching nothing at all still returns 0 rows with a suggestion and no charge — you are never billed for a run that found nothing.

Can I use this as a business email finder for local businesses?

Yes. For any category and city it crawls each business's own website and returns the MX-verified business email wherever the site exposes one, plus the mailbox it belongs to (info@, owner name, etc.). The fill rate is measured, not promised — see "How complete is the data?" above.

Measured 2026-08-08 on dentist / Austin, Texas with onlyWithWebsite: true: 55% of 55 delivered rows carried an MX-verified email, and the Cold-email list preset (onlyWithEmail plus onlyVerifiedEmail) delivered 31 leads — the whole of Austin — 100% with an MX-verified email and 94% with a phone, at $3 per 1,000 with verification included (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).

Filters — every one of them runs before you are charged

Filters here are not a tidying step applied to an invoice you have already run up: a filtered row is dropped before it is pushed to the dataset and before the charge, so filtering cuts your bill.

Filtering does not shrink your delivery. Every filter in this table deepens the candidate pool to compensate, so maxItems means "this many rows I can use", not "this many businesses considered". Filters decidable from the map data alone — onlyWithWebsite, onlyWithoutWebsite, excludeKeywords, excludeChains, skipClosed — are applied before any website is crawled, so no crawl budget is spent on a row that was never going to ship. The rest are applied batch by batch as sites are crawled, stopping the moment enough rows survive. The pool is bounded at 6x maxItems candidates so a thin area terminates instead of crawling forever; when that bound or the map itself is what stopped the run, the status message says so and gives you the counts.

FilterKeeps only
onlyWithWebsitebusinesses that list a website (recommended, see the fill-rate table)
onlyWithoutWebsitebusinesses with no site — the first-website pitch list. Mutually exclusive with the above
onlyWithEmailrows where an email was found
onlyVerifiedEmailrows whose email passed MX verification (deliverable / risky)
requirePhoneNew. rows with a phone number (from OSM, a tel: link, or the site's schema.org markup)
requireSocialNew. rows with at least one Facebook / Instagram / LinkedIn / X / YouTube profile
requireAnyContactNew. rows reachable on at least one channel — phone or email or a social profile. The loosest contactability floor there is, and the one to reach for when you do not care which channel. It matters more than it sounds: on a bare {} run of the default search (100 rows, measured 2026-08-08) exactly 50 rows carried no phone, no email and no social profile at all, and without this switch you are billed for them. Off by default for API calls and schedules (no existing run changes); pre-ticked on the Console form since 2026-08-29
excludeKeywordsdrops businesses whose name contains any of your keywords (case-insensitive). The form arrives with the list empty. Only the name is matched, so clinic cannot knock out a business on Clinic Street. Use it to strip chains, franchises or your existing customers
excludeChainsNew (2026-08-29). drops every business the map marks as a chain outlet — a brand or brand:wikidata tag, i.e. is_chain: true — leaving the independents. Decided from the map data alone, so a dropped outlet is never crawled or billed, and the status line says how many were removed. Off by default; rows you supply yourself have is_chain: null and are never dropped by it. In chain-dense categories the 6x pool bound can bite — 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged (measured 2026-08-29), so a maxItems: 10 run delivered 6 and said so; raise maxItems to deepen the pool
skipClosedNew. drops permanently-closed premises (see below)
minScore / maxScoremaxScore is new. A score ceiling is the web-design agency filter: a low score means a thin online presence, which is exactly the redesign pitch list
minRating / minReviewCountNew — read the warning below before using these

skipClosed: what it actually removes

OpenStreetMap mappers retire a business without deleting it: the primary tag moves from amenity=restaurant to disused:amenity=restaurant, so the object keeps its name and address but no longer describes an operating business. skipClosed drops elements carrying a lifecycle-prefixed primary tag (disused:, abandoned:, was:, removed:, demolished:, razed:), a disused=yes / abandoned=yes flag, opening_hours=closed, or shop=vacant. A disused: tag alongside a live tag of the same kind (a former bank that is now a café) is not treated as closed.

Measured live on 2026-08-08 in the Austin bounding box: 199 such elements, 43 of them still carrying a business name — e.g. Chago's with disused:amenity=restaurant, Corner Store with abandoned:shop=fuel. These cannot reach a normal tagged search (["amenity"="restaurant"] cannot match disused:amenity), but they do reach the business-name fallback used for unmapped categories, which is where dead businesses were being delivered as fresh leads. Verified end to end: with skipClosed: false that closed restaurant is delivered and billed; with it true it is dropped before billing.

It is off by default so existing runs, tasks and API callers are unchanged. The Apify console pre-fills it to true.

Honest warning about minRating / minReviewCount: OpenStreetMap carries no review data at all. A rating only exists when the business publishes schema.org aggregateRating on its own website — measured at roughly 3-8% of rows. So when you set a rating or review floor, rows with no rating are DROPPED, not kept: an unknown rating is not a passing rating, and nothing is ever invented to save a row. Setting either filter will cut your result count to a small fraction — an Austin restaurant run asking for 15 rows with minRating: 4.0 crawled 86 candidates looking for them and delivered 0 rows, charging nothing (re-measured 2026-08-08 after the over-fetch was generalised, so this is the deep-search result, not a shallow one). If you need ratings on every business this is the wrong source; pair it with our Google Maps Places Scraper, or use the google_maps_url on every row.

How do I run this on a schedule and get only the new businesses?

Create an Apify schedule for the run and set onlyNewBusinesses: true: the first run is the baseline, every later run delivers only businesses that were not delivered before, and a run with nothing new delivers 0 rows and bills nothing. Point a weekly schedule at "dentists in Austin" and you get the practices that appeared since the last run, and only those.

Monitoring was hardened on 2026-08-25: the memory lives in a named key-value store in your own account, entries are pruned after 90 days with the stamp refreshed on every run, each record is capped at 50,000 keys, and a schedule created before 2026-08-29 keeps its memory because excludeChains joins the key only when you set it.

How it behaves

RunWhat happens
First run of a watchBaseline. Everything the search finds is delivered, and the status message says so in those words.
Later runsOnly businesses that were not delivered before. Already-delivered ones are dropped before the crawl, so they cost nothing and are never billed.
Nothing is new0 rows, SUCCEEDED, nothing billed. The status message reads Nothing new: all N business(es) this search found were already delivered by earlier runs of this watch ... You were not charged.

What counts as the same business. The memory key is the business's OpenStreetMap object identity (node/2135639605, way/198109528) — the same identity the run already uses to dedupe in-run, so a business that renames itself or changes domain does not come back as "new". Bring-your-own-list rows have no OSM object, so they are remembered by their registered domain — the same key websiteList is already deduped on.

A row is remembered only after it has actually been delivered. Marking happens after push_data succeeds, never at the point the candidate is chosen. That ordering is the whole safety property: every filter, the maxItems trim and above all the maxTotalChargeUsd budget trim can still remove a row between the two points, and a row marked early would be recorded as delivered, skipped by every future run, and never reach you.

What identifies a watch. Your search plus the filters that shape which businesses can survive it:

  • categories / category, locations / location, countryCode, keyword
  • searchRadiusKm, centerLat, centerLon, expandNearby
  • websiteList / startUrls
  • every option in the Filters section (onlyWithWebsite, onlyWithoutWebsite, onlyWithEmail, verifyEmails, onlyVerifiedEmail, requirePhone, requireSocial, requireAnyContact, skipClosed, excludeKeywords, minScore, maxScore, minRating, minReviewCount) — and excludeChains, which joins the key only when you set it, so a schedule created before 2026-08-29 keeps the memory it already has

Change any of those and you are running a different watch, with its own independent memory — the two never share state, and the new one starts from its own baseline. Category and location lists are compared case-insensitively and order-independently, so ["gym","dentist"] and ["Dentist","GYM"] are one watch, not two.

Deliberately not part of a watch: maxItems, sortBy, outputFields, concurrency, maxPagesPerSite. Those change how many rows you get, in what order and with which columns — not which businesses exist in the search — so tuning them never resets the memory and never re-bills you for businesses you already have.

Where the memory lives. A named key-value store called local-business-leads-monitor in your own Apify account (runs execute there), one record per watch, keyed sig-<hash>. Entries older than 90 days are pruned on every write and each record is capped at 50,000 keys, so a long-running schedule cannot grow it without bound. The timestamp stored is last seen, not first, and it is refreshed on every run for businesses still returned by the search — so the 90-day prune drops businesses that have genuinely disappeared from OpenStreetMap rather than businesses that have simply been mapped for a long time (those would otherwise age out and be re-delivered, and re-billed, as "new"). If the store cannot be read the run treats itself as a first run and says so; monitoring never fails a run.

Two honest caveats.

  1. Runs take longer with it on. Already-delivered businesses have to be skipped over, so the run searches a deeper candidate pool to still fill your maxItems with genuinely new rows.
  2. New businesses appear at the speed of OpenStreetMap. This is a mapping database, not a live business registry — a new dentist shows up here when a mapper adds it. A weekly or monthly schedule matches that cadence; an hourly one will mostly report "Nothing new" and bill you nothing for the privilege.
{
"category": "dentist",
"location": "Austin, Texas",
"maxItems": 200,
"onlyWithWebsite": true,
"onlyNewBusinesses": true
}

How do I get a Slack or Discord alert when new leads land?

Put your Slack or Discord incoming-webhook URL in webhookUrl, and every run that delivers at least one row POSTs a digest to it; a run that delivers nothing sends nothing.

Slack / Discord / JSON digests shipped on 2026-08-25; the JSON payload carries the delivered count, the run and dataset links and the first 20 rows, and a webhook that fails is named in the run's status message without failing the run.

A Slack incoming webhook and a Discord webhook each get a text message. Any other URL, such as an n8n, Make or Zapier catch hook or your own endpoint, gets JSON: {actor, delivered, run_url, dataset_url, rows[:20], text}.

Pair it with onlyNewBusinesses on a schedule and the Actor is a new-business alert service on its own, with no extra automation needed just to see the rows. Delivery is best-effort: a webhook that fails is reported in the run's status message and never fails the run.

{
"category": "dentist",
"location": "Austin, Texas",
"onlyNewBusinesses": true,
"webhookUrl": "https://hooks.slack.com/services/T000/B000/XXXX"
}

Every run comes with a report

Every run that delivers at least one row also stores a one-page HTML report: the headline numbers (businesses, % with email, % verified, % with website, average lead score), the A–F lead-grade split, the top categories or cities, the email-status split and the first 100 rows, plus the same notes the status message carries. It is saved as the REPORT record of the run's key-value store — open it from the run's Output tab (record REPORT) or through the Report: link in the run's status message. It is a single self-contained HTML file (no scripts, no external assets), so it is safe to forward, attach to an email or screenshot for a client. The dataset stays the source of truth: the report summarises what was delivered and never replaces the rows, and a run that delivers nothing writes no report.

Choosing your columns and sort order

  • outputFields — pick the columns you want (e.g. ["name","email","phone","website","lead_grade"]) and the dataset carries only those. Every row still shares one identical column tuple, so CSV headers never shift mid-export, and the columns keep the documented ROW order regardless of the order you listed them in. name and attribution are always included whatever you choose — attribution because the OpenStreetMap licence (ODbL) has to travel with the data into your CSV. Column names are matched case-insensitively (and -/space count as _, so Lead Grade works). A single unknown name is logged and ignored; if none of the names you list exists, the run stops before billing rather than charging you full price for a name-only export.
  • sortByscore_desc (default, what every previous build did), name_asc, review_count_desc or rating_desc. Sorting never removes a row; all filtering already happened. Rows missing the sort value (no rating, no review count) are placed last, because a missing value is unknown rather than zero.

Output fields

73 columns on every row (61 before 2026-08-29; the twelve new ones — has_google_ads_tag, speciality, cuisine, wheelchair, operator, osm_description, brand, brand_wikidata, is_chain, city_source, osm_last_edited, osm_check_date — are appended at the END of the column order, so an existing CSV import keeps its positions). Re-measured on a 100-row dentist / Austin, Texas run of the previous build with schema defaults: all 100 rows carried the same stable column tuple, so CSV headers don't shift mid-export — and if you narrow the export with outputFields, every row still shares one identical (smaller) tuple. Fill rates are from the onlyWithWebsite: true reference run above; anything conditional says so. Fields marked New (0.1.21) show fill rates from the 100-row filter-off benchmark instead — read them accordingly, since roughly half of those rows had no website to crawl.

In the Console, the Output tab opens on the Overview view; the All columns tab shows every field listed below, in this order. Views only change what the Console table renders — a CSV, JSON or Excel export always carries every column the run produced (or exactly the ones you named in outputFields).

Identity & location — from OpenStreetMap

FieldFillDescription
name100%Business name
category100%OSM category tag (e.g. dentist, hairdresser, lawyer)
category_match100%New (0.1.21). How the row matched your category: category = matched the mapped OpenStreetMap tag; name_keyword = found by the business-name fallback search (used for unmapped terms, and as a rescue when a mapped tag returns zero places in the area). Filter on category if you only want tag-confirmed rows
address96%Street address assembled from OSM address tags
city / state / postal_code93% / 93% / 93%Address components from the OSM addr:* tags. Since 2026-08-29 a row that sits inside the searched city's outline but carries no addr:city tag gets city (and a null state) filled from the geocoder's breakdown of the searched location — zero extra requests, and city_source says which happened. For US locations the filled state is the postal abbreviation (TX, not Texas), the same form mappers write in addr:state, so one run never splits a state into two spellings. Rows admitted by expandNearby rings and radius searches (no outline) keep their nulls
city_source100% when city is setNew (2026-08-29). osm_tag = the mapper wrote addr:city; search_area = filled from the searched location because the row is inside its administrative outline; null when city is null. Appended at the end of the column order
country100%From the geocode. Empty in radius mode and on user-supplied rows (no geocode happens there)
query_category100%New. Which of your input categories produced this row — the column that makes a multi-category run readable. Null on user-supplied rows
query_location100%New. Which of your input locations produced this row (the geocoder's resolved name, or the radius description). Null on user-supplied rows
latitude / longitude100%Coordinates
osm_type / osm_id100%OpenStreetMap source identifiers
google_maps_url100%A Google Maps listing search for the business?api=1&query=<name>, <address>, URL-encoded — so one click opens Google's search for that business rather than a bare coordinate pin (a discovered row with coordinates but no address falls back to the pin). The link carries none of this actor's data and nothing is scraped from Google: whatever rating you see there is Google's, not the rating column (see the warning below). Need Google's fields at scale? Pair with our Google Maps Places Scraper
osm_url100%New (0.1.21). Link to the row's source OpenStreetMap element — instant provenance, and the place to fix bad map data
specialityvaries by categoryNew (2026-08-29). The OSM healthcare:speciality tag as a list (["orthodontics"], ["general", "paediatric"]; first 10 values) — dentists, doctors and clinics; empty for every other trade. The keyword input matches it too
cuisinevaries by categoryNew (2026-08-29). The OSM cuisine tag as a list (["pizza", "italian"]; first 10 values) — restaurants, cafes, fast food; empty elsewhere. The keyword input matches it too
wheelchairvaries by categoryNew (2026-08-29). The OSM wheelchair tag as written (yes, no, limited)
operatorvaries by categoryNew (2026-08-29). The OSM operator tag — the company running the premises, where the mapper recorded one
osm_descriptionvaries by categoryNew (2026-08-29). The mapper's free-text description tag, clamped to 150 characters
brand / brand_wikidatavaries by categoryNew (2026-08-29). The OSM brand and brand:wikidata tags (Aspen Dental / Q4807808) — dense on pharmacies, banks and fast food, sparse on independents
is_chain100% on OSM rowsNew (2026-08-29). true when either brand tag is present, false otherwise — the one-column chain flag; set excludeChains: true to drop those rows before billing (excludeKeywords still works by name). Null on user-supplied rows

Contact

FieldFillDescription
phone96%Primary phone. OSM tag first; falls back to a tel: link or JSON-LD telephone only when OSM has none
phones73%New. All tel: numbers found on the site. May include a call-tracking number — phone stays the trusted value
website100%Business website URL. A social page in OSM's website tag is routed to that social column instead, so this is always a real site
domain100%New. Bare registered domain — the field CRMs dedup on
email55%Primary contact email, chosen by mailbox ownership then deliverability
emails55%All emails for the row, primary first (previously excluded a primary that came from OSM)
contact_page_url42%New. The exact page the primary email was found on, so you can spot-check it. Null when the email came from OpenStreetMap rather than from a crawled page
email_guess0% unless opted inNew, opt-in, and deliberately not an email. With emailPatternGuess: true, a business that has a website but publishes no address anywhere we crawled gets a pattern-guessed info@<domain> here. It is never promoted into email, never counts as has_email, never earns a lead-score point and never satisfies onlyWithEmail / onlyVerifiedEmail. Nobody checked that this mailbox exists — treat it as a lead, not an address
email_guess_confidencesameNew. A statement about the domain, never the mailbox: mx_ok (the domain does run mail servers), no_mx (it does not — the guess is almost certainly dead), unchecked (no lookup completed — email verification is off, the DNS lookup itself failed, or the run's time budget stopped it)

Email quality — verification is included, not an add-on

FieldFillDescription
email_status100%deliverable / risky / undeliverable when verification is on and a candidate exists; missing when no email was found; found / missing when verifyEmails is off. Turning verification off changes the vocabulary
email_provider55%Mailbox host where identifiable (Google Workspace, Microsoft 365...). Only when verifyEmails is on and an email was found
email_domain_match55%Whether the email's domain matches the website's
email_type55%own_domain / free_mail / third_party / unknown. third_party means the address belongs to the business's marketing agency or web designer — it is deliverable but does not reach the business, and it is scored at half weight. Measured on the reference run: 12% of all harvested addresses were third-party, but only 10% of primary emails, because candidates are ranked by mailbox ownership before one is promoted
email_types55%The same classification for every entry in emails

Socials

FieldFillDescription
facebook73%Facebook profile URL
instagram55%Instagram profile URL
youtube25%New. YouTube channel URL
linkedin22%New. LinkedIn company/profile URL
twitter18%New. X/Twitter profile URL

Share buttons, tracking pixels and embedded posts are filtered out, so these are profile URLs rather than facebook.com/tr or instagram.com/p/....

Website & marketing tech

FieldFillDescription
website_platform71%CMS / site builder: WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, Custom / Next.js, Custom / React. WordPress plugins (Elementor, WP Rocket, Divi) report as WordPress, not as their own platform
website_platform_status100%New. Why website_platform is what it is — and what actually happened to the site: detected, unknown (a page loaded, nothing recognisable on it), site_blocked (the site answered but refused us — a 403, a bot-check page or a login wall), not_found (every URL that answered came back 404 or 410 — the address in the listing is dead, the host serving it is not), site_error (the site answered with a server error), no_page (it answered, but never with a readable web page), site_unreachable (nothing answered at all — DNS failure, refused connection, broken TLS or a timeout), no_website, not_crawled (crawling was off), crawl_error (our own crawler failed on that site — worth reporting). Only site_unreachable means the business has no reachable website, and only crawl_error says nothing about the site; every other value tells you the site exists and why we could not read it, so a blocked or 404 site is still a live prospect. A null platform is explained rather than unexplained
platform_version27%New. Only when the site's own meta generator names the platform and a version — never a plugin's version passed off as the platform's
site_generator42%New. The raw meta name="generator" string
tech84%New. Marketing/booking/commerce tech found on the page (Meta Pixel, Google Analytics, GTM, Calendly, NexHealth, Klaviyo, WooCommerce, Yelp Reviews, live chat...)
has_meta_pixel / has_google_analytics / has_booking_widget / has_google_ads_tagsee noteNew. Booleans derived from tech. Tri-state: true/false when the site was read, null whenever no page was read — unreachable, blocked, a bot-check interstitial, or not crawled at all. false never means "we could not check", so a no-pixel filter cannot fill up with sites nobody ever saw. has_google_ads_tag (new 2026-08-29, appended at the end of the column order) is true when the Google Ads conversion tag (googleadservices.com / gtag AW-…) is on the page — a wiring signal, not proof of live spend (see the FAQ)
website_title40%New (0.1.21). The crawled homepage's <title> — a drop-in mail-merge personalization field. Also a data tripwire: a title that clearly names a different business means OSM's website tag is wrong, and this column lets you catch that before you hit send
website_description35%New (0.1.21). The homepage's meta description — the other personalization field, and the same tripwire
mobile_viewport40%New (0.1.21). Whether the homepage declares a mobile viewport tag. A site without one predates responsive design — a concrete web-agency pitch signal
copyright_year_stale3%New (0.1.21). Flags a visibly out-of-date copyright year in the footer — a small but unambiguous neglected-site signal (flagged on 3 of the 100 benchmark rows)

Changed in this build — website_platform_status got more precise. It used to fold every crawl failure into one value, site_unreachable, which read as "this business has no working website". Most of those sites were alive and simply refused a datacenter request. The value set now separates site_blocked, not_found, site_error, no_page and crawl_error from site_unreachable, which is reserved for hosts that produced no HTTP response at all. If you have a saved filter, Make/Zapier step or script that matches website_platform_status == "site_unreachable", it will now match fewer rows — by design: the rows it stops matching are live websites. Match on the list above (or on email_status) to get the old, broader set. Pricing does not change, and no run delivers more rows than its maxItems. This build does read pages the previous one threw away (HTML served as text/plain, as XHTML, or with no Content-Type at all), so a site that used to yield no email can now yield one — with onlyWithEmail or minScore on, that can change which businesses fill your order, always within the same cap.

Business detail & reviews

FieldFillDescription
opening_hours80%From OSM; gap-filled from the site's JSON-LD only when OSM has none
rating7%New, and sparse — see the warning below. Star rating (1-5) the business publishes in its own schema.org markup
review_count7%New, sparse. Review count from the same markup
price_range24%New. schema.org priceRange (e.g. $$)

Honest warning about rating / review_count: these are not Google Maps ratings. They are only present when a business publishes aggregateRating in its own JSON-LD, and most do not. Measured on-platform: 7% of website-bearing rows for dentists in Austin (4/55) and 13% for roofing contractors in Denver (1/8); the filter-off benchmark (n=100) measured 4%. An earlier 12-row roofer sample hit 33%, so the rate swings wildly with category, city and sample size — assume under 10% and treat anything higher as luck. Do not build a workflow that needs a rating on every row. If you need star ratings and review counts for every business, this is the wrong source — pair it with our Google Maps Places Scraper; every row's google_maps_url also opens a Google Maps search for the business in one click. OpenStreetMap carries no review data at all, and this actor never invents a substitute: lead_score is a data-completeness score, not a customer rating.

Lead scoring

FieldFillDescription
lead_score100%0-100 completeness/reachability score
lead_grade100%A ≥80, B ≥65, C ≥50, D ≥35, else F
lead_tier100%hot ≥75, warm ≥50, else cold
score_breakdown100%Per-component points, plus signals_from — the list of extra signals that actually scored — so the score is auditable
has_email / has_phone / has_website100%Booleans for quick filtering

Scoring rubric (sums to a true 100, so grade A is reachable — the filter-off benchmark's top row scored the full 100): email deliverable 40 / risky 25 / present-but-unverified 15 — halved when email_type is third_party; phone 20; website 15; socials 5 each capped at 10; extra signals 5 each capped at 15, drawn from website platform detected, opening hours, marketing tech found, and a published star ratingscore_breakdown.signals_from names the ones that counted.

Provenance

FieldFillDescription
source100%OpenStreetMap for discovered rows, user_supplied for rows that came from your own websiteList
attribution100% on OSM rowsThe ODbL attribution string, so the licence travels with an exported CSV. Null on user_supplied rows — they are not OpenStreetMap data, so attaching the notice would be a false licence claim. The column is always present either way
osm_last_edited100% of OSM rows (measured 2026-08-29 on three runs: 25/25 dentists in Austin, 20/20 restaurants in Chicago)New (2026-08-29). ISO date of the element's last edit by any mapper, read from OpenStreetMap's own edit metadata (out meta) on the same request as before. Sort or filter on it to skip rows nobody has touched in years; the last column but one
osm_check_datevaries by category (measured 2026-08-29: 4/25 dentists in Austin, 3/20 restaurants in Chicago)New (2026-08-29). The mapper's check_date tag — the day someone confirmed the business on the ground (used mostly for opening hours); present on a minority of rows. The last column
enriched_from_website33%New. Lists only the fields where an OSM-overlapping value was taken from the site instead — rating, review_count, price_range (OSM carries none of these) plus phone / opening_hours / email where OSM was empty and the site's schema.org data filled the gap. It is not a full provenance map: emails, the five socials, website_platform, platform_version, site_generator and tech are always crawled from the site and are deliberately not repeated here

Example output

Real rows from a run with onlyWithWebsite: true. This is a best-case slice, not a typical one — see the measured fill rates above; on the website-filtered reference run roughly half of rows carry an email.

NamePhoneEmailStatusTypePlatformTechScore
Aloha Dental+1-512-707-7300riverside@aloha-dental.comdeliverableown_domainWordPressMeta Pixel, GTM, Yelp95 (A)
Avery Ranch Dental+1-512-246-7645smile@averyranchdental.comdeliverableown_domainWordPressGTM, reCAPTCHA95 (A)
Aviva Dental Care+1 512 852 8528dr.apurva@avivadentalcare.comdeliverableown_domainWordPressGoogle Analytics90 (A)

How to read the output

One row per business, 73 columns, always in the same order. The Console's Output tab opens on the Overview view; the other views are the same rows with a different set of columns in front:

  • Overview — the first-look table: Name, Category, City, State, Phone, Email, Email status, Website, Rating, Grade, Lead score, Map link. Read lead_grade (A–F) first, then email_status.
  • Email-ready — only what a cold-email sequence needs: the address, its deliverability grade (email_status), whose mailbox it is (email_type), and the page it was found on.
  • Business details — what the mapper wrote about the business: speciality, cuisine, brand, is_chain, opening hours, price range, and the rating / review count when the business's own site publishes one (roughly 3–8% of rows).
  • Web-agency targets — the redesign-pitch signals: site builder, mobile viewport, stale copyright year, Meta Pixel and Google Ads tag, with website_platform_status saying whether the site was actually read.
  • Map & source — where the row is (address, coordinates, Google Maps and OpenStreetMap links), how fresh it is (osm_last_edited, osm_check_date) and the ODbL attribution that has to travel with the data.
  • All columns — every column, in CSV order.

Column labels follow one vocabulary across every Flash Scrape actor: Name, Category, City, State, Phone, Email, Email status, Website, Lead score, Grade, Map link. Booleans (has_email, has_phone, has_website, is_chain) are real true/false values; numbers (rating, review_count, lead_score, latitude, longitude) are real numbers, never strings; the two OSM dates are YYYY-MM-DD.

There are no derived or duplicated columns: address is already the full one-line postal address next to its city / state / postal_code / country parts, and google_maps_url is a ready-made link, so nothing needs assembling on your side.

Exporting just one view. Views change what the Console shows, not what its Export button downloads — a Console export always carries every stored column (Apify's dataset-schema docs: views only affect the Console display). To download only a view's columns use the view parameter of the dataset-items API, e.g. https://api.apify.com/v2/datasets/<datasetId>/items?view=email_ready&format=csv (view keys: overview, email_ready, business_details, web_agency, map_source, all_columns). Leave view off, or use all_columns, for the full table; outputFields in the input narrows what is stored in the first place.

The run report and status message. Every run that delivers rows ends with the same one-line status — Done — N leads delivered for <what>. <coverage / filter notes> Report: <link>. — and stores the REPORT HTML page whose table shows exactly the Overview columns (up to 10 of them: the sparse rating and the Map link are the first to be left out when the table is full), with tiles for businesses, % with email, % verified (when email verification ran), % with website and the average lead score.

Input examples

A preset plus the two dropdowns — the shortest useful input there is:

{
"preset": "cold_email",
"categorySelect": "hair salon",
"citySelect": "Casablanca, Morocco",
"maxItems": 100
}

The preset sets onlyWithEmail and onlyVerifiedEmail; everything else stays at its default. Add any field you want and yours wins — {"preset": "web_design", "maxScore": 100} keeps maxScore: 100 and the log says which setting the preset stood down on.

One category, one city — the classic run, unchanged:

{
"category": "hair salon",
"location": "Miami, Florida",
"maxItems": 200,
"crawlEmails": true,
"onlyWithWebsite": true,
"onlyWithEmail": false,
"verifyEmails": true,
"maxPagesPerSite": 3,
"concurrency": 8
}

Several categories across several cities, contactable rows only, narrow columns:

{
"categories": ["dentist", "orthodontist"],
"locations": ["Austin, Texas", "Dallas, Texas"],
"maxItems": 200,
"onlyWithWebsite": true,
"requirePhone": true,
"excludeKeywords": ["Aspen Dental"],
"skipClosed": true,
"outputFields": ["name", "email", "phone", "website", "lead_grade", "query_location"],
"sortBy": "name_asc"
}

A 5 km sales territory around one address:

{
"category": "restaurant",
"searchRadiusKm": 5,
"centerLat": 30.2672,
"centerLon": -97.7431,
"maxItems": 150,
"onlyWithEmail": true
}

Web-design prospect list — businesses with a site, but a weak one:

{
"category": "hair salon",
"location": "Lyon, France",
"countryCode": "fr",
"onlyWithWebsite": true,
"maxScore": 45,
"skipClosed": true
}

Enrich your own list (no map data fetched at all):

{
"websiteList": ["aloha-dental.com", "averyranchdental.com", "typotes.com"],
"verifyEmails": true,
"emailPatternGuess": true
}

How much does it cost?

This Actor is pay-per-result and costs $3 per 1,000 delivered leads, which is $0.003 per lead on the free plan and less on paid plans (the Pricing tab always carries the current rate). You are charged for delivered businesses only, not for API calls or compute. There is no subscription, and new Apify users get platform free credits to test with. A scheduled pricing record raises this to $0.005 per delivered lead on 2026-09-14 (notified 2026-08-30); the Pricing tab on this page is authoritative.

Read from the live pricing on 2026-08-29 and re-read 2026-09-05: $0.003 per delivered lead on the free plan and $0.0027 (Bronze) down to $0.0021 (Diamond) on paid plans, so $5 buys 1,666 delivered leads; a change to $0.005 per result is scheduled for 2026-09-14 (notified 2026-08-30) and the Pricing tab is authoritative. For comparison, lukaskrivka/google-maps-with-contact-details charges $0.005 per place, $0.0025 per contact enrichment and $0.10 per verified email on the free plan (Store pricing read 2026-09-05), falling to $0.004 per verified email on Bronze.

$5 buys 1,666 delivered leads at that rate (1,000 after the scheduled 2026-09-14 change to $0.005 per lead). A free-plan visitor can run the form as it opens (25 dentists in Austin, at most $0.075 of leads plus a few cents of compute — well inside Apify's monthly free credit), download the CSV and judge every column before paying anything.

A run's total is simply rows delivered x the per-lead rate. Be aware what the rows contain: on the reference run with onlyWithWebsite: true, ~55% carried an email, so 500 rows ≈ 275 emailable leads (roughly 1.8x the per-lead rate per emailable lead). With the website filter off, 26% carried an email on the filter-off benchmark (roughly three-quarters on website-verified trade categories). Filters run before billing, so onlyWithEmail: true is the cheapest way to buy emails specifically — and the adaptive over-fetch delivers as close to maxItems email rows as the city allows (measured 2026-08-08: asked 25, delivered 25 in Austin; the same input delivered 19 before the pool-depth fix, and the whole city tops out at 31, which the run tells you when you ask for more).

How it compares

Read from Apify's public Store API on 2026-09-05. Nearly every alternative is a Google Maps scraper: lukaskrivka/google-maps-with-contact-details (87,957 users, 4.63 from 221 reviews, $5 per 1,000 places, with contact enrichment at $2.50/1,000 and email verification at $100/1,000 on Apify's free plan, falling to $4/1,000 on Bronze), s-r/google-maps-contact-details (166, no reviews, $4/1,000 plus $2/1,000 enrichment), leadharbor (40, no reviews, $3/1,000 MX-checked), code-node-tools (33, no reviews, $1.10/1,000) and jurassic_jove (99, no reviews, $20/1,000). Paid Apify plans get tiered discounts on several of these, this Actor included, so read every figure off the Pricing tab for the plan you are on.

This Actor (32 users on the public Store API cache, 5.0 from 3) is $3 per 1,000 with MX verification included — a change to $5 per 1,000 is scheduled for 2026-09-14 — but discovery runs on OpenStreetMap, not Google Maps, so there are no Google star ratings and van-based trades are thinly mapped.

For comparison, lukaskrivka/google-maps-with-contact-details charges $0.005 per place plus $0.0025 per contact enrichment plus $0.10 per verified email on Apify's free plan (its live Store pricing record, read 2026-09-05) — the verification event alone is $100 per 1,000 there, falling to $4 per 1,000 on Bronze. Here, MX verification, mailbox-ownership classification and lead scoring are all included in the single per-lead rate.

Where does the data come from, and what licence is it under?

Business listings come from OpenStreetMap under the Open Database License (ODbL) v1.0, and every contact detail comes from the business's own public website, which ODbL does not cover. © OpenStreetMap contributorshttps://www.openstreetmap.org/copyright. Every row carries source and attribution fields; keep them if you redistribute or publish the data, as ODbL requires attribution. The columns added on 2026-08-29 from OSM tags — speciality, cuisine, wheelchair, operator, osm_description, brand, brand_wikidata, is_chain, osm_last_edited, osm_check_date — are OSM data under the same licence (city_source and has_google_ads_tag are ours).

The website-crawled fields — email, emails, socials, website_platform, tech, rating, review_count, price_range, phones — are not OSM-derived. They come from each business's own public website and are not covered by ODbL. Those fields are site-derived on every row. The enriched_from_website column is narrower than that: it flags only the fields that OSM could have supplied but didn't, so treat the list above — not that column — as the ODbL boundary.

Frequently asked questions

How do I find local business leads with verified email addresses?

Run this Actor with a category, a city and onlyWithEmail: true, and every delivered row carries an email address; with verification left on (the default) each one is MX-checked over DNS-over-HTTPS and graded deliverable, risky or undeliverable in email_status. Measured on dentist / Austin, Texas with onlyWithWebsite: true, 55% of the 55 delivered rows carried an MX-verified email and 96% carried a phone; a roofing contractor / Denver run delivered emails on 6 of 8 rows.

Two things make those addresses usable rather than just present. email_type says whether the mailbox is the business's own_domain, a free inbox, or its marketing agency (mail to an agency mailbox never reaches the business). And onlyWithEmail and onlyVerifiedEmail run before billing, so a row without an email is dropped rather than charged, which makes them the cheapest way to buy emails specifically.

Those figures date from 2026-08-08, when the Cold-email list preset (onlyWithEmail plus onlyVerifiedEmail) delivered 31 leads for Austin dentists — the whole city that day — 100% with an MX-verified email and 94% with a phone.

This Actor reads only publicly available data: OpenStreetMap records, which are an open-data project licensed under ODbL, and each business's own public website, which is the same information anyone could read by visiting the site. There is no login, no private data, no anti-bot circumvention and no Google Maps terms-of-service exposure. Keep the attribution column with the data if you republish it, because ODbL requires attribution.

Does it work outside the US? How should I type the location?

Yes, it works anywhere OpenStreetMap covers, and you can type the location however you like. The raw string is tried first, and only if that finds nothing is it automatically re-spelled (Title Case, and a comma inserted before the trailing country word) before the run is given up on. RABAT MAROC, Rabat Maroc, casablanca morocco and Rabat, Morocco all resolve to the same place. If the run still cannot geocode, the status message lists every spelling it tried instead of a generic failure.

Verified: Rabat with countryCode mt resolves to Rabat, Western Region, Malta, while Morocco wins without the pin. The country column was fixed on 2026-08-08 — every Moroccan row used to ship a three-script run-on where Morocco belonged — and Austin, Texas and Lyon, France were unaffected.

If your city name exists in more than one country — Rabat is a city in both Morocco and Malta, Cambridge in both the UK and the US — set countryCode to an ISO 3166-1 alpha-2 code (ma, mt, us, fr, gb) to pin the geocoder. Verified: location: "Rabat" with countryCode: "mt" resolves to Rabat, Western Region, Malta; without it, Morocco wins. The pin is enforced: the fallback geocoder (Photon) has no country filter, so it is skipped whenever countryCode is set. A city that does not exist inside that country ends the run with a clear message and no charge, instead of quietly returning the same-named city elsewhere (location: "Austin" + countryCode: "ma" used to deliver Austin, Texas).

One non-US bug is fixed as of this build: the geocoder is now asked for English place names. It previously answered in the local language(s), and since the row's country column is derived from the geocoder's answer, every Moroccan row used to ship country: "Maroc ⵍⵎⵖⵔⵉⴱ المغرب" — a three-script run-on where Morocco belonged. Measured and fixed on 2026-08-08; Austin, Texas and Lyon, France are unaffected. Note that city and address still come straight from the OpenStreetMap tags a local mapper wrote, so a Moroccan row can legitimately read Témara تمارة — that is the map's own data and it is not rewritten.

How fresh is the data?

Every email, social profile, website platform and tech signal is crawled live from the business's website during the run, so that half of the row is current at run time; the OpenStreetMap half — the name, the address and the phone where the map supplies it — is only as fresh as the last mapper edit, which each row states in osm_last_edited.

That makes freshness something you can filter on rather than trust:

  • osm_last_edited — the day any mapper last changed the element, read from the map's own edit metadata. On a 163-element sample of named dentist elements in the Austin bounding box (measured 2026-08-29), 12% had not been edited since before 2020.
  • osm_check_date — the mapper's check_date tag, set when someone confirmed the business on the ground. It was present on 13% of that same sample, and the share varies by category and city.

Sort or filter on osm_last_edited to skip rows nobody has touched in years, and keep skipClosed on to drop premises mappers have retired.

How do I know whether a business email will bounce before I send to it?

Read the email_status column: with verification on (the default) every address is checked over DNS before delivery and graded deliverable, risky or undeliverable, and that verification is included in the per-lead price rather than sold as an add-on ($0.003 today, $0.005 from 2026-09-14 under a scheduled pricing record). Set onlyVerifiedEmail: true to keep only the rows that passed — deliverable and risky — and drop the ones proved dead. Turn verifyEmails off and the column falls back to found / missing; an address the run ran out of time to check also ships as found, unverified, and the status message says so.

On the 2026-08-08 reference run 55% of the 55 delivered rows carried an MX-verified email; 12% of all harvested addresses belonged to a third party (a marketing agency or web designer) but only 10% of primary emails did, because candidates are ranked by mailbox ownership before one is promoted.

Four checks run, all over DNS and none over SMTP: syntax, a mail-server (MX, or implicit-MX A record) lookup on the domain, a role-address check (info@, sales@ and the like grade risky), and a disposable-domain list (grades undeliverable). The MX lookup goes over DNS-over-HTTPS to the public Google and Cloudflare resolvers, so no mail server is ever contacted and no proxy is needed. No mailbox is ever probed — that is what keeps verification proxy-free, fast and included in the price, and it is also what it cannot see: an address on a catch-all domain, or a mailbox that was deleted while the domain kept its mail servers, can still grade deliverable. Read deliverable as "the domain accepts mail and the address is not a known-bad shape", and warm a new list with a soft first send. When the MX lookup itself fails the row grades risky rather than undeliverable, so a busy resolver never deletes a good lead.

Can it tell me whether a business is running ads?

Partly: has_google_ads_tag and has_meta_pixel tell you whether the ads or conversion tag is present on the crawled page, which is a wiring signal rather than proof of live ad spend. Each is true, false, or null when no page was read. Sites keep the tag long after a campaign ends, and a business can run ads with no on-site tag at all. For actual live creatives, feed each row's domain into our Competitor Google Ads Scraper (its queries input takes domains), which lists what a domain is currently running from Google's Ads Transparency Center.

has_google_ads_tag was appended on 2026-08-29; both flags are tri-state (null when no page was read), and the underlying tech column — Meta Pixel, Google Analytics, GTM, Calendly, Klaviyo and about 25 more — was filled on 84% of rows on the 2026-08-08 reference run (n=55).

Why is email empty on some rows?

Three distinct reasons, and email_status plus website_platform_status tell you which: the business has no website in OSM (no_website), no page could be read from the site it does have (site_blocked, not_found, site_error, no_page, or site_unreachable — and only that last one means there is nothing there to reach), or the site simply does not publish an address anywhere on the pages crawled (missing). Cloudflare-obfuscated addresses are decoded, and since 0.1.21 the crawler follows the site's own contact/about/team links — including Shopify /pages/contact and non-English slugs — within the maxPagesPerSite budget. Raising maxPagesPerSite finds a few more.

Sized on 2026-08-08: with every filter off, 26% of 100 rows carried an email because roughly half of mapped businesses list no website; with onlyWithWebsite: true it was 55% of 55, and 8 of those 55 sites refused datacenter requests (site_blocked) — a live business, not a dead one.

Does it get star ratings and review counts?

Only when a business publishes them in its own website markup, and these are never Google Maps ratings: measured at 7-33% of website-bearing rows depending on the category, and 4% on the latest filter-off benchmark. Do not build a workflow that needs a rating on every row; see the warning in the output section. OpenStreetMap has no review data, so if ratings are essential for every row, pair this with our Google Maps Places Scraper — every row's google_maps_url opens a Google Maps search for the business in one click.

Measured 2026-08-08: 4 of 55 Austin dentists (7%) published a rating, and an Austin restaurant run asking for 15 rows with minRating: 4.0 crawled 86 candidates and delivered 0 rows, charging nothing — a rating floor drops unrated rows rather than keeping them.

How do I build a list of businesses running WordPress, Wix, Squarespace or Shopify in one city?

Run the category and city with onlyWithWebsite: true and read the website_platform column, which names the CMS or site builder behind each site (WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, and Custom / Next.js or Custom / React for hand-built sites). It was filled on 71% of rows on the website-filtered reference run, and website_platform_status explains every row where it is empty.

There is no platform input to filter on, so the pattern is: pull the city, then filter the CSV on website_platform. Knowing a business runs on Wix, GoDaddy or Weebly rather than WordPress or a custom build lets you pre-qualify before picking up the phone, and combining it with has_meta_pixel: false and has_booking_widget: false narrows the list to businesses visibly under-invested in their web presence.

That 71% is the 2026-08-08 reference run (dentist / Austin, Texas, onlyWithWebsite: true, n=55); with every filter off the same city measured 34% of 100 rows, because roughly half of mapped businesses list no site to fingerprint.

Only independents? How do I leave the chains out?

Set excludeChains: true. OpenStreetMap marks a chain outlet with a brand tag (and usually brand:wikidata, both seeded from the name-suggestion-index — brand=Aspen Dental, brand:wikidata=Q4807808); that is exactly what sets the is_chain column, so the filter is the column applied before billing: a brand-tagged outlet is dropped before it is crawled, pushed or charged, and the status line and run report say how many were removed. It is a mapper-written tag, so an outlet nobody has branded on the map still comes through — pair it with excludeKeywords (["Aspen Dental"]) for a name you know. Rows you supply yourself (websiteList / startUrls) carry no brand tag (is_chain: null) and are never dropped by it.

Added 2026-08-29 and measured the same day: 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged, so a maxItems: 10 run with excludeChains delivered 6 independents and said why — raise maxItems to deepen the 6x candidate pool in chain-dense categories.

{ "category": "dentist", "location": "Austin, Texas", "excludeChains": true, "maxItems": 50 }

How much does it cost to scrape 1,000 local business leads?

1,0