Local Business Leads Scraper: Verified Business Emails & Finder
Pricing
from $2.40 / 1,000 business leads
Local Business Leads Scraper: Verified Business Emails & Finder
Local business leads scraper, local business email finder and local business email scraper: any category, any city. Verified business emails (MX-checked), phones, socials on every row - local leads with emails, no API key, no proxy, $3 per 1,000. Business email finder for agencies.
Pricing
from $2.40 / 1,000 business leads
Rating
5.0
(3)
Developer
Flash Scrape
Maintained by CommunityActor stats
6
Bookmarked
33
Total users
22
Monthly active users
4 hours ago
Last modified
Categories
Share
Local Business Leads Scraper is a pay-per-result lead-generation actor that finds local businesses in any category and any city on earth and delivers MX-verified emails, phones, social profiles, website platform and a 0-100 lead score — built on OpenStreetMap, so it needs no API key, no proxy and no login.
Try it: Find local business leads with emails by category
Local business leads scraper and business email finder — verified business emails, phones and socials for any category in any city; a local business email scraper built for agencies.
What you get per lead
One row per business, 73 stable columns, with a Google Maps link and OpenStreetMap provenance on every row. Fill rates are measured on the reference run — dentist / Austin, Texas, onlyWithWebsite: true, n=55, 2026-08-08 (google_maps_url and osm_url also measured 100% on the same day's filter-off n=100 run); the full story is under Measured, not promised below.
| Field | Filled | What it is |
|---|---|---|
| MX-verified email + whose mailbox | 55% | email graded deliverable / risky / undeliverable; email_type says own domain, free inbox or the business's marketing agency |
| Phone | 96% | OSM tag first, then the site's own tel: link |
| Social profiles | Facebook 73% · Instagram 55% · YouTube 25% · LinkedIn 22% | profile URLs the business's own site links |
| Website platform + tech | 71% | WordPress, Wix, Shopify, Squarespace… plus Meta-pixel / Google-Ads-tag flags |
| Lead score + grade | 100% | 0-100 completeness/reachability score and an A-F grade — not a customer rating |
| Google Maps link | 100% | google_maps_url opens a Google Maps search for the business in one click; nothing is scraped from Google |
| OpenStreetMap provenance | 100% | osm_url, coordinates and the ODbL attribution string |
$3 per 1,000 delivered leads ($0.003 per lead) on the free plan; paid plans pay less (Pricing tab). Every filter — onlyWithEmail, onlyVerifiedEmail, onlyWithWebsite, requirePhone, requireAnyContact and the rest — runs before you are charged, MX verification is included in that price, and a run that finds nothing charges nothing. A scheduled pricing record raises this to $0.005 per delivered lead on 2026-09-14 (notified 2026-08-30); the Pricing tab on this page is authoritative.

This Actor takes a business category and a city and returns one row per local business, 73 stable columns wide and every row scored 0-100 — with an MX-verified email, whose mailbox that email is, the phone, the social profiles and the website platform wherever the business's own site exposes them (measured fill rates below). It costs $3 per 1,000 delivered leads ($0.003 per lead on the free plan; the Pricing tab always carries the current rate, and a change to $0.005 per lead is scheduled for 2026-09-14), and email verification is part of that price rather than an add-on.
Dentists, gyms, lawyers, plumbers, roofers, salons, real estate agencies, restaurants and 50+ more curated categories, plus any term you type. No API key, no proxy, no separate scraper per niche.
Key facts:
- $3 per 1,000 delivered leads ($0.003 per lead; rising to $0.005 per lead on 2026-09-14 under a scheduled pricing record) — MX email verification, mailbox-ownership classification and lead scoring are included in the single per-lead rate; filtered rows are dropped before billing, and a run that finds nothing charges nothing.
- No API key, no proxy, no login — discovery runs on OpenStreetMap and enrichment crawls each business's own public website, from datacenter IPs.
- 95 curated categories (220+ terms), any city on earth — plus a radius search around a coordinate, or bring your own website list and skip discovery entirely.
- 73 stable columns on every row — export as CSV, JSON or Excel, or connect the dataset to your CRM via API/webhook; column headers never shift mid-export.
- Built-in monitoring and alerts —
onlyNewBusinessesturns a schedule into a new-business alert, andwebhookUrlposts a Slack / Discord / JSON digest whenever a run delivers rows.
What it does not do, stated up front so nothing here is a surprise after you have paid:
- It does not scrape Google Maps. Listings come from OpenStreetMap, so there are no Google star ratings and no Google review counts. Pair it with our Google Maps Places Scraper when you need those.
- The
ratingandreview_countcolumns are sparse. They exist only when a business publishes a rating in its own website markup, measured at about 7% of rows (4 of 55 on the 2026-08-08 reference run), and they are never a Google rating. - Van-based trades are thinly mapped. OpenStreetMap held 171 dentists in the Austin bounding box but 9 plumbers and 5 electricians. The coverage section below names which categories are dense.
- A guessed email is never sold as a verified one. A pattern-guessed address stays in
email_guessand is never promoted intoemail.
Use from an AI agent
- MCP: point Claude, ChatGPT, Cursor or any MCP client at
https://mcp.apify.com?tools=flash_scraper/local-business-leads; the tool is named after the Store slug and takes this actor's input unchanged. Keywords for the server'ssearch-actorstool: business leads, local business, lead generation, business emails, verified business emails. Tool-name spellings, payment without an Apify token and measured timings: Use it from an AI agent (MCP). - Smallest useful call (Python
apify-client; the same JSON works in the Console, the REST API and n8n/Make/Zapier):
from apify_client import ApifyClientclient = ApifyClient("<APIFY_TOKEN>")run = client.actor("flash_scraper/local-business-leads").call(run_input={"category": "dentist", "location": "Austin, Texas", "maxItems": 10, "crawlEmails": False})rows = client.dataset(run["defaultDatasetId"]).list_items().items
- Output contract: the same 73 columns on every row, headers never shift mid-export;
email_guessis never promoted intoemail. The full field list with measured fill rates is under What data you get; every run also writes a machine-readableRUN_SUMMARYrecord to its key-value store.
What you get
- Any category, any city — 220+ category terms (95 curated categories plus their aliases) mapped to exact OpenStreetMap tags, any other term attempted directly and then rescued by a business-name search, and any city on earth geocoded (
Austin, Texas,casablanca morocco,Dubai UAE). - Emails that are verified, not guessed — every address is MX-checked over DNS-over-HTTPS and graded
deliverable/risky/undeliverable. Verification is part of the price, not an upsell. - Whose mailbox it is —
own_domain, a free inbox, or the business's marketing agency. Emailing an agency mailbox never reaches the business, so this column decides whether a lead is worth a send. - Redesign pitch signals —
website_platform(WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more),mobile_viewport,copyright_year_stale,has_meta_pixel,has_google_ads_tag. A builder-tier site with no tracking is the highest-intent pitch there is. - A score you can audit — 0-100, an A-F grade, a hot/warm/cold tier, and a
score_breakdownobject showing the arithmetic that produced it. - Five ready-made Output views — Overview, Email-ready, Web-agency targets, Map & source, All columns. Pick one above the results table instead of scrolling a wall of 73 raw columns.
- Filters that cut your bill, not just the table — every filter drops the row before it is pushed and before it is charged, and the run log names each filter and how many rows it removed.
- Schedulable — only what is new —
onlyNewBusinessesturns a weekly schedule into a new-business alert: the first run is the baseline, every later run delivers only businesses that were not delivered before, and a run with nothing new delivers 0 rows and bills nothing. How it works. - One-click provenance —
google_maps_url(a Google Maps search for the business's name and address, so it opens the listing search rather than a bare map pin) andosm_urlon every discovered row, plus the ODbL attribution the licence requires you to keep.
Is there a Google Maps scraper alternative that does not scrape Google Maps?
Yes: this Actor discovers businesses on OpenStreetMap and then crawls each business's own public website, so no listing, rating or review is ever taken from Google Maps and no Google page is ever scraped. That is the whole design, not a setting you have to switch on.
On the 2026-08-08 reference run (dentist / Austin, Texas, onlyWithWebsite: true, n=55) that design delivered a phone on 96% of rows, an MX-verified email on 55% and a detected website platform on 71%, with no proxy input and no proxy cost, at $3 per 1,000 delivered leads (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).
What that buys you:
- No proxy bill and no proxy setup. There is no proxy input on this Actor and no proxy cost in a run: OpenStreetMap's public endpoints and most business websites answer datacenter IPs, so runs go direct. A minority refuse them — measured at 8 of the 55 website-bearing rows in the reference run — and those rows say
site_blockedrather than pretending the business is gone. - No Google Maps terms-of-service exposure. What you receive is open-licensed map data plus pages the businesses publish themselves.
- Fields a Maps listing does not carry. MX-verified emails, mailbox ownership (
own_domain/ free inbox / the business's marketing agency),website_platform,has_meta_pixelandmobile_viewportall come from the business's own site, and the 0-100 lead score is computed from those signals together with the phone and website the map supplies. - A Google Maps link on every discovered row anyway.
google_maps_urlis a constructed search link (?api=1&query=<name>, <address>), so you can open the real listing in one click. Nothing is read from Google to build it.
What you give up is equally concrete. There are no Google star ratings or review counts, and OpenStreetMap maps some trades thinly, which the coverage section below quantifies. If you need Google's own review data, run our Google Maps Places Scraper alongside this one.
Is this a small business leads scraper with MX verified emails?
Yes. It is a small business leads scraper for any category and any city: every row is a real local business found on OpenStreetMap, and every email is MX verified against the domain's mail servers before delivery, so what you get is verified email leads rather than guessed addresses. Local leads with emails, phones and social profiles arrive in one table, and rows removed by a filter are never billed.
Measured, not promised
Every number below is from a real run of this Actor on dentist / Austin, Texas. Inputs and dates are given so you can reproduce them.
| Run | Result |
|---|---|
| Console form untouched, pressed Save & start | Today's untouched form is dentist / Austin, Texas at a cap of 25 (the form's starting value since 2026-08-29; it was 100 before) with Website required, Skip closed and Widen to nearby areas pre-ticked and no keyword exclusions — every delivered row lists a website. No count is quoted here on purpose: the 2026-08-08 measurement of this row was taken at cap 25 with a two-chain exclusion list (Aspen Dental, Walmart) the form pre-filled at the time and no longer does, so its figure cannot be reproduced from the form as it opens now. The next row is the same search with the same website filter at cap 100 (the API default) |
Console form as it opened on 2026-08-08 (it then pre-filled excludeKeywords: ["Aspen Dental", "Walmart"], which today's form does not), Max businesses raised to 100, pressed Save & start | 55 businesses, 100% contactable — 96% with a phone, 55% with an MX-verified email. Under the 100 asked for, and the run says why: OpenStreetMap holds 171 dentist records in Austin and 55 of them are crawlable with a website inside the search area. That is the whole city, not a truncation |
Bare {} from the API (no filters at all) — 2026-08-08 | 100 businesses, 50% contactable — the honest filter-off case, see the fill-rate table below |
| Cold-email list preset — 2026-08-08 | 31 rows (every emailable dentist Austin has — the exact count varies by city and over time), 100% with an MX-verified email, 94% with a phone |
| Call list preset, cap 100 — 2026-08-08 | 56 rows, 100% with a phone — the same number on two independent runs, and the run reports it as the whole of Austin rather than a truncation |
| Web-design prospects preset, cap 100 — 2026-08-08 | 7 rows, every one scored ≤45 and reachable on at least one channel. The score ceiling alone matches 68 businesses in Austin, but only 7 of those publish a phone, email or social profile — the preset drops the other 61 rather than bill you for map pins you cannot contact |
| Full enrichment preset — 2026-08-08 | 100 rows plus 23 email_guess addresses, kept out of email and never billed as verified |
| Column stability across every run above | one identical column tuple on every row (61 columns on those runs; 73 on the current build — the 12 additions of 2026-08-29 are appended at the END, so existing CSV column positions are unchanged) — CSV headers never shift mid-export |
Every 2026-08-08 row that was started from the Console form (the preset rows included) ran with the two-chain excludeKeywords list the form pre-filled at the time and no longer does; the bare {} row did not. The counts stay as historical figures.
The gap between the Console-form rows and the bare {} row is the whole story of this Actor's honesty: the Console form pre-fills onlyWithWebsite, which is why it delivers 100% contactable businesses, while a bare API call keeps every mapped location including the ones with nothing to contact. Both numbers are published; neither is hidden. The same goes for the 55-of-100 line: a run that cannot reach your cap says so in its status message, with the counts that explain it.
Max businesses starts at 25 in the Console form — raise it once the first run looks right. The form also pre-ticks requireAnyContact (since 2026-08-29), so a first run never bills you for a business with no phone, no email and no social profile — untick it if you want every listing. At $3 per 1,000 that first run costs at most $0.075 in leads (at most $0.125 once the scheduled 2026-09-14 change to $5 per 1,000 lands) and finishes in a couple of minutes. The API default is unchanged at 100: API calls, tasks and schedules that send no maxItems keep the ceiling they always had.
Quick start
The shortest complete input is a business category and a location. Everything else has a working default:
{"category": "dentist","location": "Austin, Texas"}
That is a complete run, and it is also exactly what the form already contains. Open the Actor, press Save & start without touching anything, and you get businesses. The defaults behind it: Max businesses starts at 25 in the form (raise it once the first run looks right; an API call that sends no maxItems gets 100), website crawling on, email verification on, 3 pages per website.
The input form at a glance
Eight sections, in the order the form shows them. A first run only ever touches the first three; the rest are already tuned and safe to ignore.
| Section | What it is for |
|---|---|
| Quick start | One dropdown, preset, that fills in the rest of the form for one outreach job — see the presets below. Custom changes nothing |
| What to search | The business category and the city (dropdown or free text, or the categories / locations lists for several at once), an optional countryCode and keyword (matched against the name, and since 2026-08-29 the cuisine / healthcare:speciality tags too), maxItems — how many businesses this run may deliver, i.e. your cost ceiling; the form starts at 25, raise it once the first run looks right (API default 100) — and expandNearby, pre-ticked, which widens the search around the city when the city itself runs out (each widened row labelled in query_location) |
| Filters — businesses are dropped BEFORE you are charged | Website / email / phone / social floors, excludeKeywords (arrives empty — nothing is excluded until you type a keyword), skipClosed (pre-ticked), and the score and rating floors |
| 🔔 Monitoring | onlyNewBusinesses — put the run on a schedule and get only the businesses that were not delivered before |
| Enrichment | What gets read from each business website: emails, social profiles, phones, optional pattern-guessed addresses, plus the maxPagesPerSite and concurrency crawl knobs |
| Other ways to search | A radius around a coordinate (searchRadiusKm + centerLat / centerLon) instead of a city, or bring your own list (websiteList / startUrls) and skip OpenStreetMap discovery entirely |
| Output | outputFields — which columns you get, and in what order — and sortBy |
| 🔔 Alerts | webhookUrl — a Slack / Discord / JSON digest whenever a run delivers rows |
Presets
The Use-case preset dropdown at the top configures the rest of the form for one specific outreach job. It is optional — leave it on Custom and nothing at all changes.
| Preset | What it sets | Measured on the default search (Austin dentists) |
|---|---|---|
| Cold-email list | onlyWithEmail, onlyVerifiedEmail | 31 leads (the whole of Austin; the count varies by city and over time), 100% with an MX-verified email, 94% with a phone |
| Web-design prospects | maxScore: 45, requireAnyContact | 7 leads, all in the thin/neglected half — no published email, no marketing tech, often no mobile viewport — and all reachable. maxScore: 45 on its own matches 68 businesses; the contact floor is what removes the 61 you could not pitch to |
| Call list | requirePhone | 56 leads, 100% with a phone number |
| Full enrichment | emailPatternGuess, maxPagesPerSite: 6 (every crawl toggle is already on by default) | 100 leads plus 23 guessed email_guess addresses, kept out of email |
Anything you set yourself wins, no preset ever widens your bill, and the run log names exactly what each preset applied and what it stood down on. The three guarantees are spelled out — and tested — under Presets: the three guarantees below.
What the Output tab looks like
Six curated views ship with the Actor. Pick one above the results table; the CSV / JSON / Excel export is unaffected and always carries every column you asked for.
| View | Columns | Use it for |
|---|---|---|
| Overview (default tab) | name, category, city, state, phone, email, email_status, website, rating, lead_grade, lead_score, google_maps_url | The twelve columns that answer "is this a lead?" — plus one click to the live Google Maps listing |
| Email-ready | name, email, email_status, email_type, phone, website, contact_page_url, lead_score | Loading a cold-email sequence — deliverability grade and mailbox owner side by side |
| Business details | name, category, speciality, cuisine, brand, is_chain, opening_hours, price_range, rating, review_count, wheelchair, operator, city | What the business actually is: speciality / cuisine / brand tags, chain or independent, hours, price range, and the sparse rating / review count |
| Web-agency targets | name, website, website_platform, mobile_viewport, copyright_year_stale, has_meta_pixel, has_google_ads_tag, phone, email, website_platform_status | Building a redesign pitch list from the neglect signals |
| Map & source | name, address, google_maps_url, osm_url, latitude, longitude, osm_last_edited, osm_check_date, attribution | Verifying a row in one click, territory mapping, row freshness, keeping the ODbL attribution with the data |
| All columns | all 73, in CSV order | Everything, when you want the full table |
Three deliberate choices in those views. website, contact_page_url, google_maps_url (labelled Map link) and osm_url render as clickable links; rating, review_count, lead_score, latitude and longitude render as numbers (sortable). The tri-state flags — mobile_viewport, copyright_year_stale, has_meta_pixel, has_google_ads_tag — render as text, not as a checkbox, because they are null when the site was never crawled and a checkbox cannot tell "no mobile viewport" apart from "we never looked". has_email / has_phone / has_website / is_chain are never null, so those do get the real boolean widget.
How do I find local businesses whose website is neglected enough to pitch a redesign?
Set maxScore: 45 with requireAnyContact (the Web-design prospects preset), then read the neglect columns on the rows that come back. The score cap keeps both kinds of prospect — a business with a neglected site and a business with no site at all — so add onlyWithWebsite: true if you only want the ones that already have a site to replace. Four columns carry the pitch:
Measured 2026-08-08 on Austin dentists at cap 100: the Web-design prospects preset delivered 7 rows, every one scored 45 or below and reachable on at least one channel. The score ceiling alone matched 68 businesses; the contact floor dropped the other 61 — unbilled — rather than sell you map pins you cannot contact.
website_platform— WordPress, Wix, Squarespace, GoDaddy, Shopify, Webflow, Weebly, Duda and more. A builder-tier platform usually means a self-built site.mobile_viewport—falsemeans the homepage declares no mobile viewport tag, so the site predates responsive design.copyright_year_stale— a visibly out-of-date footer copyright year.has_meta_pixel—falsemeans nobody is measuring anything on the site.
A builder-tier site with no mobile viewport, a stale copyright year and no tracking pixel is the strongest redesign signal this Actor can give you. The Web-agency targets output view shows exactly those columns and nothing else.
Read email_type before you send. It says whether the mailbox belongs to the business, to a free inbox, or to the marketing agency that already holds the account, so you can drop the leads where you would only be pitching a competitor.
How do I find local businesses that have no website, for a web-design pitch list?
Type the trade and the city and tick Businesses with no website (onlyWithoutWebsite: true), and every delivered row is a business that lists no site at all. That is the first-website pitch list for web designers and local SEO agencies, verified at 15 of 15 rows without websites on a 15-row run.
Size the list honestly: on the filter-off benchmark of 2026-08-08 (n=100), 50 of the 52 businesses with no website had no phone, no email and no social profile either, so these rows are name, address and coordinates — which is also why nobody else can cold-email them.
Expect name, address, coordinates and occasionally a phone. With no site to crawl, no email, platform or tech enrichment is possible on these rows, which is also what keeps them uncrowded: nobody else can cold-email them either. Every row still carries google_maps_url, one click to a Google Maps search for the business's name and address.
The filter is mutually exclusive with onlyWithWebsite. Setting both stops the run with an explanatory status message before anything is charged, and nothing is billed. Like every other filter it runs before billing, and a thin city can be widened with expandNearby.
What does it do?
It turns a business category and a city into one row per local business — 73 stable columns with an MX-verified email, whose mailbox it is, phone, socials, website platform and a 0-100 lead score — at $3 per 1,000 delivered leads. Measured 2026-08-08 on dentist / Austin, Texas with the website filter on: 55 businesses, 96% with a phone, 55% with an MX-verified email (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).
This actor takes a plain-English business category (e.g. "dentist," "hair salon," "roofing contractor") and a location, maps the category to the right OpenStreetMap tags, and pulls every matching business in the area. It then crawls each business's own website — following the site's own contact/about/team links (including Shopify /pages/contact and non-English slugs) within the maxPagesPerSite budget — and extracts, from pages it has already downloaded:
- a contact email, MX-verified over DNS-over-HTTPS and graded
deliverable/risky/undeliverable - whose mailbox it is — the business's own domain, a free inbox, or its marketing agency (emailing that one never reaches the business)
- social profiles — Facebook, Instagram, LinkedIn, X/Twitter, YouTube
- which website platform it runs on — WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Joomla, Drupal, HubSpot CMS, and a "Custom / Next.js" bucket for hand-built sites
- marketing & booking tech — Meta Pixel, Google Analytics/GTM, Calendly, NexHealth, Klaviyo, WooCommerce and ~25 more
- the homepage's title and meta description — drop-in mail-merge personalization fields (and a data tripwire: a title naming a different business exposes a wrong OSM
websitetag) - web-agency pitch signals — a missing mobile viewport tag, a stale footer copyright year
- one-click source links —
google_maps_url, a Google Maps search for the business's name and address, andosm_urlto the OpenStreetMap element - a lead score 0-100, an A-F grade, and a hot/warm/cold tier
Why use it / who's it for
- Web design & marketing agencies — filter for a builder-tier platform (Wix, GoDaddy, Weebly) and
has_meta_pixel: falseto find businesses with a dated site and no tracking: the highest-intent redesign pitch there is. TheCustom / Next.jsvalue is the inverse signal — don't pitch a DIY-site rebuild to someone who already paid a developer. Build 0.1.21 adds two more neglect signals —mobile_viewport(missing = the site predates responsive design) andcopyright_year_stale— plus anonlyWithoutWebsitefilter that returns only businesses with no site at all: the first-website pitch list. - Freelancers on Fiverr/Upwork — generate a "100 dentists in Austin with verified emails" list on demand for any client vertical without building a new scraper per niche.
- SaaS sales teams — any B2B tool sold to local businesses (booking software, payment processors, review management) can target by category and city, and
has_booking_widgettells you who already has a competitor installed. - B2B lead-gen resellers — one actor covers any category, replacing dozens of niche scrapers.
- Franchise & market researchers — count and map competitor density for a category in a target city (set
onlyWithWebsite: falseto keep every mapped location, including ones with no contact details).
How to use it
The form is ready to run as it opens: hit Save & start without touching anything and you get Austin dentists that all list a website, at a cap of 25 (raise it once the first run looks right) with Website required, Skip closed and Widen to nearby areas pre-ticked and no keyword exclusions — at $3 per 1,000 that is at most $0.075 of leads (at most $0.125 after the scheduled 2026-09-14 change to $5 per 1,000). (The count that used to be quoted here was measured on 2026-08-08 at cap 25 with a chain-exclusion list the form pre-filled at the time and no longer does, so it is not repeated.)
Filtered runs used to come up short of the cap. The cause was on our side, not the map's: only the two email filters deepened the candidate pool, so any run using Website required, Phone required or Social profile required filtered a pool sized for an unfiltered run and quietly came up short. Every filter that can drop a row now deepens the pool, and the filters that can be decided from the map data alone (website present, name keywords, permanently closed) are applied before any site is crawled, so no crawl budget is spent on a row that was never going to ship. Same-day comparisons at cap 25 in Austin (Console form, 2026-08-08, which then pre-filled a two-chain exclusion list the form no longer carries) showed Website required, Phone required and Email required each filling the cap after the fix where they had come up short before; those counts are not repeated here because the form no longer reproduces them.
Max businesses is still a ceiling, not a promise — a thin area really can run out. When that happens the run now says so in its status message with the counts behind it, instead of returning fewer rows without comment. A live example, dentist in Laramie, Wyoming at cap 25: "you asked for up to 25 and 1 could be delivered. That is the whole of Laramie, Albany County, Wyoming, United States, not a silent truncation: OpenStreetMap returned 4 'dentist' record(s) there, 1 became crawlable candidate(s), and 1 survived your filter(s) (onlyWithWebsite)." The run only claims an area is exhausted when no search came back at its row ceiling and every candidate was actually checked; otherwise it says which limit it hit and what to change.
- Pick a Business category from the dropdown, or leave it on Custom and type one in plain English (e.g.
medspa,HVAC contractor,funeral home). - Pick a City from the dropdown, or leave it on Custom and type any city on earth — city and region/country (e.g.
Austin, Texas). Capitalisation and the comma are optional:RABAT MAROCandcasablanca moroccoresolve too. - Leave Website required on unless you are doing density research — see the honest note below. (Web-design agencies can flip Businesses with no website instead to get the no-site prospect list; the two filters are mutually exclusive, and setting both stops the run with an explanatory status message (nothing is billed) with a clear error before anything is charged.)
- Run the actor. It geocodes the location, pulls matching places from OpenStreetMap, then crawls each business website for contact details, tech and reviews.
- Export as CSV, JSON, or Excel — or connect the dataset to your CRM/outreach tool via API/webhook.
Presets: the three guarantees
The preset table is at the top of this page. Behind it sit three guarantees, each of them tested:
- Anything you set yourself wins. A preset only fills in a field you left at its default. Set
maxScore: 100alongside the web-design preset and you keep 100 — the run log says so explicitly (left as you set them: maxScore (you chose 100)). - A preset never widens your bill. Every preset either adds a filter (fewer rows delivered, so fewer rows charged) or turns on enrichment that adds columns to rows you were already getting. None of them clears a filter you set or raises
maxItems. - It tells you what it did. One log line names the preset and lists exactly which settings it applied and which it left alone, and the run summary names the preset too.
Presets are Console and API: send "preset": "cold_email" from the API and you get the same behaviour. Existing API callers, tasks and schedules that send no preset are completely unaffected.
Which field wins: the precedence chain
Category and location each accept three inputs. Precedence is the same for both, highest first:
| Rank | Category | Location | Why |
|---|---|---|---|
| 1 | categories (list) | locations (list) | The plural list is an explicit multi-search request — it beats everything when non-empty. |
| 2 | categorySelect (dropdown) | citySelect (dropdown) | 20 categories with hand-verified OpenStreetMap tags; 25 metros verified against the live geocoder. |
| 3 | category (free text) | location (free text) | Anything else — over a hundred more category aliases are mapped, and any city on earth geocodes. |
Both dropdowns default to "" ("Custom"), which is why every input that worked before this option existed still resolves to exactly the same search.
Three search modes
| Mode | Set | What happens |
|---|---|---|
| City (default) | location (or locations) + category (or categories) | The location is geocoded and every matching business inside its bounding box is returned. Several categories x several cities run as a matrix in one run. |
| Radius | searchRadiusKm + centerLat + centerLon | Searches a circle around a coordinate instead of a city box — sales territories, franchise catchment areas, "everything within 5 km of this address". Replaces the bounding box entirely; no geocoding happens, so country stays empty. |
| Bring your own list | websiteList (or startUrls) | OpenStreetMap discovery is skipped entirely. The actor runs only the crawl + email verification + scoring pipeline over the websites you supply. Enrich a CRM export, a conference exhibitor list, or a list you bought elsewhere. |
What happens when the city runs out of businesses?
Set expandNearby: true and the same search is widened in growing rings around the city centre until your maxItems is met or the region is genuinely exhausted. It is needed because a single city often holds fewer businesses than you asked for: measured live, "dentist" in Austin, Texas tops out near 119 rows, and Round Rock, Texas holds 21.
- The rings are sized to the city and reach up to about 150 km.
- Every widened row is labelled: its
query_locationreadswithin ~16 km of Round Rock, Williamson County, Texas, United Statesinstead of the city name, so you can always tell expansion rows apart — or filter them out afterwards. - The same dedup and every pre-charge filter apply to ring rows, and the status message reports exactly how many delivered rows came from outside the city.
- Measured (2026-08-15, discovery-only):
dentist/Round Rock, Texas/maxItems: 100delivered 21 rows without the flag, 100 with it — 79 labelled ring rows, 0 duplicates.
It is off by default for API callers (existing inputs keep a byte-identical pull and bill) and pre-ticked in the Console form. It never fires when you set your own radius (searchRadiusKm) or bring your own list, when the city pull was truncated (deepening, not widening, is the fix there — the status message tells you), or during a mirror outage.
Searching several categories and cities at once
categories and locations are the plural versions of category and location, and they cross into a matrix — ["dentist","orthodontist"] x ["Austin, Texas","Dallas, Texas"] is four searches in one run. The single fields keep working exactly as before; the plural ones take precedence when non-empty.
Three things make the matrix safe rather than a footgun:
maxItemsis the whole run's budget, not a per-search one. The budget is split between the searches and results are interleaved, so the first city cannot eat the entire quota.- The same business found by two categories is delivered — and billed — once. Deduplication is on the OpenStreetMap object identity (
osm_type+osm_id), so a clinic tagged bothdentistandorthodontistappears one time. - 25 category x location combinations is the ceiling for one run. Above it the run fails immediately with a message naming the numbers, before any network request and before any charge — OpenStreetMap's public mirrors are a free shared resource.
Every row carries query_category and query_location, so you always know which search produced it.
Bring your own list (skip discovery)
Already have the businesses and only need the emails, socials, platform and score? Put the domains in websiteList (bare domains and full URLs both work; startUrls is accepted as an alias):
{ "websiteList": ["aloha-dental.com", "https://www.averyranchdental.com", "typotes.com"] }
No map data is fetched at all. Those rows differ from discovered rows in exactly three honest ways:
sourceisuser_supplied, notOpenStreetMap;attributionis null — the rows are not OSM-derived, so stamping the ODbL notice on them would be a false licence claim. The column is still present, so a mixed export keeps one stable header row;namecomes from the site's own<title>, falling back to the bare domain when the site does not answer. Nothing is invented;latitude,longitude,osm_idandosm_urlstay empty.
Everything else — the contact crawl, MX verification, mailbox-owner classification, tech fingerprinting, scoring and every filter — behaves identically.
How complete is the data? (measured, not estimated)
On the reference run of dentist in Austin, Texas with onlyWithWebsite: true (n=55), 96% of rows carried a phone, 55% an MX-verified email, 71% a detected website platform and 100% a website. With every filter off (n=100) the same city measured 48% phone, 26% email and 34% platform, because roughly half of mapped businesses list no website to crawl.
Both reference runs were taken on 2026-08-08 (the website-filtered run and the bare API run in the Measured, not promised table above), and a roofing contractor / Denver run delivered emails on 6 of 8 rows (75%) — roughly three-quarters on website-verified trade categories.
Both reference runs are in the table below, so you can see the gaps before you pay rather than after:
onlyWithWebsite: true (n=55) | schema defaults, filter off (n=100) | |
|---|---|---|
phone | 96% | 48% |
website | 100% | 48% |
email | 55% | 26% |
website_platform | 71% | 34% |
opening_hours | 80% | 44% |
facebook | 73% | 33% |
| rows graded F | 0% | 51% |
On that same filter-off n=100 run, the crawl-derived fields measured: contact_page_url 18%, website_title 40%, website_description 35%, mobile_viewport 40%, copyright_year_stale flagged on 3 rows; google_maps_url, osm_url and attribution sat at 100%.
With onlyWithWebsite: true the email fill runs far higher than the filter-off numbers — the reference run above measured 55% for dentists, and a roofing contractor / Denver run delivered emails on 6 of 8 rows (75%): roughly three-quarters on website-verified trade categories.
Read that second column before you run. OpenStreetMap has no website for a large share of businesses, and — measured on the filter-off run — 50 of the 52 rows with no website had no phone, no email and no social profile either (the other 2 carried only an OSM phone): name, coordinates and usually an address, nothing contactable. With the filter off you pay for those rows. onlyWithWebsite and onlyWithEmail drop non-matching rows before you are charged, so they cut your bill rather than just tidying the output. Every run's status message now reports the contactable ratio it actually delivered.
Every filter now goes further: the actor over-fetches, deepening the candidate pool and — for filters that need the site crawled — crawling extra candidates in batches (up to 24× maxItems, hard-capped at 30,000 candidates — the ceiling that makes a thin area terminate) until it has maxItems surviving rows or the pool is spent. This used to apply to onlyWithEmail / onlyVerifiedEmail only, which is why onlyWithWebsite, requirePhone and requireSocial quietly returned short. Compared 2026-08-08 on dentist / Austin, Texas at cap 25, one filter at a time, through a Console form that then pre-filled a two-chain exclusion list it no longer carries: onlyWithWebsite, requirePhone and onlyWithEmail each filled the cap after the change where each had come up short before (the exact counts are not repeated because today's form cannot reproduce them). Where the pool genuinely runs out first, the status message reports the counts instead of leaving you to guess — e.g. dentist in Laramie, Wyoming returns 1 row and says the map holds 4 dentist records there, 1 of them with a website.
The default is false (not true) so that existing API callers, scheduled tasks and density-research use cases keep getting every mapped location. The Apify console pre-fills it to true.
There is also the mirror filter, onlyWithoutWebsite: keep only businesses that list no website — the prospect list for web-design agencies pitching a first site (verified: a 15-row run delivered 15/15 rows without websites). It is mutually exclusive with onlyWithWebsite; setting both stops the run with an explanatory status message (nothing is billed) with a clear error before anything is charged. Expect these rows to be name + address + coordinates (and occasionally a phone) — with no site to crawl, no email/platform enrichment is possible.
Which business categories does OpenStreetMap cover well, and which are sparse?
OpenStreetMap covers businesses with premises a mapper walks past, and is thin on trades run from a van or a home office: the Austin bounding box held 171 dentist records but 9 plumbers, 5 electricians, 0 chiropractors and about 12 roofers. This is the honest limit of an OSM-based source, and it matters more than any field:
Two dated counts from the same city: 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged chain outlets (measured 2026-08-29), and on a 163-element sample of named Austin dentists taken the same day, 12% had not been edited since before 2020 and 13% carried a mapper's on-the-ground check_date.
- Dense: businesses with premises a mapper walks past — restaurants, cafés, dentists, pharmacies, hairdressers, gyms, hotels, shops, banks, clinics. (171 dentist records in the Austin bounding box.)
- Sparse: trades run from a van or a home office — plumbers (9 in the same box), electricians (5), chiropractors (0), roofers (~12). A metro of a million people can return single digits. That is what is mapped, not a bug.
95 categories (220+ terms with aliases) are curated and mapped to exact OSM tags — build 0.1.21 added pest control, photographer, moving company, self storage, funeral home and dry cleaner. Any other term is attempted as an OSM tag directly, and common phrasings are handled (auto repair shop → shop=car_repair, landscaping company → craft=gardener, insurance agency → office=insurance). Terms with no OSM tag now fall back to a business-name keyword search of OSM, and a mapped category that returns zero tagged places in the area is rescued by the same name-keyword query. Rows found only by name are labelled category_match: "name_keyword" (tag-matched rows say category) so you can filter them out if you only trust tag-confirmed rows. Measured: pest control in Denver returned 1 labelled row on 0.1.21 where the previous build returned 0 — sparse trades stay sparse, but no longer invisible. A term matching nothing at all still returns 0 rows with a suggestion and no charge — you are never billed for a run that found nothing.
Can I use this as a business email finder for local businesses?
Yes. For any category and city it crawls each business's own website and returns the MX-verified business email wherever the site exposes one, plus the mailbox it belongs to (info@, owner name, etc.). The fill rate is measured, not promised — see "How complete is the data?" above.
Measured 2026-08-08 on dentist / Austin, Texas with onlyWithWebsite: true: 55% of 55 delivered rows carried an MX-verified email, and the Cold-email list preset (onlyWithEmail plus onlyVerifiedEmail) delivered 31 leads — the whole of Austin — 100% with an MX-verified email and 94% with a phone, at $3 per 1,000 with verification included (free-plan rate read 2026-08-29 and re-read 2026-09-05; a change to $0.005 per result is scheduled for 2026-09-14 — the Pricing tab is authoritative).
Filters — every one of them runs before you are charged
Filters here are not a tidying step applied to an invoice you have already run up: a filtered row is dropped before it is pushed to the dataset and before the charge, so filtering cuts your bill.
Filtering does not shrink your delivery. Every filter in this table deepens the candidate pool to compensate, so maxItems means "this many rows I can use", not "this many businesses considered". Filters decidable from the map data alone — onlyWithWebsite, onlyWithoutWebsite, excludeKeywords, excludeChains, skipClosed — are applied before any website is crawled, so no crawl budget is spent on a row that was never going to ship. The rest are applied batch by batch as sites are crawled, stopping the moment enough rows survive. The pool is bounded at 6x maxItems candidates so a thin area terminates instead of crawling forever; when that bound or the map itself is what stopped the run, the status message says so and gives you the counts.
| Filter | Keeps only |
|---|---|
onlyWithWebsite | businesses that list a website (recommended, see the fill-rate table) |
onlyWithoutWebsite | businesses with no site — the first-website pitch list. Mutually exclusive with the above |
onlyWithEmail | rows where an email was found |
onlyVerifiedEmail | rows whose email passed MX verification (deliverable / risky) |
requirePhone | New. rows with a phone number (from OSM, a tel: link, or the site's schema.org markup) |
requireSocial | New. rows with at least one Facebook / Instagram / LinkedIn / X / YouTube profile |
requireAnyContact | New. rows reachable on at least one channel — phone or email or a social profile. The loosest contactability floor there is, and the one to reach for when you do not care which channel. It matters more than it sounds: on a bare {} run of the default search (100 rows, measured 2026-08-08) exactly 50 rows carried no phone, no email and no social profile at all, and without this switch you are billed for them. Off by default for API calls and schedules (no existing run changes); pre-ticked on the Console form since 2026-08-29 |
excludeKeywords | drops businesses whose name contains any of your keywords (case-insensitive). The form arrives with the list empty. Only the name is matched, so clinic cannot knock out a business on Clinic Street. Use it to strip chains, franchises or your existing customers |
excludeChains | New (2026-08-29). drops every business the map marks as a chain outlet — a brand or brand:wikidata tag, i.e. is_chain: true — leaving the independents. Decided from the map data alone, so a dropped outlet is never crawled or billed, and the status line says how many were removed. Off by default; rows you supply yourself have is_chain: null and are never dropped by it. In chain-dense categories the 6x pool bound can bite — 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged (measured 2026-08-29), so a maxItems: 10 run delivered 6 and said so; raise maxItems to deepen the pool |
skipClosed | New. drops permanently-closed premises (see below) |
minScore / maxScore | maxScore is new. A score ceiling is the web-design agency filter: a low score means a thin online presence, which is exactly the redesign pitch list |
minRating / minReviewCount | New — read the warning below before using these |
skipClosed: what it actually removes
OpenStreetMap mappers retire a business without deleting it: the primary tag moves from amenity=restaurant to disused:amenity=restaurant, so the object keeps its name and address but no longer describes an operating business. skipClosed drops elements carrying a lifecycle-prefixed primary tag (disused:, abandoned:, was:, removed:, demolished:, razed:), a disused=yes / abandoned=yes flag, opening_hours=closed, or shop=vacant. A disused: tag alongside a live tag of the same kind (a former bank that is now a café) is not treated as closed.
Measured live on 2026-08-08 in the Austin bounding box: 199 such elements, 43 of them still carrying a business name — e.g. Chago's with disused:amenity=restaurant, Corner Store with abandoned:shop=fuel. These cannot reach a normal tagged search (["amenity"="restaurant"] cannot match disused:amenity), but they do reach the business-name fallback used for unmapped categories, which is where dead businesses were being delivered as fresh leads. Verified end to end: with skipClosed: false that closed restaurant is delivered and billed; with it true it is dropped before billing.
It is off by default so existing runs, tasks and API callers are unchanged. The Apify console pre-fills it to true.
Honest warning about
minRating/minReviewCount: OpenStreetMap carries no review data at all. A rating only exists when the business publishes schema.orgaggregateRatingon its own website — measured at roughly 3-8% of rows. So when you set a rating or review floor, rows with no rating are DROPPED, not kept: an unknown rating is not a passing rating, and nothing is ever invented to save a row. Setting either filter will cut your result count to a small fraction — an Austin restaurant run asking for 15 rows withminRating: 4.0crawled 86 candidates looking for them and delivered 0 rows, charging nothing (re-measured 2026-08-08 after the over-fetch was generalised, so this is the deep-search result, not a shallow one). If you need ratings on every business this is the wrong source; pair it with our Google Maps Places Scraper, or use thegoogle_maps_urlon every row.
How do I run this on a schedule and get only the new businesses?
Create an Apify schedule for the run and set onlyNewBusinesses: true: the first run is the baseline, every later run delivers only businesses that were not delivered before, and a run with nothing new delivers 0 rows and bills nothing. Point a weekly schedule at "dentists in Austin" and you get the practices that appeared since the last run, and only those.
Monitoring was hardened on 2026-08-25: the memory lives in a named key-value store in your own account, entries are pruned after 90 days with the stamp refreshed on every run, each record is capped at 50,000 keys, and a schedule created before 2026-08-29 keeps its memory because excludeChains joins the key only when you set it.
How it behaves
| Run | What happens |
|---|---|
| First run of a watch | Baseline. Everything the search finds is delivered, and the status message says so in those words. |
| Later runs | Only businesses that were not delivered before. Already-delivered ones are dropped before the crawl, so they cost nothing and are never billed. |
| Nothing is new | 0 rows, SUCCEEDED, nothing billed. The status message reads Nothing new: all N business(es) this search found were already delivered by earlier runs of this watch ... You were not charged. |
What counts as the same business. The memory key is the business's OpenStreetMap object identity (node/2135639605, way/198109528) — the same identity the run already uses to dedupe in-run, so a business that renames itself or changes domain does not come back as "new". Bring-your-own-list rows have no OSM object, so they are remembered by their registered domain — the same key websiteList is already deduped on.
A row is remembered only after it has actually been delivered. Marking happens after push_data succeeds, never at the point the candidate is chosen. That ordering is the whole safety property: every filter, the maxItems trim and above all the maxTotalChargeUsd budget trim can still remove a row between the two points, and a row marked early would be recorded as delivered, skipped by every future run, and never reach you.
What identifies a watch. Your search plus the filters that shape which businesses can survive it:
categories/category,locations/location,countryCode,keywordsearchRadiusKm,centerLat,centerLon,expandNearbywebsiteList/startUrls- every option in the Filters section (
onlyWithWebsite,onlyWithoutWebsite,onlyWithEmail,verifyEmails,onlyVerifiedEmail,requirePhone,requireSocial,requireAnyContact,skipClosed,excludeKeywords,minScore,maxScore,minRating,minReviewCount) — andexcludeChains, which joins the key only when you set it, so a schedule created before 2026-08-29 keeps the memory it already has
Change any of those and you are running a different watch, with its own independent memory — the two never share state, and the new one starts from its own baseline. Category and location lists are compared case-insensitively and order-independently, so ["gym","dentist"] and ["Dentist","GYM"] are one watch, not two.
Deliberately not part of a watch: maxItems, sortBy, outputFields, concurrency, maxPagesPerSite. Those change how many rows you get, in what order and with which columns — not which businesses exist in the search — so tuning them never resets the memory and never re-bills you for businesses you already have.
Where the memory lives. A named key-value store called local-business-leads-monitor in your own Apify account (runs execute there), one record per watch, keyed sig-<hash>. Entries older than 90 days are pruned on every write and each record is capped at 50,000 keys, so a long-running schedule cannot grow it without bound. The timestamp stored is last seen, not first, and it is refreshed on every run for businesses still returned by the search — so the 90-day prune drops businesses that have genuinely disappeared from OpenStreetMap rather than businesses that have simply been mapped for a long time (those would otherwise age out and be re-delivered, and re-billed, as "new"). If the store cannot be read the run treats itself as a first run and says so; monitoring never fails a run.
Two honest caveats.
- Runs take longer with it on. Already-delivered businesses have to be skipped over, so the run searches a deeper candidate pool to still fill your
maxItemswith genuinely new rows. - New businesses appear at the speed of OpenStreetMap. This is a mapping database, not a live business registry — a new dentist shows up here when a mapper adds it. A weekly or monthly schedule matches that cadence; an hourly one will mostly report "Nothing new" and bill you nothing for the privilege.
{"category": "dentist","location": "Austin, Texas","maxItems": 200,"onlyWithWebsite": true,"onlyNewBusinesses": true}
How do I get a Slack or Discord alert when new leads land?
Put your Slack or Discord incoming-webhook URL in webhookUrl, and every run that delivers at least one row POSTs a digest to it; a run that delivers nothing sends nothing.
Slack / Discord / JSON digests shipped on 2026-08-25; the JSON payload carries the delivered count, the run and dataset links and the first 20 rows, and a webhook that fails is named in the run's status message without failing the run.
A Slack incoming webhook and a Discord webhook each get a text message. Any other URL, such as an n8n, Make or Zapier catch hook or your own endpoint, gets JSON: {actor, delivered, run_url, dataset_url, rows[:20], text}.
Pair it with onlyNewBusinesses on a schedule and the Actor is a new-business alert service on its own, with no extra automation needed just to see the rows. Delivery is best-effort: a webhook that fails is reported in the run's status message and never fails the run.
{"category": "dentist","location": "Austin, Texas","onlyNewBusinesses": true,"webhookUrl": "https://hooks.slack.com/services/T000/B000/XXXX"}
Every run comes with a report
Every run that delivers at least one row also stores a one-page HTML report: the headline numbers (businesses, % with email, % verified, % with website, average lead score), the A–F lead-grade split, the top categories or cities, the email-status split and the first 100 rows, plus the same notes the status message carries. It is saved as the REPORT record of the run's key-value store — open it from the run's Output tab (record REPORT) or through the Report: link in the run's status message. It is a single self-contained HTML file (no scripts, no external assets), so it is safe to forward, attach to an email or screenshot for a client. The dataset stays the source of truth: the report summarises what was delivered and never replaces the rows, and a run that delivers nothing writes no report.
Choosing your columns and sort order
outputFields— pick the columns you want (e.g.["name","email","phone","website","lead_grade"]) and the dataset carries only those. Every row still shares one identical column tuple, so CSV headers never shift mid-export, and the columns keep the documented ROW order regardless of the order you listed them in.nameandattributionare always included whatever you choose —attributionbecause the OpenStreetMap licence (ODbL) has to travel with the data into your CSV. Column names are matched case-insensitively (and-/space count as_, soLead Gradeworks). A single unknown name is logged and ignored; if none of the names you list exists, the run stops before billing rather than charging you full price for a name-only export.sortBy—score_desc(default, what every previous build did),name_asc,review_count_descorrating_desc. Sorting never removes a row; all filtering already happened. Rows missing the sort value (no rating, no review count) are placed last, because a missing value is unknown rather than zero.
Output fields
73 columns on every row (61 before 2026-08-29; the twelve new ones — has_google_ads_tag, speciality, cuisine, wheelchair, operator, osm_description, brand, brand_wikidata, is_chain, city_source, osm_last_edited, osm_check_date — are appended at the END of the column order, so an existing CSV import keeps its positions). Re-measured on a 100-row dentist / Austin, Texas run of the previous build with schema defaults: all 100 rows carried the same stable column tuple, so CSV headers don't shift mid-export — and if you narrow the export with outputFields, every row still shares one identical (smaller) tuple. Fill rates are from the onlyWithWebsite: true reference run above; anything conditional says so. Fields marked New (0.1.21) show fill rates from the 100-row filter-off benchmark instead — read them accordingly, since roughly half of those rows had no website to crawl.
In the Console, the Output tab opens on the Overview view; the All columns tab shows every field listed below, in this order. Views only change what the Console table renders — a CSV, JSON or Excel export always carries every column the run produced (or exactly the ones you named in outputFields).
Identity & location — from OpenStreetMap
| Field | Fill | Description |
|---|---|---|
name | 100% | Business name |
category | 100% | OSM category tag (e.g. dentist, hairdresser, lawyer) |
category_match | 100% | New (0.1.21). How the row matched your category: category = matched the mapped OpenStreetMap tag; name_keyword = found by the business-name fallback search (used for unmapped terms, and as a rescue when a mapped tag returns zero places in the area). Filter on category if you only want tag-confirmed rows |
address | 96% | Street address assembled from OSM address tags |
city / state / postal_code | 93% / 93% / 93% | Address components from the OSM addr:* tags. Since 2026-08-29 a row that sits inside the searched city's outline but carries no addr:city tag gets city (and a null state) filled from the geocoder's breakdown of the searched location — zero extra requests, and city_source says which happened. For US locations the filled state is the postal abbreviation (TX, not Texas), the same form mappers write in addr:state, so one run never splits a state into two spellings. Rows admitted by expandNearby rings and radius searches (no outline) keep their nulls |
city_source | 100% when city is set | New (2026-08-29). osm_tag = the mapper wrote addr:city; search_area = filled from the searched location because the row is inside its administrative outline; null when city is null. Appended at the end of the column order |
country | 100% | From the geocode. Empty in radius mode and on user-supplied rows (no geocode happens there) |
query_category | 100% | New. Which of your input categories produced this row — the column that makes a multi-category run readable. Null on user-supplied rows |
query_location | 100% | New. Which of your input locations produced this row (the geocoder's resolved name, or the radius description). Null on user-supplied rows |
latitude / longitude | 100% | Coordinates |
osm_type / osm_id | 100% | OpenStreetMap source identifiers |
google_maps_url | 100% | A Google Maps listing search for the business — ?api=1&query=<name>, <address>, URL-encoded — so one click opens Google's search for that business rather than a bare coordinate pin (a discovered row with coordinates but no address falls back to the pin). The link carries none of this actor's data and nothing is scraped from Google: whatever rating you see there is Google's, not the rating column (see the warning below). Need Google's fields at scale? Pair with our Google Maps Places Scraper |
osm_url | 100% | New (0.1.21). Link to the row's source OpenStreetMap element — instant provenance, and the place to fix bad map data |
speciality | varies by category | New (2026-08-29). The OSM healthcare:speciality tag as a list (["orthodontics"], ["general", "paediatric"]; first 10 values) — dentists, doctors and clinics; empty for every other trade. The keyword input matches it too |
cuisine | varies by category | New (2026-08-29). The OSM cuisine tag as a list (["pizza", "italian"]; first 10 values) — restaurants, cafes, fast food; empty elsewhere. The keyword input matches it too |
wheelchair | varies by category | New (2026-08-29). The OSM wheelchair tag as written (yes, no, limited) |
operator | varies by category | New (2026-08-29). The OSM operator tag — the company running the premises, where the mapper recorded one |
osm_description | varies by category | New (2026-08-29). The mapper's free-text description tag, clamped to 150 characters |
brand / brand_wikidata | varies by category | New (2026-08-29). The OSM brand and brand:wikidata tags (Aspen Dental / Q4807808) — dense on pharmacies, banks and fast food, sparse on independents |
is_chain | 100% on OSM rows | New (2026-08-29). true when either brand tag is present, false otherwise — the one-column chain flag; set excludeChains: true to drop those rows before billing (excludeKeywords still works by name). Null on user-supplied rows |
Contact
| Field | Fill | Description |
|---|---|---|
phone | 96% | Primary phone. OSM tag first; falls back to a tel: link or JSON-LD telephone only when OSM has none |
phones | 73% | New. All tel: numbers found on the site. May include a call-tracking number — phone stays the trusted value |
website | 100% | Business website URL. A social page in OSM's website tag is routed to that social column instead, so this is always a real site |
domain | 100% | New. Bare registered domain — the field CRMs dedup on |
email | 55% | Primary contact email, chosen by mailbox ownership then deliverability |
emails | 55% | All emails for the row, primary first (previously excluded a primary that came from OSM) |
contact_page_url | 42% | New. The exact page the primary email was found on, so you can spot-check it. Null when the email came from OpenStreetMap rather than from a crawled page |
email_guess | 0% unless opted in | New, opt-in, and deliberately not an email. With emailPatternGuess: true, a business that has a website but publishes no address anywhere we crawled gets a pattern-guessed info@<domain> here. It is never promoted into email, never counts as has_email, never earns a lead-score point and never satisfies onlyWithEmail / onlyVerifiedEmail. Nobody checked that this mailbox exists — treat it as a lead, not an address |
email_guess_confidence | same | New. A statement about the domain, never the mailbox: mx_ok (the domain does run mail servers), no_mx (it does not — the guess is almost certainly dead), unchecked (no lookup completed — email verification is off, the DNS lookup itself failed, or the run's time budget stopped it) |
Email quality — verification is included, not an add-on
| Field | Fill | Description |
|---|---|---|
email_status | 100% | deliverable / risky / undeliverable when verification is on and a candidate exists; missing when no email was found; found / missing when verifyEmails is off. Turning verification off changes the vocabulary |
email_provider | 55% | Mailbox host where identifiable (Google Workspace, Microsoft 365...). Only when verifyEmails is on and an email was found |
email_domain_match | 55% | Whether the email's domain matches the website's |
email_type | 55% | own_domain / free_mail / third_party / unknown. third_party means the address belongs to the business's marketing agency or web designer — it is deliverable but does not reach the business, and it is scored at half weight. Measured on the reference run: 12% of all harvested addresses were third-party, but only 10% of primary emails, because candidates are ranked by mailbox ownership before one is promoted |
email_types | 55% | The same classification for every entry in emails |
Socials
| Field | Fill | Description |
|---|---|---|
facebook | 73% | Facebook profile URL |
instagram | 55% | Instagram profile URL |
youtube | 25% | New. YouTube channel URL |
linkedin | 22% | New. LinkedIn company/profile URL |
twitter | 18% | New. X/Twitter profile URL |
Share buttons, tracking pixels and embedded posts are filtered out, so these are profile URLs rather than facebook.com/tr or instagram.com/p/....
Website & marketing tech
| Field | Fill | Description |
|---|---|---|
website_platform | 71% | CMS / site builder: WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, Custom / Next.js, Custom / React. WordPress plugins (Elementor, WP Rocket, Divi) report as WordPress, not as their own platform |
website_platform_status | 100% | New. Why website_platform is what it is — and what actually happened to the site: detected, unknown (a page loaded, nothing recognisable on it), site_blocked (the site answered but refused us — a 403, a bot-check page or a login wall), not_found (every URL that answered came back 404 or 410 — the address in the listing is dead, the host serving it is not), site_error (the site answered with a server error), no_page (it answered, but never with a readable web page), site_unreachable (nothing answered at all — DNS failure, refused connection, broken TLS or a timeout), no_website, not_crawled (crawling was off), crawl_error (our own crawler failed on that site — worth reporting). Only site_unreachable means the business has no reachable website, and only crawl_error says nothing about the site; every other value tells you the site exists and why we could not read it, so a blocked or 404 site is still a live prospect. A null platform is explained rather than unexplained |
platform_version | 27% | New. Only when the site's own meta generator names the platform and a version — never a plugin's version passed off as the platform's |
site_generator | 42% | New. The raw meta name="generator" string |
tech | 84% | New. Marketing/booking/commerce tech found on the page (Meta Pixel, Google Analytics, GTM, Calendly, NexHealth, Klaviyo, WooCommerce, Yelp Reviews, live chat...) |
has_meta_pixel / has_google_analytics / has_booking_widget / has_google_ads_tag | see note | New. Booleans derived from tech. Tri-state: true/false when the site was read, null whenever no page was read — unreachable, blocked, a bot-check interstitial, or not crawled at all. false never means "we could not check", so a no-pixel filter cannot fill up with sites nobody ever saw. has_google_ads_tag (new 2026-08-29, appended at the end of the column order) is true when the Google Ads conversion tag (googleadservices.com / gtag AW-…) is on the page — a wiring signal, not proof of live spend (see the FAQ) |
website_title | 40% | New (0.1.21). The crawled homepage's <title> — a drop-in mail-merge personalization field. Also a data tripwire: a title that clearly names a different business means OSM's website tag is wrong, and this column lets you catch that before you hit send |
website_description | 35% | New (0.1.21). The homepage's meta description — the other personalization field, and the same tripwire |
mobile_viewport | 40% | New (0.1.21). Whether the homepage declares a mobile viewport tag. A site without one predates responsive design — a concrete web-agency pitch signal |
copyright_year_stale | 3% | New (0.1.21). Flags a visibly out-of-date copyright year in the footer — a small but unambiguous neglected-site signal (flagged on 3 of the 100 benchmark rows) |
Changed in this build —
website_platform_statusgot more precise. It used to fold every crawl failure into one value,site_unreachable, which read as "this business has no working website". Most of those sites were alive and simply refused a datacenter request. The value set now separatessite_blocked,not_found,site_error,no_pageandcrawl_errorfromsite_unreachable, which is reserved for hosts that produced no HTTP response at all. If you have a saved filter, Make/Zapier step or script that matcheswebsite_platform_status == "site_unreachable", it will now match fewer rows — by design: the rows it stops matching are live websites. Match on the list above (or onemail_status) to get the old, broader set. Pricing does not change, and no run delivers more rows than itsmaxItems. This build does read pages the previous one threw away (HTML served astext/plain, as XHTML, or with noContent-Typeat all), so a site that used to yield no email can now yield one — withonlyWithEmailorminScoreon, that can change which businesses fill your order, always within the same cap.
Business detail & reviews
| Field | Fill | Description |
|---|---|---|
opening_hours | 80% | From OSM; gap-filled from the site's JSON-LD only when OSM has none |
rating | 7% | New, and sparse — see the warning below. Star rating (1-5) the business publishes in its own schema.org markup |
review_count | 7% | New, sparse. Review count from the same markup |
price_range | 24% | New. schema.org priceRange (e.g. $$) |
Honest warning about
rating/review_count: these are not Google Maps ratings. They are only present when a business publishesaggregateRatingin its own JSON-LD, and most do not. Measured on-platform: 7% of website-bearing rows for dentists in Austin (4/55) and 13% for roofing contractors in Denver (1/8); the filter-off benchmark (n=100) measured 4%. An earlier 12-row roofer sample hit 33%, so the rate swings wildly with category, city and sample size — assume under 10% and treat anything higher as luck. Do not build a workflow that needs a rating on every row. If you need star ratings and review counts for every business, this is the wrong source — pair it with our Google Maps Places Scraper; every row'sgoogle_maps_urlalso opens a Google Maps search for the business in one click. OpenStreetMap carries no review data at all, and this actor never invents a substitute:lead_scoreis a data-completeness score, not a customer rating.
Lead scoring
| Field | Fill | Description |
|---|---|---|
lead_score | 100% | 0-100 completeness/reachability score |
lead_grade | 100% | A ≥80, B ≥65, C ≥50, D ≥35, else F |
lead_tier | 100% | hot ≥75, warm ≥50, else cold |
score_breakdown | 100% | Per-component points, plus signals_from — the list of extra signals that actually scored — so the score is auditable |
has_email / has_phone / has_website | 100% | Booleans for quick filtering |
Scoring rubric (sums to a true 100, so grade A is reachable — the filter-off benchmark's top row scored the full 100): email deliverable 40 / risky 25 / present-but-unverified 15 — halved when email_type is third_party; phone 20; website 15; socials 5 each capped at 10; extra signals 5 each capped at 15, drawn from website platform detected, opening hours, marketing tech found, and a published star rating — score_breakdown.signals_from names the ones that counted.
Provenance
| Field | Fill | Description |
|---|---|---|
source | 100% | OpenStreetMap for discovered rows, user_supplied for rows that came from your own websiteList |
attribution | 100% on OSM rows | The ODbL attribution string, so the licence travels with an exported CSV. Null on user_supplied rows — they are not OpenStreetMap data, so attaching the notice would be a false licence claim. The column is always present either way |
osm_last_edited | 100% of OSM rows (measured 2026-08-29 on three runs: 25/25 dentists in Austin, 20/20 restaurants in Chicago) | New (2026-08-29). ISO date of the element's last edit by any mapper, read from OpenStreetMap's own edit metadata (out meta) on the same request as before. Sort or filter on it to skip rows nobody has touched in years; the last column but one |
osm_check_date | varies by category (measured 2026-08-29: 4/25 dentists in Austin, 3/20 restaurants in Chicago) | New (2026-08-29). The mapper's check_date tag — the day someone confirmed the business on the ground (used mostly for opening hours); present on a minority of rows. The last column |
enriched_from_website | 33% | New. Lists only the fields where an OSM-overlapping value was taken from the site instead — rating, review_count, price_range (OSM carries none of these) plus phone / opening_hours / email where OSM was empty and the site's schema.org data filled the gap. It is not a full provenance map: emails, the five socials, website_platform, platform_version, site_generator and tech are always crawled from the site and are deliberately not repeated here |
Example output
Real rows from a run with onlyWithWebsite: true. This is a best-case slice, not a typical one — see the measured fill rates above; on the website-filtered reference run roughly half of rows carry an email.
| Name | Phone | Status | Type | Platform | Tech | Score | |
|---|---|---|---|---|---|---|---|
| Aloha Dental | +1-512-707-7300 | riverside@aloha-dental.com | deliverable | own_domain | WordPress | Meta Pixel, GTM, Yelp | 95 (A) |
| Avery Ranch Dental | +1-512-246-7645 | smile@averyranchdental.com | deliverable | own_domain | WordPress | GTM, reCAPTCHA | 95 (A) |
| Aviva Dental Care | +1 512 852 8528 | dr.apurva@avivadentalcare.com | deliverable | own_domain | WordPress | Google Analytics | 90 (A) |
How to read the output
One row per business, 73 columns, always in the same order. The Console's Output tab opens on the Overview view; the other views are the same rows with a different set of columns in front:
- Overview — the first-look table: Name, Category, City, State, Phone, Email, Email status, Website, Rating, Grade, Lead score, Map link. Read
lead_grade(A–F) first, thenemail_status. - Email-ready — only what a cold-email sequence needs: the address, its deliverability grade (
email_status), whose mailbox it is (email_type), and the page it was found on. - Business details — what the mapper wrote about the business:
speciality,cuisine,brand,is_chain, opening hours, price range, and the rating / review count when the business's own site publishes one (roughly 3–8% of rows). - Web-agency targets — the redesign-pitch signals: site builder, mobile viewport, stale copyright year, Meta Pixel and Google Ads tag, with
website_platform_statussaying whether the site was actually read. - Map & source — where the row is (address, coordinates, Google Maps and OpenStreetMap links), how fresh it is (
osm_last_edited,osm_check_date) and the ODbL attribution that has to travel with the data. - All columns — every column, in CSV order.
Column labels follow one vocabulary across every Flash Scrape actor: Name, Category, City, State, Phone, Email, Email status, Website, Lead score, Grade, Map link. Booleans (has_email, has_phone, has_website, is_chain) are real true/false values; numbers (rating, review_count, lead_score, latitude, longitude) are real numbers, never strings; the two OSM dates are YYYY-MM-DD.
There are no derived or duplicated columns: address is already the full one-line postal address next to its city / state / postal_code / country parts, and google_maps_url is a ready-made link, so nothing needs assembling on your side.
Exporting just one view. Views change what the Console shows, not what its Export button downloads — a Console export always carries every stored column (Apify's dataset-schema docs: views only affect the Console display). To download only a view's columns use the view parameter of the dataset-items API, e.g. https://api.apify.com/v2/datasets/<datasetId>/items?view=email_ready&format=csv (view keys: overview, email_ready, business_details, web_agency, map_source, all_columns). Leave view off, or use all_columns, for the full table; outputFields in the input narrows what is stored in the first place.
The run report and status message. Every run that delivers rows ends with the same one-line status — Done — N leads delivered for <what>. <coverage / filter notes> Report: <link>. — and stores the REPORT HTML page whose table shows exactly the Overview columns (up to 10 of them: the sparse rating and the Map link are the first to be left out when the table is full), with tiles for businesses, % with email, % verified (when email verification ran), % with website and the average lead score.
Input examples
A preset plus the two dropdowns — the shortest useful input there is:
{"preset": "cold_email","categorySelect": "hair salon","citySelect": "Casablanca, Morocco","maxItems": 100}
The preset sets onlyWithEmail and onlyVerifiedEmail; everything else stays at its default. Add any field you want and yours wins — {"preset": "web_design", "maxScore": 100} keeps maxScore: 100 and the log says which setting the preset stood down on.
One category, one city — the classic run, unchanged:
{"category": "hair salon","location": "Miami, Florida","maxItems": 200,"crawlEmails": true,"onlyWithWebsite": true,"onlyWithEmail": false,"verifyEmails": true,"maxPagesPerSite": 3,"concurrency": 8}
Several categories across several cities, contactable rows only, narrow columns:
{"categories": ["dentist", "orthodontist"],"locations": ["Austin, Texas", "Dallas, Texas"],"maxItems": 200,"onlyWithWebsite": true,"requirePhone": true,"excludeKeywords": ["Aspen Dental"],"skipClosed": true,"outputFields": ["name", "email", "phone", "website", "lead_grade", "query_location"],"sortBy": "name_asc"}
A 5 km sales territory around one address:
{"category": "restaurant","searchRadiusKm": 5,"centerLat": 30.2672,"centerLon": -97.7431,"maxItems": 150,"onlyWithEmail": true}
Web-design prospect list — businesses with a site, but a weak one:
{"category": "hair salon","location": "Lyon, France","countryCode": "fr","onlyWithWebsite": true,"maxScore": 45,"skipClosed": true}
Enrich your own list (no map data fetched at all):
{"websiteList": ["aloha-dental.com", "averyranchdental.com", "typotes.com"],"verifyEmails": true,"emailPatternGuess": true}
How much does it cost?
This Actor is pay-per-result and costs $3 per 1,000 delivered leads, which is $0.003 per lead on the free plan and less on paid plans (the Pricing tab always carries the current rate). You are charged for delivered businesses only, not for API calls or compute. There is no subscription, and new Apify users get platform free credits to test with. A scheduled pricing record raises this to $0.005 per delivered lead on 2026-09-14 (notified 2026-08-30); the Pricing tab on this page is authoritative.
Read from the live pricing on 2026-08-29 and re-read 2026-09-05: $0.003 per delivered lead on the free plan and $0.0027 (Bronze) down to $0.0021 (Diamond) on paid plans, so $5 buys 1,666 delivered leads; a change to $0.005 per result is scheduled for 2026-09-14 (notified 2026-08-30) and the Pricing tab is authoritative. For comparison, lukaskrivka/google-maps-with-contact-details charges $0.005 per place, $0.0025 per contact enrichment and $0.10 per verified email on the free plan (Store pricing read 2026-09-05), falling to $0.004 per verified email on Bronze.
$5 buys 1,666 delivered leads at that rate (1,000 after the scheduled 2026-09-14 change to $0.005 per lead). A free-plan visitor can run the form as it opens (25 dentists in Austin, at most $0.075 of leads plus a few cents of compute — well inside Apify's monthly free credit), download the CSV and judge every column before paying anything.
A run's total is simply rows delivered x the per-lead rate. Be aware what the rows contain: on the reference run with onlyWithWebsite: true, ~55% carried an email, so 500 rows ≈ 275 emailable leads (roughly 1.8x the per-lead rate per emailable lead). With the website filter off, 26% carried an email on the filter-off benchmark (roughly three-quarters on website-verified trade categories). Filters run before billing, so onlyWithEmail: true is the cheapest way to buy emails specifically — and the adaptive over-fetch delivers as close to maxItems email rows as the city allows (measured 2026-08-08: asked 25, delivered 25 in Austin; the same input delivered 19 before the pool-depth fix, and the whole city tops out at 31, which the run tells you when you ask for more).
How it compares
Read from Apify's public Store API on 2026-09-05. Nearly every alternative is a Google Maps scraper: lukaskrivka/google-maps-with-contact-details (87,957 users, 4.63 from 221 reviews, $5 per 1,000 places, with contact enrichment at $2.50/1,000 and email verification at $100/1,000 on Apify's free plan, falling to $4/1,000 on Bronze), s-r/google-maps-contact-details (166, no reviews, $4/1,000 plus $2/1,000 enrichment), leadharbor (40, no reviews, $3/1,000 MX-checked), code-node-tools (33, no reviews, $1.10/1,000) and jurassic_jove (99, no reviews, $20/1,000). Paid Apify plans get tiered discounts on several of these, this Actor included, so read every figure off the Pricing tab for the plan you are on.
This Actor (32 users on the public Store API cache, 5.0 from 3) is $3 per 1,000 with MX verification included — a change to $5 per 1,000 is scheduled for 2026-09-14 — but discovery runs on OpenStreetMap, not Google Maps, so there are no Google star ratings and van-based trades are thinly mapped.
For comparison, lukaskrivka/google-maps-with-contact-details charges $0.005 per place plus $0.0025 per contact enrichment plus $0.10 per verified email on Apify's free plan (its live Store pricing record, read 2026-09-05) — the verification event alone is $100 per 1,000 there, falling to $4 per 1,000 on Bronze. Here, MX verification, mailbox-ownership classification and lead scoring are all included in the single per-lead rate.
Where does the data come from, and what licence is it under?
Business listings come from OpenStreetMap under the Open Database License (ODbL) v1.0, and every contact detail comes from the business's own public website, which ODbL does not cover. © OpenStreetMap contributors — https://www.openstreetmap.org/copyright. Every row carries source and attribution fields; keep them if you redistribute or publish the data, as ODbL requires attribution. The columns added on 2026-08-29 from OSM tags — speciality, cuisine, wheelchair, operator, osm_description, brand, brand_wikidata, is_chain, osm_last_edited, osm_check_date — are OSM data under the same licence (city_source and has_google_ads_tag are ours).
The website-crawled fields — email, emails, socials, website_platform, tech, rating, review_count, price_range, phones — are not OSM-derived. They come from each business's own public website and are not covered by ODbL. Those fields are site-derived on every row. The enriched_from_website column is narrower than that: it flags only the fields that OSM could have supplied but didn't, so treat the list above — not that column — as the ODbL boundary.
Frequently asked questions
How do I find local business leads with verified email addresses?
Run this Actor with a category, a city and onlyWithEmail: true, and every delivered row carries an email address; with verification left on (the default) each one is MX-checked over DNS-over-HTTPS and graded deliverable, risky or undeliverable in email_status. Measured on dentist / Austin, Texas with onlyWithWebsite: true, 55% of the 55 delivered rows carried an MX-verified email and 96% carried a phone; a roofing contractor / Denver run delivered emails on 6 of 8 rows.
Two things make those addresses usable rather than just present. email_type says whether the mailbox is the business's own_domain, a free inbox, or its marketing agency (mail to an agency mailbox never reaches the business). And onlyWithEmail and onlyVerifiedEmail run before billing, so a row without an email is dropped rather than charged, which makes them the cheapest way to buy emails specifically.
Those figures date from 2026-08-08, when the Cold-email list preset (onlyWithEmail plus onlyVerifiedEmail) delivered 31 leads for Austin dentists — the whole city that day — 100% with an MX-verified email and 94% with a phone.
Is it legal to scrape local business data?
This Actor reads only publicly available data: OpenStreetMap records, which are an open-data project licensed under ODbL, and each business's own public website, which is the same information anyone could read by visiting the site. There is no login, no private data, no anti-bot circumvention and no Google Maps terms-of-service exposure. Keep the attribution column with the data if you republish it, because ODbL requires attribution.
Does it work outside the US? How should I type the location?
Yes, it works anywhere OpenStreetMap covers, and you can type the location however you like. The raw string is tried first, and only if that finds nothing is it automatically re-spelled (Title Case, and a comma inserted before the trailing country word) before the run is given up on. RABAT MAROC, Rabat Maroc, casablanca morocco and Rabat, Morocco all resolve to the same place. If the run still cannot geocode, the status message lists every spelling it tried instead of a generic failure.
Verified: Rabat with countryCode mt resolves to Rabat, Western Region, Malta, while Morocco wins without the pin. The country column was fixed on 2026-08-08 — every Moroccan row used to ship a three-script run-on where Morocco belonged — and Austin, Texas and Lyon, France were unaffected.
If your city name exists in more than one country — Rabat is a city in both Morocco and Malta, Cambridge in both the UK and the US — set countryCode to an ISO 3166-1 alpha-2 code (ma, mt, us, fr, gb) to pin the geocoder. Verified: location: "Rabat" with countryCode: "mt" resolves to Rabat, Western Region, Malta; without it, Morocco wins. The pin is enforced: the fallback geocoder (Photon) has no country filter, so it is skipped whenever countryCode is set. A city that does not exist inside that country ends the run with a clear message and no charge, instead of quietly returning the same-named city elsewhere (location: "Austin" + countryCode: "ma" used to deliver Austin, Texas).
One non-US bug is fixed as of this build: the geocoder is now asked for English place names. It previously answered in the local language(s), and since the row's country column is derived from the geocoder's answer, every Moroccan row used to ship country: "Maroc ⵍⵎⵖⵔⵉⴱ المغرب" — a three-script run-on where Morocco belonged. Measured and fixed on 2026-08-08; Austin, Texas and Lyon, France are unaffected. Note that city and address still come straight from the OpenStreetMap tags a local mapper wrote, so a Moroccan row can legitimately read Témara تمارة — that is the map's own data and it is not rewritten.
How fresh is the data?
Every email, social profile, website platform and tech signal is crawled live from the business's website during the run, so that half of the row is current at run time; the OpenStreetMap half — the name, the address and the phone where the map supplies it — is only as fresh as the last mapper edit, which each row states in osm_last_edited.
That makes freshness something you can filter on rather than trust:
osm_last_edited— the day any mapper last changed the element, read from the map's own edit metadata. On a 163-element sample of nameddentistelements in the Austin bounding box (measured 2026-08-29), 12% had not been edited since before 2020.osm_check_date— the mapper'scheck_datetag, set when someone confirmed the business on the ground. It was present on 13% of that same sample, and the share varies by category and city.
Sort or filter on osm_last_edited to skip rows nobody has touched in years, and keep skipClosed on to drop premises mappers have retired.
How do I know whether a business email will bounce before I send to it?
Read the email_status column: with verification on (the default) every address is checked over DNS before delivery and graded deliverable, risky or undeliverable, and that verification is included in the per-lead price rather than sold as an add-on ($0.003 today, $0.005 from 2026-09-14 under a scheduled pricing record). Set onlyVerifiedEmail: true to keep only the rows that passed — deliverable and risky — and drop the ones proved dead. Turn verifyEmails off and the column falls back to found / missing; an address the run ran out of time to check also ships as found, unverified, and the status message says so.
On the 2026-08-08 reference run 55% of the 55 delivered rows carried an MX-verified email; 12% of all harvested addresses belonged to a third party (a marketing agency or web designer) but only 10% of primary emails did, because candidates are ranked by mailbox ownership before one is promoted.
Four checks run, all over DNS and none over SMTP: syntax, a mail-server (MX, or implicit-MX A record) lookup on the domain, a role-address check (info@, sales@ and the like grade risky), and a disposable-domain list (grades undeliverable). The MX lookup goes over DNS-over-HTTPS to the public Google and Cloudflare resolvers, so no mail server is ever contacted and no proxy is needed. No mailbox is ever probed — that is what keeps verification proxy-free, fast and included in the price, and it is also what it cannot see: an address on a catch-all domain, or a mailbox that was deleted while the domain kept its mail servers, can still grade deliverable. Read deliverable as "the domain accepts mail and the address is not a known-bad shape", and warm a new list with a soft first send. When the MX lookup itself fails the row grades risky rather than undeliverable, so a busy resolver never deletes a good lead.
Can it tell me whether a business is running ads?
Partly: has_google_ads_tag and has_meta_pixel tell you whether the ads or conversion tag is present on the crawled page, which is a wiring signal rather than proof of live ad spend. Each is true, false, or null when no page was read. Sites keep the tag long after a campaign ends, and a business can run ads with no on-site tag at all. For actual live creatives, feed each row's domain into our Competitor Google Ads Scraper (its queries input takes domains), which lists what a domain is currently running from Google's Ads Transparency Center.
has_google_ads_tag was appended on 2026-08-29; both flags are tri-state (null when no page was read), and the underlying tech column — Meta Pixel, Google Analytics, GTM, Calendly, Klaviyo and about 25 more — was filled on 84% of rows on the 2026-08-08 reference run (n=55).
Why is email empty on some rows?
Three distinct reasons, and email_status plus website_platform_status tell you which: the business has no website in OSM (no_website), no page could be read from the site it does have (site_blocked, not_found, site_error, no_page, or site_unreachable — and only that last one means there is nothing there to reach), or the site simply does not publish an address anywhere on the pages crawled (missing). Cloudflare-obfuscated addresses are decoded, and since 0.1.21 the crawler follows the site's own contact/about/team links — including Shopify /pages/contact and non-English slugs — within the maxPagesPerSite budget. Raising maxPagesPerSite finds a few more.
Sized on 2026-08-08: with every filter off, 26% of 100 rows carried an email because roughly half of mapped businesses list no website; with onlyWithWebsite: true it was 55% of 55, and 8 of those 55 sites refused datacenter requests (site_blocked) — a live business, not a dead one.
Does it get star ratings and review counts?
Only when a business publishes them in its own website markup, and these are never Google Maps ratings: measured at 7-33% of website-bearing rows depending on the category, and 4% on the latest filter-off benchmark. Do not build a workflow that needs a rating on every row; see the warning in the output section. OpenStreetMap has no review data, so if ratings are essential for every row, pair this with our Google Maps Places Scraper — every row's google_maps_url opens a Google Maps search for the business in one click.
Measured 2026-08-08: 4 of 55 Austin dentists (7%) published a rating, and an Austin restaurant run asking for 15 rows with minRating: 4.0 crawled 86 candidates and delivered 0 rows, charging nothing — a rating floor drops unrated rows rather than keeping them.
How do I build a list of businesses running WordPress, Wix, Squarespace or Shopify in one city?
Run the category and city with onlyWithWebsite: true and read the website_platform column, which names the CMS or site builder behind each site (WordPress, Wix, Shopify, Squarespace, Webflow, GoDaddy, Weebly, Duda, Jimdo, Site123, Framer, HubSpot CMS, BigCommerce, Joomla, Drupal, and Custom / Next.js or Custom / React for hand-built sites). It was filled on 71% of rows on the website-filtered reference run, and website_platform_status explains every row where it is empty.
There is no platform input to filter on, so the pattern is: pull the city, then filter the CSV on website_platform. Knowing a business runs on Wix, GoDaddy or Weebly rather than WordPress or a custom build lets you pre-qualify before picking up the phone, and combining it with has_meta_pixel: false and has_booking_widget: false narrows the list to businesses visibly under-invested in their web presence.
That 71% is the 2026-08-08 reference run (dentist / Austin, Texas, onlyWithWebsite: true, n=55); with every filter off the same city measured 34% of 100 rows, because roughly half of mapped businesses list no site to fingerprint.
Only independents? How do I leave the chains out?
Set excludeChains: true. OpenStreetMap marks a chain outlet with a brand tag (and usually brand:wikidata, both seeded from the name-suggestion-index — brand=Aspen Dental, brand:wikidata=Q4807808); that is exactly what sets the is_chain column, so the filter is the column applied before billing: a brand-tagged outlet is dropped before it is crawled, pushed or charged, and the status line and run report say how many were removed. It is a mapper-written tag, so an outlet nobody has branded on the map still comes through — pair it with excludeKeywords (["Aspen Dental"]) for a name you know. Rows you supply yourself (websiteList / startUrls) carry no brand tag (is_chain: null) and are never dropped by it.
Added 2026-08-29 and measured the same day: 38 of the 60 pharmacies OpenStreetMap returned for Austin, Texas were brand-tagged, so a maxItems: 10 run with excludeChains delivered 6 independents and said why — raise maxItems to deepen the 6x candidate pool in chain-dense categories.
{ "category": "dentist", "location": "Austin, Texas", "excludeChains": true, "maxItems": 50 }
How much does it cost to scrape 1,000 local business leads?
1,0