# Directory Leads Scraper — Yellow Pages, BBB, Superpages & More (`inovaflow/directory-leads-scraper`) Actor

Search Yellow Pages, Superpages, BBB, Yellow Pages Canada and Thumbtack by keyword and location and get one deduplicated lead per business: phone, address, website, rating, years in business, plus emails and social profiles from its website. No login, dataset-only, MCP-ready.

- **URL**: https://apify.com/inovaflow/directory-leads-scraper.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** Lead generation, Business, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 leads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

If you sell to local businesses, the classic directories still hold prospects you will not find anywhere else: Yellow Pages and Superpages list every plumber, dentist and law office with a phone and a website, BBB adds the accredited ones with a letter grade, Yellow Pages Canada covers the north, and Thumbtack shows which service pros are actually winning jobs. **Directory Leads Scraper** searches all of them at once — `dentists in Boston, MA`, `roofers in Toronto, ON` — merges the same business found in several directories into **one lead**, and then visits each business website to add **emails, extra phone numbers and social profiles**.

- **Agencies & freelancers** — a list of every restaurant / clinic / contractor in a city, with phone, website and email, from sources Google Maps scrapers do not cover.
- **Sales teams (SDRs)** — territory prospecting by trade and city, deduplicated, straight into your CRM.
- **Local marketers & directory builders** — citation audits: which directories list a business, under which categories, with which phone.
- **Founders & researchers** — market maps of a niche with tenure (years in business), reputation and BBB grade side by side.

### What does Directory Leads Scraper do?

It is a **business directory scraper and lead extractor**: it runs your searches on **yellowpages.com, superpages.com, bbb.org, yellowpages.ca and thumbtack.com** (the directories that fit the location's country, or the ones you pick), pages through the results, and groups listings that share a phone number, website or name + street into one business. Every business then gets its website crawled (homepage plus contact / about pages) for **email addresses, additional phone numbers and Facebook, Instagram, LinkedIn, X, YouTube, TikTok and WhatsApp links**, a **lead score** and a **contact status**. No login, no API key, no browser extension — dataset only, MCP-ready.

### Why use this directory leads scraper?

- **Five directories, one row per business.** A dentist listed on Yellow Pages, Superpages and BBB comes back once, with Yellow Pages' phone and website, Superpages' years in business and BBB's A+ grade — and `sources[]` pointing at all three listings.
- **Leads, not raw listings.** Emails and socials are included in every run; the contact fee applies **only when a website actually yielded one**.
- **Fields the directories really show.** Rating, review count, years in business, open status, coordinates, BBB accreditation, Thumbtack hires — real or `null`, never guessed.
- **Filters that save money:** only businesses with a website / phone / email, minimum rating, minimum reviews, exclude sponsored placements. Filtered-out businesses are never charged.
- **CRM-ready.** Download CSV / Excel / JSON, or grab the auto-generated `LEADS.csv` — one row per lead with the primary email up front.
- **Runs anywhere.** Schedule weekly territory refreshes, call it from the API, from Zapier / Make / n8n, or from an AI agent through MCP.

### What data does it extract?

| Field | Description |
| --- | --- |
| `name`, `category`, `categories` | Business name and every category the directories list it under |
| `address`, `street`, `city`, `state`, `postalCode`, `countryCode` | Structured address (`US` or `CA`) |
| `phone`, `phoneE164` | Phone number as listed and in international format |
| `website`, `domain` | Business website (directory self-links are never reported as a website) |
| `emails`, `primaryEmail` | Email addresses found on the website (same-domain, role addresses like `info@` first) |
| `phonesFromWebsite` | Additional phone numbers found on the website |
| `socials` | `facebook`, `instagram`, `linkedin`, `twitter`, `youtube`, `tiktok`, `pinterest`, `whatsapp` profile links |
| `rating`, `reviewsCount` | Star rating and number of reviews on the primary directory |
| `yearsInBusiness` | Years in business (Yellow Pages / Superpages badge, Thumbtack founding year) |
| `bbbRating` | BBB letter grade (`A+` … `F`) when the business is on BBB |
| `extras` | Directory-specific facts: `bbbAccredited`, `bbbRatingScore`, `hires`, `badge`, `employees`, `foundedYear`, `responseTimeHours`, `serviceArea`, `otherPhones` |
| `description`, `openStatus`, `lat`, `lng` | Listing snippet, open/closed status and coordinates where shown |
| `directory`, `directoryUrl`, `directoryId` | The primary listing this row was built from |
| `directories`, `sources` | Every directory (and listing URL) the business was merged from |
| `leadScore` | 0–100 reachability score (email, phone, website, socials, reputation, tenure, multi-directory presence) |
| `contactStatus`, `websitePagesCrawled` | `ok`, `no_contacts`, `no_website`, `unreachable`, `skipped` |
| `isAd`, `query`, `scrapedAt` | Provenance |

### How to scrape business directory leads

1. Open the Actor and click **Try for free**.
2. Under **What to search** type one search per line with the place in it, e.g. `plumbers in Austin, TX`, `wedding photographers in Toronto, ON`. Or type only the trade (`plumbers`) and fill **Location** (`Austin, TX`).
3. Leave **Directories** empty to search every directory that covers the country, or pick specific ones. Set **Max listings per search and directory** (default 50).
4. Turn on **Only leads with an email address** if you pay only for emailable leads; **Only businesses with a phone number** removes Thumbtack-only rows.
5. Click **Start**. Leads stream into the **Output** tab once the directory search is complete; a 4-directory search with 50 listings each takes 2–4 minutes including contact extraction.
6. Export as **CSV, Excel, JSON** or open `LEADS.csv` and import it into HubSpot, Pipedrive, Lemlist, Instantly or Google Sheets.

#### Pasting directory URLs

Under **Advanced → Directory search URLs** you can paste search or category pages copied from any of the five directories (for example `https://www.yellowpages.com/search?search_terms=plumbers&geo_location_terms=Austin%2C+TX` or `https://www.bbb.org/search?find_country=USA&find_text=roofing&find_loc=Denver%2C+CO`). Each is paged like a search.

#### How to cover a whole city

Yellow Pages and Superpages return up to a few thousand listings per search; BBB a few hundred. To cover a metro, split the search by neighbourhood or suburb (`dentists in Cambridge, MA`, `dentists in Brookline, MA`) — the same business found by two searches is delivered once.

### How much does it cost to scrape directory leads?

Pay-per-event, no subscription: **$0.01 per lead delivered** plus **$0.01 when the business website yielded at least one email or social profile**, and a small run-start fee. So 100 fully enriched leads cost at most **$2** — and less in practice, because businesses without contacts are charged only the base price. Businesses you filter out, duplicates across directories and searches, and failed searches are free. Apify's free plan is enough for a few hundred leads a month.

### Input

Only **What to search** (or a pasted directory URL) is required. Everything else has sensible defaults — see the **Input** tab. Example:

```json
{
    "searchQueries": ["dentists in Boston, MA", "orthodontists in Boston, MA"],
    "maxPerQuery": 100,
    "onlyWithWebsite": true,
    "minRating": 4
}
```

### Output

One dataset item per business:

```json
{
    "name": "Family Dental Care",
    "category": "Dentists",
    "categories": ["Dentists", "Cosmetic Dentistry"],
    "address": "1444 Dorchester Avenue, Boston, MA 02122",
    "street": "1444 Dorchester Avenue",
    "city": "Boston",
    "state": "MA",
    "postalCode": "02122",
    "countryCode": "US",
    "phone": "(857) 375-7616",
    "phoneE164": "+18573757616",
    "website": "https://www.familydentalcarema.com/",
    "domain": "familydentalcarema.com",
    "emails": ["info@familydentalcarema.com"],
    "primaryEmail": "info@familydentalcarema.com",
    "phonesFromWebsite": ["(857) 375-7616"],
    "socials": { "facebook": "https://www.facebook.com/familydentalcarema", "instagram": null, "linkedin": null, "twitter": null, "youtube": null, "tiktok": null, "pinterest": null, "whatsapp": null },
    "rating": null,
    "reviewsCount": null,
    "yearsInBusiness": 37,
    "bbbRating": null,
    "description": "New Patients Welcome - We Speak Spanish, Portuguese, Vietnamese, Tamil, Russian …",
    "openStatus": "closed now",
    "lat": null,
    "lng": null,
    "directory": "yellowpages",
    "directoryUrl": "https://www.yellowpages.com/boston-ma/mip/family-dental-care-523622",
    "directoryId": "523622",
    "directories": ["yellowpages", "superpages"],
    "sources": ["https://www.yellowpages.com/boston-ma/mip/family-dental-care-523622", "https://www.superpages.com/boston-ma/bpp/family-dental-care-523622"],
    "extras": {},
    "contactStatus": "ok",
    "websitePagesCrawled": 2,
    "leadScore": 88,
    "isAd": false,
    "query": "dentists in Boston, MA",
    "scrapedAt": "2026-09-26T11:20:14.000Z"
}
```

The key-value store holds `LEADS.csv` (first 5,000 rows) and `OUTPUT`, a run summary with leads delivered, listings per directory, how many were merged, filtered out, pages fetched and any notice about skipped searches.

### Tips

- **Location format.** `City, ST` (`Denver, CO`) or `City, Province` (`Calgary, AB`) works best: it picks the right country's directories and lets Thumbtack form its category URL.
- **Thumbtack categories** use the service name as a slug (`house cleaning` → `/ma/boston/house-cleaning/`); an unknown service simply returns no Thumbtack rows.
- **Sponsored listings** (`isAd: true`) are real businesses that paid for placement — keep them unless you audit organic rank.

### FAQ

**Which directories are supported?** yellowpages.com, superpages.com, bbb.org (US & Canada), yellowpages.ca and thumbtack.com. Yelp, Houzz, Angi, TripAdvisor and Manta are not — they block automated reads or serve JavaScript challenges, and this Actor never uses a browser, a login or a bypass service.

**Why does a BBB-only row have no website?** BBB's search results do not include one, and its profile pages are behind a challenge. When the same business is on Yellow Pages or Superpages the website comes from there.

**Why has a Thumbtack row no phone?** Thumbtack does not show phone numbers or websites; those rows carry rating, reviews, hires and tenure and merge with other directories by name and city.

**Is a business ever charged twice?** No. Listings are merged before delivery; a business that appears in three directories and two searches is one lead.

**Do I need a proxy?** Apify datacenter proxies (the default) are enough for BBB, Yellow Pages Canada and Thumbtack. Yellow Pages and Superpages are read through the Actor's own connection.

**Can an AI agent use it?** Yes — run it through Apify's MCP server or the API; the flat rows with `directoryUrl`, `sources[]` and `scrapedAt` are made for that.

### Support

Open an issue on the Actor's page or write to the Inovaflow team via the Apify Console. Related Actors: [Google Maps Leads Scraper](https://apify.com/inovaflow/google-maps-leads-scraper), [Local Business Email Finder](https://apify.com/inovaflow/local-business-email-finder).

# Actor input Schema

## `searchQueries` (type: `array`):

One search per line, with the place in it — e.g. `dentists in Boston, MA`, `plumbers in Austin, TX`, `roofing contractors in Toronto, ON`. Or type only the trade (`dentists`) and set "Location" below once for all of them.

## `location` (type: `string`):

City and state/province used for every search that does not name one — e.g. `Boston, MA` or `Toronto, ON`. US and Canada are covered.

## `directories` (type: `array`):

Which directories to search. Leave empty to search every directory that covers the location's country (US: Yellow Pages, Superpages, BBB, Thumbtack; Canada: Yellow Pages Canada, BBB). The same business found in several directories is delivered once.

## `maxPerQuery` (type: `integer`):

How many listings to collect from each directory for each search (30 per page on Yellow Pages / Superpages, 15 on BBB, 35 on Yellow Pages Canada, one page of ~8 on Thumbtack). Listings that turn out to be the same business are merged, so the number of delivered leads is usually lower.

## `enrichContacts` (type: `boolean`):

Visits every business website (home + contact/about pages) and pulls out email addresses, extra phone numbers and Facebook / Instagram / LinkedIn / X / YouTube / TikTok / WhatsApp links. Charged only when something was found.

## `onlyWithEmail` (type: `boolean`):

Deliver (and pay for) only businesses where at least one email address was found on the website.

## `onlyWithWebsite` (type: `boolean`):

Skip businesses that list no website in any directory.

## `onlyWithPhone` (type: `boolean`):

Skip businesses that list no phone number in any directory (Thumbtack never shows one).

## `minRating` (type: `integer`):

Skip businesses rated below this many stars (1–5) in the directory. Businesses without a star rating are skipped too. Leave empty for no minimum.

## `minReviews` (type: `integer`):

Skip businesses with fewer directory reviews than this. Businesses without a review count are skipped too.

## `includeAds` (type: `boolean`):

Keep businesses that appear as paid/sponsored placements on the directory page (they are real local businesses too). They are flagged with `isAd: true`.

## `startUrls` (type: `array`):

Paste search or category pages copied from yellowpages.com, superpages.com, bbb.org/search, yellowpages.ca or thumbtack.com (e.g. `https://www.yellowpages.com/search?search_terms=plumbers&geo_location_terms=Austin%2C+TX`). Each is paged like a search.

## `maxContactPages` (type: `integer`):

How many pages of each business website to read when extracting contacts (homepage + contact-like pages).

## `maxConcurrency` (type: `integer`):

How many directory pages and business websites are fetched in parallel.

## `proxyConfiguration` (type: `object`):

Used for BBB, Yellow Pages Canada and Thumbtack pages; Apify datacenter proxies are the default and work for all shipped directories. Yellow Pages and Superpages are read through the Actor's own connection (their edge blocks proxy ranges) and fall back to this proxy only if that is blocked.

## Actor input object example

```json
{
  "searchQueries": [
    "dentists in Boston, MA"
  ],
  "directories": [],
  "maxPerQuery": 15,
  "enrichContacts": true,
  "onlyWithEmail": false,
  "onlyWithWebsite": false,
  "onlyWithPhone": false,
  "includeAds": true,
  "maxContactPages": 3,
  "maxConcurrency": 8,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `leads` (type: `string`):

One row per business with directory details plus emails, phones and social profiles from its website.

## `csv` (type: `string`):

Spreadsheet / CRM-ready CSV of the leads (first 5,000 rows).

## `summary` (type: `string`):

Counts: leads delivered, with email, with any contact, listings per directory, merged, filtered out, pages fetched, blocked requests.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "dentists in Boston, MA"
    ],
    "maxPerQuery": 15,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/directory-leads-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["dentists in Boston, MA"],
    "maxPerQuery": 15,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/directory-leads-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "dentists in Boston, MA"
  ],
  "maxPerQuery": 15,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call inovaflow/directory-leads-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/directory-leads-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/AliQfST5X0eT7RpKR/builds/01ZYzFEQyMmZma7S6/openapi.json
