# YellowPages Scraper - Business Leads, Phones, Emails & Websites (`rel8ble/yellowpages-scraper`) Actor

Extract US business leads from YellowPages.com: name, phone, full address, website, categories, ratings, hours, years in business, and optional email. Search any keyword in any city. Fast, cheap, HTTP-only - no browser.

- **URL**: https://apify.com/rel8ble/yellowpages-scraper.md
- **Developed by:** [Giovanni Rich](https://apify.com/rel8ble) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YellowPages Scraper - Business Leads, Phone Numbers & Emails

This **Yellow Pages scraper** turns any YellowPages.com search into a clean **business leads** list: business name, phone number, full street address (split into street / city / state / ZIP), website, categories, ratings, opening hours and years in business. Turn on `includeDetails` to also get **business emails** (when listed), description, payment methods, services, social links and GPS coordinates - a simple YellowPages API for lead generation.

Search any keyword in any number of US cities in one run: `["plumber", "roofer"]` x `["Charlotte, NC", "Austin, TX", "10001"]`. It runs over plain HTTP with **no headless browser**, rotates proxy sessions automatically when YellowPages' Cloudflare edge pushes back, and fetches result pages in parallel - fast and cheap.

### How to use

1. Enter one or more **search terms** (e.g. `plumber`, `dentist`) and **locations** (e.g. `Charlotte, NC` or a ZIP code), or paste YellowPages search URLs. Set `maxResults` and tick `includeDetails` if you want emails.
2. Click **Start** and wait for the run to finish (a 20-result test takes well under a minute).
3. Download your leads as **JSON, CSV, Excel** or HTML from the Output tab, or fetch them via the Apify API.

### What you get

- **Contact**: phone, website, full address, and with `includeDetails` the email address, extra phone/fax numbers and social links (Facebook, Instagram, ...).
- **Address split for CRMs**: `street`, `city`, `state`, `zip` as separate fields. Service-area businesses ("Serving the Charlotte area") are flagged with `isServiceAreaOnly`.
- **Reputation**: YellowPages star rating and review count, TripAdvisor rating and review count (restaurants, hotels, attractions), price range.
- **Business profile**: categories, opening hours (`Mo-Fr 08:00-17:00` format), open-now status, years in business, years on YellowPages, claimed/unclaimed listing, review snippet, thumbnail.
- **IDs and links**: YellowPages URL and `ypid` so you can dedupe and track the same business over time.
- **Pagination**: 30 listings per page, pages fetched in parallel until `maxResults`. Duplicates are removed.

### Use cases

- Cold outreach and lead generation for agencies, SaaS and B2B sales (local service businesses: contractors, dentists, lawyers, restaurants, salons, ...)
- Market sizing: how many plumbers / dentists / gyms are in a city, and how established they are
- Enriching a CRM with phone, address, website and category
- Finding businesses **without a website** or with an **unclaimed listing** (great prospects for web and marketing agencies)

### Input

| Field | Default | Description |
|---|---|---|
| `searchTerms` | - | What you'd type into the YellowPages "Find" box: `"plumber"`, `"dentist"`, `"pizza"` |
| `locations` | - | US city + state or ZIP: `"Charlotte, NC"`, `"Austin, TX"`, `"10001"`. Every term is searched in every location. |
| `startUrls` | - | Optional: paste YellowPages search URLs directly (keeps any sorting/filters you set on the site) |
| `maxResults` | 100 | Businesses per term x location (0 = everything YellowPages has, up to ~3,000) |
| `includeDetails` | false | Adds email, description, hours, payment methods, services, social links, founded year, coordinates. One extra request per business. |
| `includeSponsored` | false | Include the paid ad cards YellowPages shows above the results (often from other areas) |
| `proxyConfiguration` | Apify Proxy | Switch to RESIDENTIAL (US) if you see many blocks |

#### Input example

```json
{
    "searchTerms": ["plumber", "hvac contractor"],
    "locations": ["Charlotte, NC", "Austin, TX"],
    "maxResults": 200,
    "includeDetails": true
}
```

### Output example

One dataset item per business. This one is from a real cloud run: "roofing contractors" in Charlotte, NC with `includeDetails: true` (long text shortened).

```json
{
    "searchTerm": "roofing contractors",
    "location": "Charlotte, NC",
    "position": 5,
    "name": "Roofing Contractors of America",
    "ypid": "520015737",
    "categories": ["Roofing Contractors", "General Contractors", "Fire & Water Damage Restoration"],
    "rating": 5,
    "reviewCount": 4,
    "tripAdvisorRating": null,
    "tripAdvisorReviewCount": null,
    "priceRange": null,
    "phone": "(844) 251-9995",
    "address": "10130 Perimeter Pkwy, Charlotte, NC 28216",
    "street": "10130 Perimeter Pkwy",
    "city": "Charlotte",
    "state": "NC",
    "zip": "28216",
    "website": "http://roofingcontractorsofamerica.localsearch.com",
    "yellowPagesUrl": "https://www.yellowpages.com/charlotte-nc/mip/roofing-contractors-of-america-520015737?lid=1001643842969",
    "openingHours": ["Mo-Su"],
    "openStatus": "open 24 hours",
    "yearsInBusiness": null,
    "yearsWithYellowPages": null,
    "snippet": "I just want to see that I had my hesitations about getting my roof done... RCA was amazing.",
    "thumbnail": "https://i4.ypcdn.com/blob/fbe8d98e053d2241e122b33a51931e5458eef6e6_130x130_crop.jpg",
    "isServiceAreaOnly": false,
    "isSponsored": false,
    "isClaimed": true,
    "searchPage": 1,
    "scrapedAt": "2026-09-24T05:30:26.958Z",
    "email": "admin@rcacarolina.com",
    "description": "What we Promise? One thing is for sure that when we will finish the restoration, you will be surprised to see your home or office...",
    "paymentMethods": ["check", "financing available", "insurance", "amex", "visa", "discover", "mastercard", "debit"],
    "foundedYear": null,
    "latitude": 35.347534,
    "longitude": -80.85175,
    "services": ["Storm Damage, Fire Damage, Water Damage, Mold Remediation, Roof Repair, ..."],
    "socialLinks": ["https://www.facebook.com/RCACarolina/", "https://twitter.com/RCACarolinas", "https://www.instagram.com/rcacarolina/"],
    "extraPhones": ["Phone: (704) 251-9995"],
    "aka": null,
    "image": "https://i4.ypcdn.com/blob/fbe8d98e053d2241e122b33a51931e5458eef6e6"
}
```

Restaurants also get `tripAdvisorRating` / `tripAdvisorReviewCount` (e.g. a Chicago pizza place: 4.5 from 2,387 TripAdvisor reviews).

The dataset has two ready-made views: **Overview** (photo, name, phone, address, website, email, categories, rating) and **Contacts (CRM export)** (flat name / phone / email / website / street / city / state / ZIP).

#### Field fill rates (real cloud run: 235 businesses, dentists and plumbers in Denver and Austin)

| Field | Fill rate |
|---|---|
| name, phone, address, categories, YellowPages URL | 100% |
| city / state / ZIP | 98% (the rest are service-area businesses) |
| website | 94% |
| opening hours | 86% |
| review snippet | 83% |
| years in business | 57% |
| YellowPages rating + review count | 46% (only businesses that have YP reviews) |
| TripAdvisor rating | 0% for trades; ~60% for restaurants |

With `includeDetails` (21 detail pages): email 57%, payment methods 86%, hours 95%, description / coordinates 81%, founded year 76%, social links 38%.

### Pricing

**Pay per result: $1.50 per 1,000 results** (businesses saved to the dataset).

- 1,000 business leads = **$1.50**
- 10,000 business leads = **$15**

Duplicates are removed before saving, so you only pay for unique businesses. If you set a maximum cost per run, the scraper stops cleanly when it reaches it. The Apify free plan includes $5 of monthly credit, enough to try it out.

### Integrations

- **Apify API**: run the actor and read results over REST, or with the official JavaScript and Python API clients.
- **Make, Zapier, n8n**: trigger runs and push new leads into your CRM, email tool or spreadsheet.
- **Google Sheets**: export the dataset straight into a sheet.
- **Webhooks**: get notified (or kick off your own pipeline) when a run finishes.
- **MCP for AI agents**: through the Apify MCP server (https://mcp.apify.com), Claude, ChatGPT, Cursor and other agents can call this scraper and use the leads directly.

### Limits (read before large runs)

- **United States only.** YellowPages.com covers US businesses. (Canada, Australia and other countries have separate Yellow Pages sites that this actor does not cover.)
- **Up to ~3,000 results per search** (YellowPages serves at most ~100 pages of 30). For a whole metro area, split by suburb or ZIP code.
- **Email needs `includeDetails`** and only exists when the business published one on YellowPages. Detail mode opens one extra page per business, so it's slower and uses more proxy traffic.
- **Cloudflare**: YellowPages sits behind Cloudflare and challenges a large share of automated requests. The scraper uses the client profile that passes most often and retries challenged pages on a fresh proxy session automatically (in testing: every search page got through, typically after 1-4 retries). **For `includeDetails` runs or big jobs, use the RESIDENTIAL proxy group**: in testing it loaded 100% of pages, versus ~75% of detail pages on datacenter proxies. A detail page that still fails is saved without the extra fields (marked with `detailsError`), so you never lose the listing.
- **Review texts** are not included in this version (only the snippet shown on the results page).

### FAQ

**Does it need a browser?**
No. It makes plain HTTPS requests and reads the listing HTML and the schema.org data YellowPages embeds. That's why it's fast and cheap.

**Which proxy should I use?**
Start with the default Apify Proxy. If the log shows lots of "Request blocked - received 403" retries, switch the proxy group to RESIDENTIAL (US).

**Why did a search return fewer results than `maxResults`?**
YellowPages had fewer listings for that term and location, or a page gave up after all retries. The `RUN_SUMMARY` record in the key-value store shows, per search, how many results YellowPages reports, how many were collected and why it stopped.

**Can I use my own filters or sorting?**
Yes: set them on yellowpages.com, copy the search URL and paste it into `startUrls`.

**Is it legal to scrape YellowPages?**
This actor only collects publicly listed business information. Some listings include names or emails of sole proprietors; you're responsible for using the data in line with YellowPages' terms and privacy and anti-spam laws (e.g. CAN-SPAM, GDPR, CCPA). This is not legal advice - if you're unsure, check with a lawyer.

**How does it avoid blocks?**
Every request goes through Apify Proxy with a desktop browser fingerprint. When YellowPages/Cloudflare answers with a 403, 429 or challenge page, the scraper retires that proxy session and retries on a fresh IP (up to `maxRequestRetries`, default 12). For big or `includeDetails` runs, the RESIDENTIAL (US) proxy group gets through most reliably.

**What are the limits?**
US businesses only, up to ~3,000 results per search (about 100 pages of 30), emails only when the business published one and `includeDetails` is on, and no full review texts. See the Limits section above.

### How it works (for developers)

Each search page is fetched over HTTP with a desktop browser fingerprint and parsed with Cheerio: the result cards give name, phone, address, categories, ratings and badges, and the page's schema.org `LocalBusiness` JSON-LD adds structured address parts and opening hours. The total result count on page 1 is used to fan out the remaining page URLs in parallel. Detail pages are read from their schema.org JSON-LD (email, geo, payment, services, hours) with HTML fallbacks, including decoding Cloudflare-obfuscated email links. Blocked responses (403/429/challenge pages) retire the session and retry on a new proxy IP.

Run it locally:

```bash
npm install
npm test                                  # parser tests on saved fixtures
node src/main.js                          # input in storage/key_value_stores/default/INPUT.json
```

# Actor input Schema

## `searchTerms` (type: `array`):

What you would type into the YellowPages "Find" box, one per line. Examples: "plumber", "dentist", "roofing contractors", "pizza". Every term is searched in every location. If empty (and no search URLs are given), defaults to "plumber".

## `locations` (type: `array`):

US city + state or ZIP code, one per line. Examples: "Charlotte, NC", "Austin, TX", "10001". If empty, defaults to "Charlotte, NC".

## `startUrls` (type: `array`):

Optional alternative to search terms: paste YellowPages.com search URLs (keeps any filters/sorting you set on the site). Example: https://www.yellowpages.com/search?search\_terms=dentist\&geo\_location\_terms=Austin%2C+TX

## `maxResults` (type: `integer`):

Maximum businesses to collect for each term x location search (YellowPages shows 30 per page). Example: 100. Set 0 for everything YellowPages has (up to ~3,000 per search).

## `includeDetails` (type: `boolean`):

Open each business page to add the email address (when listed), full opening hours, description, payment methods, services, social links, year founded and GPS coordinates. Adds one request per business, so runs are slower.

## `includeSponsored` (type: `boolean`):

YellowPages shows a few paid ad cards above the results, often from other areas or categories. Off = only real search results; on = ads are included and flagged with isSponsored.

## `maxConcurrency` (type: `integer`):

How many pages are fetched in parallel. Default 5 is a good balance; higher is faster but triggers more blocks.

## `maxRequestRetries` (type: `integer`):

How many times a failed or blocked page is retried. Each retry uses a fresh proxy session (new IP). Default 12.

## `proxyConfiguration` (type: `object`):

Apify Proxy is required. Datacenter (default) works for most search runs; switch to the RESIDENTIAL group (country US) for includeDetails or large runs if you see many blocks.

## Actor input object example

```json
{
  "searchTerms": [
    "plumber"
  ],
  "locations": [
    "Charlotte, NC"
  ],
  "maxResults": 20,
  "includeDetails": false,
  "includeSponsored": false,
  "maxConcurrency": 5,
  "maxRequestRetries": 12,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `businesses` (type: `string`):

All businesses found, with phone, address, website, categories and ratings.

## `summary` (type: `string`):

Per-search counts, YellowPages totals and stop reasons.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "plumber"
    ],
    "locations": [
        "Charlotte, NC"
    ],
    "maxResults": 20,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("rel8ble/yellowpages-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["plumber"],
    "locations": ["Charlotte, NC"],
    "maxResults": 20,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("rel8ble/yellowpages-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "plumber"
  ],
  "locations": [
    "Charlotte, NC"
  ],
  "maxResults": 20,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call rel8ble/yellowpages-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rel8ble/yellowpages-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4q7AvZ8tpAgmPHkWJ/builds/WIQoONpLYSUwaIypE/openapi.json
