# Yellow Pages US Business Scraper (`scrapyx/yellowpages-us-scraper`) Actor

US business listings from YellowPages.com by city and category: name, phone, street address, website, categories, years in business, Yellow Pages and TripAdvisor ratings, and whether the listing is paid. Lead lists for any US city.

- **URL**: https://apify.com/scrapyx/yellowpages-us-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Yellow Pages US Business Scraper

US business listings from **YellowPages.com** by city and category — for
example every plumber in Austin, TX or every dentist in San Diego, CA. For
each business: **name, phone, street address, city and ZIP, website,
categories, years in business, Yellow Pages rating and review count,
TripAdvisor rating and review count**, open-now status, a description
snippet, and whether it is a **paid listing**.

Reads Yellow Pages' public city/category pages. No login, no browser.

### What it is for

- **Local lead generation** — phone, address and website for every business
  of a type in a city.
- **Market mapping** — how many businesses of a kind, how long they've
  operated, how they're rated.
- **Advertiser research** — which businesses pay for placement, and how high.

### Input

| field | what it does |
| --- | --- |
| `locations` | `"Austin, TX"`, `"New York, NY"`, … |
| `categories` | `plumbers`, `restaurants`, `dentists`, `auto repair service`, … |
| `sortBy` | `best_match` (Yellow Pages' default), `name`, `rating`. |
| `maxItems` | Businesses per city × category (default 60; 30 per page). |

### Four things worth knowing before you trust the data

#### 1. A city Yellow Pages doesn't know is silently replaced

Asking for plumbers in a misspelled city returned "Best 18 Plumbers in
**Winters, TX**" — a different real town, no error. This Actor checks that the
page is about the city you asked for; if not, it returns **no rows** for that
search and tells you which city Yellow Pages answered for.

#### 2. "Best match" is largely a paid ranking

For plumbers in Austin, **17 of the first 30** results are paying subscribers,
listed above the free listings — separately from the sponsored ads box, which
this Actor leaves out entirely. Every row has **`isPaidListing`**, and each
search's summary counts them. (In some categories — New York restaurants —
all 30 were free.) Use `sortBy: name` or `rating` for a non-paid order.

#### 3. The total changes while you page

Page 1 said "of 491"; page 17 said "of 511". The Actor keeps going until a
page has no results rather than trusting the first total, and reports both.

#### 4. Ratings are stored as words

Yellow Pages encodes its stars as CSS words ("four half"); TripAdvisor's
rating travels separately. Both are converted to numbers (`ypRating: 4.5`,
`tripAdvisorRating: 4.5`) with their review counts.

### Output

```json
{
  "recordType": "BUSINESS",
  "name": "Paul & Jimmy's Restaurant",
  "phone": "(212) 475-9540",
  "streetAddress": "123 E 18th St",
  "locality": "New York, NY 10003",
  "website": "https://www.paulandjimmys.com",
  "categories": ["Restaurants", "Italian Restaurants", "Bars"],
  "yearsInBusiness": 76,
  "ypRating": 4,
  "ypReviewCount": 1,
  "tripAdvisorRating": 4.5,
  "tripAdvisorReviewCount": 77,
  "isPaidListing": false
}
```

### Limits

- Yellow Pages' **search box** results are not used — its robots.txt
  disallows them. City × category pages are the permitted route and cover the
  same businesses.
- Business detail pages are not fetched; the category page already carries
  the fields above.

# Actor input Schema

## `locations` (type: `array`):

US cities as 'City, ST' (e.g. 'Austin, TX', 'San Diego, CA'). A city Yellow Pages doesn't recognise is silently replaced by another town there — this Actor detects that and returns no rows for it.

## `categories` (type: `array`):

Yellow Pages category names as in its URLs: 'plumbers', 'restaurants', 'dentists', 'auto repair service', 'real estate agents', 'electricians'. Every city × category pair is one search.

## `sortBy` (type: `string`):

The orders measured to work.

## `maxItems` (type: `integer`):

30 per page. 0 = all.

## `maxConcurrency` (type: `integer`):

Searches in parallel.

## `minRequestInterval` (type: `number`):

Global pacing.

## `proxyConfiguration` (type: `object`):

Off by default, by measurement: Apify's own servers were served on every request, while Apify's datacenter proxy was refused by Yellow Pages' Cloudflare on every request. Residential is the fallback if you are ever refused.

## Actor input object example

```json
{
  "locations": [
    "Austin, TX",
    "New York, NY"
  ],
  "categories": [
    "plumbers",
    "dentists"
  ],
  "sortBy": "best_match",
  "maxItems": 60,
  "maxConcurrency": 3,
  "minRequestInterval": 1.5,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "locations": [
        "Austin, TX",
        "New York, NY"
    ],
    "categories": [
        "plumbers",
        "dentists"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/yellowpages-us-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "locations": [
        "Austin, TX",
        "New York, NY",
    ],
    "categories": [
        "plumbers",
        "dentists",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/yellowpages-us-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "locations": [
    "Austin, TX",
    "New York, NY"
  ],
  "categories": [
    "plumbers",
    "dentists"
  ]
}' |
apify call scrapyx/yellowpages-us-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/yellowpages-us-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/e4bDmkwGx3AncVIFa/builds/TcVY5zzOVLgYwRTay/openapi.json
