# Yellow Pages Indonesia Business & Lead Scraper (`fanndev/yellowpages-id-business-scraper`) Actor

Extract Indonesian business listings from Yellow Pages ID: name, phone, email, website, full address and GPS coordinates. Six modes - keyword search, category listing, company profiles, B2B products, market-density insights per province, and a province/city directory. No API key, no login.

- **URL**: https://apify.com/fanndev/yellowpages-id-business-scraper.md
- **Developed by:** [Faisal Ahdan naufal](https://apify.com/fanndev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Yellow Pages Indonesia Business & Lead Scraper

Extract Indonesian business listings from Yellow Pages ID — name, phone, email, website, full
postal address and GPS coordinates — plus B2B products, per-province market density and a
province/city directory. Pure HTTP, no browser, no API key, no login.

***

### Modes

| Mode | Input | Emits | What it is for |
| --- | --- | --- | --- |
| `search_businesses` | `queries` (+ `region` / `city`) | `BUSINESS` | The general lead builder. Keyword search across the whole directory. |
| `category_businesses` | `queries` = category slugs (+ `city`) | `BUSINESS` | The curated "top suppliers" block of a `/places/<slug>` page. Small but hand-ranked. |
| `company_details` | `companyUrls` | `COMPANY` | Full profile: labelled phone numbers, founding year, legal form, categories. |
| `search_products` | `queries` | `PRODUCT` | B2B product offers with price, seller and the seller's real shop URL. |
| `market_insights` | `queries` | `MARKET_INSIGHT` | Listing counts per province and city for a keyword, plus related keywords. One request. |
| `geo_directory` | `provinceCodes` | `GEO_NODE` | Provinces and their cities — the seed list for a systematic nationwide crawl. |

Set `enrichWithDetails: true` on the two business modes to merge each company's profile page into
its listing row (one extra request per company).

#### Why one actor with modes rather than six actors

Every mode hits the same Django application, through the same host choice, with the same
pagination ceiling and the same record envelope. The host choice (below) is the fragile part of
this scraper — splitting it across six actors would mean six copies to re-verify whenever the
site's WAF configuration changes. Modes that emit different shapes stay separable through
`recordType`, which the dataset schema exposes as its own named views.

***

### How it gets the data

`www.yellowpages.id` puts a **Cloudflare Managed Challenge** on exactly the three paths that carry
the directory data:

| Path | yellowpages.id | yoys.id |
| --- | --- | --- |
| `/` , `/product-*` , `/citymap-*` | 200 | 200 |
| `/profil-*` (company profiles) | **403 `cf-mitigated: challenge`** | 200 |
| `/places/*` (categories) | **403 `cf-mitigated: challenge`** | 200 |
| `/listing/*` (search) | **403 `cf-mitigated: challenge`** | 200 |

All 18 `curl_cffi` TLS profiles — `chrome99_android` through `chrome136`, `safari15_5`–`18_0`,
`edge99`/`edge101`, `firefox133`/`firefox135` — return the **identical** 403 on the gated paths
while the homepage returns 200 from the same IP and the same session. A homepage warm-up issues no
clearance cookie. That makes it a path-scoped WAF *rule*, not a JA3 fingerprint gate, and it is not
solvable HTTP-only.

`www.yoys.id` is the same application serving the same Indonesian dataset — YP Media Ltd runs both
brands off one backend. Verified by resolving identical IDs on both hosts: `product-…_52399`
returns company `492034` on each, and `profil-551113` is the same company on each, byte-comparable
apart from brand strings. yoys.id applies no challenge on any path.

**So the actor reads from `yoys.id` and writes the matching `www.yellowpages.id` URL into
`yellowpagesUrl` on every record**, so results stay citable against the brand you asked for. If
yoys.id ever acquires the same rule the client fails over to yellowpages.id automatically and
rotates through the TLS ladder before giving up.

No IP-based blocking was observed, so **the proxy is off by default** — turn on Apify Proxy only
for very high volume.

***

### Known limits

These are properties of the site, measured rather than assumed. They are the things that will
surprise you if you do not know them.

- **250 rows per query, hard.** The result header advertises figures like "10000 hasil", but the
  backend serves only 10 pages of 25. Page 11 answers **HTTP 200 with a fully rendered, empty
  page** — not a 404 — so anything that pages on status code alone will loop silently. To go
  deeper, slice the same keyword by `region` or `city`; each slice gets its own 250. Run
  `market_insights` first to see where the volume actually is.
- **The free-text location box is a decoy.** The site's own `l=` parameter returns zero results for
  every value while echoing the place name in the page title. This actor ignores it and uses the
  facet parameters (`adm` / `cty`) that the sidebar actually uses.
- **Promoted listings ignore the region filter.** A sponsored entry from another province can
  appear in a filtered result set. Filter on `address.region` downstream if that matters.
- **Two kinds of entry share the directory.** `entryType: "COMPANY"` has a claimed profile page and
  a `companyId`; `entryType: "PHONE_LISTING"` is a phone-book record with a number but no profile
  page, so `companyId` is null and enrichment skips it.
- **Profile pages carry no coordinates.** Geo comes from listing pages only, so enrichment
  preserves the listing's `geo` rather than overwriting it.
- **Emails are rare.** Most listings publish none. Where one exists it is behind Cloudflare's email
  obfuscation, which this actor decodes.
- **Province codes are the site's own**, not Indonesia's official BPS numbering (`02` is Bali, not
  North Sumatra). Unused codes return a 200 page titled "Aceh" with zero cities; those are reported
  as `hollow_response` errors rather than pushed as empty provinces.
- **Missing companies return 302 → `/410`**, not 404. Reported as `not_found`.
- Fields are passed through as the site publishes them, so occasional upstream junk (an email in a
  `website` field, `foundedYear: "10"`) appears verbatim rather than being silently cleaned.

***

### Output

Every record carries the house envelope — `_input`, `_source`, `_scrapedAt`, `recordType` — on top
of the payload. Failures become `ERROR` rows with `_error` / `_errorDetail` instead of vanishing,
and truncated queries carry a `_warning`.

```json
{
  "_input": "percetakan",
  "_source": "S1-listing-search",
  "_scrapedAt": "2026-09-20T15:23:18Z",
  "recordType": "BUSINESS",
  "entryId": "365207",
  "entryType": "COMPANY",
  "companyId": "365207",
  "name": "Percetakan Raja Setting",
  "description": "Rajasetting merupakan Percetakan Offset dan Digital Printing…",
  "phone": "+62 81572606669",
  "emails": [],
  "website": "https://rajasetting.com",
  "address": {
    "street": "Jl. Babakan H. Tamim No. 25",
    "locality": "Bandung",
    "region": "Jawa Barat",
    "postalCode": "40125",
    "country": "Indonesia"
  },
  "geo": { "latitude": -6.90525, "longitude": 107.647461 },
  "detailUrl": "https://www.yoys.id/profil-365207-percetakan-raja-setting-bandung.html",
  "yellowpagesUrl": "https://www.yellowpages.id/profil-365207-percetakan-raja-setting-bandung.html"
}
```

Measured completeness on a 20-row `percetakan` / Jawa Barat run: phone 20/20, website 20/20,
coordinates 20/20, postcode 19/20, region 16/20, street 15/20.

***

### Recipes

**Build a city lead list**

```json
{ "mode": "search_businesses", "queries": ["percetakan"], "city": "Bandung",
  "enrichWithDetails": true, "maxItemsPerQuery": 250 }
```

**Find where the volume is before spending requests**

```json
{ "mode": "market_insights", "queries": ["bank"] }
```

Then feed the returned `regions[].filterValue` back in as `region` to collect 250 rows per
province instead of 250 nationwide.

**Nationwide sweep** — run `geo_directory` once for the city list, then iterate
`search_businesses` over `city` values.

***

### Stack

`curl_cffi` (TLS impersonation, `chrome124` primary with a 5-profile ladder) + `selectolax`
for parsing. Extraction follows JSON-LD first, HTML selectors only where JSON-LD does not reach.
Listing rows merge both layers per field, because neither is complete on its own.

Retries are exponential with jitter; `429`/`5xx` back off, challenges rotate TLS profile then host.
A shape change raises `UnexpectedShape` and fails the run loudly rather than emptying the dataset.

# Actor input Schema

## `mode` (type: `string`):

What to extract. Business search is the general-purpose lead builder; the others target one surface each.

## `queries` (type: `array`):

One or more search keywords (e.g. "bank", "hotel", "percetakan"). For Category listing mode use the slug from a /places/<slug> URL. Required for every mode except Company details and Geo directory.

## `companyUrls` (type: `array`):

Company details mode only. Accepts a full https://www.yellowpages.id/profil-<id>-<slug>.html URL, the equivalent yoys.id URL, or a bare numeric company ID.

## `region` (type: `string`):

Business search mode. Province name as the site spells it, e.g. "Jawa Barat", "Bali", "Daerah Khusus Ibukota Jakarta". Run Market insights first to see the exact values and their hit counts. Note that promoted listings are returned regardless of this filter.

## `city` (type: `string`):

City name, e.g. "Bandung", "Kota Medan". Used as a query filter in Business search and as a URL segment in Category listing.

## `provinceCodes` (type: `array`):

Geo directory mode. The site uses its own two-digit codes, not Indonesia's official BPS numbering - e.g. 01 Aceh, 02 Bali, 04 Jakarta, 07 Central Java, 30 West Java. Leave empty to crawl all 34 codes that exist; unused codes return an empty page and are reported as errors.

## `maxItemsPerQuery` (type: `integer`):

Cap on records per keyword. The site itself serves at most 250 rows per query (10 pages x 25), so higher values will not return more - split the query by region or city instead.

## `enrichWithDetails` (type: `boolean`):

Business search and Category listing modes: fetch each company's profile page to add founding year, legal form, labelled phone numbers and categories. Costs one extra request per company.

## `maxConcurrency` (type: `integer`):

Parallel requests for profile fetches and the geo directory.

## `proxyConfiguration` (type: `object`):

Optional. This target applies no IP-based blocking - the same address is served on every open path - so runs work fine with no proxy and that is the default. Turn on Apify Proxy (Residential recommended) only if you run very high volume and want to spread the load.

## Actor input object example

```json
{
  "mode": "search_businesses",
  "queries": [
    "bank"
  ],
  "companyUrls": [
    "https://www.yellowpages.id/profil-185272-pt-bank-mandiri-persero-tbk-jakarta.html"
  ],
  "maxItemsPerQuery": 250,
  "enrichWithDetails": false,
  "maxConcurrency": 4,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Every business, company, product, insight, geo and error record produced by this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "bank"
    ],
    "companyUrls": [
        "https://www.yellowpages.id/profil-185272-pt-bank-mandiri-persero-tbk-jakarta.html"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fanndev/yellowpages-id-business-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["bank"],
    "companyUrls": ["https://www.yellowpages.id/profil-185272-pt-bank-mandiri-persero-tbk-jakarta.html"],
}

# Run the Actor and wait for it to finish
run = client.actor("fanndev/yellowpages-id-business-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "bank"
  ],
  "companyUrls": [
    "https://www.yellowpages.id/profil-185272-pt-bank-mandiri-persero-tbk-jakarta.html"
  ]
}' |
apify call fanndev/yellowpages-id-business-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fanndev/yellowpages-id-business-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/85Ljy4A0ej4Hr2f2u/builds/7TttdCazaqBfvJ6L8/openapi.json
