# Facebook Business Contact Extractor (`fanndev/facebook-business-contact-extractor`) Actor

Pull public contact details from Facebook business pages at scale: email, phone, website, address, category, price range, services and followers. Feed it a list of pages, or discover them by keyword and country through the Ad Library, since Facebook's page search is closed to logged-out clients.

- **URL**: https://apify.com/fanndev/facebook-business-contact-extractor.md
- **Developed by:** [Faisal Ahdan naufal](https://apify.com/fanndev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Facebook Business Contact Extractor

Pull the public contact details businesses put on their own Facebook pages:
email, phone, website, street address, category, price range, services, opening
status, follower count and rating.

Feed it a list of pages, or let it **find** the pages by keyword and country.

No login, no cookies to paste, no proxy required.

### How keyword discovery works, and why

Facebook's own page search is closed to signed-out clients. `/search/pages/`,
`/search/places/`, `/search/top/` — every vertical returns **HTTP 404**. There is
no public "find pages matching X" endpoint any more.

The Ad Library is the exception: it *is* keyword-addressable, it is
country-scoped, and every result names its advertiser's page. So that is where
discovery runs. One useful side effect: every business it finds is currently
spending money on advertising, which is a sharper filter than a keyword match.

A live run for `"dentist"` in the US surfaced **53 distinct business pages** in
one pass.

If you already have a list, skip discovery and pass `pages` instead.

### What you get

| Field | Example |
| --- | --- |
| `pageName`, `primaryCategory` | `BARBUTO \| New York NY`, `Italian restaurant` |
| `email`, `phone` | `info@barbutonyc.com`, `+1 212-924-9700` |
| `website` | `http://www.barbutonyc.com/` (unwrapped from Facebook's redirect) |
| `address` | `113 horatio street, New York, NY, United States, 10014` |
| `priceRange`, `services` | `Price range · ££`, `["Dine in"]` |
| `hoursLabel` | `Closed now` |
| `followers`, `recommendPercent`, `reviewCount` | `6987`, `94.0`, `1236` |
| `contactScore` | `4` — how many of the four channels are published |
| `discoveredVia`, `discoveryQuery` | `ad-library`, `dentist @ US` |

With **`includeTransparency`** on, each page also gives `pageCreationDate`
(`17 February 2010`) and `isRunningAds` — useful for telling an established
business from a week-old one, and a prospect already buying ads from one that
is not.

### Input example

```json
{
  "searchKeywords": ["dental implants", "orthodontist"],
  "countries": ["US", "GB"],
  "maxPages": 300,
  "requireAnyContact": true,
  "includeTransparency": true,
  "exportFormats": ["xlsx"]
}
```

Or, with a list you already have:

```json
{
  "pages": ["BarbutoNYC", "https://www.facebook.com/gramercytavern"],
  "requireEmail": true
}
```

### Honest limits

- **Opening hours are not published.** Facebook shows logged-out visitors only
  `Open now` / `Closed now` — there is no weekly schedule in the payload. The
  field is named `hoursLabel` rather than `hours` so nobody mistakes it for one.
- **A null contact field means the business publishes none**, not that the
  parser missed it. That distinction is the whole point of a lead list, so the
  actor never falls back to guessing a value from nearby text. Addresses found
  in the page's *description* rather than its contact fields are reported
  separately as `emailsFromDescription`.
- **Missing slugs are reported as errors.** Facebook answers HTTP 200 for pages
  that do not exist, so a status-code check would return a blank row for every
  typo. This actor checks the payload size and gives you a `page_not_found`
  error row instead.
- **These pages are big** — 8 to 18 MB each. That is why `concurrency` defaults
  to 4. Raising it speeds a run up and costs memory and bandwidth.
- Follower counts are abbreviated by Facebook; the parsed number is approximate
  and the raw label is kept beside it.

# Actor input Schema

## `searchKeywords` (type: `array`):

Find pages instead of listing them. Facebook's own page search is closed to logged-out clients (every /search/ path returns 404), so discovery runs through the Ad Library: it is keyword-addressable and every result names its advertiser. A useful side effect is that the businesses it finds are all currently spending money on ads.

## `countries` (type: `array`):

ISO country codes the Ad Library search runs in. Ignored when you supply pages directly.

## `pages` (type: `array`):

Extract contacts from a list you already have: vanity slugs ('BarbutoNYC') or full Facebook page URLs.

## `startUrls` (type: `array`):

Page URLs, for pasting a list straight out of another actor's dataset.

## `maxPages` (type: `integer`):

Caps how many business pages are opened, across both direct input and discovery.

## `includeTransparency` (type: `boolean`):

Adds the page creation date and whether the business is currently running ads - both useful for qualifying a lead. Costs one extra request per page.

## `requireEmail` (type: `boolean`):

Drop pages that publish no email address.

## `requirePhone` (type: `boolean`):

Drop pages that publish no phone number.

## `requireWebsite` (type: `boolean`):

Drop pages that publish no website.

## `requireAnyContact` (type: `boolean`):

Drop pages that publish no email, phone, website or address at all.

## `minFollowers` (type: `integer`):

Pages whose follower count could not be read are dropped too, since an unknown count cannot be shown to clear the bar.

## `categoryContains` (type: `string`):

Case-insensitive substring match against the page's self-declared categories, e.g. 'restaurant', 'dental'.

## `contentLanguages` (type: `array`):

Two-letter language codes. Narrows the Ad Library discovery search to ads in these languages.

## `concurrency` (type: `integer`):

About tabs are large - 8 to 18 MB each - so raising this speeds the run up but costs memory and bandwidth.

## `emitSummary` (type: `boolean`):

Append one RUN\_SUMMARY record: how many pages were read, and how many publish an email, phone, website or address.

## `exportFormats` (type: `array`):

Also write the results to the key-value store in these formats. The dataset is always produced regardless.

## `proxyConfiguration` (type: `object`):

Optional. Page about tabs apply no TLS or IP gate and this actor does not paginate through Facebook's GraphQL, so the direct connection works on the Apify platform too. Keyword discovery reads the Ad Library's rendered first page only, which is not rate-limited. Use a proxy to spread a very large run across IPs.

## Actor input object example

```json
{
  "searchKeywords": [
    "dentist"
  ],
  "countries": [
    "US"
  ],
  "maxPages": 100,
  "includeTransparency": false,
  "requireEmail": false,
  "requirePhone": false,
  "requireWebsite": false,
  "requireAnyContact": false,
  "contentLanguages": [],
  "concurrency": 4,
  "emitSummary": true,
  "exportFormats": [],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

One contact row per Facebook page, plus discovery rows, the run summary and error rows.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchKeywords": [
        "dentist"
    ],
    "countries": [
        "US"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fanndev/facebook-business-contact-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchKeywords": ["dentist"],
    "countries": ["US"],
}

# Run the Actor and wait for it to finish
run = client.actor("fanndev/facebook-business-contact-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchKeywords": [
    "dentist"
  ],
  "countries": [
    "US"
  ]
}' |
apify call fanndev/facebook-business-contact-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fanndev/facebook-business-contact-extractor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZBav7jRtK0jfEg2hS/builds/oDlyIBKPAog5e1Ipw/openapi.json
