# PagesJaunes Business Data — Leads & Contacts (`b2b_leads/pagesjaunes-real-time-data-scraper`) Actor

Collect live PagesJaunes business data across France: search by activity and city, get phones, emails, websites, ratings, opening hours, and legal ids (SIRET/SIREN). Real-time streaming to your dataset with optional webhooks. Free plan exports 2 records; paid plans are unlimited.

- **URL**: https://apify.com/b2b\_leads/pagesjaunes-real-time-data-scraper.md
- **Developed by:** [Emmanuel](https://apify.com/b2b_leads) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $15.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PagesJaunes Real-Time Data

Live PagesJaunes business directory intelligence: professional search across every French city, full listing details, public reviews, and lead enrichment (phones, websites, emails, socials, legal identifiers). Checkbox features — enable only what you need. Structured JSON streamed to your dataset in real time.

> **Free plan notice:** on the Apify Free plan this Actor exports a small sample of results per run (2 records). Upgrade to a paid Apify plan for full, unlimited exports.

### Who is this for

Built for agencies, local-SEO teams, sales teams, and CRM owners who work with French businesses:

- **Lead generation** — pull plumbers, dentists, lawyers, restaurants, garages… with phones, websites, and emails ready for outreach.
- **Local SEO & citation audits** — track which businesses in a city carry ratings, review counts, and working websites.
- **Market & competitor research** — size any trade in any postcode: how many players, how well reviewed, who has a site.
- **CRM & enrichment pipelines** — stream records straight into Sheets, Airtable, HubSpot, Pipedrive via webhooks.
- **Franchise & network monitoring** — watch listings by city for presence, ratings, and contact accuracy.
- **AI agents & MCP workflows** — feed live directory answers into your agent (see below).

### Feature matrix

| Feature | What it does | Default |
|---|---|---|
| 🔎 **Professional & business search** | Multi-task search: pair an activity ("plombier", "restaurant", "avocat"…) with any French location ("Paris 75015", "Lyon", "Bordeaux 33000"). Each task collects its businesses across result pages — every business on every page is a row. | **On** |
| 📋 **Listing Details** | Turn numeric listing ids or profile URLs into full records: phone, opening hours, payment methods, coordinates, legal identifiers. | Off |
| 🔗 **Scrape By URL** | Paste any PagesJaunes listing URL (collects every business on that page) or profile URL (one full record). | Off |
| 🎯 **Lead details** | Enriches **every** row in place: phone, website, emails (including from the business's own website), socials, opening hours, geo, SIRET/SIREN, legal form. Enrich, never filter — every business is still exported. | **On** |
| ⭐ **Reviews** | Public customer reviews for businesses requested via Listing Details — one row per review. | Off |
| 🌐 **Webhooks** | JSON or Slack. Every record is POSTed to your URL the moment it is collected — CRM, Zapier, Make, n8n. | Off |

Features combine freely: search + lead details is the classic lead-gen run; Listing Details + Reviews is the deep-dive run; Scrape By URL re-collects a page you already know.

### Input reference

#### 🔎 Professional & business search

| Field | Type | Default | Notes |
|---|---|---|---|
| `enableSearch` | boolean | `true` | Master switch for search. |
| `searchTasks` | array | 2 demo rows | One row per search: `activity`, `location`, optional `maxResults`. |
| `maxResultsPerTask` | integer | `10` | Global default cap per task; override per task with `maxResults`. |
| `maxPagesPerTask` | integer | *empty* | Optional page-depth cap (20 businesses per page). Empty = keep going until this task's businesses are all collected or the directory's last page. |

Example `searchTasks`:

```json
[
    { "activity": "plombier", "location": "Paris 75015", "maxResults": 10 },
    { "activity": "restaurant", "location": "Lyon 69002", "maxResults": 10 },
    { "activity": "dentiste", "location": "Marseille 13008", "maxResults": 10 }
]
```

Activities work best in French ("plombier", "coiffeur", "garage automobile", "avocat", "kine"). Locations accept city, postcode, or arrondissement forms — adding the postcode sharpens arrondissement-level searches.

#### 📋 Listing details

| Field | Type | Default | Notes |
|---|---|---|---|
| `enableListingDetails` | boolean | `false` | Master switch. |
| `listingIds` | string\[] | `[]` | Numeric ids, e.g. `64391412` — the number after `/pros/` in any profile URL. |
| `listingUrls` | string\[] | `[]` | Full profile URLs. |

Every requested id/URL produces exactly one row (found or not), so counts stay predictable.

#### 🔗 Scrape By URL

| Field | Type | Default | Notes |
|---|---|---|---|
| `enableScrapeByUrl` | boolean | `false` | Master switch. |
| `scrapeUrls` | string\[] | `[]` | Listing page URLs (one row per business card) or profile URLs (one full record). |

Only `pagesjaunes.fr` URLs are accepted; anything else fails validation with a clear error.

#### 🎯 Lead details

| Field | Type | Default | Notes |
|---|---|---|---|
| `enableLeadDetails` | boolean | `true` | Enrich **every** business row in place. Never filters — output count stays predictable. |

What enrichment adds, where publicly available:

- **Phone** and **website** from the full listing.
- **Emails** — from the listing *and* from the business's own website (homepage plus its contact/legal pages).
- **Socials** — Facebook, Instagram, LinkedIn, X, YouTube, TikTok, Pinterest.
- **Opening hours** (structured per day), **payment methods**.
- **Coordinates** (`lat`, `lng`).
- **Legal identifiers** — SIRET, SIREN, legal form, workforce.

Enrichment runs with bounded parallelism (see `concurrency` below) and adds a little extra time per business. A business whose details cannot be read right now is still exported with its search-card data.

#### ⭐ Reviews

| Field | Type | Default | Notes |
|---|---|---|---|
| `enableReviews` | boolean | `false` | One row per review; requires Listing Details. |
| `maxReviewsPerBusiness` | integer | `10` | Cap per business (1–20). |

#### ⚙️ Output & limits

| Field | Type | Default | Notes |
|---|---|---|---|
| `maxItems` | integer | `10000` | Global dataset-row cap across all features. The run also respects your max total charge for the run. |
| `concurrency` | integer | `3` | Parallel detail lookups (1–8). Higher is faster; 3 balances speed and reliability. |
| `delayBetweenRequestsMs` | integer | `250` | Polite pacing between page fetches (0–5000 ms). |
| `webhookUrl` | string | `""` | Optional real-time POST per record. |
| `webhookFormat` | select | `json` | `json` (full record) or `slack` (message payload). |
| `proxyConfiguration` | object | Residential FR | Apify residential proxy, France by default. Change the country only if you know why. |

### Output field reference

One dataset row per business (search), per listing, per review, or per scraped URL — identified by `featureType`:

| `featureType` | Meaning |
|---|---|
| `search` | A business found by activity + location search. |
| `listing_details` | A full profile record for a requested id/URL. |
| `reviews` | One public review (linked to its business via `proId`). |
| `scrape_by_url` | A record collected directly from a pasted URL. |

Key fields on business rows:

| Field | Description |
|---|---|
| `name`, `proId`, `profileUrl` | Identity: business name, listing id, canonical profile link. |
| `phone`, `website`, `emails` | Contact channels where publicly listed. |
| `socials` | Facebook / Instagram / LinkedIn / X / YouTube / TikTok / Pinterest profiles. |
| `street`, `city`, `zipCode`, `location` | Address; `location` is joined, ready for CRM import. |
| `geo` | `{ lat, lng }` coordinates. |
| `rating`, `ratingSource`, `reviewCount` | Rating value, whether it comes from PagesJaunes or Google, and review count. |
| `categories` | Business categories (e.g. `Plumber`, `Electrician`). |
| `openingHours` | Structured weekly hours: `[{ days: [...], opens, closes }]`. |
| `paymentAccepted` | Payment methods when listed. |
| `siret`, `siren`, `legalForm`, `workforce` | French legal identifiers from the listing's legal block. |
| `searchActivity`, `searchLocation`, `searchTaskLabel`, `position` | Which task found this row and where it ranked. |
| `leadDetails`, `detailsFetched`, `hasWebsite` | Enrichment flags for this row. |
| `scrapedAt` | ISO timestamp. |

Review rows add `businessName`, `reviewIndex`, `rating`, `date`, `author`, `text`. Listing-details and scrape rows carry the same business fields; scrape rows also include a `data` object with extras.

### Webhook guide

Set `webhookUrl` and every record is POSTed as JSON the moment it is collected — in addition to the dataset, never instead of it.

- **JSON format** — the full record object, identical to the dataset row.
- **Slack format** — a formatted message with name, rating, location, phone, website, and a link to the listing.

Works with Slack incoming webhooks, Discord, Zapier, Make, n8n, or your own URL. Delivery is best-effort: a failing webhook never interrupts the run or the dataset writes.

### Using with AI agents (MCP)

The Actor pairs naturally with Apify's MCP server, so an AI agent can answer questions with live directory data. Example questions an agent can now answer:

> "Find the 10 best-reviewed plumbers in Paris 15th with a website and a phone number."

> "How many dentists practice in Bordeaux and what share has a rating above 4.5?"

> "List pizzerias in Lyon with their opening hours and Google ratings for a competitor scan."

### FAQ

**Do I need my own proxies?**
No. The Actor runs on Apify residential proxies (France) by default, configured in the Connection section.

**How current is the data?**
Every run collects live from the directory at run time — what you get reflects the listings as they exist when you press Start.

**What happens on the free plan?**
Free-plan runs export a small sample (2 records) so you can verify output quality before upgrading. Paid plans are uncapped.

**Why do some rows have no phone or website?**
The directory only shows what businesses publish. Enrichment adds what is publicly available and always exports the row regardless — nothing is filtered out.

**Why do some enriched rows have no email?**
Many French small businesses simply do not publish an email. When the listing has a website, the Actor also checks that site's contact pages — but if neither publishes one, there is nothing to collect.

**How fast is a run?**
A 10-row run with search on and lead details off finishes in well under a minute. Lead details add a little extra time per business (parallel detail lookups plus a check of the business's own website when it has one); `concurrency` trades speed against reliability.

**Can I collect the same business twice?**
No — rows are deduplicated by listing id across pages and tasks, so a business re-ranked onto a later page never produces a duplicate row.

**Why did my run stop before reaching the requested number of results?**
Two reasons are possible, and the log always names the one that applies:

- **Charge limit reached** — the run's maximum cost was hit. Every row that was paid for is in the dataset. Raise the run's maximum cost (or the Actor's minimum, if you published this yourself) to collect more in one run, or split the work across several runs.
- **Results could not be saved** — the dataset could not be written to after repeated attempts. Nothing is kept without being stored, so the run ends with a failure message and the storage error is reported in the log and in the `OUTPUT` summary (`errors`). Retry the run; raise the run timeout if the job was cut short by it.

A long run also needs a long enough timeout: check the run's timeout when starting large jobs (the Actor requests a long default, but a run started with a short timeout is cut off by the platform).

### Delivery

The default dataset holds every record; the per-run summary (totals, feature flags, charge-limit status, save-failure status, paywall status) is stored in the key-value store under `OUTPUT`.

# Actor input Schema

## `enableSearch` (type: `boolean`):

Search the PagesJaunes directory by activity and location. Enabled by default.

## `searchTasks` (type: `array`):

Primary input. Add one row per search: pair an activity (e.g. "plombier", "restaurant", "dentiste") with a French location (e.g. "Paris 75015", "Lyon", "Bordeaux 33000").

## `maxResultsPerTask` (type: `integer`):

Default maximum businesses per search task. Override per task in the list above.

## `maxPagesPerTask` (type: `integer`):

Optional safety cap on pages per task (20 businesses per page). Leave empty to keep going until this task's businesses are all collected or the directory's last page is reached.

## `enableListingDetails` (type: `boolean`):

Fetch full profile records for specific business ids or URLs.

## `listingIds` (type: `array`):

Numeric listing ids (one per line). Find them in any profile URL after /pros/.

## `listingUrls` (type: `array`):

Full PagesJaunes profile URLs (one per line).

## `enableScrapeByUrl` (type: `boolean`):

Collect directly from listing or profile URLs instead of searching.

## `scrapeUrls` (type: `array`):

PagesJaunes listing or profile URLs (one per line).

## `enableLeadDetails` (type: `boolean`):

Enrich every business with lead details (phone, website, emails, socials, opening hours, geo, SIRET/SIREN, legal form) when publicly available. Every business is still exported even when no details are found — adds a little extra time per business.

## `enableReviews` (type: `boolean`):

Collect public customer reviews for businesses requested via Listing Details.

## `maxReviewsPerBusiness` (type: `integer`):

Cap on reviews collected per business (1–20).

## `maxItems` (type: `integer`):

Global cap on total dataset rows across all features. Set high for large runs (e.g. 10000+). Note: the run also respects your max total charge for this run.

## `concurrency` (type: `integer`):

How many detail lookups to run in parallel (1–8). Higher is faster; 3 is a good default.

## `delayBetweenRequestsMs` (type: `integer`):

Polite pacing between directory fetches (0–5000 ms). Higher values are slower but gentler.

## `webhookUrl` (type: `string`):

Optional. Every record is always saved to the run's dataset — this webhook is an ADDITIONAL real-time push. When set, each new record is also POSTed to this URL (CRM, Slack incoming webhook, Zapier, Make, Google Sheets).

## `webhookFormat` (type: `string`):

json = full record object; slack = Slack-friendly message payload.

## `proxyConfiguration` (type: `object`):

Apify residential proxy (FR) is enabled by default for reliable collection.

## Actor input object example

```json
{
  "enableSearch": true,
  "searchTasks": [
    {
      "activity": "plombier",
      "location": "Paris 75015",
      "maxResults": 10
    },
    {
      "activity": "restaurant",
      "location": "Lyon 69002",
      "maxResults": 10
    }
  ],
  "maxResultsPerTask": 10,
  "enableListingDetails": false,
  "listingIds": [],
  "listingUrls": [],
  "enableScrapeByUrl": false,
  "scrapeUrls": [],
  "enableLeadDetails": true,
  "enableReviews": false,
  "maxReviewsPerBusiness": 10,
  "maxItems": 10000,
  "concurrency": 3,
  "delayBetweenRequestsMs": 250,
  "webhookUrl": "",
  "webhookFormat": "json",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "FR"
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset with every row from all enabled features in this run.

## `search` (type: `string`):

Businesses from activity + location search.

## `runSummary` (type: `string`):

Per-run metadata: totals, lead-details flag, spending-limit status, and the paywall object (detected, isPaying, pricingTier, limited, blocked, freeTierMaxItems).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "enableSearch": true,
    "searchTasks": [
        {
            "activity": "plombier",
            "location": "Paris 75015",
            "maxResults": 10
        },
        {
            "activity": "restaurant",
            "location": "Lyon 69002",
            "maxResults": 10
        }
    ],
    "maxResultsPerTask": 10,
    "enableLeadDetails": true,
    "maxReviewsPerBusiness": 10,
    "maxItems": 10000,
    "concurrency": 3,
    "delayBetweenRequestsMs": 250,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "FR"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("b2b_leads/pagesjaunes-real-time-data-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "enableSearch": True,
    "searchTasks": [
        {
            "activity": "plombier",
            "location": "Paris 75015",
            "maxResults": 10,
        },
        {
            "activity": "restaurant",
            "location": "Lyon 69002",
            "maxResults": 10,
        },
    ],
    "maxResultsPerTask": 10,
    "enableLeadDetails": True,
    "maxReviewsPerBusiness": 10,
    "maxItems": 10000,
    "concurrency": 3,
    "delayBetweenRequestsMs": 250,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "FR",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("b2b_leads/pagesjaunes-real-time-data-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "enableSearch": true,
  "searchTasks": [
    {
      "activity": "plombier",
      "location": "Paris 75015",
      "maxResults": 10
    },
    {
      "activity": "restaurant",
      "location": "Lyon 69002",
      "maxResults": 10
    }
  ],
  "maxResultsPerTask": 10,
  "enableLeadDetails": true,
  "maxReviewsPerBusiness": 10,
  "maxItems": 10000,
  "concurrency": 3,
  "delayBetweenRequestsMs": 250,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "FR"
  }
}' |
apify call b2b_leads/pagesjaunes-real-time-data-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,b2b_leads/pagesjaunes-real-time-data-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/UAv4q6e2IDRGF78mO/builds/C3E4hbU2CCMEfjMoh/openapi.json
