PagesJaunes Business Data — Leads & Contacts avatar

PagesJaunes Business Data — Leads & Contacts

Pricing

from $15.00 / 1,000 results

Go to Apify Store
PagesJaunes Business Data — Leads & Contacts

PagesJaunes Business Data — Leads & Contacts

Collect live PagesJaunes business data across France: search by activity and city, get phones, emails, websites, ratings, opening hours, and legal ids (SIRET/SIREN). Real-time streaming to your dataset with optional webhooks. Free plan exports 2 records; paid plans are unlimited.

Pricing

from $15.00 / 1,000 results

Rating

0.0

(0)

Developer

Emmanuel

Emmanuel

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

20 hours ago

Last modified

Categories

Share

PagesJaunes Real-Time Data

Live PagesJaunes business directory intelligence: professional search across every French city, full listing details, public reviews, and lead enrichment (phones, websites, emails, socials, legal identifiers). Checkbox features — enable only what you need. Structured JSON streamed to your dataset in real time.

Free plan notice: on the Apify Free plan this Actor exports a small sample of results per run (2 records). Upgrade to a paid Apify plan for full, unlimited exports.

Who is this for

Built for agencies, local-SEO teams, sales teams, and CRM owners who work with French businesses:

  • Lead generation — pull plumbers, dentists, lawyers, restaurants, garages… with phones, websites, and emails ready for outreach.
  • Local SEO & citation audits — track which businesses in a city carry ratings, review counts, and working websites.
  • Market & competitor research — size any trade in any postcode: how many players, how well reviewed, who has a site.
  • CRM & enrichment pipelines — stream records straight into Sheets, Airtable, HubSpot, Pipedrive via webhooks.
  • Franchise & network monitoring — watch listings by city for presence, ratings, and contact accuracy.
  • AI agents & MCP workflows — feed live directory answers into your agent (see below).

Feature matrix

FeatureWhat it doesDefault
🔎 Professional & business searchMulti-task search: pair an activity ("plombier", "restaurant", "avocat"…) with any French location ("Paris 75015", "Lyon", "Bordeaux 33000"). Each task collects its businesses across result pages — every business on every page is a row.On
📋 Listing DetailsTurn numeric listing ids or profile URLs into full records: phone, opening hours, payment methods, coordinates, legal identifiers.Off
🔗 Scrape By URLPaste any PagesJaunes listing URL (collects every business on that page) or profile URL (one full record).Off
🎯 Lead detailsEnriches every row in place: phone, website, emails (including from the business's own website), socials, opening hours, geo, SIRET/SIREN, legal form. Enrich, never filter — every business is still exported.On
⭐ ReviewsPublic customer reviews for businesses requested via Listing Details — one row per review.Off
🌐 WebhooksJSON or Slack. Every record is POSTed to your URL the moment it is collected — CRM, Zapier, Make, n8n.Off

Features combine freely: search + lead details is the classic lead-gen run; Listing Details + Reviews is the deep-dive run; Scrape By URL re-collects a page you already know.

Input reference

FieldTypeDefaultNotes
enableSearchbooleantrueMaster switch for search.
searchTasksarray2 demo rowsOne row per search: activity, location, optional maxResults.
maxResultsPerTaskinteger10Global default cap per task; override per task with maxResults.
maxPagesPerTaskintegeremptyOptional page-depth cap (20 businesses per page). Empty = keep going until this task's businesses are all collected or the directory's last page.

Example searchTasks:

[
{ "activity": "plombier", "location": "Paris 75015", "maxResults": 10 },
{ "activity": "restaurant", "location": "Lyon 69002", "maxResults": 10 },
{ "activity": "dentiste", "location": "Marseille 13008", "maxResults": 10 }
]

Activities work best in French ("plombier", "coiffeur", "garage automobile", "avocat", "kine"). Locations accept city, postcode, or arrondissement forms — adding the postcode sharpens arrondissement-level searches.

📋 Listing details

FieldTypeDefaultNotes
enableListingDetailsbooleanfalseMaster switch.
listingIdsstring[][]Numeric ids, e.g. 64391412 — the number after /pros/ in any profile URL.
listingUrlsstring[][]Full profile URLs.

Every requested id/URL produces exactly one row (found or not), so counts stay predictable.

🔗 Scrape By URL

FieldTypeDefaultNotes
enableScrapeByUrlbooleanfalseMaster switch.
scrapeUrlsstring[][]Listing page URLs (one row per business card) or profile URLs (one full record).

Only pagesjaunes.fr URLs are accepted; anything else fails validation with a clear error.

🎯 Lead details

FieldTypeDefaultNotes
enableLeadDetailsbooleantrueEnrich every business row in place. Never filters — output count stays predictable.

What enrichment adds, where publicly available:

  • Phone and website from the full listing.
  • Emails — from the listing and from the business's own website (homepage plus its contact/legal pages).
  • Socials — Facebook, Instagram, LinkedIn, X, YouTube, TikTok, Pinterest.
  • Opening hours (structured per day), payment methods.
  • Coordinates (lat, lng).
  • Legal identifiers — SIRET, SIREN, legal form, workforce.

Enrichment runs with bounded parallelism (see concurrency below) and adds a little extra time per business. A business whose details cannot be read right now is still exported with its search-card data.

⭐ Reviews

FieldTypeDefaultNotes
enableReviewsbooleanfalseOne row per review; requires Listing Details.
maxReviewsPerBusinessinteger10Cap per business (1–20).

⚙️ Output & limits

FieldTypeDefaultNotes
maxItemsinteger10000Global dataset-row cap across all features. The run also respects your max total charge for the run.
concurrencyinteger3Parallel detail lookups (1–8). Higher is faster; 3 balances speed and reliability.
delayBetweenRequestsMsinteger250Polite pacing between page fetches (0–5000 ms).
webhookUrlstring""Optional real-time POST per record.
webhookFormatselectjsonjson (full record) or slack (message payload).
proxyConfigurationobjectResidential FRApify residential proxy, France by default. Change the country only if you know why.

Output field reference

One dataset row per business (search), per listing, per review, or per scraped URL — identified by featureType:

featureTypeMeaning
searchA business found by activity + location search.
listing_detailsA full profile record for a requested id/URL.
reviewsOne public review (linked to its business via proId).
scrape_by_urlA record collected directly from a pasted URL.

Key fields on business rows:

FieldDescription
name, proId, profileUrlIdentity: business name, listing id, canonical profile link.
phone, website, emailsContact channels where publicly listed.
socialsFacebook / Instagram / LinkedIn / X / YouTube / TikTok / Pinterest profiles.
street, city, zipCode, locationAddress; location is joined, ready for CRM import.
geo{ lat, lng } coordinates.
rating, ratingSource, reviewCountRating value, whether it comes from PagesJaunes or Google, and review count.
categoriesBusiness categories (e.g. Plumber, Electrician).
openingHoursStructured weekly hours: [{ days: [...], opens, closes }].
paymentAcceptedPayment methods when listed.
siret, siren, legalForm, workforceFrench legal identifiers from the listing's legal block.
searchActivity, searchLocation, searchTaskLabel, positionWhich task found this row and where it ranked.
leadDetails, detailsFetched, hasWebsiteEnrichment flags for this row.
scrapedAtISO timestamp.

Review rows add businessName, reviewIndex, rating, date, author, text. Listing-details and scrape rows carry the same business fields; scrape rows also include a data object with extras.

Webhook guide

Set webhookUrl and every record is POSTed as JSON the moment it is collected — in addition to the dataset, never instead of it.

  • JSON format — the full record object, identical to the dataset row.
  • Slack format — a formatted message with name, rating, location, phone, website, and a link to the listing.

Works with Slack incoming webhooks, Discord, Zapier, Make, n8n, or your own URL. Delivery is best-effort: a failing webhook never interrupts the run or the dataset writes.

Using with AI agents (MCP)

The Actor pairs naturally with Apify's MCP server, so an AI agent can answer questions with live directory data. Example questions an agent can now answer:

"Find the 10 best-reviewed plumbers in Paris 15th with a website and a phone number."

"How many dentists practice in Bordeaux and what share has a rating above 4.5?"

"List pizzerias in Lyon with their opening hours and Google ratings for a competitor scan."

FAQ

Do I need my own proxies? No. The Actor runs on Apify residential proxies (France) by default, configured in the Connection section.

How current is the data? Every run collects live from the directory at run time — what you get reflects the listings as they exist when you press Start.

What happens on the free plan? Free-plan runs export a small sample (2 records) so you can verify output quality before upgrading. Paid plans are uncapped.

Why do some rows have no phone or website? The directory only shows what businesses publish. Enrichment adds what is publicly available and always exports the row regardless — nothing is filtered out.

Why do some enriched rows have no email? Many French small businesses simply do not publish an email. When the listing has a website, the Actor also checks that site's contact pages — but if neither publishes one, there is nothing to collect.

How fast is a run? A 10-row run with search on and lead details off finishes in well under a minute. Lead details add a little extra time per business (parallel detail lookups plus a check of the business's own website when it has one); concurrency trades speed against reliability.

Can I collect the same business twice? No — rows are deduplicated by listing id across pages and tasks, so a business re-ranked onto a later page never produces a duplicate row.

Why did my run stop before reaching the requested number of results? Two reasons are possible, and the log always names the one that applies:

  • Charge limit reached — the run's maximum cost was hit. Every row that was paid for is in the dataset. Raise the run's maximum cost (or the Actor's minimum, if you published this yourself) to collect more in one run, or split the work across several runs.
  • Results could not be saved — the dataset could not be written to after repeated attempts. Nothing is kept without being stored, so the run ends with a failure message and the storage error is reported in the log and in the OUTPUT summary (errors). Retry the run; raise the run timeout if the job was cut short by it.

A long run also needs a long enough timeout: check the run's timeout when starting large jobs (the Actor requests a long default, but a run started with a short timeout is cut off by the platform).

Delivery

The default dataset holds every record; the per-run summary (totals, feature flags, charge-limit status, save-failure status, paywall status) is stored in the key-value store under OUTPUT.