# Pagine Gialle Businesses Scraper (`nice_dev/paginegialle-businesses-scraper`) Actor

Scrape Italian business leads from PagineGialle.it: many activities × places per run, phones (mobile, WhatsApp), e-mail (+ company websites), website, VAT, address, GPS, hours, rating. New-businesses-only monitoring, free lead filters, auto split of searches capped at 200.

- **URL**: https://apify.com/nice\_dev/paginegialle-businesses-scraper.md
- **Developed by:** [Nice Dev](https://apify.com/nice_dev) (community)
- **Categories:** Lead generation, MCP servers
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.77 / 1,000 businesses

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### 🏢 What is Pagine Gialle Businesses Scraper?

**Pagine Gialle Businesses Scraper** extracts **business leads from [PagineGialle.it](https://www.paginegialle.it)**, the Italian Yellow Pages: **company name, phone numbers (mobile and WhatsApp flagged), e-mail, website, VAT number (P.IVA), full address with postal code, province and region, GPS coordinates, opening hours, rating and description** for any activity in any Italian place.

Type one or several **activities** (`ristoranti`, `idraulici`, `avvocati`) and one or several **places** (`milano`, `20159`, `MI`, `lombardia`), click **Start**, and download the businesses in JSON, CSV or Excel. No login, nothing to configure: about **1,000 businesses in 2 minutes**, e-mail and GPS included, without opening every business page. It can also visit each **company website to find more e-mail addresses**, and return **only the businesses added since your last run**.

### 📋 What data can you extract from Pagine Gialle?

One item per business, 73 fields:

| Category | What you get |
| --- | --- |
| 🏷️ **Business** | name, shop sign, registered name, the site's id and link, paid / premium / free listing, head office of a branch — `Ristorante La Pobbia 1850` |
| 🗂️ **Category** | the activity shown, its code, every category the business is filed under — `Ristoranti e trattorie` |
| 📍 **Address and map** | full address, street, postal code, city, province, region, GPS coordinates, distance from the place searched — `Via Gallarate, 92, 20151 Milano (MI)` |
| 📞 **Phones** | every phone number, also in international format (`+39…`), mobile and WhatsApp numbers flagged, fax — `347 3765549` |
| 📧 **E-mail and web** | e-mail addresses (and those found on the company website), website, Facebook, Instagram and other social links — `lapobbia@lapobbia.com` |
| 🧾 **Company data** | VAT number (P.IVA), tax code (codice fiscale) — `10122980963` |
| 📝 **Description** | short and full description, highlights, services, payment methods |
| 🕒 **Opening hours** | hours day by day, the same in one line for spreadsheets, open 24/7, open at the time of the scrape |
| ⭐ **Rating** | average rating, number of reviews, the reviews themselves — `4.1` from 64 reviews |
| 🖼️ **Photos and links** | logo, photos, number of branches, TheFork booking, menu and booking links of restaurants |
| ✅ **Quick filters** | yes / no columns (has a phone, a mobile, an e-mail, a website, WhatsApp, social pages) and the list of ways to reach the business |
| 🔎 **Search** | the search the business came from, the sub-area actually read, how the site understood the place, results page and rank, time of the scrape |

Every field, with an example, is listed in the **Output** section below.

With **Extract details** on, one more request per business adds `taxCode` (codice fiscale), `services`, `paymentMethods`, `reviews` (up to 20), the social links of the page, `menuUrl`, `bookingUrl` and all photos (and, on category pages, `legalName` and `categoryCode`, which those pages do not print).

### ✅ Why use Pagine Gialle Businesses Scraper?

- 🚀 **Fast**: about 1,000 businesses in 2 minutes — e-mail, website, VAT number, GPS, opening hours and rating come with the search results, so 100 leads take about 15 seconds.
- 🗂️ **Several searches in one run**: activities × places (3 × 4 = 12 searches), with a **cap per search** so that one big city cannot eat the whole budget.
- 🧩 **No 200-result wall**: Pagine Gialle never shows more than 200 results per search. When a search is capped, the Actor automatically re-runs it for every sub-area the site lists (region → provinces → municipalities → districts) and deduplicates.
- 🔔 **Monitoring built in**: tick **New businesses only**, schedule the Actor, and each run returns (and charges) only businesses it has never delivered before.
- 📧 **More e-mails**: optionally visits each company website (contact / legal pages) and adds the addresses published there.
- 🎯 **Lead filters**: with a phone, e-mail or website, minimum rating, open now, paid or free listings, category codes, words in the name, city or category — filtered businesses are not saved; a business you keep costs its normal price, each business a filter drops is charged a check fee (see pricing). **Open now** is sent to Pagine Gialle itself, so the Actor reads only the businesses open at that moment (measured: 53 plumbers of a town → 4 open, one page instead of three).
- 🧭 **Honest place matching**: if the site does not recognise a place it silently answers with nationwide results; the Actor detects it (`searchLocationMatch`) and skips such searches instead of selling you unrelated leads.
- 📞 **Ready for dialers and CRMs**: phones in E.164 (`+39…`), mobile and WhatsApp numbers flagged, `contactChannels`, flat opening hours for CSV, optional removal of empty fields.
- 💾 **Cheap**: $0.80 per 1,000 businesses (e-mail, website, GPS, hours included), and the proxy is included in the price.
- 🔌 API, scheduling, monitoring, integrations (Make, Zapier, n8n, Google Sheets…) and JSON/CSV/Excel export via the Apify platform.

### 🚀 How to scrape Pagine Gialle

1. Create a free Apify account.
2. Open **Pagine Gialle Businesses Scraper**, type an activity in **What** (in Italian, as on the site: `ristoranti`, `dentisti`, `agenzie immobiliari`) and a place in **Where** (`roma`, `Milano Quartiere Isola`, `00122`, `TO`, `provincia di torino`, `toscana`) — add more in **More activities** and **More places**.
3. Or paste Pagine Gialle URLs in **Start URLs**: a search, a category page (`https://www.paginegialle.it/lombardia/milano/ristoranti.html`) or single business pages.
4. Set **Max businesses** (100 by default, 0 = no limit) and **Max businesses per search**, optionally the **Filters**, **Extract details**, **Find e-mails on company websites** or **New businesses only**, and click **Start**.
5. Download the dataset in JSON, CSV, Excel or via API.

### 💰 How much does it cost to scrape Pagine Gialle?

This Actor uses **pay per event** pricing: **$0.80 per 1,000 businesses** — plus **$0.002 per run start** at the Actor's default memory (20 cents per 100 runs; the platform counts that event once per gigabyte, so a run you give more memory pays proportionally more). Higher Apify plans pay less per business: $0.79 (Bronze), $0.78 (Silver) and $0.77 (Gold and above) per 1,000. **Extract details** adds **$0.70 per 1,000 businesses read from their own page** ($1.50 per 1,000 in all); a pasted business URL is always read from its page, so it counts as one too. A business whose page is gone is saved from the results page and not charged for it. **Find e-mails on company websites** adds **$2.00 per 1,000 businesses whose website gave a new e-mail address** (businesses whose website gives nothing new are not charged for it). Businesses skipped by **New businesses only** are never charged. The **Filters** are applied by the Actor on every business it reads: a business you keep costs its normal price, a business a filter drops costs **$0.90 per 1,000 businesses dropped** on the results page by any of them (*Only businesses with a phone*, *an e-mail* or *a website*, *Minimum rating*, *Only businesses open right now*, *Listing types*, *Category codes*, *Category contains*, *City contains*, *Name contains*, *Exclude names containing*), and **$0.90 per 1,000 business pages or websites read and dropped** by a filter the results page does not give (*Only businesses with a website* or *Category codes* on category-page results, any filter on a pasted business URL, *Only businesses with an e-mail* judged on the website visit of **Find e-mails on company websites**) — both $0.88 on Bronze, $0.86 on Silver and $0.84 on Gold and above.
Example: 5,000 businesses ≈ $4.00 ($7.50 with details); a weekly refresh of new plumbers in Roma costs only the new ones. Platform usage (compute, proxy) is included in the price.

### ⚙️ Input

```json
{
    "what": "ristoranti",
    "searchQueries": ["pizzerie"],
    "where": "milano",
    "locations": ["roma"],
    "maxItems": 1000,
    "maxItemsPerQuery": 250,
    "extractDetails": false,
    "enrichEmails": false,
    "requirePhone": true,
    "minRating": 4
}
```

Or with your own URLs, only the new businesses since the last run:

```json
{
    "startUrls": [
        { "url": "https://www.paginegialle.it/lombardia/milano/dentisti.html" },
        { "url": "https://www.paginegialle.it/ricerca/avvocati/torino" },
        { "url": "https://www.paginegialle.it/ristoranteantico1850" }
    ],
    "maxItems": 0,
    "onlyNew": true,
    "stateKey": "weekly-leads"
}
```

| Field                                                                    | Type                                    | Default                                     | Notes                                                                                                       |
| ------------------------------------------------------------------------ | --------------------------------------- | ------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| `what` / `searchQueries`                                                 | string / string\[]                       | `ristoranti` / `[]`                         | Activities, in Italian; crossed with every place                                                            |
| `where` / `locations`                                                    | string / string\[]                       | `milano` / `[]`                             | Municipality, `City Quartiere X`, postal code, province code, `provincia di X` or region; empty = all Italy |
| `searches`                                                               | string\[]                                | `[]`                                        | `what \| where` pairs that must not be crossed with the lists                                               |
| `startUrls`                                                              | array                                   | `[]`                                        | Pagine Gialle search, category or business page URLs                                                        |
| `maxItems`                                                               | integer                                 | `100`                                       | 0 = unlimited                                                                                               |
| `maxItemsPerQuery`                                                       | integer                                 | `0`                                         | Cap of each search, its sub-areas included; 0 = none                                                        |
| `extractDetails`                                                         | boolean                                 | `false`                                     | Visit each business page (+1 request per business)                                                          |
| `enrichEmails`                                                           | boolean                                 | `false`                                     | Visit each company website for e-mail addresses (separate event)                                            |
| `expandLocations`                                                        | boolean                                 | `true`                                      | Split capped searches by sub-area                                                                           |
| `sortBy`                                                                 | `relevance` / `distance` / `popularity` | `relevance`                                 | Site order (matters for capped searches); `distance` also fills the `distance` field                        |
| `requirePhone` / `requireEmail` / `requireWebsite`                       | boolean                                 | `false`                                     | Lead filters                                                                                                |
| `minRating`                                                              | number                                  | `0`                                         | 1-5; `1` = rated businesses only                                                                            |
| `openNowOnly`                                                            | boolean                                 | `false`                                     | Asks the site for its "Aperto ora" results, then re-checks the hours (Italian time)                         |
| `listingTypes`                                                           | `paid` / `premium` / `free` list        | `[]`                                        | Keep these listing types only                                                                               |
| `categoryCodes`                                                          | string\[]                                | `[]`                                        | Pagine Gialle category codes (`007585100`)                                                                  |
| `categoryContains` / `cityContains` / `nameContains` / `excludeKeywords` | string\[]                                | `[]`                                        | Words, case and accents ignored                                                                             |
| `onlyNew` / `stateKey` / `resetState`                                    | boolean / string / boolean              | `false` / `default` / `false`               | Monitoring memory                                                                                           |
| `excludeEmptyFields`                                                     | boolean                                 | `false`                                     | Drop null and empty values from the items                                                                   |
| `proxyConfiguration`                                                     | object                                  | Apify Proxy (included in the price)         | Advanced                                                                                                    |
| `maxConcurrency` / `maxRequestsPerMinute` / `maxRequestRetries`          | integer                                 | `4` / `120` / `5`                           | Advanced                                                                                                    |
| `skipUnrecognizedLocation`                                               | boolean                                 | `true`                                      | Skip searches whose place the site did not recognise                                                        |
| `debugLog`                                                               | boolean                                 | `false`                                     | Verbose log                                                                                                 |

### 📦 Output

One business, shortened (every column is listed below):

```json
{
    "id": "7859A46A-D3E7-81C8-E050-020A3E743887",
    "url": "https://www.paginegialle.it/ristoranteantico1850",
    "name": "Ristorante La Pobbia 1850",
    "listingType": "paid",
    "schemaType": "Restaurant",
    "category": "Ristoranti e trattorie",
    "categoryCode": "027585100",
    "categoryCodes": ["007585100", "007502600"],
    "address": "Via Gallarate, 92, 20151 Milano (MI)",
    "postalCode": "20151",
    "city": "Milano",
    "province": "MI",
    "region": "Lombardia",
    "country": "IT",
    "latitude": 45.4961,
    "longitude": 9.13435,
    "distance": 1194,
    "phone": "347 3765549",
    "phones": ["347 3765549", "02 38006641"],
    "phonesE164": ["+393473765549", "+390238006641"],
    "mobilePhones": ["347 3765549"],
    "whatsappPhones": ["347 3765549"],
    "email": "lapobbia@lapobbia.com",
    "emails": ["lapobbia@lapobbia.com"],
    "websiteEmails": [],
    "website": "http://lapobbia.com",
    "websiteDomain": "lapobbia.com",
    "socialLinks": ["https://www.instagram.com/lapobbia1850/?hl=it"],
    "instagramUrl": "https://www.instagram.com/lapobbia1850/?hl=it",
    "vatNumber": "10122980963",
    "openingHours": { "mon": ["12:30 - 14:30"], "tue": ["12:30 - 14:30", "19:30 - 22:00"] },
    "openingHoursText": "mon 12:30-14:30; tue 12:30-14:30, 19:30-22:00",
    "isOpenNow": false,
    "ratingAverage": 4.1,
    "ratingCount": 64,
    "theForkUrl": "https://module.thefork.com/it_IT/module/840549-a0e5e/83153-798",
    "menuUrl": "https://www.instagram.com/lapobbia1850/?hl=it",
    "bookingUrl": "https://lapobbia.com/prenota/",
    "hasMobilePhone": true,
    "contactChannels": ["phone", "mobile", "whatsapp", "email", "website", "social"],
    "searchWhat": "ristoranti",
    "searchWhere": "Milano Quartiere Certosa",
    "searchLocation": "milano",
    "searchLocationMatch": "nearby",
    "position": 7,
    "scrapedAt": "2026-09-17T09:30:00.000Z"
}
```

You can download the dataset in JSON, CSV or Excel, or fetch it via API.

#### All 73 fields

| Fields | What you get |
| --- | --- |
| `id`, `parentId`, `url`, `name`, `brand`, `legalName`, `listingType`, `schemaType` | **Business** — `7859A46A-…`, id of the head office for a branch, `https://www.paginegialle.it/ristoranteantico1850`, `Ristorante La Pobbia 1850`, the shop sign when the listing prints one (`null` here), `Ristorante La Pobbia 1850`, `paid`, `Restaurant` |
| `category`, `categoryCode`, `categoryCodes`, `categories` | **Category** — `Ristoranti e trattorie`, `027585100` (the code of that category), `["007585100", "007502600"]` (every category the business is filed under), `["RISTORANTI"]` |
| `address`, `street`, `postalCode`, `city`, `province`, `region`, `country` | **Address** — `Via Gallarate, 92, 20151 Milano (MI)`, `Via Gallarate, 92`, `20151`, `Milano`, `MI`, `Lombardia`, `IT` |
| `latitude`, `longitude`, `distance` | **Map** — `45.4961`, `9.13435`, `1194` (metres from the centre of the place searched, with **Sort by: distance**) |
| `phone`, `phones`, `phonesE164`, `mobilePhones`, `whatsappPhones`, `fax` | **Phones** — `347 3765549`, `["347 3765549", "02 38006641"]`, `["+393473765549", …]`, `["347 3765549"]`, `["347 3765549"]` |
| `email`, `emails`, `websiteEmails`, `emailSource` | **E-mail** — `lapobbia@lapobbia.com`, all addresses, those found on the company website, the page they were found on |
| `website`, `websiteDomain`, `links`, `socialLinks`, `facebookUrl`, `instagramUrl` | **Web** — `https://www.lapobbia.com`, `lapobbia.com`, other published links, Facebook / Instagram / LinkedIn… |
| `vatNumber`, `taxCode` | **Company data** — `10122980963`, `NZGRRT68T20F205L` |
| `shortDescription`, `description`, `highlights`, `services`, `paymentMethods` | **Description** — texts of the listing, `["Tavoli all'aperto"]`, `["cucina milanese"]`, `["carte di credito"]` |
| `openingHours`, `openingHoursText`, `is24h`, `isOpenNow` | **Opening hours** — `{ "mon": ["12:30 - 14:30"] }`, `mon 12:30-14:30; tue 12:30-14:30, 19:30-22:00`, `false`, `true` |
| `ratingAverage`, `ratingCount`, `reviews` | **Rating** — `4.1`, `64`, up to 20 reviews |
| `logo`, `images`, `branchCount`, `theForkUrl`, `menuUrl`, `bookingUrl` | **Photos and links** — image URLs, number of branches, TheFork booking widget, menu and booking links of restaurants |
| `hasPhone`, `hasMobilePhone`, `hasEmail`, `hasWebsite`, `hasWhatsapp`, `hasSocial`, `contactChannels` | **Quick filters** — booleans and `["phone", "mobile", "email", "website"]` for quick filtering |
| `searchWhat`, `searchWhere`, `searchLocation`, `searchLocationMatch`, `searchUrl`, `page`, `position` | **Search** — the search the business came from, the sub-area actually read, how the site resolved the place (`exact`, `postalCode`, `nearby`, `notFound`), rank in the search |
| `scrapedAt` | ISO timestamp |

### 💡 Tips

#### How to get more results

A city search is capped at 200 by the site; keep **Expand locations** on for large cities, or search by postal code / district yourself. Set **Max businesses** to `0` for no limit.

A search skipped with the warning "place not recognised" means the site did not know the `where` value; check the spelling (Italian names: `Firenze`, not `Florence`) or use the postal code.

#### How to reduce costs

The price is per business, so the levers are **Max businesses**, **Max businesses per search**, the filters (a filtered-out business is not saved, only its check fee is charged; a business you keep pays no check fee) and **New businesses only** for recurring runs (you never pay twice for the same business). The default mode already returns e-mail, website, GPS and hours. Turn **Extract details** on only when you need the tax code, reviews, services, social links or menu / booking links. It makes runs slower and adds $0.70 per 1,000 businesses.

#### Open now, distance and free listings

**Open now** asks Pagine Gialle for its own "Aperto ora" list — far fewer pages, and it also brings the 24/7 emergency listings the normal pages hide. **Sort by: distance** fills `distance` (metres from the centre of the place you searched), so you can sort your leads by proximity; without that sort the field stays empty.

Free listings (`listingType: free`) usually carry a phone, an address and a VAT number but no e-mail or website: use the filters when you need contactable leads, or **Find e-mails on company websites**.

#### Several searches in one run

Fill **More activities** (`searchQueries`) and / or **More places** (`locations`): the Actor runs one search per activity × place (3 activities × 4 places = 12 searches, up to 500 per run). `what` and `where` still work and are added to the lists; `searches` lines (`what | where`) are added as they are. A business found by several searches is saved — and charged — once. Set **Max businesses per search** (`maxItemsPerQuery`) to give every search its own cap (the automatic sub-areas of a search count for it): without it the first searches can use up the whole `maxItems` budget. Pasted search and category URLs are searches of their own, with the same cap.

#### Monitoring: only the new businesses

Tick **New businesses only** (`onlyNew`) and schedule the Actor. The first run returns everything; each later run skips the businesses already delivered: they are not saved, not charged, and their business page is not even opened. The memory lives in a named key-value store of your account (`paginegialle-businesses-scraper-seen`, up to 150,000 businesses per key) and is only updated with businesses that really reached the dataset, so a failed run never hides anything. Give each schedule its own **Memory key** (`stateKey`), and tick **Reset the memory** once to start over. Pagine Gialle has no "newest first" order: with a low **Max businesses**, the next run simply returns the next businesses you do not have yet.

#### Find e-mails on company websites

With `enrichEmails` on, the Actor reads each company's own website (home page, then up to 2 contact / legal notice pages linked from it) and adds the business e-mail addresses published there to `websiteEmails` and `emails` (`emailSource` = the page). It works on every business that has a website: in the default mode (search results) and with **Extract details**; category pages do not show the website, so combine them with **Extract details**. Facebook / Instagram pages given as a website are not visited. With **Only businesses with an e-mail**, a business is kept when its website gives one. The websites are read directly by the Actor, not through the proxy, at most 3 pages and 20 seconds per site: a site that does not answer, refuses the visit or shows an anti-bot page simply gives no address, and the run summary counts these cases (`2/40 websites gave an e-mail (others: 20 HTTP 403, 10 timeout, 7 no address on the pages read)`).

### 🔌 Integrations and API

Call the Actor via the Apify API, the JavaScript or Python clients, or connect it with integrations and webhooks (Make, Zapier, n8n, Google Sheets, Slack, Airtable…). Schedule it with **New businesses only** to feed a CRM every week. The dataset can be fetched as JSON, CSV or Excel from any tool.

### 🤖 Use with AI agents (MCP)

AI agents (Claude, ChatGPT, Cursor…) can find and run this Actor through the [Apify MCP server](https://mcp.apify.com), billed to their Apify account like any run. It returns one item per business listed on Pagine Gialle. Actor id: `nice_dev/paginegialle-businesses-scraper`; MCP server with this Actor only: `https://mcp.apify.com/?tools=fetch-actor-details,nice_dev/paginegialle-businesses-scraper`.

Smallest input, for a cheap first call:

```json
{
    "what": "ristoranti",
    "where": "milano",
    "maxItems": 10
}
```

Key output fields: `url`, `name`, `category`, `phone`, `email`, `website`, `address`, `city`.

Cost: $0.80 per 1,000 businesses plus $0.002 per run start at the default memory ($0.77 per 1,000 on the Gold plan and above); **Extract details**, **Find e-mails on company websites** and the filters cost extra, see the pricing section above. Cap each call with `maxItems` and, through the API, with the run option `maxTotalChargeUsd`.

### ❓ FAQ

#### Is it legal to scrape Pagine Gialle?

The Actor only reads what Pagine Gialle shows publicly to any anonymous visitor (company name, professional phone numbers and e-mail addresses, VAT number, address, opening hours). It logs in to nothing and bypasses no access control or captcha. Results can contain personal data — a sole trader's listing carries their name — which is protected by GDPR: do not store it without a legitimate reason. You are responsible for using the data in compliance with Pagine Gialle's Terms of Use and applicable law. This Actor is not affiliated with Pagine Gialle or Italiaonline.

**Where do the website e-mail addresses come from?** With `enrichEmails` on, the Actor reads the pages each company publishes on its own website (home page, contact page, legal notice) and returns the business contact addresses shown there — what any visitor sees. It does not log in, guess addresses, read contact forms or bypass any protection. Some of these addresses are personal data under GDPR (e.g. `firstname.lastname@company.com`): you are responsible for having a lawful basis before using them, for informing the people concerned and for honouring opt-outs. Rules for B2B e-mail prospecting differ by country (opt-out in some, prior consent in others).

#### Why did I get 0 results?

- The activity is not one the site knows: type it **in Italian, as on the site** (`idraulici`, not `plumbers`). Some English words happen to work (`restaurants` gives the restaurants), most do not: an unknown activity gives an empty results page.
- The place was not recognised (see the next question) and `skipUnrecognizedLocation` is on, so the search was skipped — the log says so.
- The filters removed every business (the run summary says "N filtered out"), or **New businesses only** skipped businesses already delivered ("already delivered by a previous run").
- A run that saves nothing **and** had requests failing for good ends as **failed**, so a site change or a block never looks like an empty search. Check the `FAILED_REQUESTS` record in the key-value store.

#### What does "place not recognised" mean?

When the site does not know the `where` value, it silently answers with results from all of Italy. The Actor detects this (`searchLocationMatch: notFound`) and, by default, skips the search with a warning instead of saving unrelated leads. Check the spelling (Italian names: `Firenze`, not `Florence`; `Milano`, not `Milan`), use the postal code, the province code (`MI`) or a district as listed on the site (`Milano Quartiere Isola`). Set `skipUnrecognizedLocation` to `false` to keep the nationwide results anyway.

#### Does it need a login or a proxy?

No login. The proxy is included in the price: leave the default setting (the residential proxy cannot be chosen in the input). A request the site turns away is retried at once on a new proxy session, up to 10 times on top of the retries, and a page turned away on 3 proxy sessions in a row is retried through an Italian residential proxy — that page only, at no extra cost to you. With your own proxy URLs, only your proxies are used. Without a proxy, a request turned away is retried after a pause of 5 seconds, doubled at each retry up to 150 seconds. The proxy setting applies to the Pagine Gialle pages only: with **Find e-mails on company websites**, the company websites are read directly, not through the proxy.

#### Is the data safe to open in Excel or to show on a web page?

Business names, descriptions and highlights are the businesses' own words, copied as they are. A text can begin with `-`, `+`, `=` or `@`, and a phone number is a text (`347 3765549`, `+39 011 …`): Excel and Google Sheets may read such a cell of a CSV file as a formula or as a number, and would drop the leading `+` of `phonesE164`. The Actor leaves the text as it is, so that the JSON and the API give the real value: when you open a CSV, import these columns as text.

Every field that holds a URL — `url`, `website`, `links`, `socialLinks`, `facebookUrl`, `instagramUrl`, `logo`, `images`, `theForkUrl`, `menuUrl`, `bookingUrl` — is either a complete `http(s)` address or `null`: no other kind of link ever comes out, never one with a user name or password (`https://bank.it@other.site`) or a bare IP address, and an address the site gives broken is dropped rather than shortened. An e-mail is a plain address (letters, digits and `. _ + -`), never one with a quote, an angle bracket or a `?`. Texts are not escaped, though: on a web page, escape every field like any text written by a stranger.

#### Known limitations

- The site caps every search at 200 results: **Expand locations** works around it by searching each sub-area, but a nationwide search (empty `where`) only yields the top 200 unless you paste a region. On small towns the sub-areas mostly repeat the same businesses. A capped search is always read in the site's order first; with a cap (`maxItems`, `maxItemsPerQuery`) and filters or **New businesses only**, the last businesses of the cap may come from its sub-areas, each in the site's order for that sub-area.
- Category pages show no website, VAT number or category code: `requireWebsite` and `categoryCodes` need **Extract details** there (the business page gives the website, the VAT number and the main category code).
- `premium` listings are recognised on search results only (category pages show them as `paid`); `theForkUrl` is read on category pages and HTML fallback pages only; on those pages `ratingAverage` is the stars the card shows, rounded to the half star.
- `onlyNew` remembers business ids, not their content: a business whose phone changed is not returned again. Two runs sharing the same `stateKey` at the same time may both return the same new business.
- Macro categories (`categories`) are missing on some searches: the site does not always send them. `taxCode`, `services`, `paymentMethods`, `reviews`, `menuUrl` and `bookingUrl` come from the business page: they are empty when **Extract details** is off (`legalName` comes with every search result; on category pages it needs **Extract details** too). `name` is the name Pagine Gialle prints — the shop sign when the listing has one — and `legalName` the registered name.
- No filter by publication date: a directory listing has no date.

**A run the platform stops without warning** (out of memory, run timeout)

- Resurrect it: it goes on from where it stood at most a minute before the stop. What it had read since is read again, and the businesses already saved are skipped: none is delivered or charged twice, and **Max businesses** still counts them.
- With **New businesses only**, the memory is saved once a minute: resurrect the stopped run and the businesses it had saved meanwhile join the memory; leave it stopped for good, and the next run may return up to a minute of them once more.

#### Something doesn't work?

The last line of the log counts the businesses saved, filtered out, already delivered by a previous run and no longer on Pagine Gialle (a pasted business page that was removed), and the requests that failed after every retry. Those requests and the removed pages are listed, with the reason, in the `FAILED_REQUESTS` record of the run's key-value store. A run that saved nothing and had failed requests fails, and its last message gives the cause (a pasted category page that does not exist says so, instead of "run it again"). A run that saved some businesses fails too when at least as many requests failed for good as were read: a green run with a short dataset would hide an outage. One failed request among many is only a warning.

If Pagine Gialle changes its pages, you are told instead of paying for blank rows. A results page that counts businesses but gives none the Actor can read is an error (listed in `FAILED_REQUESTS`), never a quiet "0 results"; so is a first results page that no longer says how many businesses the search has (the Actor would otherwise read that one page only). If the first 20 businesses read all lack their city, category, phone, GPS position, street, postal code or province — or, on search results, their region, category code or legal name, or, when business pages are read, the legal name those pages give — the run saves nothing more, stops and fails, and its last message names the missing field: at most those first businesses are charged. A business a filter drops because one of those fields is missing (`requirePhone`, `categoryCodes`…) counts among those 20.

### 🛟 Support

Open an issue in the **Issues** tab with a link to your run: the log shows every search, the number of results the site reported, the reason of every skipped search, and the `FAILED_REQUESTS` record of the key-value store lists the URLs that failed and why.

# Actor input Schema

## `what` (type: `string`):

Activity, category or business name **in Italian**, as typed in the site's *cosa* box: `ristoranti`, `idraulici`, `avvocati`, `parrucchieri`, `agenzie immobiliari`.

## `searchQueries` (type: `array`):

Extra activities, one per line (`dentisti`, `idraulici`). Every activity is searched in every place of *Where* + *More places*: 3 activities × 4 places = 12 searches in one run, each with its own *Max businesses per search*.

## `where` (type: `string`):

Italian municipality (`milano`, `reggio emilia`), district (`Milano Quartiere Isola`), postal code (`20159`), province code (`MI`, `provincia di milano`) or region (`lombardia`). Leave empty for the whole country (top 200 results only, unless *Expand locations* is on and you paste a region).

## `locations` (type: `array`):

Extra places, one per line (`roma`, `00122`, `TO`). Crossed with every activity (see *More activities*), within the run's limit of searches (max 500).

## `searches` (type: `array`):

Searches that must NOT be crossed with the lists above, one per line, as `what | where` (e.g. `idraulici | roma`, `dentisti | 00122`). They run in the same dataset, deduplicated by business id.

## `startUrls` (type: `array`):

Pagine Gialle search URLs (`https://www.paginegialle.it/ricerca/idraulici/roma`), category pages (`https://www.paginegialle.it/lombardia/milano/ristoranti.html`, with or without `/p-N` or a district) or single business pages (`https://www.paginegialle.it/ristoranteantico1850`). Used in addition to the fields above. Max 1 000 URLs, of which at most 500 searches or category pages (a run reads at most 500 searches in all).

## `maxItems` (type: `integer`):

Maximum number of businesses to save in the whole run (after deduplication and filters). 0 = no limit. For reference: the site returns at most 200 results per search; with *Expand locations* on, a big city goes far beyond it (Milano lists 73 districts, each searched up to 200).

## `maxItemsPerQuery` (type: `integer`):

Cap of EACH search (activity × place, `what | where` pair, or start URL), its automatic sub-areas included — so the first search cannot use the whole *Max businesses* budget. 0 = no per-search cap.

## `extractDetails` (type: `boolean`):

Off (default): 25 businesses per request with phones, e-mail, website, VAT number, GPS, hours, rating, description already included. On: one extra request per business to add the tax code (codice fiscale), services, payment methods, reviews, social links and photos — and, on category pages, the legal name and the category code. Charged as a separate `business-details` event ($0.70 per 1,000 businesses).

## `enrichEmails` (type: `boolean`):

Visits each business's own website (home page + contact / legal pages, 3 pages max) and adds the e-mail addresses published there (`websiteEmails`, also merged into `emails`). Slower. Charged as a separate `email-enriched` event, only when the website gives an address Pagine Gialle does not show. Category pages carry no website: combine them with *Extract details*. The websites are read directly, not through the proxy; the run summary says why the others gave no address.

## `expandLocations` (type: `boolean`):

When a search hits the site's cap of 200 results, automatically re-run it for every sub-area the site lists (region → provinces → municipalities → districts) and deduplicate. Costs one extra page per capped search.

## `sortBy` (type: `string`):

Order of the site's results (matters when a search is capped at 200): relevance (site default), distance from the place centre, or popularity. Sorting by distance is also what makes the site fill the `distance` field of every item. Category pages keep the site's own order.

## `requirePhone` (type: `boolean`):

Skip businesses without a phone number. Filtered businesses are not saved and not counted in the caps.

## `requireEmail` (type: `boolean`):

Skip businesses without an e-mail address. With *Find e-mails on company websites* on, a business is kept when its website gives one.

## `requireWebsite` (type: `boolean`):

Skip businesses without a website. Category pages do not show the website: with them, turn *Extract details* on.

## `minRating` (type: `number`):

Keep businesses rated at least this (1-5, e.g. `4` or `4.5`); unrated businesses are skipped. 0 = off. Use `1` to keep rated businesses only.

## `openNowOnly` (type: `boolean`):

Keep only businesses open at the moment they are scraped. The Actor asks the site for its own "Aperto ora" results (far fewer pages to read) and checks the hours again on every item (Italian time). Businesses with no hours and no 24/7 flag are skipped. A pasted search URL that already carries the site's "Aperto ora" filter switches this on.

## `listingTypes` (type: `array`):

Keep only these listing types: `paid` (advertisers, richest data), `premium` (top advertisers — recognised on search results only; category pages show them as paid), `free`. Empty = all.

## `categoryCodes` (type: `array`):

Keep businesses filed under one of these Pagine Gialle category codes (the `categoryCode` / `categoryCodes` output fields, e.g. `007585100` = restaurants). Search results give every code of a business; a category page gives none — turn *Extract details* on there, the business page gives its main code; a pasted business URL gives its main code.

## `categoryContains` (type: `array`):

Keep businesses whose category contains one of these words (`pizzeria`, `sushi`).

## `cityContains` (type: `array`):

Keep businesses whose municipality contains one of these words — useful after a province or region search (`monza`, `sesto`).

## `nameContains` (type: `array`):

Keep businesses whose name contains one of these words (`srl`, `studio`).

## `excludeKeywords` (type: `array`):

Skip businesses whose name contains one of these words (`mcdonald`, `franchising`).

## `onlyNew` (type: `boolean`):

Skip (and do not charge) the businesses that a previous run with the same *Memory key* already delivered. Schedule the same input every week to receive only the businesses added since. The memory lives in the named key-value store `paginegialle-businesses-scraper-seen` of your account.

## `stateKey` (type: `string`):

Name of the memory used by *New businesses only* (letters, digits, `_`, `-`). Give each schedule its own key, e.g. `dentisti-roma`.

## `resetState` (type: `boolean`):

Forget the businesses remembered under *Memory key* before this run (it then returns everything again).

## `excludeEmptyFields` (type: `boolean`):

Remove null values, empty texts, empty lists and empty objects from every business: a compact JSON, smaller exports. Off: every business has all the fields (stable CSV columns).

## `proxyConfiguration` (type: `object`):

Apify Proxy or your own proxies, for the Pagine Gialle pages (company websites visited by *Find e-mails* are read directly). Keep the default: it is included in the price. The residential Apify proxy is not available in this Actor.

## `maxConcurrency` (type: `integer`):

How many pages the Actor reads at the same time. Default: 4.

## `maxRequestsPerMinute` (type: `integer`):

Upper limit on the Actor's request rate. Default: 120; lower it to run more slowly.

## `maxRequestRetries` (type: `integer`):

Retries per page when a request fails. Behind a proxy, a page the site turns away is also retried on a new proxy session up to 10 times without using up these retries. Default: 5.

## `skipUnrecognizedLocation` (type: `boolean`):

When the site does not recognise the *where* value it silently returns nationwide results. On (default): such searches are skipped with a warning instead of saving unrelated businesses.

## `debugLog` (type: `boolean`):

Verbose logging (one line per page and per skipped business).

## Actor input object example

```json
{
  "what": "ristoranti",
  "searchQueries": [],
  "where": "milano",
  "locations": [],
  "searches": [],
  "startUrls": [],
  "maxItems": 100,
  "maxItemsPerQuery": 0,
  "extractDetails": false,
  "enrichEmails": false,
  "expandLocations": true,
  "sortBy": "relevance",
  "requirePhone": false,
  "requireEmail": false,
  "requireWebsite": false,
  "minRating": 0,
  "openNowOnly": false,
  "listingTypes": [],
  "categoryCodes": [],
  "categoryContains": [],
  "cityContains": [],
  "nameContains": [],
  "excludeKeywords": [],
  "onlyNew": false,
  "stateKey": "default",
  "resetState": false,
  "excludeEmptyFields": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxConcurrency": 4,
  "maxRequestsPerMinute": 120,
  "maxRequestRetries": 5,
  "skipUnrecognizedLocation": true,
  "debugLog": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "what": "ristoranti",
    "where": "milano",
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("nice_dev/paginegialle-businesses-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "what": "ristoranti",
    "where": "milano",
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("nice_dev/paginegialle-businesses-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "what": "ristoranti",
  "where": "milano",
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call nice_dev/paginegialle-businesses-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nice_dev/paginegialle-businesses-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/IuS4je4WEMefpKhto/builds/h6UdZMw9C77W4UCX4/openapi.json
