PagineGialle Scraper — Italian Business Leads avatar

PagineGialle Scraper — Italian Business Leads

Pricing

from $7.00 / 1,000 business records

Go to Apify Store
PagineGialle Scraper — Italian Business Leads

PagineGialle Scraper — Italian Business Leads

Extract Italian business leads from PagineGialle.it with emails, phone numbers, WhatsApp contacts, websites, locations, ratings, social profiles, flat output, automatic deduplication, and per-query or global result limits.

Pricing

from $7.00 / 1,000 business records

Rating

4.3

(2)

Developer

Emiliano Mastragostino

Emiliano Mastragostino

Maintained by Community

Actor stats

2

Bookmarked

134

Total users

9

Monthly active users

13 hours

Issues response

6 days ago

Last modified

Share

PagineGialle Scraper — Italian Business Leads Scraper (paginegialle.it) 🇮🇹

🇮🇹 Sviluppato in Italia per il mercato italiano.
🇬🇧 Developed in Italy for the Italian market.


Quick Navigation / Scorciatoie


🇮🇹 Versione Italiana

Estrai lead aziendali strutturati da PagineGialle.it, la directory italiana di Pagine Gialle.

Inserisci categorie e località, incolla URL specifici dei risultati di ricerca di paginegialle.it, oppure combina le due modalità. L’Actor restituisce record aziendali pronti per attività di lead generation, con deduplicazione automatica, indicatori sulla qualità dei contatti, tracciamento delle query e limiti globali o per singola ricerca.

Progettato per esecuzioni lunghe e affidabili, può elaborare numerose categorie, località, CAP e URL di ricerca diretti in una sola esecuzione, mantenendo pulito il dataset finale. Quando la deduplicazione è attiva, vengono addebitati solo i record aziendali unici effettivamente salvati nel dataset.

L’Actor è pensato come generatore di dataset di lead aziendali italiani, non come scraper generico. Ogni riga salvata è quindi strutturata per facilitare analisi, filtri e workflow successivi.

Perché usare questo scraper di Pagine Gialle?

Usa l’Actor per:

  • Creare liste di lead B2B italiani.
  • Raccogliere indirizzi email, numeri di telefono, contatti WhatsApp, siti web e profili social disponibili pubblicamente.
  • Analizzare mercati locali, concorrenti e fornitori di servizi.
  • Esportare dati aziendali in CSV, Excel, Google Sheets, CRM o workflow Apify.
  • Eseguire estrazioni di grandi dimensioni basate su più ricerche, senza che i record duplicati incidano sui limiti o sul budget di fatturazione.

Funzionalità principali

  • Input semplice e combinabile — inserisci categorie e località, URL diretti dei risultati di ricerca di PagineGialle oppure entrambi. L’Actor prepara automaticamente gli URL di partenza.
  • Adatto a esecuzioni lunghe — combina numerose categorie, località, CAP e URL in un’unica esecuzione, con limiti per query, limiti globali, deduplicazione e riepiloghi finali.
  • Deduplicazione automatica — attiva per impostazione predefinita in ogni esecuzione e basata sugli identificatori univoci dei profili aziendali di PagineGialle.
  • Fatturazione successiva alla deduplicazione — quando la deduplicazione è attiva, i duplicati ignorati non vengono salvati nel dataset e non vengono addebitati come risultati.
  • Output semplice e pronto all’uso — un record aziendale per ogni riga del dataset, ideale per CSV, Excel, Google Sheets e importazioni nei CRM.
  • Indicatori sulla qualità dei contatti — filtra rapidamente i record in base alla presenza di email, telefono, cellulare, WhatsApp, sito web, profili social o alla possibilità complessiva di contattare l’azienda.
  • Limiti flessibili sui risultati — controlla il numero massimo di aziende salvate nell’intera esecuzione o per ciascuna query.
  • Estrazione affidabile — paginazione automatica, rotazione dei proxy e mitigazione dei blocchi.

Come estrarre dati da PagineGialle

  1. Inserisci una o più categorie aziendali e località, incolla gli URL dei risultati di ricerca di PagineGialle oppure combina le due opzioni.
  2. Configura facoltativamente maxItems, maxItemsPerQuery e deduplicateResults nella sezione tecnica.
  3. Avvia l’Actor dalla Console Apify o tramite API.
  4. Scarica i lead aziendali ottenuti in JSON, CSV, Excel, XML o in un altro formato supportato dal dataset.

🇬🇧 English Version

Extract structured business leads from PagineGialle.it, Italy’s Yellow Pages directory.

Enter categories and locations, paste specific paginegialle.it search-result URLs, or combine both input methods. The Actor returns lead-ready business records with automatic deduplication, contactability flags, query-source metadata, and global or per-query result limits.

Designed for reliable long-running jobs, it can process many categories, locations, Italian postal codes (CAPs), and direct search URLs in a single run while keeping the final dataset clean. When deduplication is enabled, only unique business records successfully saved to the dataset are billed.

The Actor is designed as a dataset builder for Italian business leads rather than as a generic scraper. Each saved row is therefore structured for analysis, filtering, and downstream workflows.

Why use this Pagine Gialle scraper?

Use the Actor to:

  • Build Italian B2B lead lists.
  • Collect publicly listed email addresses, phone numbers, WhatsApp contacts, websites, and social profiles.
  • Research local markets, competitors, and service providers.
  • Export business data to CSV, Excel, Google Sheets, CRMs, or Apify workflows.
  • Run large multi-query extractions without duplicate records counting toward your limits or billing budget.

Key features

  • Simple, combinable input — provide categories and locations, direct PagineGialle search URLs, or both. No mode selection is required.
  • Designed for long runs — combine many categories, locations, postcodes, and URLs in a single run, with per-query limits, global limits, deduplication, and run summaries.
  • Automatic deduplication — enabled by default within each run and based on reliable PagineGialle profile identifiers.
  • Billing after deduplication — when deduplication is enabled, skipped duplicates are not saved to the dataset and are not billed as results.
  • Lead-ready output — each dataset row contains one business record, suitable for CSV, Excel, Google Sheets, and CRM imports.
  • Contactability flags — quickly filter records by the presence of email, phone, mobile phone, WhatsApp, website, social profiles, or overall contactability.
  • Flexible result limits — control the maximum number of businesses saved across the entire run or for each individual query.
  • Reliable extraction — automatic pagination, proxy rotation, and blocking mitigation.

How to scrape PagineGialle

  1. Enter one or more business categories and locations, paste direct PagineGialle search-result URLs, or combine both options.
  2. Optionally configure maxItems, maxItemsPerQuery, and deduplicateResults in the technical setup.
  3. Run the Actor from Apify Console or through the API.
  4. Download the resulting business leads in JSON, CSV, Excel, XML, or another dataset-supported format.

⚙️ Technical Documentation & API Reference

Input Configuration

FieldTypeDescription
categoriesarrayBusiness categories in Italian, such as "dentisti", "avvocati" or "ristoranti". Must be used together with locations, unless searchUrls is provided.
locationsarrayLocations expressed as city names, province codes, or CAPs/postcodes, for example "Roma", "MI" or "00121". Must be used together with categories, unless searchUrls is provided.
searchUrlsarrayPagineGialle search-result URLs. Can be used on their own or together with categories and locations.
maxItemsintegerMaximum number of businesses saved across the entire run. Use 0 for no global limit. Default: 0.
maxItemsPerQueryintegerMaximum number of businesses saved for each category/location pair or direct search URL. Use 0 for no per-query limit. Default: 0.
deduplicateResultsbooleanRemoves duplicate PagineGialle profiles within the current run. Default: true.

Valid input patterns are:

  • categories + locations
  • searchUrls
  • categories + locations + searchUrls

Each category is combined with each location. For example, two categories and three locations generate six separate searches. Overlapping results are deduplicated by default. For searchUrls, use search-result URLs copied from your browser.

Input Examples

{
"searchUrls": ["https://www.paginegialle.it/ricerca/avvocati/milano"]
}
{
"categories": ["dentisti"],
"locations": ["Roma", "00121"]
}
{
"categories": ["dentisti", "commercialisti"],
"locations": ["Roma", "FI", "00121"],
"searchUrls": ["https://www.paginegialle.it/ricerca/avvocati/milano"],
"maxItems": 700,
"maxItemsPerQuery": 100,
"deduplicateResults": true
}

Extracted Data Fields

Data GroupExample Fields
Business detailscompany_name, business_category, description
Contact dataemail, emails (array), phone, phones (array), secondary_phone, whatsapp, whatsapps
Online presencewebsite, website_domain, facebook_url, instagram_url, tiktok_url, logo_url
Locationaddress, postcode (CAP), city, province, region, country, latitude, longitude
Ratingsrating_average, rating_count
Lead-quality flagshas_email, has_phone, has_mobile_phone, has_whatsapp, has_website, has_social, is_contactable
Query attributionquery_category, query_location, query_input_type, query_strategy, query_id, query_search_url
Source metadataprofile_url, profile_id, source, search_url, source_page_number, scraped_at

Note: Scalar fields may be null when PagineGialle does not provide a value. Collection fields are returned as arrays and may be empty.

  • has_mobile_phone uses a conservative heuristic for identifying Italian mobile numbers and may return false when a number cannot be identified with sufficient confidence.
  • is_contactable is true when the business has at least one email address, phone number, WhatsApp number, or website. A social profile alone does not set this flag.

Output Example (JSON)

{
"company_name": "Monti Studio",
"business_category": "Studi commercialisti",
"description": "Lo STUDIO MONTI fornisce ai propri clienti servizi professionali in materia fiscale e del lavoro.",
"website": "[https://www.montistudio.eu](https://www.montistudio.eu)",
"website_domain": "montistudio.eu",
"email": "info@montistudio.it",
"emails": ["info@montistudio.it"],
"phone": "06 5812270",
"phones": ["06 5812270", "333 7430257"],
"secondary_phone": "333 7430257",
"whatsapp": "333 7430257",
"whatsapps": ["333 7430257"],
"address": "Via Costanza Baudana Vaccolini, 5",
"postcode": "00153",
"city": "Roma",
"province": "RM",
"region": "Lazio",
"country": "Italy",
"latitude": 41.87813,
"longitude": 12.46669,
"rating_average": 5,
"rating_count": 3,
"profile_url": "[https://www.paginegialle.it/montistudioroma](https://www.paginegialle.it/montistudioroma)",
"profile_id": "0ec6b49e-c36e-41f4-9221-79797f606906",
"logo_url": "[https://img.italiaonline.it/0WO5p000003g9yRGAQ/IOL4YOU_LOGO_1647186095662.gif](https://img.italiaonline.it/0WO5p000003g9yRGAQ/IOL4YOU_LOGO_1647186095662.gif)",
"facebook_url": "[https://www.facebook.com/montistudio.roma/](https://www.facebook.com/montistudio.roma/)",
"instagram_url": null,
"tiktok_url": null,
"has_email": true,
"has_phone": true,
"has_mobile_phone": true,
"has_whatsapp": true,
"has_website": true,
"has_social": true,
"is_contactable": true,
"contact_channels": ["email", "phone", "whatsapp", "website"],
"query_category": "commercialisti",
"query_location": "Roma",
"query_input_type": "category_location",
"query_strategy": "category_location",
"query_id": "category_location:commercialisti:roma",
"query_search_url": "[https://www.paginegialle.it/ricerca/commercialisti/Roma](https://www.paginegialle.it/ricerca/commercialisti/Roma)",
"source": "paginegialle.it",
"search_url": "[https://www.paginegialle.it/ricerca/commercialisti/Roma](https://www.paginegialle.it/ricerca/commercialisti/Roma)",
"source_page_number": 1,
"scraped_at": "2026-07-02T10:00:00.000Z"
}

Deduplication, Limits, and Billing

Deduplication is enabled by default and applies within a single run. The Actor identifies duplicates using reliable PagineGialle profile identifiers: profile_id and normalized profile_url.

The first matching record is saved. Any matching records found later are skipped rather than merged. The Actor does not deduplicate records by company name, phone number, address, email, or website domain. Records without a reliable profile identifier cannot be safely deduplicated and are therefore treated as unique.

When deduplication is enabled, skipped duplicates:

  • are not saved to the dataset;
  • are not billed as saved results;
  • do not count toward maxItems;
  • do not count toward maxItemsPerQuery.

Billing therefore applies only to final business records that remain after deduplication and limit checks and are successfully saved to the dataset.

  • maxItems controls the total number of businesses saved across the entire run.
  • maxItemsPerQuery applies separately to each category/location pair and each direct search URL. It is useful for distributing coverage more evenly across multiple searches.
  • Deduplication does not persist across separate Actor runs. If deduplicateResults is disabled, duplicate records may be saved, counted toward limits, and billed normally.

🇮🇹 Limiti per gli utenti del piano gratuito Apify

Per gli utenti con un piano Apify gratuito, questo Actor restituisce al massimo 100 risultati per esecuzione e 200 risultati al giorno. Il conteggio giornaliero si azzera a mezzanotte UTC e considera solo i record unici salvati.

Quando il limite viene raggiunto, l’esecuzione si conclude regolarmente con i risultati già raccolti e lo spiega nel messaggio di stato. Con un piano Apify a pagamento non si applica alcun limite: valgono solo i tuoi maxItems e maxItemsPerQuery.

🇬🇧 Limits for Apify free-plan users

For users on a free Apify plan, this Actor returns at most 100 results per run and 200 results per day. The daily counter resets at midnight UTC and counts only final unique saved records.

When a limit is reached, the run finishes normally with the results collected so far and explains the limit in its status message. On a paid Apify plan no such cap applies — only your own maxItems and maxItemsPerQuery do.


Run Summary and Deduplication State

The default dataset contains business records only. Run metadata is stored separately in the default key-value store:

  • RUN_SUMMARY — query counts, saved records, duplicates removed, contact coverage, limits, and failure statistics.
  • DEDUPE_STATE — a snapshot of the deduplication state for the current run.

These records are not added to the dataset, do not appear as business leads, and are not billed as dataset results. DEDUPE_STATE is useful for monitoring long-running jobs.


Tips for Better Coverage

  • Use multiple CAPs (italian postcodes) to cover a large city more precisely.
  • Paste specific filtered PagineGialle search URLs when you need full control over the original search.
  • Combine generated category/location searches with direct URLs in the same run.
  • Use maxItemsPerQuery to distribute results more evenly across multiple searches.
  • Use maxItems to control the total number of saved records and your overall budget.
  • Keep deduplicateResults enabled for large runs involving overlapping locations or multiple geographic areas.

A broad PagineGialle search may return only a limited number of results. Using multiple postcodes or specific filtered search URLs can increase coverage. When searches overlap, deduplication prevents the same PagineGialle profiles from being saved and billed more than once within the same run.


🛠️ Maintenance and Support

This Actor is actively maintained.