PagineGialle Scraper — Italian Business Leads
Pricing
from $7.00 / 1,000 business records
PagineGialle Scraper — Italian Business Leads
Extract Italian business leads from PagineGialle.it with emails, phone numbers, WhatsApp contacts, websites, locations, ratings, social profiles, flat output, automatic deduplication, and per-query or global result limits.
Pricing
from $7.00 / 1,000 business records
Rating
4.3
(2)
Developer
Emiliano Mastragostino
Maintained by CommunityActor stats
2
Bookmarked
134
Total users
9
Monthly active users
13 hours
Issues response
6 days ago
Last modified
Categories
Share
PagineGialle Scraper — Italian Business Leads Scraper (paginegialle.it) 🇮🇹
🇮🇹 Sviluppato in Italia per il mercato italiano.
🇬🇧 Developed in Italy for the Italian market.
Quick Navigation / Scorciatoie
- 🇮🇹 Versione Italiana
- 🇬🇧 English Version
- ⚙️ Technical Setup & API Docs (Input/Output)
- 🛠️ Support & Contacts
🇮🇹 Versione Italiana
Estrai lead aziendali strutturati da PagineGialle.it, la directory italiana di Pagine Gialle.
Inserisci categorie e località, incolla URL specifici dei risultati di ricerca di paginegialle.it, oppure combina le due modalità. L’Actor restituisce record aziendali pronti per attività di lead generation, con deduplicazione automatica, indicatori sulla qualità dei contatti, tracciamento delle query e limiti globali o per singola ricerca.
Progettato per esecuzioni lunghe e affidabili, può elaborare numerose categorie, località, CAP e URL di ricerca diretti in una sola esecuzione, mantenendo pulito il dataset finale. Quando la deduplicazione è attiva, vengono addebitati solo i record aziendali unici effettivamente salvati nel dataset.
L’Actor è pensato come generatore di dataset di lead aziendali italiani, non come scraper generico. Ogni riga salvata è quindi strutturata per facilitare analisi, filtri e workflow successivi.
Perché usare questo scraper di Pagine Gialle?
Usa l’Actor per:
- Creare liste di lead B2B italiani.
- Raccogliere indirizzi email, numeri di telefono, contatti WhatsApp, siti web e profili social disponibili pubblicamente.
- Analizzare mercati locali, concorrenti e fornitori di servizi.
- Esportare dati aziendali in CSV, Excel, Google Sheets, CRM o workflow Apify.
- Eseguire estrazioni di grandi dimensioni basate su più ricerche, senza che i record duplicati incidano sui limiti o sul budget di fatturazione.
Funzionalità principali
- Input semplice e combinabile — inserisci categorie e località, URL diretti dei risultati di ricerca di PagineGialle oppure entrambi. L’Actor prepara automaticamente gli URL di partenza.
- Adatto a esecuzioni lunghe — combina numerose categorie, località, CAP e URL in un’unica esecuzione, con limiti per query, limiti globali, deduplicazione e riepiloghi finali.
- Deduplicazione automatica — attiva per impostazione predefinita in ogni esecuzione e basata sugli identificatori univoci dei profili aziendali di PagineGialle.
- Fatturazione successiva alla deduplicazione — quando la deduplicazione è attiva, i duplicati ignorati non vengono salvati nel dataset e non vengono addebitati come risultati.
- Output semplice e pronto all’uso — un record aziendale per ogni riga del dataset, ideale per CSV, Excel, Google Sheets e importazioni nei CRM.
- Indicatori sulla qualità dei contatti — filtra rapidamente i record in base alla presenza di email, telefono, cellulare, WhatsApp, sito web, profili social o alla possibilità complessiva di contattare l’azienda.
- Limiti flessibili sui risultati — controlla il numero massimo di aziende salvate nell’intera esecuzione o per ciascuna query.
- Estrazione affidabile — paginazione automatica, rotazione dei proxy e mitigazione dei blocchi.
Come estrarre dati da PagineGialle
- Inserisci una o più categorie aziendali e località, incolla gli URL dei risultati di ricerca di PagineGialle oppure combina le due opzioni.
- Configura facoltativamente
maxItems,maxItemsPerQueryededuplicateResultsnella sezione tecnica. - Avvia l’Actor dalla Console Apify o tramite API.
- Scarica i lead aziendali ottenuti in JSON, CSV, Excel, XML o in un altro formato supportato dal dataset.
🇬🇧 English Version
Extract structured business leads from PagineGialle.it, Italy’s Yellow Pages directory.
Enter categories and locations, paste specific paginegialle.it search-result URLs, or combine both input methods. The Actor returns lead-ready business records with automatic deduplication, contactability flags, query-source metadata, and global or per-query result limits.
Designed for reliable long-running jobs, it can process many categories, locations, Italian postal codes (CAPs), and direct search URLs in a single run while keeping the final dataset clean. When deduplication is enabled, only unique business records successfully saved to the dataset are billed.
The Actor is designed as a dataset builder for Italian business leads rather than as a generic scraper. Each saved row is therefore structured for analysis, filtering, and downstream workflows.
Why use this Pagine Gialle scraper?
Use the Actor to:
- Build Italian B2B lead lists.
- Collect publicly listed email addresses, phone numbers, WhatsApp contacts, websites, and social profiles.
- Research local markets, competitors, and service providers.
- Export business data to CSV, Excel, Google Sheets, CRMs, or Apify workflows.
- Run large multi-query extractions without duplicate records counting toward your limits or billing budget.
Key features
- Simple, combinable input — provide categories and locations, direct PagineGialle search URLs, or both. No mode selection is required.
- Designed for long runs — combine many categories, locations, postcodes, and URLs in a single run, with per-query limits, global limits, deduplication, and run summaries.
- Automatic deduplication — enabled by default within each run and based on reliable PagineGialle profile identifiers.
- Billing after deduplication — when deduplication is enabled, skipped duplicates are not saved to the dataset and are not billed as results.
- Lead-ready output — each dataset row contains one business record, suitable for CSV, Excel, Google Sheets, and CRM imports.
- Contactability flags — quickly filter records by the presence of email, phone, mobile phone, WhatsApp, website, social profiles, or overall contactability.
- Flexible result limits — control the maximum number of businesses saved across the entire run or for each individual query.
- Reliable extraction — automatic pagination, proxy rotation, and blocking mitigation.
How to scrape PagineGialle
- Enter one or more business categories and locations, paste direct PagineGialle search-result URLs, or combine both options.
- Optionally configure
maxItems,maxItemsPerQuery, anddeduplicateResultsin the technical setup. - Run the Actor from Apify Console or through the API.
- Download the resulting business leads in JSON, CSV, Excel, XML, or another dataset-supported format.
⚙️ Technical Documentation & API Reference
Input Configuration
| Field | Type | Description |
|---|---|---|
categories | array | Business categories in Italian, such as "dentisti", "avvocati" or "ristoranti". Must be used together with locations, unless searchUrls is provided. |
locations | array | Locations expressed as city names, province codes, or CAPs/postcodes, for example "Roma", "MI" or "00121". Must be used together with categories, unless searchUrls is provided. |
searchUrls | array | PagineGialle search-result URLs. Can be used on their own or together with categories and locations. |
maxItems | integer | Maximum number of businesses saved across the entire run. Use 0 for no global limit. Default: 0. |
maxItemsPerQuery | integer | Maximum number of businesses saved for each category/location pair or direct search URL. Use 0 for no per-query limit. Default: 0. |
deduplicateResults | boolean | Removes duplicate PagineGialle profiles within the current run. Default: true. |
Valid input patterns are:
categories+locationssearchUrlscategories+locations+searchUrls
Each category is combined with each location. For example, two categories and three locations generate six separate searches. Overlapping results are deduplicated by default. For searchUrls, use search-result URLs copied from your browser.
Input Examples
{"searchUrls": ["https://www.paginegialle.it/ricerca/avvocati/milano"]}
{"categories": ["dentisti"],"locations": ["Roma", "00121"]}
{"categories": ["dentisti", "commercialisti"],"locations": ["Roma", "FI", "00121"],"searchUrls": ["https://www.paginegialle.it/ricerca/avvocati/milano"],"maxItems": 700,"maxItemsPerQuery": 100,"deduplicateResults": true}
Extracted Data Fields
| Data Group | Example Fields |
|---|---|
| Business details | company_name, business_category, description |
| Contact data | email, emails (array), phone, phones (array), secondary_phone, whatsapp, whatsapps |
| Online presence | website, website_domain, facebook_url, instagram_url, tiktok_url, logo_url |
| Location | address, postcode (CAP), city, province, region, country, latitude, longitude |
| Ratings | rating_average, rating_count |
| Lead-quality flags | has_email, has_phone, has_mobile_phone, has_whatsapp, has_website, has_social, is_contactable |
| Query attribution | query_category, query_location, query_input_type, query_strategy, query_id, query_search_url |
| Source metadata | profile_url, profile_id, source, search_url, source_page_number, scraped_at |
Note: Scalar fields may be null when PagineGialle does not provide a value. Collection fields are returned as arrays and may be empty.
has_mobile_phoneuses a conservative heuristic for identifying Italian mobile numbers and may returnfalsewhen a number cannot be identified with sufficient confidence.is_contactableistruewhen the business has at least one email address, phone number, WhatsApp number, or website. A social profile alone does not set this flag.
Output Example (JSON)
{"company_name": "Monti Studio","business_category": "Studi commercialisti","description": "Lo STUDIO MONTI fornisce ai propri clienti servizi professionali in materia fiscale e del lavoro.","website": "[https://www.montistudio.eu](https://www.montistudio.eu)","website_domain": "montistudio.eu","email": "info@montistudio.it","emails": ["info@montistudio.it"],"phone": "06 5812270","phones": ["06 5812270", "333 7430257"],"secondary_phone": "333 7430257","whatsapp": "333 7430257","whatsapps": ["333 7430257"],"address": "Via Costanza Baudana Vaccolini, 5","postcode": "00153","city": "Roma","province": "RM","region": "Lazio","country": "Italy","latitude": 41.87813,"longitude": 12.46669,"rating_average": 5,"rating_count": 3,"profile_url": "[https://www.paginegialle.it/montistudioroma](https://www.paginegialle.it/montistudioroma)","profile_id": "0ec6b49e-c36e-41f4-9221-79797f606906","logo_url": "[https://img.italiaonline.it/0WO5p000003g9yRGAQ/IOL4YOU_LOGO_1647186095662.gif](https://img.italiaonline.it/0WO5p000003g9yRGAQ/IOL4YOU_LOGO_1647186095662.gif)","facebook_url": "[https://www.facebook.com/montistudio.roma/](https://www.facebook.com/montistudio.roma/)","instagram_url": null,"tiktok_url": null,"has_email": true,"has_phone": true,"has_mobile_phone": true,"has_whatsapp": true,"has_website": true,"has_social": true,"is_contactable": true,"contact_channels": ["email", "phone", "whatsapp", "website"],"query_category": "commercialisti","query_location": "Roma","query_input_type": "category_location","query_strategy": "category_location","query_id": "category_location:commercialisti:roma","query_search_url": "[https://www.paginegialle.it/ricerca/commercialisti/Roma](https://www.paginegialle.it/ricerca/commercialisti/Roma)","source": "paginegialle.it","search_url": "[https://www.paginegialle.it/ricerca/commercialisti/Roma](https://www.paginegialle.it/ricerca/commercialisti/Roma)","source_page_number": 1,"scraped_at": "2026-07-02T10:00:00.000Z"}
Deduplication, Limits, and Billing
Deduplication is enabled by default and applies within a single run. The Actor identifies duplicates using reliable PagineGialle profile identifiers: profile_id and normalized profile_url.
The first matching record is saved. Any matching records found later are skipped rather than merged. The Actor does not deduplicate records by company name, phone number, address, email, or website domain. Records without a reliable profile identifier cannot be safely deduplicated and are therefore treated as unique.
When deduplication is enabled, skipped duplicates:
- are not saved to the dataset;
- are not billed as saved results;
- do not count toward
maxItems; - do not count toward
maxItemsPerQuery.
Billing therefore applies only to final business records that remain after deduplication and limit checks and are successfully saved to the dataset.
maxItemscontrols the total number of businesses saved across the entire run.maxItemsPerQueryapplies separately to each category/location pair and each direct search URL. It is useful for distributing coverage more evenly across multiple searches.- Deduplication does not persist across separate Actor runs. If
deduplicateResultsis disabled, duplicate records may be saved, counted toward limits, and billed normally.
🇮🇹 Limiti per gli utenti del piano gratuito Apify
Per gli utenti con un piano Apify gratuito, questo Actor restituisce al massimo 100 risultati per esecuzione e 200 risultati al giorno. Il conteggio giornaliero si azzera a mezzanotte UTC e considera solo i record unici salvati.
Quando il limite viene raggiunto, l’esecuzione si conclude regolarmente con i risultati già raccolti e lo spiega nel messaggio di stato. Con un piano Apify a pagamento non si applica alcun limite: valgono solo i tuoi maxItems e maxItemsPerQuery.
🇬🇧 Limits for Apify free-plan users
For users on a free Apify plan, this Actor returns at most 100 results per run and 200 results per day. The daily counter resets at midnight UTC and counts only final unique saved records.
When a limit is reached, the run finishes normally with the results collected so far and explains the limit in its status message. On a paid Apify plan no such cap applies — only your own maxItems and maxItemsPerQuery do.
Run Summary and Deduplication State
The default dataset contains business records only. Run metadata is stored separately in the default key-value store:
RUN_SUMMARY— query counts, saved records, duplicates removed, contact coverage, limits, and failure statistics.DEDUPE_STATE— a snapshot of the deduplication state for the current run.
These records are not added to the dataset, do not appear as business leads, and are not billed as dataset results. DEDUPE_STATE is useful for monitoring long-running jobs.
Tips for Better Coverage
- Use multiple CAPs (italian postcodes) to cover a large city more precisely.
- Paste specific filtered PagineGialle search URLs when you need full control over the original search.
- Combine generated category/location searches with direct URLs in the same run.
- Use
maxItemsPerQueryto distribute results more evenly across multiple searches. - Use
maxItemsto control the total number of saved records and your overall budget. - Keep
deduplicateResultsenabled for large runs involving overlapping locations or multiple geographic areas.
A broad PagineGialle search may return only a limited number of results. Using multiple postcodes or specific filtered search URLs can increase coverage. When searches overlap, deduplication prevents the same PagineGialle profiles from being saved and billed more than once within the same run.
🛠️ Maintenance and Support
This Actor is actively maintained.