# Deep Site Audit - Per-Page + Company Dataset

**Use case:** 

Crawl an entire website deeply and get a complete contact dataset. Depth 4 with output_mode=both returns a row per crawled page - primary_email, emails, email_count, primary_phone, phone_numbers, current_url and page_title - plus one deduplicated summary row per site aggregating all emails and phone_numbers. verify_emails scores each address via email_status and deliverable_email_count.

## Input

```json
{
  "Urls": [
    "https://www.debian.org/"
  ],
  "Depth": 4,
  "Total_num": 100,
  "Lock_domain": true,
  "render_mode": "auto",
  "output_mode": "both",
  "only_with_contacts": false,
  "verify_emails": true,
  "enrich_social_profiles": [],
  "Concurrency": 5,
  "Max_urls_per_depth": 200,
  "url_include_patterns": [],
  "url_exclude_patterns": []
}
```

## Output

```json
{
  "type": {
    "label": "Type",
    "format": "text"
  },
  "company_name": {
    "label": "Company Name",
    "format": "text"
  },
  "domain": {
    "label": "Domain",
    "format": "text"
  },
  "category": {
    "label": "Category",
    "format": "text"
  },
  "city": {
    "label": "City",
    "format": "text"
  },
  "region": {
    "label": "Region/State",
    "format": "text"
  },
  "country": {
    "label": "Country",
    "format": "text"
  },
  "address": {
    "label": "Address",
    "format": "text"
  },
  "latitude": {
    "label": "Latitude",
    "format": "text"
  },
  "longitude": {
    "label": "Longitude",
    "format": "text"
  },
  "opening_hours": {
    "label": "Opening Hours",
    "format": "text"
  },
  "primary_email": {
    "label": "Primary Email",
    "format": "text"
  },
  "email_count": {
    "label": "# Emails",
    "format": "number"
  },
  "deliverable_email_count": {
    "label": "# Deliverable Emails",
    "format": "number"
  },
  "emails": {
    "label": "Emails",
    "format": "array"
  },
  "email_status": {
    "label": "Email Verification",
    "format": "object"
  },
  "primary_phone": {
    "label": "Primary Phone",
    "format": "text"
  },
  "phone_count": {
    "label": "# Phones",
    "format": "number"
  },
  "phone_numbers": {
    "label": "Phone Numbers",
    "format": "array"
  },
  "uncertain_phone_numbers": {
    "label": "Uncertain Phone Numbers",
    "format": "array"
  },
  "instagram": {
    "label": "Instagram",
    "format": "array"
  },
  "instagram_profiles": {
    "label": "Instagram Profiles (followers, verified…)",
    "format": "array"
  },
  "tiktok": {
    "label": "TikTok",
    "format": "array"
  },
  "tiktok_profiles": {
    "label": "TikTok Profiles (followers, verified…)",
    "format": "array"
  },
  "youtube": {
    "label": "YouTube",
    "format": "array"
  },
  "youtube_profiles": {
    "label": "YouTube Profiles (subscribers…)",
    "format": "array"
  },
  "twitter": {
    "label": "Twitter/X",
    "format": "array"
  },
  "twitter_profiles": {
    "label": "Twitter/X Profiles (followers…)",
    "format": "array"
  },
  "facebook": {
    "label": "Facebook",
    "format": "array"
  },
  "facebook_profiles": {
    "label": "Facebook Profiles (followers…)",
    "format": "array"
  },
  "linkedin": {
    "label": "LinkedIn",
    "format": "array"
  },
  "linkedin_profiles": {
    "label": "LinkedIn Company (followers, employees, industry…)",
    "format": "array"
  },
  "pinterest": {
    "label": "Pinterest",
    "format": "array"
  },
  "github": {
    "label": "GitHub",
    "format": "array"
  },
  "reddit": {
    "label": "Reddit",
    "format": "array"
  },
  "snapchat": {
    "label": "Snapchat",
    "format": "array"
  },
  "whatsapp": {
    "label": "WhatsApp",
    "format": "array"
  },
  "telegram": {
    "label": "Telegram",
    "format": "array"
  },
  "medium": {
    "label": "Medium",
    "format": "array"
  },
  "discord": {
    "label": "Discord",
    "format": "array"
  },
  "threads": {
    "label": "Threads",
    "format": "array"
  },
  "page_title": {
    "label": "Page Title",
    "format": "text"
  },
  "current_url": {
    "label": "Current URL",
    "format": "link"
  },
  "start_url": {
    "label": "Start URL",
    "format": "link"
  },
  "pages_crawled": {
    "label": "Pages Crawled",
    "format": "number"
  },
  "scraped_urls": {
    "label": "Scraped URLs",
    "format": "array"
  },
  "render_method": {
    "label": "Render Method",
    "format": "text"
  },
  "depth": {
    "label": "Depth",
    "format": "number"
  },
  "referrer_url": {
    "label": "Referrer URL",
    "format": "link"
  }
}
```

## About this Actor

This example demonstrates how to use [Contact Info Scraper](https://apify.com/delicious_zebu/contact-info-scraper.md) with a specific input configuration. Visit the [Actor detail page](https://apify.com/delicious_zebu/contact-info-scraper.md) to learn more, explore other use cases, and run it yourself.


## How to integrate an Actor?

This Task's input is already configured above. Use it as-is rather than inventing a new one.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For full API examples (JavaScript, Python, CLI, MCP, OpenAPI), see this Task's Actor page: https://apify.com/delicious_zebu/contact-info-scraper.md

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).
