# Website Lead Qualifier: Emails, Tech Stack & Sales Signals (`toolsheder/website-lead-qualifier`) Actor

Reads each business website and lists what an agency could sell it: no Meta pixel or analytics, no HTTPS, not mobile-ready, no online booking, outdated WordPress, missing SEO basics. Adds emails with MX check, phones and socials. Takes a URL list or a Google Maps Scraper dataset.

- **URL**: https://apify.com/toolsheder/website-lead-qualifier.md
- **Developed by:** [Kostas Skutulas](https://apify.com/toolsheder) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 analysed websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Lead Qualifier: Emails, Tech Stack & Sales Signals

Website Lead Qualifier is a lead generation tool for agencies and web
designers. It reads business websites and returns one flat row per business
with emails, phones, the tech stack, marketing tags such as Meta Pixel and
Google Ads, and short codes for what an agency could sell the business, such as
`no-meta-pixel`, `no-online-booking`, `no-https` and `outdated-wordpress`. It
takes a list of websites or a Google Maps Scraper dataset, and does not search
for businesses itself.

### What it finds

For each website it records:

- **Contacts**: emails (with an MX check), phones in international format,
  WhatsApp, and links to Facebook, Instagram, LinkedIn, X, YouTube and TikTok.
- **Tech stack**: platform (WordPress, WooCommerce, Shopify, Wix, Squarespace,
  Webflow and a dozen more), WordPress version, and whether the site is a shop.
- **Marketing tags**: Google Analytics 4, the old Universal Analytics, Google
  Tag Manager, Google Ads, Meta Pixel, TikTok, LinkedIn, Hotjar and Clarity.
  Tags are read from the page and from inside its Tag Manager container, where
  agencies put most of them.
- **Widgets**: cookie consent tool, live chat, online booking system and
  newsletter form, each named by vendor.
- **Basics**: HTTPS, mobile viewport, title, meta description, noindex,
  structured data, copyright year, server response time and accessibility
  signals.
- **Opportunities**: the codes in the table below, for filtering the list.

### How to use

#### With Google Maps Scraper

1. Run [Google Maps Scraper](https://apify.com/compass/crawler-google-places)
   for your niche and city.
2. Paste the run's Dataset ID into this actor's "Dataset ID" field and start
   it. You can also add this actor as an integration on the Maps run. It then
   reads the dataset by itself when the Maps run finishes.

Each row keeps the business name, category, city, rating, review count and the
phone number that Google Maps lists, next to what was found on the website.
Businesses without a website are included as `no-website` and are not charged.

#### With your own list

Paste domains or URLs into "Websites", one per line. Tracking parameters are
removed and the rest of the URL is kept, so a branch page on a chain's site is
read as that branch page.

#### Filtering the results

Sort by `opportunityCount`, or filter `opportunities` by the code for the
service you sell. Check `emailHasMx` before you send email to the addresses in
`emails`.

### Opportunity codes

| Code | Raised when | Who sells the fix |
|---|---|---|
| `no-website` | The dataset row has no website at all | Web design |
| `no-own-website` | The "website" is a Facebook, Instagram, Booksy, Wolt, Linktree or similar page, or the domain forwards to one | Web design |
| `website-down` | The site did not answer, or its home page returned an error | Web design, hosting |
| `parked-domain` | The domain shows a registrar's parking or for-sale page | Web design |
| `placeholder-page` | "Coming soon", "under construction" or a server's default page, with little else on it | Web design |
| `no-https` | The site is served over plain HTTP | Web design, hosting |
| `https-broken` | HTTPS failed on its certificate and the site only answered over HTTP | Web design, hosting |
| `no-mobile-viewport` | No `width=device-width` viewport: the site was not built for phones | Web design |
| `outdated-copyright` | The latest copyright year is two or more years old | Web design |
| `outdated-wordpress` | WordPress is three or more feature releases behind the current one | WordPress care |
| `site-builder` | Built on Wix, Squarespace, Webflow, Weebly, GoDaddy, Jimdo, Duda or Tilda | Web design |
| `slow-server` | The home page took more than 2.5 s to start answering | Hosting, performance |
| `accessibility-issues` | Four or more of eight machine-readable accessibility signals failed | Accessibility |
| `noindex` | The page tells search engines not to index it | SEO |
| `missing-title` | No `<title>` | SEO |
| `no-meta-description` | No meta description | SEO |
| `no-structured-data` | No schema.org Organization or LocalBusiness data | SEO |
| `no-analytics` | No analytics of any kind: Google, Tag Manager, Matomo, Plausible, Hotjar or Clarity | Analytics |
| `legacy-analytics-only` | Only Universal Analytics, which stopped collecting data in July 2024. A GA4 tag that the old tag loads as a connected tag counts as GA4 | Analytics |
| `no-meta-pixel` | No Meta Pixel on the page, in Tag Manager or in Shopify's Meta app | Paid social |
| `no-google-ads-tag` | No Google Ads tag on the page, in Tag Manager or in Shopify's Google app | PPC |
| `no-cookie-consent` | Tracking tags run but no consent banner was found, in the EU, EEA, UK or Switzerland | Privacy compliance |
| `no-online-booking` | No booking system, for a business whose Maps category books appointments (salons, clinics, restaurants, gyms, hotels...) | Booking software |
| `no-live-chat` | No chat widget and no Messenger, WhatsApp, Viber or Telegram link | Chat, chatbots |
| `no-contact-email` | No email address on the site or its contact page | n/a |
| `no-social-links` | No link to any social profile | Social media |

A code is raised only on evidence. When something could not be read, its code
is left off rather than guessed. A wrong `no-meta-pixel` would send an ads
agency to a business that already has a pixel.

### Measured on 155 real businesses

Two samples, one run each, with default settings.

| | 130 Lithuanian shops and showrooms | 25 UK dentists, Berlin restaurants, Austin salons |
|---|---|---|
| Analysed | 128 (2 refused access) | 18 (3 without a site of their own, 3 down or refusing access) |
| Run time | 24 s at 10 sites at once | 9 s |
| Pages opened per site | 1.2 on average | 1.8 |
| Email found | 98% | 78% |
| Phone found | 96% | 100% |
| Tags found only inside Tag Manager | 40% of sites | 44% |
| Booking system found | n/a | 44% |
| Opportunities per site | 3.8 | 4.6 |

The Tag Manager row is the reason the container is read. On four sites in ten,
a pixel or an ads tag that is not in the page source was loaded from the
container. Reading the HTML alone would have reported those businesses as not
advertising.

Consent findings were checked by hand. Twelve sites flagged `no-cookie-consent`
were opened in a browser, and nine had no banner. The other three had banners
the actor did not detect at the time (Borlabs Cookie 3, a banner text stored as
escaped JSON, and Wix's own banner). All three cases are now handled.

### What HTML cannot show

- **Cookie banners on site builders.** Wix, Squarespace and the other builders
  draw their own banner after the page loads, and it leaves no trace in the
  HTML. Consent is not judged on these sites: `cookieConsent` is left empty and
  `no-cookie-consent` is never raised.
- **Tags added in a builder's own marketing settings**, such as the "Facebook
  Pixel" integration of a Wix site, load the same way. On site builders,
  `no-meta-pixel` is likely but not certain.
- **Booking forms a business built itself** are not recognised. Only the
  seventy-odd booking systems and plugins listed in the source code are. This
  is why `no-online-booking` is raised only for categories that book
  appointments.
- **Speed** is the server's response time. It shows slow hosting. It is not a
  Lighthouse score and does not measure how heavy the page is.
- **Accessibility** is eight machine-readable signals, not a full audit. See
  the recommendations below.

### Pricing

$4 per 1,000 websites analysed: $0.004 for each site that was reached and read.

Rows for businesses with no website or with only a social or booking page, and
rows for sites that did not answer, are recorded free of charge. When several
Google Maps listings share one website (branches, or a clinic listed once per
doctor), the site is read and charged once and the result is copied to each
listing.

Apify adds its standard start fee of $0.00005 per run. Compute is included.

### Input and output

#### Input

```json
{
  "datasetId": "YOUR_GOOGLE_MAPS_SCRAPER_DATASET_ID",
  "websites": ["kirpykla.lt", "https://www.example-dental.co.uk/"],
  "maxRequestsPerSite": 5,
  "concurrency": 10
}
```

#### Output

A real row from the measured run, shortened:

```json
{
  "name": "Thornley Park Dental",
  "category": "Dentist",
  "city": "Manchester",
  "countryCode": "GB",
  "website": "https://thornleyparkdental.com/",
  "status": "ok",
  "opportunities": ["no-online-booking", "no-live-chat"],
  "opportunityCount": 2,
  "emails": ["info@thornleyparkdental.com"],
  "emailHasMx": true,
  "phones": ["+441613368036"],
  "instagram": "https://www.instagram.com/thornleyparkdental/",
  "platform": "wordpress",
  "https": true,
  "responseMs": 455,
  "mobileViewport": true,
  "copyrightYear": 2026,
  "schemaBusiness": true,
  "googleAnalytics": true,
  "googleTagManager": true,
  "googleAdsTag": true,
  "metaPixel": true,
  "foundInTagManager": ["Google Analytics 4", "Google Ads", "Meta Pixel"],
  "cookieConsent": true,
  "cookieConsentTool": "Borlabs Cookie",
  "liveChat": false,
  "onlineBooking": false,
  "a11yChecksFailed": 3,
  "requestsMade": 1,
  "elapsedMs": 1250
}
```

This practice runs Google Ads and a Meta Pixel, both loaded through Tag
Manager, so its page source shows neither. It is not a lead for an ads agency.
Its two codes are for booking software and live chat.

Every row has a `status`. `ok` and `parked` rows are analyses and are charged.
`no-website`, `no-own-website`, `unreachable`, `blocked` and `invalid` rows are
not. Rows copied from an earlier listing with the same website carry
`sameWebsiteAs` and are not charged again.

### Recommendations

- **Filter on the code you sell.** `opportunityCount` is for sorting. To build
  a list for one service, filter `opportunities` by that service's code.
- **Check `emailHasMx` before sending.** An address on a domain with no mail
  server bounces, and bounces harm the sending domain's reputation.
- **Use `a11yChecksFailed` as a signal only.** It counts eight checks that can
  be read from the source, such as missing alternative text, unlabelled inputs
  and blocked zoom. Failures cluster there, but most of WCAG needs a rendered
  page and a person testing with a keyboard, so it is not a compliance verdict.
- **Do not raise `concurrency` to finish a big list sooner.** Pages on one site
  are opened one at a time with a pause between them, to keep the load on small
  business sites low.

### FAQ

**Does it find businesses as well?**
No. It analyses the websites it is given. Get the businesses from Google Maps
Scraper or from your own list.

**Why is a site `blocked`?**
It answered with 401, 403, 407 or 429, usually from a firewall that blocks
automated visitors. The row records the refusal and is not charged.

**Does it respect robots.txt?**
Yes, including a group written for `website-lead-qualifier` by name. The home
page is requested once, and robots.txt is read before any further page.

**Why is a pixel reported that I cannot see in the page source?**
It is loaded by Google Tag Manager or by a Shopify app. `foundInTagManager`
names the tags that were found only in the container.

### Related tools

- [Lead Qualifier Chrome extension](https://itneeds.lt/lead-qualifier/): the
  same checks for the one site open in your browser, in one click, with a draft
  first email.
- [Shop Intel](https://apify.com/toolsheder/shop-intel), for online shops:
  platform, catalogue photo and dimension readiness, contacts.
- [Photo to 3D](https://apify.com/toolsheder/photo-to-3d): product photos to
  web-ready 3D models at real size.
- Made by [ITneeds](https://itneeds.lt/).

# Actor input Schema

## `websites` (type: `array`):

One business website per line. Domains or full URLs both work; tracking parameters are dropped. A Facebook page, Instagram profile, Booksy or Wolt page given instead of a website is recorded as "no-own-website" and not charged. Leave empty when you give a dataset below.

## `datasetId` (type: `string`):

Read the websites from a dataset instead of, or as well as, the list above. Google Maps Scraper output works as it is: each row keeps its business name, category, city, rating and Maps phone next to the analysis. Any dataset with a website, domain or url field works too. When this actor runs as an integration after another actor, its dataset is used automatically.

## `maxWebsites` (type: `integer`):

Stop after this many rows of input, to try the actor on the start of a large dataset. Leave empty to process everything.

## `maxRequestsPerSite` (type: `integer`):

The most pages opened on one website: the one given, then its contact or imprint page if the first had no email. The cap limits both the cost and the load on small business sites.

## `concurrency` (type: `integer`):

How many different websites are read in parallel. Each one is still visited one page at a time, with a pause between pages.

## Actor input object example

```json
{
  "websites": [
    "https://www.apify.com",
    "https://wordpress.org",
    "https://www.wix.com"
  ],
  "maxRequestsPerSite": 5,
  "concurrency": 10
}
```

# Actor output Schema

## `leads` (type: `string`):

One row per business: what an agency could sell it, how to reach it, and the facts behind both.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "https://www.apify.com",
        "https://wordpress.org",
        "https://www.wix.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("toolsheder/website-lead-qualifier").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "https://www.apify.com",
        "https://wordpress.org",
        "https://www.wix.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("toolsheder/website-lead-qualifier").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "https://www.apify.com",
    "https://wordpress.org",
    "https://www.wix.com"
  ]
}' |
apify call toolsheder/website-lead-qualifier --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,toolsheder/website-lead-qualifier"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cHfbsDsgVMKCxVRZT/builds/Djq7htzLh9YS0R85P/openapi.json
