# Local Business Website Lead Qualifier (`weio/local-business-website-lead-qualifier`) Actor

Lead qualification for web designers: score local business websites 0-100 on how much they need a new site (no mobile layout, broken HTTPS, outdated code, SEO spam). Enrich a Google Maps scraper dataset, or chain it after your scrape via Integrations. Role emails, phones and socials included.

- **URL**: https://apify.com/weio/local-business-website-lead-qualifier.md
- **Developed by:** [Weio, Inc.](https://apify.com/weio) (community)
- **Categories:** Lead generation, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 site qualifieds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Give it a list of local business websites, or a dataset you already have from a places or business-directory
scraper, and it tells you which businesses most need a new website, with a **0-100 "needs a new website" score**
and the plain reasons behind it. Built for web designers, agencies and freelancers who sell website work to local
businesses and want to contact the right ones first.

Lists of businesses *without* a website are easy to get. This actor finds the ones that *have* a website that is
hurting them: no mobile layout, a broken certificate, code from another decade, or a homepage stuffed with casino
spam by a hacker.

### How to use

1. Paste domains or URLs into **Websites**, or pick a dataset in **Dataset with websites** (for example the output
   of a places or directory scraper you ran on Apify; the actor reads each item's website field).
2. Optional: set **Only output sites scoring at least** (for example 40) so you only get, and only pay for, likely leads.
3. Optional: tick **Phone screenshot** to get a picture of how each site looks on a phone.
4. Start the run, open the **Leads** view, sort by score and export to CSV, Excel or JSON, or read the dataset
   through the API.

**Run it automatically after your places scraper.** In the task of your Google Maps (or other places) scraper, open
**Integrations**, add this actor and keep the prefilled input `"datasetId": "{{resource.defaultDatasetId}}"`. Each time
the scraper finishes, this actor qualifies the businesses in that run's dataset that list a website (others are free).

### What each row contains

| field | meaning |
|---|---|
| `needs_new_website_score` | 0-100, the sum of the points below (capped at 100) |
| `score_reasons` | which signals fired |
| `signals` | every signal, true/false |
| `mobile_viewport` | the page declares a phone layout |
| `https_bare`, `https_www` | certificate status (valid / incomplete\_chain / expired / wrong\_name / self\_signed / no\_https ...) and days left |
| `cms`, `technologies`, `jquery_version` | what the site is built with |
| `newest_copyright_year` | newest copyright year found on the homepage |
| `spam_terms`, `spam_links` | casino/pharma spam phrases and links to spam sites injected into the homepage, if any |
| `role_emails`, `phones`, `social`, `contact_page` | public business contact details from the site itself |
| `inputs` | every input of yours that points at this website (join the result back to your own list) |
| `no_own_website` | true when the "website" was a page on a platform (Facebook, Yelp, Google, a free-builder subdomain) |
| `phone_screenshot_url` | link to the phone screenshot, when you asked for one |

Example row (`info.cern.ch`, from the prefilled input):

```json
{
  "input": "info.cern.ch",
  "inputs": ["info.cern.ch"],
  "domain": "info.cern.ch",
  "final_url": "https://info.cern.ch/",
  "needs_new_website_score": 30,
  "score_reasons": ["no_mobile_viewport"],
  "mobile_viewport": false,
  "https_bare": {"status": "valid", "days_left": 52},
  "https_www": {"status": "unresolved", "days_left": null},
  "cms": null,
  "technologies": [],
  "newest_copyright_year": null,
  "spam_terms": [],
  "role_emails": [],
  "phones": [],
  "social": {},
  "no_own_website": false,
  "error": null
}
```

#### How the score is built

| signal | points |
|---|---|
| no mobile viewport tag (page renders desktop-width on phones) | 30 |
| no valid HTTPS certificate on either the bare or www host | 20 |
| HTTPS broken on one of the two hosts, or an incomplete certificate chain | 5 |
| fixed page width of 700 px or more (HTML width attribute, or CSS width on a page without a viewport tag) | 10 |
| newest copyright year is 3 or more years old ("2015 - present" counts as current) | 10 |
| jQuery 1.x loaded | 10 |
| Flash or Silverlight object | 10 |
| `<font>` tags, or table-based layout on a page without a mobile viewport | 5 |
| an HTTPS page still loads http:// scripts, styles or images | 5 |
| SEO spam injected into the homepage (likely hacked) | 15 |

The spam signal counts spam phrases, links to spam sites and hidden spam text. It ignores the spam family that is
the business's own trade, so a pharmacy, a casino or a lender is not called hacked for describing what it sells.

### Input

- **Websites**: domains or URLs, one per line, and/or
- **Dataset with websites**: a dataset in your Apify account. The actor reads the `website` (or `websiteUrl`,
  `webUrl`, `site`, `domain`) field of each item, so the output of most places and directory scrapers works as is.
  A listing's own page (its Google Maps or Yelp `url`) is never used as the business's website: items without a
  website, and places marked permanently closed, are skipped for free.
- **Only output sites scoring at least**: skip low scorers (they are not output and not charged).
- **Phone screenshot** (opt-in, off by default): also saves a first-screen phone screenshot (390x844) of each
  qualified site and adds `phone_screenshot_url` to the row. Useful as visual proof in a pitch: a desktop-only site
  shows up squeezed and unreadable. The link works for anyone who has it and lasts as long as your plan keeps the
  run's storage, so download the images you want to keep.

Example input:

```json
{
  "websites": ["zingermans.com", "info.cern.ch"],
  "datasetId": "<your dataset ID, optional>",
  "minScore": 40,
  "phoneScreenshot": false
}
```

### Pricing

Pay per event: **$0.01 per qualified site** ($10 per 1,000). You pay once per website domain: if your list has
40 locations of one chain that all link to the same site, that site is checked and charged once, and its row lists
every distinct input that points at it (`inputs`).

Free: sites that are unreachable or not a valid address, sites that answer with a bot check instead of their page,
sites whose robots.txt asks our crawler (user agent WeioBot) not to read the homepage,
"websites" that are a page on a platform (Facebook, Instagram, Yelp, Google, a free-builder subdomain; marked
`no_own_website`, which is often the best lead of all), and sites below your minimum score.

Opt-in phone screenshot: **$0.003 per screenshot saved**. When the site shows a bot check, answers with an error
status, or the capture fails, no screenshot is saved, the row says why in `phone_screenshot_note`, and it is not
charged.

Example: 1,000 places from a directory scraper, of which 700 have a website on 650 distinct domains, 50 of those
are unreachable and you set a minimum score of 40 that 200 sites reach: you pay for 200 sites, $2.00. The run stops
when your maximum charge is reached, and a run that is restarted by the platform does not charge again for sites
it already delivered.

### FAQ and limits

- **Does it render the page?** No. The score reads the homepage HTML and the certificates. A site that has a
  viewport tag but still overflows on a phone can score low; the phone screenshot shows you the real thing.
- **Certificate checks that time out** are never counted against a site.
- **Does it use Google?** No. The actor never fetches Google, Google Maps or social network pages. It reads the
  website addresses you give it and fetches only those businesses' own homepages.
- **Big lists:** up to 10,000 websites per run. For lists over a few thousand, give the run a longer timeout in
  the run options.
- **Related actors from Weio:**
  [Local Business Website Audit](https://apify.com/weio/local-business-website-audit) gives the full per-site issue
  list (no score) for sites you already care about, and
  [Bulk Lighthouse Audit](https://apify.com/weio/lighthouse-mobile-median-audit) runs mobile Lighthouse on the
  leads you shortlist here.

### Data and responsibility

- Only public pages are read: the homepage of each site, and a TLS handshake for the certificate.
- Only role email addresses (info@, sales@, office@ ...) are returned. Personal-name addresses are dropped.
- You are responsible for having the right to collect and use the input data and for how you contact the
  businesses (for example CAN-SPAM in the US: real sender, postal address, working opt-out).
- A business that wants its site left out can email sales@weio.ai (opt-out requests are honoured; the address is
  used for nothing else here) and it will be added to the actor's skip list.

Built and maintained by Weio, Inc. The actor was written and is operated by AI agents (Anthropic Claude), with a
person accountable at Weio.

# Actor input Schema

## `websites` (type: `array`):

Domains or URLs, one per line. Leave empty if you use a dataset below. Inputs on the same domain are checked and charged once. Up to 10,000 websites per run.

## `datasetId` (type: `string`):

A dataset you already hold, e.g. the output of a places or business-directory scraper you ran. The actor reads each item's website field (website, websiteUrl, webUrl, site, domain). Items without a website and permanently closed places are skipped for free; a listing's own page (Google Maps, Yelp) is never used as the website. You are responsible for having the right to collect and use that data.

## `minScore` (type: `integer`):

Skip sites whose 'needs a new website' score is lower; skipped sites are not output and not charged. 0 = output all. Try 40 to keep only likely leads.

## `phoneScreenshot` (type: `boolean`):

Also save a first-screen phone screenshot (390x844) of each qualified site and add its link to the row. Charged as a separate phone-screenshot event ($0.003) only when the image is saved. Off by default.

## `concurrency` (type: `integer`):

How many sites are checked at the same time (1-32).

## Actor input object example

```json
{
  "websites": [
    "zingermans.com",
    "info.cern.ch"
  ],
  "minScore": 0,
  "phoneScreenshot": false,
  "concurrency": 16
}
```

# Actor output Schema

## `siteRows` (type: `string`):

The default dataset, Leads view: one row per website.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "zingermans.com",
        "info.cern.ch"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("weio/local-business-website-lead-qualifier").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "zingermans.com",
        "info.cern.ch",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("weio/local-business-website-lead-qualifier").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "zingermans.com",
    "info.cern.ch"
  ]
}' |
apify call weio/local-business-website-lead-qualifier --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,weio/local-business-website-lead-qualifier"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/oaXx1C1FCsQcqznrE/builds/PB6dmpn5HpJr4lwVV/openapi.json
