# Website Sales Signals — Tech Stack, Contacts, SEO & Gaps (`tildekai/website-sales-signals`) Actor

For each website: tech stack (1,178 rules), role emails and phones, company identity, hosting, SEO and mail checks, and missing tools with evidence, lead score and sales hooks.

- **URL**: https://apify.com/tildekai/website-sales-signals.md
- **Developed by:** [Attila Kis](https://apify.com/tildekai) (community)
- **Categories:** Lead generation, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 analyzed websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Sales Signals

Give the Actor a list of websites. For each website it returns what the site uses, how to contact the company, and what the site does not have. The result is made for sales prospecting: each missing item is a possible reason to contact the company.

All detection is rule-based. No LLM, no browser, no proxy, no owned database. The output is company data; see [Personal data](#personal-data) for the limits.

### Output detail

`compact` (default) is made for lists and CRM import: `summary`, `platform`, `company`, `contact`, `technologySummary`, `toolBudget`, `domainInfo`, `dns`, `opportunities`, `salesHooks`, `flags`, `targetMatch`, `leadScore`, `checkSummary`, `dataQuality`. About half the size of the full result.

`full` adds the detail blocks: `technologies` with evidence and versions, `features`, `checks`, `privacy`, `pagesRead`.

`summary` is one short text for people and AI agents, for example:

> Csavarker Kft. (HU). Platform: UNAS. Contact: info@csavarker.hu, +36 30 643 0000. Tools: Hotjar. Tool budget: small\_business. Lead score: 94 (tier A). Top opportunities: No live chat or chatbot detected; …

A parked domain (for sale, registrar placeholder) gets `flags.isParkedDomain`, lead score 0 and no opportunities.

### Sample output

Part of a real compact result (Apify cloud run, 2026-09-30):

```json
{
  "website": "csavarker.hu",
  "status": "complete",
  "summary": "Csavarker Kft. (HU). Platform: UNAS. Contact: info@csavarker.hu, +36 30 643 0000. Tools: Hotjar. Tool budget: small_business. Lead score: 94 (tier A). Top opportunities: No live chat or chatbot detected; ...",
  "platform": "UNAS",
  "company": {
    "name": "Csavarker Kft.",
    "legalName": "Csavarker Kft.",
    "country": {"code": "HU", "confidence": "high"},
    "registrationIds": [{"type": "tax_number", "country": "HU", "value": "11098539-2-06", "pageKind": "contact"}]
  },
  "contact": {
    "primaryEmail": "info@csavarker.hu",
    "primaryPhone": "+36 30 643 0000",
    "socialProfiles": {"facebook": "https://www.facebook.com/csavarker", "tiktok": "https://www.tiktok.com/@csavarkerkft"}
  },
  "hosting": {"hostedOn": "Rackforest Zrt.", "cdn": null, "originHidden": false, "dnsProvider": "Forpsi"},
  "opportunities": [
    {"id": "no_live_chat", "category": "chat_support", "title": "No live chat or chatbot detected", "confidence": "high"}
  ],
  "salesHooks": [
    {"category": "chat_support", "opportunityId": "no_live_chat", "confidence": "high",
     "text": "Csavarker Kft. has no chat on the website, so visitors with a question must call or write an email."}
  ],
  "leadScore": {"score": 94, "tier": "A", "parts": {"contactability": 30, "opportunity": 35, "dataQuality": 15}}
}
```

### What you get for each website

| Block | Content |
|---|---|
| `technologies` | 1,178 product rules in 64 categories: CMS, shop platform, analytics, advertising pixels, chat, chatbot, email marketing, booking, payments, consent tools, reviews, recruiting and more. Each result has evidence. |
| `contact` | Role email addresses (`info@`, `sales@`, `support@`, …), phone numbers in E.164 format, each with the page where it was found, contact form, company social profiles. |
| `company` | Name, logo URL, legal name and legal form, company registration and tax numbers, address, country, industry hints, founding year. |
| `features` | Page types that the site links to, calls to action, forms, selling and hiring signals, content freshness, trust signals. |
| `checks` | About 45 checks: on-page SEO, indexability, structured data, robots.txt, sitemap, llms.txt, performance, security headers, accessibility. |
| `privacy` | Trackers, consent tools, policy pages. |
| `dns` | Mail provider, SPF, DMARC and services that DNS TXT records verify (for example Microsoft 365, Stripe, Atlassian). |
| `hosting` | Where the site runs: IP, network owner (for example Hetzner, OVHcloud), CDN, platform (Shopify behind Cloudflare), DNS provider. A CDN can hide the real host; `originHidden` says so. |
| `opportunities` | What the site does not have, by category, with evidence and a confidence level. |
| `salesHooks` | One ready sentence per opportunity category, for a sales message. |
| `targetMatch` | Which of your target products or categories the site uses. |
| `leadScore` | A 0–100 sorting score with tier A–D and the points of each part. |
| `toolBudget` | Spend level from the paid tools that were detected: `enterprise`, `mid_market`, `small_business`, `minimal`. A lower bound. |
| `domainInfo` | Domain registration date, expiry date, registrar, and the TLS certificate issuer and expiry. |
| `adLibraryLinks` | Links to the Google Ads Transparency Center and the Meta Ad Library, to check if the company runs ads. |
| `flags` | 25 yes/no columns (`hasLiveChat`, `hasAdPixel`, `dmarcEnforced`, …) for spreadsheet and CRM import. |

The word lists cover 17 languages. `Kontakt`, `Contactez-nous`, `Contacto` and `Yhteystiedot` are all contact pages.

### Input

```json
{
  "websites": ["apify.com", "https://www.allbirds.com"],
  "focus": ["chat_support", "seo"],
  "maxPagesPerWebsite": 4
}
```

| Field | Default | Meaning |
|---|---|---|
| `websites` | — | Up to 5000 domains or URLs. One result per host name. The Actor always starts at the home page. |
| `sourceDatasetId`, `sourceField` | —, `website` | Read the websites from a dataset of another Actor, for example a Google Maps export. Give `websites`, a dataset, or both. |
| `skipDatasetId` | — | The dataset of an earlier run. Websites in it are skipped and not charged. For scheduled runs on a list that grows. |
| `outputDetail` | `compact` | `compact` or `full`; see below. |
| `focus` | all | Keeps only these categories in `opportunities`. The other blocks do not change. |
| `maxPagesPerWebsite` | 4 | Home page plus the best contact, legal notice, about and pricing pages. When these pages give no role email and no phone, the Actor reads up to 2 more pages to find a contact page. The value 1 reads the home page only. |
| `includeDns` | true | MX, SPF, DMARC, TXT and hosting lookups. |
| `includeEvidence` | true | Adds the matched URL, header or text to each technology and opportunity. |
| `maxConcurrency` | 8 | Websites in parallel. Requests to one host are sequential. With the default 1 GB memory, 1,000 websites take about 25 minutes. |
| `requestDelayMs` | 300 | Minimum time between two requests to one host. |
| `includeDomainInfo` | true | One RDAP request and one TLS handshake per website. |
| `targetTechnologies` | none | Product names (`Klaviyo`) or category ids (`chat`). Fills `targetMatch`. |
| `onlyTargetMatches` | false | Keep only websites that use a target. |
| `onlyWithOpportunities` | false | Keep only websites with an opportunity in the selected categories. |
| `requireEmail`, `requirePhone` | false | Keep only websites with a role email or a phone number. |
| `countries` | all | Keep only these two-letter country codes. |
| `minLeadScore` | 0 | Keep only websites with at least this score. |

A website that a filter removes is not stored. It is charged with the lower `website-filtered` event, because the Actor did the same work. The run summary counts the removed websites by reason.

Example: find users of a competitor with a contact address.

```json
{
  "websites": ["shop-one.com", "shop-two.com"],
  "targetTechnologies": ["Klaviyo", "Mailchimp"],
  "onlyTargetMatches": true,
  "requireEmail": true
}
```

Opportunity categories: `chat_support`, `seo`, `advertising`, `email_marketing`, `analytics`, `web_development`, `ecommerce`, `booking`, `reviews`, `privacy_compliance`, `accessibility`, `email_security`, `localization`, `recruiting`, `ai_readiness`.

### How to read "not detected"

The Actor reads HTML. It does not run JavaScript. A tool that a tag manager, a JavaScript app or a shop platform loads at run time can be invisible.

For this reason each opportunity has a `confidence`:

| Value | Meaning |
|---|---|
| `high` | The site is server-rendered and has no tag manager or app loader. A missing tool is very probably missing. |
| `medium` | The site has a tag manager, a JavaScript framework or a hosted platform. The tool can load at run time. |
| `low` | The result is partial, the page is client-rendered, or the rule is weak by nature. |

`dataQuality.absenceConfidence` gives the same value for the whole website. Use `high` results for automatic outreach and check `medium` results by hand.

Measured on 30 vendor websites (a vendor uses its own product): 27 of 31 expected technologies were found. All 4 misses were on websites with `medium` confidence.

### Personal data

The Actor is built to output company data only.

- An email address is output only when its name is a role word (`info`, `sales`, `support`, `office`, … in 17 languages). Other addresses are counted in `nonRoleEmailsNotOutput` and not output.
- Addresses in an obscured spelling are read and marked with `obfuscated: true`: `info [at] acme [dot] com`, `info (at) acme (dot) com`, the word "at" in other languages, split markup, hidden decoy text, reversed text, data attributes and simple script concatenation. You can filter these out.
- Encoded addresses (Cloudflare email protection, TYPO3 mail encryption) are not decoded. `encodedEmailNotDecoded` reports them.
- LinkedIn person profiles, Facebook `profile.php` and `people/` links, and share links are dropped.
- Names of persons in structured data (`founder`, `employee`, `author`, reviews) are never read.
- Addresses of other organisations are dropped, also the mail link of the hosting provider on a legal notice page. A mail link on a contact page is kept, because some companies use another domain for mail.

Limits, found in a review of 220 small business outputs (2026-09-30):

- A sole trader or a partnership can use a personal name as the business name (for example `Barbara Heinze`, `… GbR`). The Actor cannot separate these.
- Text that the site publishes about itself is copied: the page title, the meta description (`company.description`) and short evidence snippets. This text can name persons that the site names, for example a chef or an editor.

Phone numbers are output because the site publishes them as a business contact.

An obscured spelling is a sign that the site owner does not want automatic collection. Check the rules for unsolicited business email in your country before you use these addresses.

### Behaviour on the web

- User-Agent: `WebsiteSalesSignalsBot/0.1`.
- The Actor obeys `robots.txt`. A website that disallows its home page gets the status `blocked`.
- No login, no CAPTCHA or challenge bypass, no proxy rotation.
- Per website: `robots.txt`, home page, one HTTP redirect check, sitemap, `llms.txt`, up to three more pages, up to two pages of the contact search, three DNS mail queries, and up to five DNS queries for hosting (two of them to the free Team Cymru IP-to-ASN service).

### Status and charging

| `status` | Meaning | Charged |
|---|---|---|
| `complete` | Home page and the selected pages were read | yes |
| `partial` | Home page was read; another request failed | yes |
| `blocked` | robots.txt or the server refused the request | no |
| `unreachable` | DNS, TLS or network failure | no |
| `invalid` | The input is not a public website address | no |
| `error` | The analysis failed; this is a defect of the Actor | no |

Pay-per-event prices: `website-analyzed` $0.01 for a stored result (= $10 per 1,000 websites) and `website-filtered` $0.002 for a website that your filters removed. Blocked, unreachable and invalid websites are free. The run summary is in the key-value store record `RUN_SUMMARY`.

With a spending limit on the run (for example the monthly free credit), the Actor stops when the limit is reached. You pay only for the stored results, and `RUN_SUMMARY` shows `stoppedByChargeLimit: true`.

# Actor input Schema

## `websites` (type: `array`):

Up to 5000 domains or URLs, for example acme.com or https://www.acme.com. One result per host name. The Actor always starts at the home page. Give this list, a source dataset, or both.

## `sourceDatasetId` (type: `string`):

A dataset of another Actor, for example a Google Maps export. The Actor reads one website from each item.

## `skipDatasetId` (type: `string`):

The dataset of an earlier run of this Actor. Websites that this dataset has are skipped and not charged. Use it for scheduled runs on a list that grows.

## `sourceField` (type: `string`):

Name of the dataset field that holds the website address.

## `outputDetail` (type: `string`):

compact: company, contacts, technology summary, opportunities, sales hooks, lead score and yes/no columns. full: adds evidence for each technology, all page features, all checks and privacy signals.

## `focus` (type: `array`):

Keeps only these categories in the opportunities list. Empty keeps all. The other output blocks do not change.

## `maxPagesPerWebsite` (type: `integer`):

Home page plus the best contact, legal notice, about and pricing pages. When these pages give no role email and no phone number, the Actor reads up to 2 more pages to find a contact page. The value 1 reads the home page only. Robots.txt, sitemap and llms.txt are extra requests.

## `includeDns` (type: `boolean`):

MX, SPF and DMARC records, services that TXT records verify, and hosting: IP, network owner, CDN and name servers.

## `includeEvidence` (type: `boolean`):

Adds the matched URL, header or text to each technology and opportunity.

## `maxConcurrency` (type: `integer`):

Requests to one host are always sequential.

## `requestDelayMs` (type: `integer`):

The Actor sends its own User-Agent and obeys robots.txt.

## `targetTechnologies` (type: `array`):

Product names (for example Klaviyo, Intercom) or category ids (for example chat, email\_marketing). Each website gets targetMatch: which of these it uses. Unknown names are listed in the run summary.

## `onlyTargetMatches` (type: `boolean`):

Other websites are not stored. They are charged with the lower website-filtered event.

## `onlyWithOpportunities` (type: `boolean`):

Uses the opportunity categories that you selected.

## `requireEmail` (type: `boolean`):

Filter on contact.emails.

## `requirePhone` (type: `boolean`):

Filter on contact.phones.

## `countries` (type: `array`):

Two-letter country codes, for example HU, DE. A website with an unknown country is removed. Empty keeps all.

## `minLeadScore` (type: `integer`):

0 keeps all. The score formula is in the output documentation.

## `includeDomainInfo` (type: `boolean`):

One RDAP request and one TLS handshake per website. Many country domains (for example .hu, .de) publish no registration dates.

## Actor input object example

```json
{
  "websites": [
    "apify.com"
  ],
  "sourceField": "website",
  "outputDetail": "compact",
  "focus": [],
  "maxPagesPerWebsite": 4,
  "includeDns": true,
  "includeEvidence": true,
  "maxConcurrency": 8,
  "requestDelayMs": 300,
  "targetTechnologies": [],
  "onlyTargetMatches": false,
  "onlyWithOpportunities": false,
  "requireEmail": false,
  "requirePhone": false,
  "countries": [],
  "minLeadScore": 0,
  "includeDomainInfo": true
}
```

# Actor output Schema

## `websites` (type: `string`):

One item per website: status, company, contact, technologies, features, checks, privacy, DNS and opportunities.

## `overviewCsv` (type: `string`):

One row per website with the main columns. Text columns come from websites: import them as text.

## `summary` (type: `string`):

Counts by status, requests, bytes and error codes.

## `crmCsv` (type: `string`):

One row per website: company, contacts, lead score, tool budget, target match and yes/no columns.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tildekai/website-sales-signals").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": ["apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("tildekai/website-sales-signals").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "apify.com"
  ]
}' |
apify call tildekai/website-sales-signals --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tildekai/website-sales-signals"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/s86lHDnAPaErfyUU0/builds/DpQdA9tSDKy392RB6/openapi.json
