# Website Tech Stack & Domain Intelligence (DNS, WHOIS, SSL) (`zhucl1006/website-tech-stack-domain-intelligence`) Actor

Bulk-check domains: detect CMS, e-commerce, analytics and 7,000+ technologies, plus email provider (MX/SPF/DMARC), RDAP WHOIS registration & expiry, SSL certificate expiry and security headers. Homepage only, robots.txt respected.

- **URL**: https://apify.com/zhucl1006/website-tech-stack-domain-intelligence.md
- **Developed by:** [leo zhu](https://apify.com/zhucl1006) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.01 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Tech Stack & Domain Intelligence (DNS, WHOIS, SSL)

Paste a list of domains and get, for each one, a **clean lead-qualification report**:

- 🧩 **Technologies** – CMS, e-commerce platform, analytics, tag managers, CDN, frameworks, payment and marketing tools (7,600+ fingerprints)
- ✉️ **Email & DNS** – MX records, **email provider** (Google Workspace, Microsoft 365, Proofpoint…), **SPF** and **DMARC policy**, SaaS verification records (TXT)
- 📅 **WHOIS via RDAP** – registrar, **registration date, expiry date, domain age**
- 🔒 **SSL/TLS certificate** – issuer, validity, **days until expiry**
- 🛡️ **Security headers** – HSTS, CSP, X-Frame-Options… with an A–F grade
- 🔗 Page title, description, language and official **social profiles**

One row per domain, ready for CSV/Excel, CRM import or API use.

**Keywords:** website technology lookup, tech stack checker, what CMS is this site using, detect Shopify / WordPress / WooCommerce sites, technographic data, bulk domain lookup, email provider lookup (MX), SPF and DMARC checker, WHOIS / RDAP domain age and expiry, SSL certificate expiry checker, security headers scanner, lead enrichment, low-cost alternative to BuiltWith-style or Wappalyzer-style technology lookups (pay per domain, no subscription).

### Who is it for?

- **Sales & lead generation** – find shops running Shopify or WooCommerce, companies on HubSpot, sites using a competitor's product.
- **Agencies** – qualify prospects: outdated CMS, poor security headers, SSL about to expire, no DMARC.
- **Security & IT** – quick external posture check of your own or your clients' domains.
- **Domain investors / brand protection** – registration and expiry dates in bulk.

#### Example workflows

- Upload 5,000 e-commerce domains, keep rows where `ecommerce` contains `Shopify`, export to CSV for outreach.
- Enrich a CRM export with `emailProviders` (Google Workspace vs Microsoft 365) to tailor messaging.
- Weekly scheduled run on your clients' domains: alert when `ssl.daysUntilExpiry` < 21 or `dns.dmarcPolicy` is missing.

### How it works (and what it does not do)

- Only the **homepage** of each host is fetched (plus a few internal pages in *Deep check* mode).
- **robots.txt is respected** (RFC 9309). If a site disallows crawling, the page is not fetched; DNS, RDAP and certificate checks still run because they do not crawl the site.
- **No JavaScript execution, no proxies, no CAPTCHA or anti-bot bypass.** Sites that present a bot challenge are reported as `blocked`.
- Polite defaults: identified user agent, request timeouts, limited concurrency, sequential requests per site.

### Input

| Field | Default | Description |
|---|---|---|
| `domains` | – | Domains or URLs. Paths are ignored; duplicates are removed and never charged twice. |
| `deepCheck` | `false` | Analyse up to `deepCheckMaxPages` internal pages (about, pricing, contact…) and check `security.txt`, sitemap, `ads.txt`, `llms.txt`. |
| `deepCheckMaxPages` | `3` | 1–10 extra pages per domain. |
| `includeTechnologies` / `includeDns` / `includeWhois` / `includeSsl` | `true` | Turn individual checks off if you don't need them. |
| `maxConcurrency` | `10` | Domains processed in parallel (1–50). |
| `requestTimeoutSecs` | `15` | Timeout per request. |

```json
{ "domains": ["apify.com", "shopify.com", "wordpress.org"], "deepCheck": false }
```

### Output example (abbreviated)

```json
{
  "domain": "wordpress.org",
  "status": "ok",
  "httpStatus": 200,
  "title": "Blog Tool, Publishing Platform, and CMS – WordPress.org",
  "cms": ["WordPress"],
  "technologyCount": 13,
  "websiteTechnologies": "Google Font API, Google Tag Manager, Gutenberg, HSTS, MySQL, Nginx, Open Graph, PHP, ...",
  "technologies": [{"name": "WordPress", "version": null, "categories": ["CMS", "Blogs"], "confidence": 100, "detectedBy": ["dom", "html", "meta"]}],
  "dns": {"mx": ["10 smtp1-dca.wordpress.org"], "hasSpf": true, "hasDmarc": true, "dmarcPolicy": "reject", "emailProviders": []},
  "whois": {"registrar": "MarkMonitor Inc.", "createdDate": "2003-03-28", "expiryDate": "2035-03-28", "domainAgeDays": 8584},
  "ssl": {"valid": true, "issuer": "Let's Encrypt", "validTo": "2026-12-22", "daysUntilExpiry": 86},
  "securityHeaders": {"grade": "D", "score": 2, "server": "nginx"},
  "socialProfiles": {"x": "https://twitter.com/WordPress", "linkedin": "https://www.linkedin.com/company/wordpress"}
}
```

`status` values: `ok`, `httpError` (site answered 4xx/5xx), `blocked` (bot challenge – not bypassed), `robotsDisallowed`, `robotsUnreachable` (robots.txt returned 5xx/429 – page not fetched, per RFC 9309), `unreachable`, `invalid`, `error`.

`websiteTechnologies` are seen on the website itself; `dnsInferredServices` are SaaS tools inferred only from DNS records (e.g. TXT domain-verification entries) or the certificate issuer.

### Pricing

Pay per event: a small start fee per run, one **domain-checked** event per domain in the dataset, and an extra **deep-check** event per domain when *Deep check* is enabled. Invalid inputs and duplicates are not charged. Example: 1,000 domains without deep check = 0.005 + 1,000 x 0.003 = about **USD 3.01**. Set a maximum cost per run and the Actor stops when it is reached.

### FAQ

**How is this different from browser-extension tech detectors?** It runs in bulk (thousands of domains per run), adds DNS/email, WHOIS/RDAP, SSL and security-header data in the same row, and exports to CSV/Excel/JSON or your CRM via the Apify API.

**Does it bypass Cloudflare or bot protection?** No. Blocked sites are reported as `blocked` and still get DNS, WHOIS and SSL data.

**Can I schedule it?** Yes - use Apify schedules and integrations (Google Sheets, Slack, webhooks, Zapier/Make).

### Limitations

- Technologies that only appear after JavaScript runs may be missed (no browser is used).
- RDAP is not offered by every country-code TLD (e.g. `.de`, `.jp`); those rows carry `whois.available = false` with the reason.
- Results reflect what the site served to an identified bot at check time.

### Credits and licences

- Technology fingerprints come from the community project **[enthec/webappanalyzer](https://github.com/enthec/webappanalyzer)**, licensed under the **GNU General Public License v3.0**. This Actor contains its own matching engine; the fingerprint data is downloaded unmodified (unused fields removed) at build time together with its licence. Each record carries a `rulesAttribution` field.
- WHOIS data comes from the registries' public **RDAP** services located through the IANA bootstrap registry.
- This Actor is not affiliated with BuiltWith, Wappalyzer or any of the detected vendors. All product names are trademarks of their respective owners.

# Actor input Schema

## `domains` (type: `array`):

Domains (example.com) or URLs (https://www.example.com/page). Paths are ignored - the homepage of each host is analysed. Duplicates are removed automatically and are not charged twice.

## `deepCheck` (type: `boolean`):

Also analyse a few internal pages (about, pricing, contact ...) for more technologies and check security.txt, sitemap, ads.txt and llms.txt. Charged as an extra 'deep-check' event per domain.

## `deepCheckMaxPages` (type: `integer`):

How many internal pages to analyse per domain in deep-check mode (robots.txt is always respected).

## `includeTechnologies` (type: `boolean`):

Fingerprint CMS, frameworks, analytics, CDN, e-commerce etc.

## `includeDns` (type: `boolean`):

Look up A/AAAA/MX/NS/TXT records, SPF and DMARC policy, and infer the email provider.

## `includeWhois` (type: `boolean`):

Registrar, creation date, expiry date and domain age from the official RDAP service of the TLD (not available for every ccTLD).

## `includeSsl` (type: `boolean`):

Issuer, validity and days until expiry of the certificate on port 443.

## `maxConcurrency` (type: `integer`):

Domains processed in parallel.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP/TLS/RDAP request.

## Actor input object example

```json
{
  "domains": [
    "apify.com",
    "shopify.com",
    "wordpress.org"
  ],
  "deepCheck": false,
  "deepCheckMaxPages": 3,
  "includeTechnologies": true,
  "includeDns": true,
  "includeWhois": true,
  "includeSsl": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 15
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "apify.com",
        "shopify.com",
        "wordpress.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("zhucl1006/website-tech-stack-domain-intelligence").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "apify.com",
        "shopify.com",
        "wordpress.org",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("zhucl1006/website-tech-stack-domain-intelligence").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "apify.com",
    "shopify.com",
    "wordpress.org"
  ]
}' |
apify call zhucl1006/website-tech-stack-domain-intelligence --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zhucl1006/website-tech-stack-domain-intelligence"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/w82bz2cn7fKSy2dkm/builds/4MrrbTRTQLRbYnjzc/openapi.json
