# Website Tech Stack Detector + DNS & WHOIS Lookup (`prevailing_glow/tech-stack-detector`) Actor

Detect the tech stack of any website: CMS, ecommerce platform, analytics, CDN and 7,600+ technologies, plus email provider, DNS host and WHOIS (RDAP) data.

- **URL**: https://apify.com/prevailing\_glow/tech-stack-detector.md
- **Developed by:** [Dave West](https://apify.com/prevailing_glow) (community)
- **Categories:** Lead generation, SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 domain analyses

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Tech Stack Detector + DNS & WHOIS Lookup

Find out **what any website is built with**: CMS, ecommerce platform, analytics and tag managers, CDN and hosting, JavaScript frameworks, payment providers, marketing tools and **7,600+ other technologies**. Each domain also gets **SEO signals** (HTTPS, HTTP/2, robots.txt, sitemap), its **email provider** (Google Workspace, Microsoft 365, ...), **DNS host**, **SaaS verification records** and **domain WHOIS data via RDAP** (registrar, creation and expiry dates).

Paste a list of domains and get one clean row per domain, ready for CSV, Excel, Google Sheets or your CRM. One plain HTTP request per homepage, no browser, so it is fast and cheap: **about $3 per 1,000 domains** with everything switched on.

### Use cases

- **Lead generation by technology**: build lists of stores on Shopify, WooCommerce, Magento or BigCommerce, sites using HubSpot, Klaviyo or Intercom, or companies on Google Workspace vs. Microsoft 365. Use the `onlyWithTech` filter and pay only for matching domains.
- **Sales prospecting**: qualify accounts before outreach. Which CMS, which analytics, which payment provider, how old is the domain, which SaaS tools have they verified their domain with?
- **Competitor research**: see the stack behind competitors' sites (platform, CDN, A/B testing, reviews, chat, consent tools) and how it changes over time.
- **SEO audits at scale**: check HTTPS, HTTP/2, robots.txt, sitemap presence and Googlebot access for hundreds of sites at once. For a deep page-by-page audit of one site, use [Sitemap SEO Audit & Monitor](https://apify.com/prevailing_glow/sitemap-seo-audit-monitor).
- **Market research**: measure the market share of platforms in a country or niche with the run summary (technology counts and most common stacks).
- **Domain portfolio and security checks**: registrar, expiry dates, EPP status, DMARC policy, SPF senders.

### What is detected

| Area | Examples | Output fields |
|---|---|---|
| CMS | WordPress, Drupal, Joomla, Webflow, Wix, Squarespace, HubSpot CMS, Contentful, Sanity, Ghost | `cms` |
| Ecommerce platform | Shopify, WooCommerce, Magento, BigCommerce, PrestaShop, Shopware, Squarespace Commerce | `ecommerce` |
| Analytics & tag managers | Google Analytics, Google Tag Manager, Matomo, Hotjar, Plausible, New Relic | `analytics`, `tagManagers` |
| CDN & hosting | Cloudflare, Fastly, Akamai, Amazon CloudFront, Vercel, Netlify, WP Engine, Kinsta | `cdn`, `hosting` |
| Frameworks & languages | React, Next.js, Vue, Nuxt, Angular, Svelte, Astro, Laravel, PHP, Node.js | `jsFrameworks`, `webFrameworks`, `programmingLanguages` |
| Marketing & sales | Klaviyo, Mailchimp, HubSpot, Marketo, Intercom, Zendesk, Drift | `marketingAutomation`, `liveChat` |
| Payments | Stripe, PayPal, Shop Pay, Klarna, Afterpay, Apple Pay | `payments` |
| More | advertising pixels, cookie consent, A/B testing, reviews, page builders, WordPress plugins, Shopify apps, security headers | `advertising`, `cookieConsent`, `abTesting`, `reviews`, `pageBuilders`, `themesAndPlugins`, `security` |
| SEO signals | HTTPS, HTTP/2, HTTP/3 advertised, redirects, robots.txt, sitemap URL, Googlebot access, title, meta description, language | `https`, `http2`, `robotsTxtPresent`, `sitemapUrl`, ... |
| DNS | A/AAAA, MX → email provider, NS → DNS host, SPF senders, TXT verification records (Google, Microsoft 365, Meta, Stripe, Atlassian, OpenAI, ...), DMARC policy | `emailProvider`, `dnsHost`, `spfServices`, `saasVerifications`, `dmarcPolicy` |
| WHOIS (RDAP) | registrar, created / updated / expiry dates, domain age, EPP status | `registrar`, `domainCreatedAt`, `domainExpiresAt`, `domainAgeYears`, `domainStatus` |

Every technology comes with its **categories**, the **version** where the site reveals it (e.g. `WordPress 6.6.2`, `Nginx 1.25.3`), a **confidence** score (0-100) and **detectedBy** (the evidence: headers, cookies, meta, script URLs, HTML, DOM, implied, or DNS/IP). Detection uses response headers, cookies, meta tags, script and stylesheet URLs, inline scripts and the HTML/DOM of the homepage.

Privacy: the RDAP lookup returns registrar-level data only. **Registrant names, emails, phone numbers and addresses are never output**, even when a registry publishes them.

### Quick start

1. Paste your domains (one per line; `example.com`, `www.example.com` and full URLs all work).
2. Keep **DNS + WHOIS enrichment** on if you want email provider, DNS host and registration data.
3. Optional: add `Shopify` (or any technology) under **Only output sites using these technologies** to get a lead list.
4. Run, then download the dataset as CSV/Excel/JSON or open the **HTML report**.

Leave the domain list empty for a free preview on five example sites (nothing is charged).

### Input example

```json
{
    "domains": ["ruggable.com", "wordpress.org", "https://www.coolblue.nl", "stripe.com"],
    "enrichDnsWhois": true,
    "checkSeoFiles": true,
    "onlyWithTech": [],
    "onlyWithCategory": [],
    "minConfidence": 50,
    "maxConcurrency": 10
}
```

Lead list of Shopify stores that use Klaviyo:

```json
{ "domains": ["..."], "onlyWithTech": ["Klaviyo"], "onlyWithCategory": ["Ecommerce"] }
```

Within one filter list any entry may match (OR); when both lists are given, both must match (AND). Technology names are matched exactly and case-insensitively, and `onlyWithTech` also matches email providers, DNS hosts and SaaS records (e.g. `Google Workspace`, `HubSpot`, `Stripe`).

### Output

One flat row per domain. `technologies` is an array of objects; all other list fields are comma-separated strings, so CSV and Excel exports stay readable. Real example (shortened `technologies`):

```json
{
  "domain": "ruggable.com",
  "finalUrl": "https://ruggable.com/",
  "statusCode": 200,
  "blocked": false,
  "technologies": [
    { "name": "Cloudflare", "categories": ["CDN"], "version": null, "confidence": 100, "detectedBy": ["headers"] },
    { "name": "Contentful", "categories": ["CMS"], "version": null, "confidence": 100, "detectedBy": ["html"] },
    { "name": "Next.js", "categories": ["JavaScript frameworks", "Web frameworks"], "version": null, "confidence": 100, "detectedBy": ["headers"] },
    { "name": "Shopify", "categories": ["Ecommerce", "CMS"], "version": null, "confidence": 50, "detectedBy": ["dom"] }
  ],
  "technologyNames": "Cloudflare, Cloudflare Bot Management, Cloudflare Browser Insights, Contentful, Google Tag Manager, HSTS, Next.js, Node.js, Open Graph, React, Swiper, Vercel, Webpack, Shopify",
  "technologyCount": 14,
  "cms": "Contentful, Shopify",
  "ecommerce": "Shopify",
  "tagManagers": "Google Tag Manager",
  "cdn": "Cloudflare",
  "title": "Washable Rugs & Washable Area Rugs by Ruggable | Ruggable US",
  "https": true,
  "http2": true,
  "robotsTxtPresent": true,
  "sitemapPresent": true,
  "sitemapUrl": "https://ruggable.com/sitemap.xml",
  "emailProvider": "Mimecast",
  "dnsHost": "Cloudflare",
  "saasVerifications": "Airtable, Anthropic, Apple, Atlassian, Google Search Console, Jamf, Klaviyo, Microsoft 365, Notion, Shopify, Smartsheet, Zoom",
  "dmarcPolicy": "quarantine",
  "registrar": "GoDaddy.com, LLC",
  "domainCreatedAt": "2009-11-20T02:32:02.000Z",
  "domainExpiresAt": "2031-11-20T02:32:02.000Z",
  "domainAgeYears": 16.9,
  "chargedEvents": ["domain-analyzed", "dns-rdap-enrichment"]
}
```

More real results from a test run (October 2026):

| Domain | Platform | CDN / hosting | Email provider |
|---|---|---|---|
| allbirds.com | Shopify | Cloudflare | Microsoft 365 |
| techcrunch.com | WordPress (WordPress VIP, Yoast SEO) | - | Mimecast |
| porterandyork.com | WordPress + WooCommerce, Elementor | Cloudflare, WP Engine | Google Workspace |
| huel.com | Shopify + Sanity, Next.js | Vercel, Imgix | Google Workspace |
| wix.com | Wix, React | Google Cloud CDN | Google Workspace |
| mozilla.org | Wagtail | - | Google Workspace |
| coolblue.nl | custom (Emotion) | Amazon CloudFront | Google Workspace |

Dataset views: **Overview**, **Technologies**, **SEO signals**, **DNS & WHOIS** and **Blocked & failed**.

The key-value store also contains:

- `SUMMARY`: technology counts and market share across all domains, most common stacks (e.g. `Cloudflare + Shopify`), category breakdowns, email providers, DNS hosts, SaaS records, blocked and failed domains, charges.
- `OUTPUT.html`: a one-page HTML report with the same numbers and a domain table.

### Blocked sites

Some big-brand websites sit behind bot protection (Cloudflare, Akamai, DataDome, AWS WAF, Vercel, ...) and answer a plain request with a challenge page, HTTP 403 or HTTP 429. Those rows are returned with `blocked: true` and a `blockedReason`, still include **DNS, email provider and WHOIS data** and the detections that the response headers reveal (e.g. `Cloudflare`, `DataDome`), and are **not charged** as an analysed domain. Hosted platforms are also recognised from the site's IP address (Shopify, Squarespace, Wix, Vercel, Netlify, GitHub Pages), so a rate-limited Shopify store still shows up as `Shopify` (`detectedBy: ["dns"]`, confidence 75) and matches a `Shopify` filter.

Expect a share of large consumer brands to be partly blocked, and some busy Shopify stores to answer with HTTP 429 when many requests come from cloud servers. Small and medium sites are rarely affected. The Actor does not use proxies or headless browsers to get around protections.

### Pricing

Pay per event, no subscription:

| Event | Price | When |
|---|---|---|
| Domain analysed (`domain-analyzed`) | **$0.002** | per domain row whose homepage was fetched and analysed (technologies + SEO signals) |
| DNS + WHOIS enrichment (`dns-rdap-enrichment`) | **$0.001** | per domain when enrichment is on and DNS returned records |

- 1,000 domains with everything on: **about $3**. Without enrichment: about $2.
- **Free**: domains that do not resolve or never answer (DNS failure, timeout, connection or TLS error), blocked/challenged homepages (only the enrichment is charged when it returned data), domains removed by your filters, and preview runs with an empty domain list.
- The run stops cleanly when your **maximum total charge per run** is reached, so you never pay more than you set.

### FAQ

**How accurate is the detection?** It uses an open, MIT-licensed database of 7,600+ technology fingerprints and checks headers, cookies, meta tags, script URLs, inline scripts and the HTML of the homepage. Platforms (CMS, ecommerce, CDN, server) are detected very reliably. Tools that are injected later by JavaScript (for example analytics loaded through a tag manager) can be missed because the page is not rendered in a browser.

**Why only the homepage?** The stack is almost always site-wide and the homepage reveals it. One request per site keeps runs fast and cheap.

**Some fields are empty for a domain. Why?** Empty means "not detected": the site may not use that kind of tool, or it is loaded in a way that a plain HTTP request cannot see. `expiresAt` is missing for some country-code domains whose registry does not publish it, and some TLDs (for example `.de`, `.so`, `.io`) have no public RDAP server, which is noted in `enrichmentError`.

**Is it polite to the websites?** Yes. Each site gets at most three small requests (robots.txt, homepage, sitemap check) with an honest User-Agent; robots.txt is respected by default, and the default concurrency is 10 different sites in parallel.

**Can I filter for several technologies?** Yes: `onlyWithTech: ["Shopify", "WooCommerce", "BigCommerce"]` keeps a domain if any of them is found. Add `onlyWithCategory` to require a category as well.

**Can I run it on a schedule or via API?** Yes, like every Apify Actor: schedule it in the Console, call it from the API, or connect it to Make, Zapier, n8n or Google Sheets.

**Does it return personal data?** No. Only technical and registrar-level information is collected; registrant contact data from WHOIS/RDAP is deliberately dropped.

### Credits and licenses

Technology fingerprints and categories come from [projectdiscovery/wappalyzergo](https://github.com/projectdiscovery/wappalyzergo) v0.3.3 (MIT License, Copyright (c) 2021 ProjectDiscovery, Inc.). The data is vendored unmodified in `data/` with its license (`data/LICENSE.md`, `data/NOTICE.md`). Registration data comes from the registries' public RDAP services via the IANA bootstrap registry.

### Support

Found a wrong or missing detection? Open an issue on the Actor's Issues tab with the domain and what you expected.

# Actor input Schema

## `domains` (type: `array`):

One domain or URL per line, e.g. <code>example.com</code> or <code>https://www.example.com/shop</code>. Only the homepage of each site is analysed. Duplicates are removed. Each analysed domain is one billable event.

## `enrichDnsWhois` (type: `boolean`):

Look up A/AAAA, MX (email provider), NS (DNS host), TXT/SPF (SaaS verification records), DMARC and the domain's RDAP registration data (registrar, created/expiry dates, status). Billed as a separate small event per domain when it returns data.

## `checkSeoFiles` (type: `boolean`):

Check robots.txt presence, the sitemap URL (from robots.txt or /sitemap.xml) and whether Googlebot may crawl the homepage. Adds 1-2 tiny requests per domain; included in the domain price.

## `onlyWithTech` (type: `array`):

Keep only domains where at least one of these is detected, e.g. <code>Shopify</code>, <code>WooCommerce</code>, <code>HubSpot</code>, <code>Google Workspace</code>. Exact names, case-insensitive; also matches email provider, DNS host and SaaS records. Filtered-out domains are not stored and not charged.

## `onlyWithCategory` (type: `array`):

Keep only domains with at least one technology in these categories, e.g. <code>Ecommerce</code>, <code>CMS</code>, <code>Live chat</code>, <code>Marketing automation</code>. Combined with the technology filter using AND.

## `minConfidence` (type: `integer`):

Hide detections below this confidence (0-100). Most detections are 100; a few weak signals are 25-50.

## `maxConcurrency` (type: `integer`):

How many domains are analysed in parallel. Every domain is a different website, and each gets at most a few sequential requests.

## `requestTimeoutSecs` (type: `integer`):

Maximum time for one HTTP request including redirects.

## `respectRobotsTxt` (type: `boolean`):

Skip the homepage request when the site's robots.txt disallows it for this bot (the row is still returned with DNS/RDAP data and is not charged as analysed).

## `userAgent` (type: `string`):

Custom User-Agent header. Default: an honest bot User-Agent that links to this Actor.

## Actor input object example

```json
{
  "domains": [
    "wordpress.org",
    "ruggable.com",
    "porterandyork.com",
    "ghost.org",
    "mozilla.org"
  ],
  "enrichDnsWhois": true,
  "checkSeoFiles": true,
  "minConfidence": 50,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20,
  "respectRobotsTxt": true
}
```

# Actor output Schema

## `report` (type: `string`):

Most common technologies, stacks, categories, email providers and a domain table.

## `overview` (type: `string`):

One row per domain: CMS, ecommerce, analytics, CDN, email provider and all technology names.

## `results` (type: `string`):

Full per-domain rows incl. technologies with versions, SEO signals, DNS and RDAP.

## `summary` (type: `string`):

Technology counts across all domains, most common stacks, blocked and failed domains.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "wordpress.org",
        "ruggable.com",
        "porterandyork.com",
        "ghost.org",
        "mozilla.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("prevailing_glow/tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "wordpress.org",
        "ruggable.com",
        "porterandyork.com",
        "ghost.org",
        "mozilla.org",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("prevailing_glow/tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "wordpress.org",
    "ruggable.com",
    "porterandyork.com",
    "ghost.org",
    "mozilla.org"
  ]
}' |
apify call prevailing_glow/tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,prevailing_glow/tech-stack-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Szh7WroWv3ZbJfLQP/builds/smhDABJuuLaOZm8MZ/openapi.json
