# Website Tech Stack Detector & Monitor - Wappalyzer Alternative (`neverempty/website-technology-detector`) Actor

For lead qualification, sales prospecting and competitor tracking: the tech stack of 1 to 1,000 websites per run - CMS, ecommerce, analytics, ads, CDN, hosting, frameworks, payments - with version and the evidence that matched. Monitor: only sites whose stack changed. Unreadable sites are free.

- **URL**: https://apify.com/neverempty/website-technology-detector.md
- **Developed by:** [NeverEmpty](https://apify.com/neverempty) (community)
- **Categories:** Lead generation, Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $21.00 / 1,000 site analyzeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Tech Stack Detector & Monitor - Wappalyzer Alternative

Give it a list of websites, get back **what each one is built with**: CMS, ecommerce platform, analytics, tag managers, advertising and retargeting pixels, marketing automation and CRM, CDN, hosting, web server, programming language, JavaScript frameworks, payment processors, live chat, email providers and more, one clean JSON row per site.

It works like a Wappalyzer or BuiltWith lookup, but as an API you can run on a list of 1 to 1,000 sites at a time, and it can **monitor** your list: after the first run, it returns only the sites that **added or removed a technology**.

Built for: lead qualification and enrichment (find every store on Shopify, every site on WordPress, every company using HubSpot), sales prospecting by technology, competitor tracking (notice when a competitor moves to a new platform, analytics tool or payment provider), agency audits, and market research on technology adoption.

- **3,588 technology fingerprints** from the open Wappalyzer project (the last release under the MIT license), matched against the page's HTTP headers, cookies, HTML, `<meta>` tags, `<script src>` URLs, inline scripts and styles, visible text, DOM elements and the domain's DNS records. 3,213 of them can be detected without a browser.
- **Evidence for every detection.** Each technology comes with its categories, version (when the site reveals it), confidence and `detectedBy`, the signals that matched (for example `header:server`, `cookie:_shopify_y`, `meta:generator`, `scriptSrc`, `dns:MX`, `impliedBy:WordPress`).
- **Ready-made columns** for the categories people filter on most: `cms`, `ecommerce`, `analytics`, `tagManagers`, `advertising`, `marketingAutomation`, `cdn`, `hosting`, `webServers`, `programmingLanguages`, `javascriptFrameworks`, `paymentProcessors`, `liveChat`, `emailServices`, `security`, plus every category in `categories`.
- **Monitor mode** returns only sites whose technology list changed, with `addedTechnologies` and `removedTechnologies`. A change is reported only after it has been seen in **two runs in a row** (and in two reads within each run). Tags that come and go (A/B tests, rotating ad tags, bot-protection cookies that are set on some visits only) are therefore usually not reported, and once a technology has been seen to change and change back it is ignored for that site. A tag that happens to show up in two runs in a row can still be reported once. Changes seen for the first time are listed in `unconfirmedChanges`.
- **Nothing is charged for a site that could not be analyzed.** A site that does not exist, refuses automated reading, shows a check page, needs a sign-in, or where no fingerprint matched comes back as a free row that says why.
- **Respects each site's robots.txt** (RFC 9309) for every URL and every redirect, and never signs in, solves check pages or rotates proxies to get around a refusal.
- **Reads real pages, not a cached database.** Every run fetches the site as it is now (up to 2 MB of HTML per page).

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `urls` | array of strings | (empty: the example `https://wordpress.org/` is used) | Websites to analyze, 1 to 1,000 per run. A plain domain such as `example.com` is read as `https://example.com/`. Each line is one page, usually the home page. The same URL given twice is read and charged once. |
| `includeDns` | boolean | `true` | Also read the domain's public MX, TXT, NS, SOA and CNAME records (email provider, DNS host, verified services). |
| `onlyChanges` | boolean | `false` | Monitor mode: return only sites where a technology was added or removed since the last run with the same `watchName`, plus sites new to the watch. |
| `watchName` | string | (none) | Name of the remembered state (letters, digits, `.`, `-`, `_`; up to 40). Setting it fills `changeType`, `addedTechnologies` and `removedTechnologies`. With monitor mode on and no name, `default` is used. |
| `resetMonitoringState` | boolean | `false` | Forget what this watch remembered, so every site is returned as a first check. |
| `maxConcurrency` | integer | `4` | Sites read in parallel, 1 to 8. |
| `requestTimeoutSecs` | integer | `20` | Time limit for one request, 5 to 60 seconds. HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a site that does not answer in time is asked once more. |

```json
{
  "urls": ["wordpress.org", "https://www.allbirds.com/", "vercel.com"]
}
```

### Output

One row per site. A real row from 2026-09-25 (the `technologies` list and `categories` are shortened here to 3 and 4 entries; the full row had 12 technologies):

```json
{
  "status": "ok",
  "position": 1,
  "url": "https://www.allbirds.com/",
  "finalUrl": "https://www.allbirds.com/",
  "httpStatus": 200,
  "contentType": "text/html; charset=utf-8",
  "redirects": [],
  "technologyCount": 12,
  "technologyNames": [
    "Apple iCloud Mail",
    "Apple Pay",
    "Cloudflare",
    "DocuSign",
    "Google Tag Manager",
    "HSTS",
    "HTTP/3",
    "Microsoft 365",
    "Open Graph",
    "PayPal",
    "Priority Hints",
    "Shopify"
  ],
  "technologies": [
    {
      "name": "Cloudflare",
      "slug": "cloudflare",
      "categories": [
        "CDN"
      ],
      "version": null,
      "confidence": 100,
      "website": "http://www.cloudflare.com",
      "detectedBy": [
        "header:cf-cache-status",
        "header:cf-ray",
        "header:server"
      ]
    },
    {
      "name": "Microsoft 365",
      "slug": "microsoft-365",
      "categories": [
        "Webmail",
        "Email"
      ],
      "version": null,
      "confidence": 100,
      "website": "https://www.microsoft.com/microsoft-365",
      "detectedBy": [
        "dns:MX"
      ]
    },
    {
      "name": "Shopify",
      "slug": "shopify",
      "categories": [
        "Ecommerce"
      ],
      "version": null,
      "confidence": 100,
      "website": "http://shopify.com",
      "detectedBy": [
        "cookie:_shopify_s",
        "cookie:_shopify_y",
        "meta:shopify-digital-wallet",
        "scriptSrc"
      ]
    }
  ],
  "categories": {
    "Webmail": [
      "Apple iCloud Mail",
      "Microsoft 365"
    ],
    "Payment processors": [
      "Apple Pay",
      "PayPal"
    ],
    "CDN": [
      "Cloudflare"
    ],
    "Miscellaneous": [
      "DocuSign",
      "HTTP/3",
      "Open Graph"
    ]
  },
  "cms": [],
  "ecommerce": [
    "Shopify"
  ],
  "analytics": [],
  "tagManagers": [
    "Google Tag Manager"
  ],
  "advertising": [],
  "marketingAutomation": [],
  "cdn": [
    "Cloudflare"
  ],
  "hosting": [],
  "webServers": [],
  "programmingLanguages": [],
  "javascriptFrameworks": [],
  "paymentProcessors": [
    "Apple Pay",
    "PayPal"
  ],
  "liveChat": [],
  "emailServices": [
    "Apple iCloud Mail",
    "Microsoft 365"
  ],
  "security": [
    "HSTS"
  ],
  "pageTitle": "Allbirds: Comfortable, Sustainable Shoes & Apparel",
  "server": "cloudflare",
  "poweredBy": null,
  "htmlBytesRead": 665022,
  "htmlTruncated": false,
  "dnsChecked": true,
  "dnsFailedTypes": [],
  "unconfirmedChanges": null,
  "changeType": null,
  "addedTechnologies": null,
  "removedTechnologies": null,
  "previousCheckedAt": null,
  "watchName": null,
  "note": null,
  "fetchedAt": "2026-09-24T18:08:55.288Z"
}
```

#### Columns

| Column | What it holds |
|---|---|
| `technologies` | Every detected technology: `name`, `slug`, `categories`, `version` (null when the site does not reveal it), `confidence` (0 to 100), `website` and `detectedBy`. |
| `technologyNames`, `technologyCount` | The names alone, and how many. |
| `categories` | Category name to the technologies in it, for every category found. |
| `cms` … `security` | Technology names in commonly used categories (see the list above). `ecommerce` lists platforms only; the generic "Cart Functionality" marker (a link to a cart) stays in `technologies` and `categories`. |
| `pageTitle`, `server`, `poweredBy` | The page's `<title>` and the raw `Server` and `X-Powered-By` headers. |
| `finalUrl`, `httpStatus`, `redirects` | Where the URL ended up after redirects, and each redirect step. |
| `htmlBytesRead`, `htmlTruncated` | How much HTML was read; `true` when the page was larger than 2 MB and only its first 2 MB were read. |
| `dnsChecked`, `dnsFailedTypes` | Whether DNS records were part of the check, and which record types (for example `TXT`) could not be read this time even after asking again (an empty list when all answered; a record type that does not exist counts as answered). |
| `changeType`, `addedTechnologies`, `removedTechnologies`, `previousCheckedAt` | Monitor mode: `first-check`, `new`, `changed` or `unchanged`, and the confirmed changes (seen in this run and the run before). |
| `unconfirmedChanges` | Monitor mode: changes seen for the first time in this run, such as `+Stripe` or `-Nginx`. They are reported as changes if the next run sees them too; if they go back, the technology is ignored for that site from then on. |

#### Rows that are not charged

| `status` | Meaning |
|---|---|
| `no-technologies-detected` | The page was read, but no fingerprint matched. |
| `robots-disallowed` | The site's robots.txt does not allow automated reading of this address; it was not requested. |
| `robots-unreachable` | The site's robots.txt answered 5xx or could not be reached; following RFC 9309 the page was not requested. |
| `login-required` | HTTP 401, or the page redirects to a sign-in page. |
| `blocked` | The site refused this reader (HTTP 403, 429 after retries, 451) or showed a CAPTCHA or browser-check page. A check page is never asked again and no technology is claimed for it. |
| `not-found` | HTTP 404 or 410. |
| `http-error` | Another HTTP answer that is not a page, or a broken redirect. |
| `unreachable` | The domain does not resolve, or the site did not answer. |
| `unreadable` | 5xx after retries, an answer that kept stopping before the end of the page, or more than 8 redirects. |
| `not-html` | The address is a PDF, image, JSON or other file, not a web page. |
| `bad-input` | Not an http(s) URL, a URL with a user name or password, or an address in a private network. |
| `no-change` | Monitor mode: nothing changed (one summary row), or a change could not be confirmed by a second read. |
| `budget-reached` | The run reached the maximum total charge you set; the rest was not requested. |

These rows keep the same columns as a technology row, with `note` explaining the reason.

### What it can and cannot see

- It reads the page the way a plain HTTP client does and **does not run the page's JavaScript**. 375 of the 3,588 fingerprints can only be seen in a browser (JavaScript globals, background requests) or need the TLS certificate; those technologies are not reported, and a missing technology is not proof that the site does not use it.
- Tools that the page loads later from another script are not seen. The most common case is Google Analytics loaded through Google Tag Manager: in a check of 8 sites against a real browser on 2026-09-25, every technology the row reported that the browser could confirm was there, and the ones missed were Google Analytics (5 sites) and Google Tag Manager (2 sites) loaded that way.
- Some technologies come from DNS alone, for example `Cloudflare` from `dns:NS` means the domain's DNS is hosted at Cloudflare, which usually but not always means the site is served through it. `detectedBy` tells you which signal it was.
- One URL is one page. Tools that appear only on a checkout, blog or login page are found when you give that page's URL.
- DNS TXT records often hold verification codes (for example for Google, Apple, Stripe or DocuSign). They show that the service was set up for the domain; they are marked `dns:TXT` in `detectedBy` so you can keep or drop them. Turn off `includeDns` to use the page only.
- In monitor mode, when a page is larger than 2 MB or some DNS records could not be read, removals are not counted for that check (what was not read may still contain the technology) and the remembered list keeps them.
- In monitor mode a real change appears one run after it happens (it has to be seen twice). Schedule the watch as often as you need that delay to be short.
- Do not run two schedules with the same `watchName` at the same time: both compare against the same remembered state, so the same change can be returned by both. Use one `watchName` per list and per schedule.
- If a monitored page or its DNS could only be read partly on the first check, the next fully read check quietly updates the starting point instead of reporting the difference as a change.

### Pricing

Pay per event:

- **Site analyzed**: charged for each row with `status: "ok"` (at least one technology detected).
- **Run start**: charged once per run that returns at least one site. In monitor mode it is charged once per run that read and compared at least one site, whether or not anything changed, because the check itself is the service. It is not charged when no site could be analyzed or the input could not be used.

A run whose maximum total charge has no room for the run start fee plus one site requests nothing and is charged nothing. The current prices are on the Pricing tab.

### Calling it from code

```bash
curl -X POST "https://api.apify.com/v2/acts/neverempty~website-technology-detector/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["wordpress.org", "vercel.com"]}'
```

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('neverempty/website-technology-detector').call({ urls: ['wordpress.org', 'vercel.com'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const site of items) console.log(site.url, site.cms, site.ecommerce, site.analytics);
```

To track a list of competitors or accounts, save a task with `onlyChanges: true` and a `watchName`, and schedule it daily or weekly. Each run then returns only the sites whose stack changed.

### Fingerprint data and license

The technology fingerprints are taken from the Wappalyzer project at its last commit released under the MIT license (25 January 2023), Copyright 2008 Wappalyzer, used under the MIT license. Nothing from later, differently licensed releases is used. This Actor is not affiliated with, endorsed by or sponsored by Wappalyzer or BuiltWith; the names are used only to describe what this Actor is comparable to.

### Limits

- Up to 1,000 URLs per run, one page per URL, up to 2 MB of HTML per page.
- Sites behind a sign-in, a CAPTCHA or a browser check, and sites whose robots.txt disallows automated reading, are not analyzed (free rows).
- The fingerprints date from January 2023, so technologies launched after that are not in the list.

### Support

Found a wrong or missing detection, or need another column? Open an issue on the Issues tab with the URL and what you expected.

# Actor input Schema

## `urls` (type: `array`):

Websites to analyze, one per line (1 to 1,000 per run). A domain such as example.com or a URL without http:// or https:// is read as https://. Each line is one page (usually the home page); the same URL given twice is read and charged once. If this is empty, the example site https://wordpress.org/ is used.

## `includeDns` (type: `boolean`):

Also read the domain's public MX, TXT, NS, SOA and CNAME records, which reveal services such as Google Workspace, Microsoft 365, SendGrid or the DNS host. Detections from DNS are marked dns:MX, dns:TXT and so on in detectedBy (TXT verification records show that a service was set up for the domain, not that it is visible on the page). Turn off to use the page only.

## `onlyChanges` (type: `boolean`):

Return only sites where a technology was added or removed since the last run with the same watch name (plus sites new to the watch). The first run returns every site as the starting point. A change is reported once it has been seen in two runs in a row (and in two reads within each run); changes seen for the first time are listed in unconfirmedChanges. A technology that comes and goes (for example an A/B-tested tag) is not reported and is ignored for that site from then on. A run in which nothing changed returns a free row saying so and charges only the run start fee.

## `watchName` (type: `string`):

Name of the remembered state used to compare runs (letters, digits, dot, dash, underscore; up to 40). Setting it (or turning on monitor mode) fills changeType, addedTechnologies and removedTechnologies. Use a different name for each list of sites you track on its own schedule. With monitor mode on and no name, the name "default" is used.

## `resetMonitoringState` (type: `boolean`):

Start this watch over: forget the remembered technologies before this run, so every site is returned as a first check.

## `maxConcurrency` (type: `integer`):

How many sites are read in parallel (1 to 8).

## `requestTimeoutSecs` (type: `integer`):

How long to wait for one site to answer (5 to 60 seconds). HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a site that does not answer in time is asked once more.

## Actor input object example

```json
{
  "urls": [
    "https://wordpress.org/",
    "https://www.allbirds.com/"
  ],
  "includeDns": true,
  "onlyChanges": false,
  "resetMonitoringState": false,
  "maxConcurrency": 4,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

One row per site: every detected technology with its categories, version, confidence and the signals that matched (header, cookie, meta tag, script URL, HTML, DOM, DNS), grouped by category, plus ready columns for CMS, ecommerce, analytics, advertising, CDN, hosting, frameworks, payments, live chat and email. In monitor mode, the technologies added and removed. A site that robots.txt does not allow, that needs a sign-in, that shows a check page or that cannot be read comes back as a free row that says why.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://wordpress.org/",
        "https://www.allbirds.com/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neverempty/website-technology-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://wordpress.org/",
        "https://www.allbirds.com/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("neverempty/website-technology-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://wordpress.org/",
    "https://www.allbirds.com/"
  ]
}' |
apify call neverempty/website-technology-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neverempty/website-technology-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WR32pzqh9kTZaebdt/builds/WrxbdYDSTNiPzUWgs/openapi.json
