# Tech Stack Detector - Wappalyzer & BuiltWith Alternative (`zenomastro/tech-stack-detector`) Actor

Detect the CMS, e-commerce platform, analytics, tag managers, frameworks, CDN and thousands of other technologies on any list of websites using open-source fingerprints, with category, version and confidence for each. From $3 per 1,000 websites, platform usage included.

- **URL**: https://apify.com/zenomastro/tech-stack-detector.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** Developer tools, SEO tools, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.40 / 1,000 analysed websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Tech Stack Detector do?

Tech Stack Detector finds the technologies behind any list of websites, like Wappalyzer or BuiltWith but in bulk and through an API. Paste URLs or bare domains and get one row per website with the CMS, e-commerce platform, analytics and tag managers, JavaScript frameworks, web server, CDN, payment processors, marketing tools, email provider and more. Each technology comes with its categories, detected version and a confidence score.

Detection uses the open-source Wappalyzer fingerprint set (about 3,600 technologies in 108 categories, last MIT-licensed release). It matches HTTP headers, cookies, meta tags, script URLs, inline scripts and styles, HTML and DOM patterns, plus public DNS records (MX, TXT, NS, CNAME, SOA). Each website gets a single, lightweight HTTP request. No browser, no login, no proxies to manage.

### Who is it for?

- **Sales and lead generation teams** who want to qualify or segment accounts by technology, e.g. every Shopify store or every company using HubSpot in a prospect list.
- **Agencies and freelancers** looking for migration or redesign leads, e.g. sites still on an old WordPress or Magento version.
- **SEO and marketing teams** auditing analytics, tag managers, CDNs and performance tools across client or competitor sites.
- **Developers and data teams** who need a Wappalyzer-style API for enrichment pipelines, CRMs or AI agents.
- **Market researchers** measuring technology adoption across thousands of domains.

### Output fields

| Field | Description |
|---|---|
| `url` | Normalized website URL that was analysed (bare domains get `https://`). |
| `finalUrl` | URL after redirects. |
| `status` | `COMPLETE`, `PARTIAL`, `VALID_EMPTY`, `INVALID_INPUT` or `UPSTREAM_FAILED` (see below). |
| `httpStatus` | HTTP status code of the final response. |
| `technologies` | List of `{name, slug, categories, version, confidence, website, detectedBy}`. |
| `technologyNames` | Technology names only, handy for CSV/Excel filters. |
| `categories` | Summary: category name to technologies, e.g. `{"Ecommerce": ["Shopify"]}`. |
| `categoryNames` | Category names found on the website. |
| `technologyCount` | Number of detected technologies. |
| `headers` | Technology-relevant response headers (`server`, `x-powered-by`, `via`, `x-cache`, `content-type`...). |
| `pageTitle` | Title of the analysed page. |
| `domain`, `redirectCount`, `htmlTruncated`, `responseTimeMs`, `error`, `analyzedAt` | Diagnostics. |

**Row status values:**

- `COMPLETE`: the page loaded and at least one technology was detected. Charged.
- `PARTIAL`: technologies were detected but part of the evidence was missing, e.g. the HTML was larger than the size limit, the response was not HTML, or a redirect was not followed by choice. Charged.
- `VALID_EMPTY`: the page loaded but no known fingerprint matched. Free.
- `INVALID_INPUT`: the entry is not a public http(s) URL or domain. Free.
- `UPSTREAM_FAILED`: timeout, DNS error, too many redirects or an HTTP 4xx/5xx answer, for example a bot wall. Header and cookie fingerprints are still listed when available. Free.

### Pricing at a glance

| | Price per 1,000 analysed websites |
|---|---:|
| **This Actor** | **$3.00** |
| Median of comparable Apify Store tech-stack Actors with a per-result price (October 2026) | $5.25 |

**Cost example:** you upload 1,000 domains. 850 return technologies (`COMPLETE` or `PARTIAL`), 100 have no detectable technology (`VALID_EMPTY`) and 50 are offline (`UPSTREAM_FAILED`). You pay 850 × $0.003 = **$2.55**, plus the run start fee of $0.0001 per GB of run memory. Empty, failed and invalid rows cost nothing.

**Subscriber discounts:** on a paid Apify plan you pay less per website: Bronze −10%, Silver −15%, Gold and higher −20% (**$2.40 per 1,000 websites** on Gold).

**Apify platform usage (compute, storage, data transfer) is included.** You pay only the per-website price plus the small start fee.

Use the run's *maximum charge* setting to cap spending. The Actor stops cleanly when the limit is reached.

### How to use

1. Open the Actor and paste your websites into **Websites (URLs or domains)**, one per line. `example.com`, `www.example.com` and `https://example.com/shop` all work.
2. Optionally change **Parallel websites**, **Request timeout**, **Follow redirects**, **Check DNS records** or **Minimum confidence**.
3. Click **Start**. The three prefilled websites finish in a few seconds.
4. Open the **Output** tab to see the results table, or export as JSON, CSV, Excel or HTML. Filter `status = COMPLETE` and `technologyNames` to build lead lists.

#### Input example

```json
{
  "urls": ["wordpress.org", "https://www.allbirds.com/", "https://apify.com/"],
  "maxConcurrency": 5,
  "requestTimeoutSecs": 15,
  "followRedirects": true,
  "includeDnsRecords": true,
  "minConfidence": 0
}
```

### Output example

```json
{
  "recordType": "site",
  "url": "https://wordpress.org/",
  "finalUrl": "https://wordpress.org/",
  "domain": "wordpress.org",
  "status": "COMPLETE",
  "httpStatus": 200,
  "technologyCount": 12,
  "technologies": [
    {
      "name": "WordPress",
      "slug": "wordpress",
      "categories": ["CMS", "Blogs"],
      "version": "7.2",
      "confidence": 100,
      "website": "https://wordpress.org",
      "detectedBy": ["headers", "html", "meta", "scriptSrc"]
    },
    {
      "name": "Nginx",
      "slug": "nginx",
      "categories": ["Web servers", "Reverse proxies"],
      "version": null,
      "confidence": 100,
      "website": "http://nginx.org/en",
      "detectedBy": ["headers"]
    }
  ],
  "technologyNames": ["WordPress", "MySQL", "PHP", "C", "Gutenberg", "Nginx", "Google Tag Manager", "..."],
  "categories": {
    "CMS": ["WordPress"],
    "Tag managers": ["Google Tag Manager"],
    "Web servers": ["Nginx"]
  },
  "headers": { "server": "nginx", "content-type": "text/html; charset=UTF-8" },
  "pageTitle": "Blog Tool, Publishing Platform, and CMS – WordPress.org",
  "responseTimeMs": 1402,
  "error": null
}
```

`detectedBy` shows which evidence matched: `headers`, `cookies`, `meta`, `scriptSrc`, `scripts`, `html`, `dom`, `css`, `text`, `url`, `dns`, or `implied` when a technology is implied by another one (for example WordPress implies PHP and MySQL).

### API

One HTTP call runs the Actor and returns the results directly:

```bash
curl -X POST "https://api.apify.com/v2/acts/zenomastro~tech-stack-detector/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["wordpress.org", "https://www.allbirds.com/"]}'
```

The same works from the Apify Python/JavaScript clients, Make, n8n, Zapier, Clay-style enrichment flows, LangChain and LlamaIndex integrations, or on a schedule.

### Use with AI agents (MCP)

Claude, ChatGPT, Cursor, VS Code, n8n AI agents and any MCP client can run this Actor through the official Apify MCP server. Add `https://mcp.apify.com?tools=zenomastro/tech-stack-detector` to your client, then ask for example:

> *"Which of these 50 domains run Shopify, and which analytics and payment tools do they use?"*

The agent fills the input from the field descriptions and receives clean JSON it can reason over.

### Why use this Actor?

Detect the CMS, e-commerce platform, analytics, tag managers, frameworks, CDN and thousands of other technologies on any list of websites using open-source fingerprints, with category, version and confidence for each. From $3 per 1,000 websites, platform usage included.

### Features

- **Websites (URLs or domains)** — Websites to analyse, one per line. Full URLs (https://www.example.com/shop) or bare domains (example.com) both work; bare domains are tried over HTTPS first, then HTTP. Duplicates are skipped.
- **Parallel websites** — How many different websites are analysed at the same time. Each website receives only one page request (plus redirects), so this stays polite even at the maximum.
- **Request timeout (seconds)** — Maximum seconds to wait for each website response before it is reported as UPSTREAM\_FAILED (not charged).
- **Follow redirects** — Follow HTTP redirects (e.g. example.com to www.example.com) and analyse the final page. When off, only the first response headers and cookies are analysed.
- **Check DNS records** — Also match public DNS records (MX, TXT, NS, CNAME, SOA) to detect email providers, DNS hosts and verification-based services such as Google Workspace or Microsoft 365.
- **Minimum confidence (%)** — Only report technologies whose fingerprint confidence is at least this value (0-100). 0 keeps every detection; 50 removes weak single-signal guesses.
- **Maximum websites per run** — Safety cap on unique websites analysed in one run, e.g. 1000. Use it together with the run's maximum charge to keep cost predictable.
- **Maximum redirect hops** — Maximum number of redirects followed per website before it is reported as UPSTREAM\_FAILED.
- **Retries** — How many times a website is retried after a network error, timeout or HTTP 429/502/503/504 (with a short backoff).
- **Maximum HTML size (KB)** — Only the first part of very large pages is downloaded and analysed. Rows whose HTML was cut are marked PARTIAL.
- **User agent** — HTTP User-Agent header sent to websites. Leave empty to use the default browser-compatible user agent that identifies this Actor.

### Use cases

- Sales lead qualification by technology.
- Competitor tech stack research.
- Agency prospecting and migration targeting.
- Website and marketing-stack audits.

### Related tools

- [Schema Markup Extractor - JSON-LD, Microdata & RDFa](https://apify.com/zenomastro/schema-markup-extractor-pro)
- [Website Screenshot API - Full Page PNG, JPEG & PDF](https://apify.com/zenomastro/website-screenshot-pro)
- [Sitemap URL Extractor](https://apify.com/zenomastro/sitemap-url-extractor-pro)

### Example input

```json
{
  "urls": [
    "https://wordpress.org/",
    "https://www.allbirds.com/",
    "https://apify.com/"
  ],
  "maxConcurrency": 5,
  "requestTimeoutSecs": 15,
  "followRedirects": true,
  "includeDnsRecords": true,
  "minConfidence": 0
}
```

### Pricing & cost control

The primary event costs **$0.003000 per analysed website** (about **$3.00 per 1,000** successful primary events).
Only successful primary events are intentionally billed by this Actor; summary/status rows add context without adding primary-event charges.

Use the bounded input limits and filters to keep both event charges and platform usage predictable.

### FAQ

**What is this Actor for?**\
It is designed for sales lead qualification by technology, competitor tech stack research, agency prospecting and migration targeting.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

### More questions

**Does it run JavaScript like a browser?**\
No. The Actor reads the server response (headers, cookies, HTML, scripts, DOM and DNS), which is fast and cheap and finds the large majority of CMS, e-commerce, analytics, tag manager, CDN and server technologies. Technologies that only appear after client-side JavaScript runs, and fingerprints based on JavaScript globals, are not evaluated.

**How accurate is the confidence score?**\
Each fingerprint carries a weight from the open-source dataset; matches are added up to 100. Most detections are 100. Use **Minimum confidence** (e.g. 50) to drop weak single-signal guesses.

**Why is a website `UPSTREAM_FAILED`?**\
The site timed out, does not resolve, or answered with an error such as 403 from a bot-protection wall. These rows are never charged. Retry later or with a longer timeout.

**Is this the Wappalyzer or BuiltWith API?**\
No. It is an independent alternative that uses the last MIT-licensed Wappalyzer open-source fingerprints. It is not affiliated with or endorsed by Wappalyzer or BuiltWith.

### Fingerprint data and license

Technology fingerprints and the matching engine are vendored unmodified from the npm package `wappalyzer@6.10.54`, the last release published under the MIT License (later releases moved to GPL-3.0 and are not used). The MIT license notice and provenance details ship with the Actor in `vendor/wappalyzer/LICENSE` and `vendor/wappalyzer/NOTICE.md`.

### Responsible use

The Actor sends a single page request per website, with polite concurrency, timeouts and an identifying user agent. It reads only publicly served technical signals, never logs in, and does not collect personal data. Private, local and reserved network addresses are blocked. Follow the terms of the websites you analyse.

# Changelog

This Actor's version history is a separate document: https://apify.com/zenomastro/tech-stack-detector/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Websites to analyse, one per line. Full URLs (https://www.example.com/shop) or bare domains (example.com) both work; bare domains are tried over HTTPS first, then HTTP. Duplicates are skipped.

## `maxConcurrency` (type: `integer`):

How many different websites are analysed at the same time. Each website receives only one page request (plus redirects), so this stays polite even at the maximum.

## `requestTimeoutSecs` (type: `integer`):

Maximum seconds to wait for each website response before it is reported as UPSTREAM\_FAILED (not charged).

## `followRedirects` (type: `boolean`):

Follow HTTP redirects (e.g. example.com to www.example.com) and analyse the final page. When off, only the first response headers and cookies are analysed.

## `includeDnsRecords` (type: `boolean`):

Also match public DNS records (MX, TXT, NS, CNAME, SOA) to detect email providers, DNS hosts and verification-based services such as Google Workspace or Microsoft 365.

## `minConfidence` (type: `integer`):

Only report technologies whose fingerprint confidence is at least this value (0-100). 0 keeps every detection; 50 removes weak single-signal guesses.

## `maxItems` (type: `integer`):

Safety cap on unique websites analysed in one run, e.g. 1000. Use it together with the run's maximum charge to keep cost predictable.

## `maxRedirects` (type: `integer`):

Maximum number of redirects followed per website before it is reported as UPSTREAM\_FAILED.

## `retries` (type: `integer`):

How many times a website is retried after a network error, timeout or HTTP 429/502/503/504 (with a short backoff).

## `maxHtmlKb` (type: `integer`):

Only the first part of very large pages is downloaded and analysed. Rows whose HTML was cut are marked PARTIAL.

## `userAgent` (type: `string`):

HTTP User-Agent header sent to websites. Leave empty to use the default browser-compatible user agent that identifies this Actor.

## Actor input object example

```json
{
  "urls": [
    "https://wordpress.org/",
    "https://www.allbirds.com/",
    "https://apify.com/"
  ],
  "maxConcurrency": 5,
  "requestTimeoutSecs": 15,
  "followRedirects": true,
  "includeDnsRecords": true,
  "minConfidence": 0,
  "maxItems": 1000,
  "maxRedirects": 6,
  "retries": 1,
  "maxHtmlKb": 2048
}
```

# Actor output Schema

## `overview` (type: `string`):

Main fields of every result, ready to preview or export as CSV, Excel or JSON.

## `results` (type: `string`):

Every result with all fields, for APIs, AI agents and integrations.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://wordpress.org/",
        "https://www.allbirds.com/",
        "https://apify.com/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://wordpress.org/",
        "https://www.allbirds.com/",
        "https://apify.com/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://wordpress.org/",
    "https://www.allbirds.com/",
    "https://apify.com/"
  ]
}' |
apify call zenomastro/tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/tech-stack-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fY5QG4NEbizvn1ENF/builds/RBVKuKA80lDM29pkV/openapi.json
