# Tech Stack Detector – Bulk Wappalyzer Alternative (`oldjard/tech-stack-detector`) Actor

Detect the tech stack of any list of websites: CMS, ecommerce platform, analytics, ad pixels, CDN, hosting, JS frameworks, payment and chat tools, from 7,600+ fingerprints. One spreadsheet row per site. $4 per 1,000 sites, failed sites free.

- **URL**: https://apify.com/oldjard/tech-stack-detector.md
- **Developed by:** [Joshua White](https://apify.com/oldjard) (community)
- **Categories:** Lead generation, Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Tech Stack Detector – Bulk Wappalyzer Alternative

**Tech Stack Detector** finds out **what any website is built with**. Paste a list of domains and get one spreadsheet
row per site: CMS, ecommerce platform, analytics and ad pixels, CDN, hosting, JavaScript frameworks, payment providers
and chat widgets, matched against 7,600+ Wappalyzer-format fingerprints. A pay-as-you-go **bulk tech stack lookup**
and **Wappalyzer / BuiltWith alternative**: no subscription, no API key.

**Try it in one click:** the input is prefilled with three sites. Press **Start** and you'll have results in under a
minute. Apify's free plan covers about 1,200 sites a month.

### What data does the tech stack lookup return?

| domain | cms | ecommerce | javascriptFrameworks | cdn / hosting | technologies |
|---|---|---|---|---|---|
| wordpress.org | WordPress 7.2 | | | | 13 |
| allbirds.com | Shopify | Shopify | | Cloudflare | 10 |
| nextjs.org | | | Next.js, React | Vercel | 10 |

Plus 21 category columns (analytics, advertising, live chat, payment processors, tag managers, web servers…), a
`technologies` array with version, confidence and the signal that found each one, and a clear error code for sites
that couldn't be scanned (not charged).

### How to detect a website's technology in 3 steps

1. Paste domains or URLs into **Websites to analyse**, one per line (or comma-separated). A whole spreadsheet column
   works; duplicates, blank lines and `www.` variants are cleaned up, and typos like `htps://` are fixed.
2. Leave **Deep scan** off for the fast, cheap scan, or turn it on to also catch analytics, ad pixels and chat tools
   that only load with JavaScript.
3. Click **Start**, then download CSV, Excel or JSON, or call it from the API, Make, Zapier or n8n.

### How much does it cost to look up a tech stack?

**$4 per 1,000 sites** scanned, so Apify's $5 monthly free credit covers about 1,250 sites with the fast scan. Deep scan adds **$3 per 1,000** sites opened in the browser. Sites that fail (no DNS,
timeouts, blocked, bot checks) are **free**. Scanning 500 domains with the fast scan costs about $2. Set a maximum
cost per run, and the actor stops cleanly when it's reached.

### Who uses a bulk technology lookup?

- **Lead lists.** Which of these 5,000 stores run Shopify? Which sites use HubSpot, Klaviyo or Intercom? Filter a
  prospect list down to the companies that use (or don't use) a given tool.
- **Competitor research.** See the analytics, A/B testing, chat and payment tools behind competitors' sites.
- **Market sizing.** Count how many sites in a niche use WordPress vs. Webflow vs. Wix.
- **Agencies and freelancers.** Qualify prospects before a call: "they're on WooCommerce with an old jQuery".
- **Security and IT inventory.** Detect server software and framework versions across your own domains.

### Fast scan or deep scan?

| | Fast scan (default) | Deep scan (`browserMode: always`) |
|---|---|---|
| How | One normal HTTP request per site | The fast scan, then the page is opened in headless Chrome |
| Finds | CMS, ecommerce platform, web server, CDN, hosting, most frameworks | All of that, plus tools that only load at runtime: analytics, ad pixels, chat, A/B testing, JS libraries |
| Technologies per site (our test set) | median 10 | median 16 |
| Speed on Apify | about 40 sites a minute | about 10 sites in 2–2.5 minutes (2 GB) |

**Auto** mode sits in between: it opens the browser only for sites where the fast scan is blocked or finds fewer than
5 technologies.

### Input example

```json
{
  "domains": ["wordpress.org", "allbirds.com", "https://www.bbc.co.uk"],
  "browserMode": "off",
  "minConfidence": 0
}
```

### Output

One row per site. The `technologyNames` text and the category columns paste cleanly into a spreadsheet, and the
`technologies` array has the full detail.

```json
{
  "input": "wordpress.org",
  "domain": "wordpress.org",
  "url": "https://wordpress.org/",
  "statusCode": 200,
  "scanMethod": "http",
  "title": "Blog Tool, Publishing Platform, and CMS – WordPress.org",
  "technologyCount": 13,
  "technologyNames": "Google Tag Manager, Gutenberg 24.1.0, HSTS, MySQL, Nginx, Open Graph, PHP, WordPress 7.2, ...",
  "technologies": [
    { "name": "WordPress", "version": "7.2", "categories": ["CMS", "Blogs"], "confidence": 100, "website": "https://wordpress.org", "detectedBy": "header" }
  ],
  "cms": "WordPress 7.2, WordPress Block Editor, WordPress Site Editor",
  "ecommerce": "",
  "analytics": "",
  "tagManagers": "Google Tag Manager",
  "webServers": "Nginx",
  "programmingLanguages": "PHP",
  "databases": "MySQL",
  "cdn": "",
  "hosting": "",
  "error": null,
  "errorCode": null,
  "scannedAt": "2026-10-05T16:13:02.114Z"
}
```

**Category columns:** `cms`, `ecommerce`, `analytics`, `tagManagers`, `advertising`, `marketingAutomation`,
`liveChat`, `paymentProcessors`, `javascriptFrameworks`, `javascriptLibraries`, `webFrameworks`,
`programmingLanguages`, `databases`, `webServers`, `cdn`, `hosting`, `security`, `cookieCompliance`, `reviews`, `fonts`, `other`.

**`detectedBy`** says which signal identified the technology: `header`, `cookie`, `html`, `script`, `css`, `meta`,
`dom`, `js` (deep scan only) or `implied` (for example, WordPress implies PHP).

**Sites that can't be scanned** still get a row, with `error` and `errorCode` filled in (`DNS`, `TIMEOUT`,
`HTTP_BLOCKED`, `BOT_CHALLENGE`, `HTTP_ERROR`, `TOO_MANY_REDIRECTS`, `ROBOTS_DISALLOWED`, `TLS`, `NETWORK`).
`BOT_CHALLENGE` means the site answered with a bot check or waiting-room page (Cloudflare, Akamai, DataDome and
similar) instead of its homepage. Error rows have no technologies (`technologyCount` is 0), because what we saw was the
block page, not the site; the error text names the bot wall when we recognise it. If the fast scan was blocked, the
error suggests trying Deep scan **Auto**, which often gets through. **You aren't charged for these rows.**

**Input lines that aren't web addresses** (`allbirds`, `not a url`) also get a row, with `errorCode: "INVALID_INPUT"`
and the line echoed in `input` and `domain`, so every line you pasted comes back. They're free, and the status counts
them: "Analysed 2 of 4 entries; 2 weren't web addresses".

The run's `OUTPUT` record in the key-value store has a summary: how many sites were scanned, loaded, failed or
skipped, and which input lines weren't valid domains.

### How it works, and how it behaves on the web

- Each site gets **one visit to its public homepage** (or the URL you gave), plus a request for its `robots.txt`. No
  crawling, no logins, nothing behind a paywall, and no personal data collected.
- It **respects robots.txt**. If a site's robots.txt disallows its homepage for automated visitors, the site is skipped
  and reported as `ROBOTS_DISALLOWED`.
- It identifies itself honestly with the user agent `TechStackDetectorBot`.
- Detection uses the open-source Wappalyzer-format fingerprints maintained by ProjectDiscovery
  ([wappalyzergo](https://github.com/projectdiscovery/wappalyzergo), MIT licence), matched against response headers,
  cookies, HTML, script URLs, inline scripts and styles, meta tags and DOM selectors, and, in deep scan, JavaScript
  globals in the live page. Fingerprints come from wappalyzergo, last refreshed 2026-10-04 (refreshed by hand, not on a
  schedule).

### Accuracy

On a test set of 38 well-known sites with publicly known stacks, the fast scan found **85%** of the expected core
technologies (CMS, platform, framework), and the deep scan found **92%**. What it can't see:

- **Back-end-only technology** that leaves no trace in the page (for example a Django site with no forms on its
  homepage).
- **Sites that block automated visitors** entirely (HTTP 403 to every client).
- **Styles in external CSS files.** CSS-only fingerprints (Tailwind CSS, for example) are matched only against styles
  inlined in the page.

As with any fingerprint-based detector, low-confidence hints can occasionally be wrong. Set **Minimum confidence** to
50 or more if you want only strong matches.

### Ready-made examples

Each one opens this actor with the input already filled in. Click **Try** to run it, or change the input to fit your own list.

- [Is it WordPress? Bulk CMS checker](https://apify.com/oldjard/tech-stack-detector/examples/is-it-wordpress-cms-checker)
- [Check which Shopify stores use Klaviyo](https://apify.com/oldjard/tech-stack-detector/examples/shopify-stores-using-klaviyo)
- [Which ecommerce platform does a store use?](https://apify.com/oldjard/tech-stack-detector/examples/ecommerce-platform-checker)
- [Tech stack of SaaS competitors](https://apify.com/oldjard/tech-stack-detector/examples/saas-competitor-tech-stack)
- [Check Google Analytics and Tag Manager on websites](https://apify.com/oldjard/tech-stack-detector/examples/google-analytics-tag-manager-check)
- [Detect Next.js and React websites](https://apify.com/oldjard/tech-stack-detector/examples/detect-nextjs-react-sites)

### More tools from oldjard

- [Sitemap URL Extractor](https://apify.com/oldjard/sitemap-url-extractor): every URL on a website, for RAG and SEO.
- [Shopify Products Scraper & Price Monitor](https://apify.com/oldjard/shopify-products-price-monitor): catalogs and price changes from any Shopify store.
- [Workday, Greenhouse, Lever & Ashby Jobs Scraper](https://apify.com/oldjard/ats-career-site-jobs): every open job from company career sites.
- [Bulk Website Screenshot & URL to PDF](https://apify.com/oldjard/screenshot-pdf): screenshots and PDFs of any list of pages.
- [AI Web Scraper (your own key)](https://apify.com/oldjard/ai-web-scraper): describe fields in English, get JSON.
- [Website Change Monitor](https://apify.com/oldjard/website-change-monitor): a before/after diff by webhook, Slack or Discord when a page changes.
- [Company Registry Lookup](https://apify.com/oldjard/company-registry-lookup): UK Companies House, Spain, France, Finland and Norway in one schema.
- [UK & EU Public Tenders](https://apify.com/oldjard/uk-eu-public-tenders): Find a Tender and TED notices in one table, with daily only-new alerts.

### Use it from an AI agent or the API

- **Minimal input:** `{"domains": ["example.com"]}`. Everything else has a sensible default.
- **Cost:** $0.004 per website analysed, plus $0.003 per site opened in the browser (`browserMode` `auto` or
  `always`). Sites that fail to load are free. 100 sites with the fast scan ≈ $0.40.
- **Run time (our runs):** 10 sites with the fast scan in 6–24 s. Deep scan works at the default 1 GB but opens one
  page at a time there; give the run 2 GB or more to render several sites at once.
- **Results:** the default dataset, one row per site (a `technologies` array plus category columns); the run summary
  is the `OUTPUT` record in the key-value store.
- Works over the Apify MCP server (`search-actors`, then `call-actor`) and is eligible for agentic payments (x402).

### FAQ

**Is this the official Wappalyzer or BuiltWith?** No. It's an independent tool that uses an open-source,
Wappalyzer-compatible fingerprint set. It isn't affiliated with Wappalyzer or BuiltWith.

**How is this different from BuiltWith?** BuiltWith keeps a historical database you search by technology. This actor
checks the live site right now, from a list you give it, and charges per site with no subscription.

**Does it look up historical technology?** No, it reports what a site runs right now.

**Can I scan specific pages, not just homepages?** Yes, give a full URL such as `https://example.com/pricing`.

**How many domains per run?** As many as you like. Large lists run at up to 50 sites in parallel (5 in the browser).

**Can I run it on a schedule?** Yes. Save your input as a task and schedule it in Apify to track stack changes over
time.

### Feedback

Found a technology it misses or gets wrong? Open an issue on the Issues tab with the domain and what you expected. It
helps us improve detection for everyone.

# Actor input Schema

## `domains` (type: `array`):

Domains or URLs, one per line (e.g. shopify.com or https://www.example.com/shop). You can paste a whole spreadsheet column; duplicates and blank lines are removed automatically. A bare domain scans its homepage.

## `browserMode` (type: `string`):

off (default): one HTTP request per site; fastest and cheapest. Finds the CMS, ecommerce platform, server, CDN and most frameworks. auto: also opens the site in headless Chrome when the plain request is blocked or finds fewer than 5 technologies. always: opens every site in Chrome, which also finds analytics, ad pixels, chat and A/B tools (about 50% more technologies). Each browser-opened site adds a $0.003 browser-render charge. Works at 1 GB memory (one page at a time); 2 GB or more is faster.

## `minConfidence` (type: `integer`):

Hide technologies detected with lower confidence than this. 0 shows everything; 50 hides weak single hints.

## `maxConcurrency` (type: `integer`):

How many sites to scan at the same time. Each site still gets only one visit (plus robots.txt). At most 5 pages render in the browser at once.

## `requestTimeoutSecs` (type: `integer`):

Give up on a site that hasn't answered within this many seconds. It's reported as a timeout and not charged.

## `canaryExpectations` (type: `object`):

Internal monitoring only; leave empty. Map of domain to technologies that must be found, e.g. {"wordpress.org": \["WordPress"]}. The run fails if more than 20% of the checks miss.

## Actor input object example

```json
{
  "domains": [
    "wordpress.org",
    "allbirds.com",
    "https://www.bbc.co.uk"
  ],
  "browserMode": "off",
  "minConfidence": 0,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset, one row per website: input, domain, url, statusCode, scanMethod (http/browser), title, technologyCount, technologyNames, a technologies array (name, version, confidence, categories, evidence), category columns (cms, ecommerce, analytics, javascriptFrameworks, cdn, hosting, ...), error and errorCode. Read it with GET {url}?format=json or csv.

## `summary` (type: `string`):

Run summary (JSON): sites scanned, loaded, failed and rendered in a browser, invalid inputs, duplicates removed, and why the run stopped early (if it did).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "wordpress.org",
        "allbirds.com",
        "nextjs.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("oldjard/tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "wordpress.org",
        "allbirds.com",
        "nextjs.org",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("oldjard/tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "wordpress.org",
    "allbirds.com",
    "nextjs.org"
  ]
}' |
apify call oldjard/tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,oldjard/tech-stack-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8NOWRpcrgcRcV3xQj/builds/c8pumaZzjlsQF9lyf/openapi.json
