# Tech Stack & Framework Detector (CMS, Next.js, React) (`thodor/tech-stack-detector`) Actor

Detect the exact frontend framework (Next.js, React, Nuxt, Vue) and CMS (Shopify, WordPress, Headless) powering any website. Bulk-analyze thousands of domains to build B2B lead lists based on tech stacks.

- **URL**: https://apify.com/thodor/tech-stack-detector.md
- **Developed by:** [Thodor](https://apify.com/thodor) (community)
- **Categories:** Lead generation, SEO tools, Other
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 2 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

A **tech stack and framework detector** tool. Paste a list of domains, get back the exact frontend framework and CMS each one runs, plus the CDN, analytics, and marketing stack, one row per domain, as CSV, Excel, or JSON. One domain or 100,000 in a single run, and it works as a single-site **CMS checker** too.

- ⚛️ **Frontend frameworks**: React, Next.js, Vue, Nuxt, Svelte, SvelteKit, Astro, Remix, Gatsby, Angular
- 🎨 **CSS and styling**: Tailwind CSS, Bootstrap, Material UI
- 🛍️ **CMS and commerce**: WordPress, Shopify (standard and headless), Webflow, Wix, Ghost, Magento

A **Wappalyzer alternative** and **BuiltWith alternative** without the subscription: no monthly minimum, no 2-technology filter cap, no API credits that [expire after 60 days](https://www.wappalyzer.com/pricing/). Detection runs on the actively maintained open-source successor to Wappalyzer's database, about 7,500 technologies, not the frozen pre-2023 copy most alternatives still ship, plus a dedicated heuristic layer for frameworks that strip their markers in production builds.

### 📋 How to check the tech stack of a whole list of domains

1. Paste your domains into **Start URLs or domains**, one per line.
2. Click **Start**.
3. Open the **Output** tab and click **Export** for CSV, Excel, JSON, or HTML.

1,000 domains take 5 to 10 minutes, 10,000 about an hour and a half, 100,000 run overnight.

### 🎁 So what do you get?

| | | |
| --- | --- | --- |
| **🏷️ Type:** *CMS, shop, builder, framework* | **⚛️ Framework:** *Next.js, React, Astro* | **🧱 CMS:** *WordPress 6.9.4, Shopify, Wix* |
| **🌐 CDN:** *Cloudflare, Fastly, Akamai* | **📊 Analytics:** *GA4, Plausible, Mixpanel* | **📣 Marketing:** *Klaviyo, HubSpot, GTM* |
| **🧾 Full breakdown:** *every technology found* | **💸 Pricing per tech:** *low / mid / high / freemium* | **#️⃣ Tech count:** *18 on techcrunch.com* |

### ⚖️ Compared to BuiltWith, Wappalyzer, and WhatCMS

| | This actor | Wappalyzer Pro | BuiltWith Basic | WhatCMS |
|---|---|---|---|---|
| 🆓 Free tier | ✅ ~1,200 analyses per month, Apify's $5 free credit | ⚠️ Browser extension only | ⚠️ Single lookups | ⚠️ Single lookups |
| 💰 Billing | ✅ Pay per result | ❌ Subscription | ❌ Subscription | ⚠️ Per lookup or subscription |
| ⚛️ Modern frameworks (Next.js, Astro, SvelteKit) | ✅ Dedicated heuristic layer | ⚠️ Partial | ⚠️ Partial | ❌ CMS only |
| ⏳ Credit expiration | ✅ None | ❌ 60 days | n/a | n/a |
| 🔎 Technologies you can filter on | ✅ Unlimited | ✅ Unlimited | ❌ Capped at 2 | ❌ CMS only |
| 🗄️ Database | ✅ Maintained fork, updated weekly | ⚠️ Closed since 2023 | ⚠️ Closed | ⚠️ Closed |

### 🎯 Three things people run this for

| | How |
|---|---|
| 🛍️ **Headless commerce detection** | A headless store leaks both layers into one row: `cms` says Shopify while the `breakdown` lists Next.js and Tailwind CSS. Filter for both and you have the stores with real engineering budgets, the exact list performance and dev-tool vendors want |
| 🏗️ **Agency migration lead gen** | Every site on a legacy stack in your market is a modernization pitch: Joomla, Drupal 7, and Magento 1 by CMS name and version, AngularJS and Gatsby by framework. Then pull contacts with the [Email Scraper](https://apify.com/thodor/apify-email-scraper-tool) |
| 📇 **SaaS lead qualification** | Feed inbound domains through the API and route them by stack: enterprise-priced tools (`high`, `poa` pricing tags) to sales, Shopify stores to the commerce team, freemium stacks to self-serve |

### 📥 Input

```json
{
  "start_urls": [
    { "url": "shopify.com" },
    { "url": "https://www.nytimes.com" },
    { "url": "techcrunch.com" }
  ]
}
```

- `start_urls`: 1 to 100,000 domains. Everything is normalized to the homepage, lowercased, `www.` stripped, so `www.x.com` and `x.com` become one row. Subdomains stay distinct: `docs.example.com` and `example.com` are two rows

That is the whole input.

### 📤 Output

One row per domain that returned a result. Failed fetches (DNS error, 4xx, 5xx, timeout) push no row and are not billed.

![Tech stack detector output example: dataset table with one row per domain showing type, CMS, framework, CDN, analytics, marketing tools and full tech breakdown](https://api.apify.com/v2/key-value-stores/LHcvkclm26dcJvwP1/records/output-example-cms-detector.png)

A real row, from a live run on TechCrunch:

```json
{
    "domain": "techcrunch.com",
    "url_checked": "https://techcrunch.com/",
    "type": "CMS",
    "cms": { "name": "WordPress", "version": "6.9.4" },
    "framework": "WordPress",
    "cdn": [],
    "analytics": [],
    "marketing": ["Google Tag Manager", "Sailthru"],
    "breakdown": [
        { "name": "MySQL",     "version": null,    "categories": ["Databases"],             "pricing": [] },
        { "name": "React",     "version": null,    "categories": ["JavaScript frameworks"], "pricing": [] },
        { "name": "Sailthru",  "version": null,    "categories": ["Marketing automation"],  "pricing": ["poa"] },
        { "name": "WordPress", "version": "6.9.4", "categories": ["CMS", "Blogs"],          "pricing": ["low", "recurring", "freemium"] }
    ],
    "tech_count": 18
    // HIDDEN: 14 more breakdown entries (Nginx, PHP, Yoast SEO Premium, ...)
}
```

A pure framework site (no CMS on top) comes back as `type: "Framework"` with the framework named. A site where nothing matches is `type: "Unknown"` with an empty breakdown; hand-built static pages and heavily stripped SPAs land there.

> ⚠️ **No JavaScript is run.** Pages are read the way a search engine reads them, which is what keeps a run fast and cheap, and on ~95% of sites the platform shows up in the source anyway. Tools that only exist after JavaScript executes (the Adobe enterprise stack, Drift, Cloudflare Zaraz) are invisible, and JS-rendered sites may return `Unknown`, though framework traces like `/_next/` usually still give you `type: "Framework"` with the right name.

#### Fields

| Field | Meaning |
|---|---|
| `domain`, `url_checked` | Canonical hostname and the exact homepage fetched |
| `type` | `CMS`, `Ecommerce`, `Website builder`, `Blog`, `Framework`, or `Unknown` (site reached, nothing in the HTML matched). Always set when the fetch succeeded |
| `cms` | `{ name, version }`, or `null` when no CMS-tier match. `version` is `null` when the site strips it |
| `framework` | Plain-string "what runs this site". When a CMS is detected it names the CMS; on headless builds the frontend framework sits in `breakdown` |
| `cdn`, `analytics`, `marketing` | Three ready-to-filter lists |
| `breakdown` | Every technology found: `name`, `version`, `categories`, `pricing`. This is where you filter for React, Tailwind CSS, or Sanity |
| `breakdown[].pricing` | Cost tier (`low` under $100/mo, `mid`, `high` over $1k/mo, `poa`) and billing model (`freemium`, `recurring`, `onetime`, `payg`). About 60% of technologies carry pricing data |
| `tech_count` | Unique technologies detected. Sort on it to find the most stack-rich domains |

### 🧬 The technology database

In August 2023, Wappalyzer closed its open-source rules. The [enthec/webappanalyzer](https://github.com/enthec/webappanalyzer) project picked up where it left off: open, public, roughly weekly updates, about 7,500 technologies. Most "Wappalyzer alternative" tools still ship the frozen pre-2023 database, which misses Next.js App Router, modern Shopify themes, and everything added since. This actor reads the maintained database directly.

On top of it sits a curation pass, because the raw data has known quirks ([source is public](https://github.com/Polluxs/apify/tree/master/apify-cms-detector)):

- A heuristic layer catches Next.js, Nuxt, React, Vue, Svelte, Astro, Remix, and Gatsby on production builds that strip the standard markers, via traces like `/_next/`, `__NUXT__`, and `data-astro-`
- The generic "Cart Functionality" fingerprint no longer steals the CMS slot from Shopify and Stripe
- Post-2023 category renumbering is remapped, so marketing tools land in `marketing` instead of wrong buckets
- Some 25 label fixes: Amazon S3 is object storage, not a CDN; Leadfeeder is B2B retargeting, not analytics; Datadog is APM
- Sites that cloak (paywall for browsers, clean page for crawlers) are refetched with a search-engine user agent

Independent testing across ~2,000 sites puts every major tool's CMS accuracy at 87 to 93%; this actor sits in that band, and the database freshness is what wins the modern-stack edge cases.

### ⚙️ Use it as a tech stack detection API

Every run is an HTTP endpoint: POST the same JSON as the form and the rows come back in the response body.

#### Python

```python
import requests

resp = requests.post(
    "https://api.apify.com/v2/acts/thodor~tech-stack-detector/run-sync-get-dataset-items",
    params={"token": "YOUR_APIFY_TOKEN"},
    json={"start_urls": [{"url": "shopify.com"}, {"url": "techcrunch.com"}]},
)

for row in resp.json():
    stack = {t["name"] for t in row["breakdown"]}
    headless = row["cms"] and "Next.js" in stack
    print(row["domain"], row["type"], "headless!" if headless else "")
```

#### Node.js

```javascript
import axios from "axios";

const { data } = await axios.post(
  "https://api.apify.com/v2/acts/thodor~tech-stack-detector/run-sync-get-dataset-items",
  { start_urls: [{ url: "shopify.com" }] },
  { params: { token: process.env.APIFY_TOKEN } }
);

console.log(data[0].domain, data[0].type, data[0].tech_count);
```

#### curl

```bash
curl -X POST "https://api.apify.com/v2/acts/thodor~tech-stack-detector/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"start_urls":[{"url":"shopify.com"}]}'
```

#### Clay, n8n, Make, Zapier

For tools that only send GET requests, like a Clay HTTP enrichment column, Apify's `?method=POST` trick turns the same endpoint into a GET URL:

```
GET https://api.apify.com/v2/acts/thodor~tech-stack-detector/run-sync-get-dataset-items
  ?token=APIFY_TOKEN&method=POST&start_urls[0][url]={{Domain}}
```

That appends `cms`, `framework`, `cdn`, and `marketing` to every row of a domain list without Clay's Explorer tier. n8n, Make, and Zapier take the plain POST from the examples above.

> 💡 **Tip:** no need to write the JSON by hand. Fill in the form on the Input tab, switch the editor from **Form** to **JSON**, and copy the result into your code.

### 💰 How much does bulk tech stack detection cost?

Billing is per successful detection, at the rate on the price card on this page: one row pushed, one result billed. Failed fetches are free, and `maxItems` on a run is respected, so you are never charged for more than you asked for.

### ❓ FAQ

**What about Cloudflare-protected sites?**
Most work. Sites behind the hardest JavaScript challenges sometimes return `Unknown`.

**Why is `cms` null for a site?**
The site either has no CMS (check `breakdown` for the framework) or hides it. If another tool sees a CMS this one misses, open a ticket with the URL.

**Why wasn't a tool detected even though I can see it on the site?**
A technology is only reported when detection is certain: strong signals fire alone, weak ones need a second confirmation. A few misses on obscure tools, far fewer false positives.

**Why does a re-platformed site still show its old platform?**
Migrations rarely strip every legacy marker. Old paths and tags linger, and typically fade over 6 to 18 months.

**Can a site fake its tech stack?**
Yes, every detection signal is public and can be faked; one researcher [tricked Wappalyzer into reporting 1,929 technologies](https://lab.julienverneaut.com/wappalyzer-1900-technologies) on one page. Treat the data as a signal, not gospel.

**Does it detect headless CMSes like Contentful, Sanity, or Strapi?**
When they leak their CDN domains or API endpoints, yes. A fully proxied backend shows only the frontend framework.

### 🛟 Support

A detection that's wrong, or a field missing? Message me in the Issues tab with the URL and what you expected; both false positives and false negatives usually turn into a one-line fingerprint or curation fix in the next build. I'm a solo dev, so don't hesitate.

Need contacts for the domains you just classified? The [Email Scraper](https://apify.com/thodor/apify-email-scraper-tool) pulls email addresses off the same list.

- Thodor

# Actor input Schema

## `start_urls` (type: `array`):

Websites to check. Each entry can be a bare domain (example.com, www.example.com) or any URL on the site. The Actor always inspects the site's homepage.

## Actor input object example

```json
{
  "start_urls": [
    {
      "url": "https://apify.com"
    }
  ]
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "start_urls": [
        {
            "url": "https://apify.com"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("thodor/tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "start_urls": [{ "url": "https://apify.com" }] }

# Run the Actor and wait for it to finish
run = client.actor("thodor/tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "start_urls": [
    {
      "url": "https://apify.com"
    }
  ]
}' |
apify call thodor/tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thodor/tech-stack-detector"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/YzcmXu9EAMKR4oMXu/builds/OjbFkDV61Ic9WEvkI/openapi.json
