# Tech Stack Detector: BuiltWith & Wappalyzer Alternative (`conserving_celerytop/website-tech-stack-detector`) Actor

Tech stack detector for any list of websites: the CMS, ecommerce platform, analytics, frameworks, CDN, hosting and payment tools. A BuiltWith and Wappalyzer alternative at $2 per 1,000 websites; sites that fail are free. Email provider, company profile and change tracking included. No login.

- **URL**: https://apify.com/conserving\_celerytop/website-tech-stack-detector.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Stats:** 2 total users, 1 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.60 / 1,000 website analyzeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Tech Stack Detector

Website Tech Stack Detector finds the technologies behind a list of websites: the CMS, ecommerce platform, analytics and tag managers, web and JavaScript frameworks, CDN, hosting and payment processors. Give it domains or URLs and it returns one row per website with every technology it recognised, its version when the page shows one, a confidence score and its categories. It is an alternative to BuiltWith and Wappalyzer lookups for lead lists, market research and competitor checks.

Each website gets one homepage request, and only when the site's robots.txt allows it. The Actor reads the response headers, cookie names (never cookie values), meta tags, script and stylesheet addresses and the HTML, and matches them against 7,628 open technology fingerprints from the [webappanalyzer](https://github.com/enthec/webappanalyzer) project (GPL-3.0).

### What the tech stack lookup returns

- **Technologies**: name, slug, version, confidence (0 to 100), categories and the technology's website.
- **Quick columns**: `cms`, `ecommerce`, `analytics`, `frameworks`, `cdn`, `hosting`, `paymentProcessors`, `tagManagers`, each a list of technology names, ready for a spreadsheet.
- **Email setup**: the company's email provider (for example Google Workspace or Microsoft 365), its email security gateway, the services allowed to send its email (for example HubSpot, SendGrid, Zendesk), its DMARC policy and the services it has verified its domain with (for example Atlassian, Stripe, Google Search Console). Read from public DNS records, with no extra request to the website.
- **Company profile**: page title, description and language, the company name, legal name and logo the homepage publishes, and links to the company's own LinkedIn, X, Facebook, Instagram, YouTube, GitHub and TikTok pages. Personal profiles are left out.
- **Changes since the last run**: turn on change tracking, run on a schedule, and each row says which technologies were added or removed since the previous run and whether the email provider changed.
- **Status for every website**: `status` is `ok` when the homepage loaded, or says why it did not (for example `robots_disallowed`, `blocked`, `timeout`, `domain_not_found`), with a plain-language `error`.

### How to find the tech stack of a website, step by step

1. Open the Actor and go to the **Input** tab.
2. In **Websites**, enter one website per line, as a domain (`example.com`) or a URL (`https://www.example.com/pricing`). The Actor always checks the homepage of each site.
3. Set **Maximum websites** to the number of sites you want checked. Duplicates count once.
4. Optional: pick **Categories** (for example CMS and Ecommerce) to return only those technologies, or turn off **Include versions**, **Email provider and DNS** or **Company profile and social pages**.
5. Optional: turn on **Track changes since the last run** and add a schedule (for example weekly). The first run saves the baseline; later runs fill the `changes` field.
6. Click **Start**. A run with the three example websites takes a few seconds.
7. Open the **Output** tab. The **Overview** view shows the quick columns, **All technologies** the full list, **Company and email** the company and email fields, and **Changes** what changed since the last run. Download the results as CSV, Excel or JSON, or read them through the Apify API.

### How much does it cost to detect a website's tech stack?

You pay per website analyzed: one `site-analyzed` event, with the email, company and change fields included at no extra cost, for each website whose homepage loaded. Websites that could not be read (robots.txt does not allow it, the site blocked the request, the domain does not exist, the request timed out) return a row with the reason and are **not charged**. Invalid entries and duplicates are free.

**Worked example:** at $2.00 per 1,000 websites analyzed, a list of 500 domains where 460 homepages load and 40 fail costs 460 x $0.002 = **$0.92**. The 40 failed websites cost nothing.

Set a spending limit on the run and the Actor stops starting new websites when the next one would go over it. The run log and the `STATS` record say how many websites were left unchecked.

### Input example

```json
{
    "websites": ["pypi.org", "crates.io", "www.wikipedia.org"],
    "maxSites": 3,
    "includeVersions": true,
    "categories": [],
    "includeEmailDns": true,
    "includeCompanyProfile": true,
    "monitorChanges": false,
    "maxConcurrency": 5
}
```

| Field | Meaning |
| --- | --- |
| `websites` | Domains or URLs, one per entry. |
| `maxSites` | Check at most this many websites (1 to 10,000). |
| `includeVersions` | Return version numbers when the page shows them. |
| `categories` | Return only technologies in these categories. Empty returns all. |
| `includeEmailDns` | Add the email provider, gateway, senders, DMARC policy and verified services. On by default. |
| `includeCompanyProfile` | Add page title, description, language, company name and logo, and company social pages. On by default. |
| `monitorChanges` | Compare with the previous run and fill `changes`. Off by default. |
| `monitorStoreName` | Key-value store in your account that keeps the last result per website (default `tech-stack-monitor`). Use one name per watchlist. |
| `maxConcurrency` | Websites checked at the same time (1 to 10). |

### Output example

A real row from a test run (technologies shortened):

```json
{
    "inputUrl": "crates.io",
    "finalUrl": "https://crates.io/",
    "domain": "crates.io",
    "httpStatus": 200,
    "status": "ok",
    "technologies": [
        { "name": "SvelteKit", "slug": "sveltekit", "version": null, "confidence": 100, "categories": ["UI frameworks"], "website": "https://kit.svelte.dev" },
        { "name": "Varnish", "slug": "varnish", "version": null, "confidence": 100, "categories": ["Caching"], "website": "https://www.varnish-cache.org" }
    ],
    "cms": [],
    "ecommerce": [],
    "analytics": [],
    "frameworks": ["Svelte"],
    "cdn": [],
    "hosting": [],
    "paymentProcessors": [],
    "tagManagers": [],
    "emailDomain": "crates.io",
    "emailProvider": "Mailgun",
    "emailSecurityGateway": null,
    "emailSenders": ["Mailgun"],
    "mxHosts": ["mxa.mailgun.org", "mxb.mailgun.org"],
    "hasSpf": true,
    "dmarcPolicy": "none",
    "verifiedServices": [],
    "dnsError": null,
    "pageTitle": "crates.io: Rust Package Registry",
    "metaDescription": "crates.io serves as a central registry for sharing crates, which are packages or libraries written in Rust that you can use to enhance your projects",
    "language": "en",
    "company": null,
    "socialProfiles": { "linkedin": null, "x": null, "facebook": null, "instagram": null, "youtube": null, "github": null, "tiktok": null },
    "changes": { "isFirstCheck": false, "previousCheckedAt": "2026-09-19T16:30:12.004Z", "changed": false, "added": [], "removed": [], "emailProviderBefore": null },
    "checkedAt": "2026-09-26T16:33:45.741Z",
    "charged": true,
    "error": null
}
```

The quick columns map to these fingerprint categories: `cms` (CMS), `ecommerce` (Ecommerce), `analytics` (Analytics), `frameworks` (Web frameworks, JavaScript frameworks), `cdn` (CDN), `hosting` (Hosting, PaaS), `paymentProcessors` (Payment processors), `tagManagers` (Tag managers).

### What "no JavaScript rendering" means

The Actor reads the page the server sends. It does not run the page's JavaScript in a browser. Technologies that are visible in that first response (headers, cookies, meta tags, script addresses, HTML markup) are found. Technologies that only appear after scripts run, for example a chat widget or an analytics tool loaded by a tag manager, may be missed. Fingerprint patterns that need a browser (JavaScript variables, network calls made later) are not used. This keeps the Actor fast and cheap and means one request per website.

### FAQ

**Is it legal to check a website's technologies?**
The Actor reads only public homepages, one request per site, without logging in, and only when the site's robots.txt allows the homepage for all user agents. It does not collect personal data: cookie values are never read into the output, and no names, email addresses or phone numbers are stored. The email fields describe the company's mail setup (which provider and services it uses), not any mailbox. Social links are kept only for company pages; personal profiles are skipped. You are responsible for how you use the results; check the terms of the sites on your list.

**Is this Actor affiliated with BuiltWith, Wappalyzer or Enthec?**
No. It is not affiliated with, endorsed by or sponsored by BuiltWith, Wappalyzer or Enthec. The names describe what the Actor does. The fingerprints come from the open webappanalyzer project under the GPL-3.0 licence; see the NOTICE file.

**Why does a website show `robots_disallowed`?**
Its robots.txt does not allow the homepage for all user agents, so the Actor did not request it. The row is free.

**Why does a website show `blocked`?**
The site answered with an access-denied or challenge page (for example HTTP 403 or a bot check). The Actor does not retry or work around it. The row is free.

**Why is a technology I know about missing?**
It may load only through JavaScript (see above), or the open fingerprint set may not describe it. The confidence score shows how sure a match is; 100 means at least one pattern matched fully.

**How many websites can I check?**
Up to 10,000 per run. The Actor checks up to 10 websites at a time and each website gets at most one robots.txt request and one homepage request (plus redirects and up to two retries after a network error or a server error).

**How does change tracking work?**
With **Track changes since the last run** on, each analyzed website's technology list is saved in a named key-value store in your own Apify account. The next run compares against it and writes `added`, `removed` and `changed`. Only websites that loaded update the saved result, so a site that is down for one run keeps its old baseline. If you change **Categories**, a new baseline starts for that filter.

**Why is `emailProvider` "Other" or null?**
"Other" means the domain receives mail on a server that is not a known hosted provider (often its own server or its web host). null means the domain has no mail server, or an email security gateway hides the provider; `emailSecurityGateway` then names the gateway.

**What if a page is very large?**
The first 3 MB are checked and the row says so in `error`.

# Actor input Schema

## `websites` (type: `array`):

Enter the websites to check, one per line, as a domain (example.com) or a URL. Each site’s homepage is read once.

## `maxSites` (type: `integer`):

Check at most this many websites from the list, in list order. Duplicates count once.

## `includeVersions` (type: `boolean`):

Return the version number when the page shows it (for example WordPress 6.5).

## `categories` (type: `array`):

Return only technologies in these categories. Leave empty to return every category.

## `includeEmailDns` (type: `boolean`):

Add the company's email provider (for example Google Workspace or Microsoft 365), email security gateway, the services allowed to send its email (for example HubSpot, SendGrid, Zendesk), its DMARC policy and the services it has verified the domain with. Read from public DNS records, no extra web request.

## `includeCompanyProfile` (type: `boolean`):

Add the page title, description, language, the company name, legal name and logo the homepage publishes, and links to the company's own LinkedIn, X, Facebook, Instagram, YouTube, GitHub and TikTok pages. Personal profiles are left out.

## `monitorChanges` (type: `boolean`):

Compare each website with the result saved by the previous run and add a "changes" field (added, removed, email provider change). The first run saves the baseline. Same price per website.

## `monitorStoreName` (type: `string`):

Name of the key-value store in your account that keeps the last result per website. Use a different name for each watchlist.

## `maxConcurrency` (type: `integer`):

Check this many websites at the same time. Each website gets at most one homepage request and one robots.txt request.

## Actor input object example

```json
{
  "websites": [
    "pypi.org",
    "crates.io",
    "www.wikipedia.org"
  ],
  "maxSites": 3,
  "includeVersions": true,
  "categories": [],
  "includeEmailDns": true,
  "includeCompanyProfile": true,
  "monitorChanges": false,
  "monitorStoreName": "tech-stack-monitor",
  "maxConcurrency": 5
}
```

# Actor output Schema

## `overview` (type: `string`):

inputUrl, domain, httpStatus, status, cms, ecommerce, analytics, frameworks, cdn, hosting, paymentProcessors, tagManagers, checkedAt and error.

## `technologies` (type: `string`):

The full technologies list per website: name, slug, version, confidence, categories and website.

## `company` (type: `string`):

Company name, logo, social pages, email provider, email senders, DMARC policy and verified services.

## `changes` (type: `string`):

With change tracking on: technologies added and removed per website.

## `stats` (type: `string`):

JSON with websites checked and analyzed, statuses, requests, retries, charges, peak memory and the fingerprint source commit.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "pypi.org",
        "crates.io",
        "www.wikipedia.org"
    ],
    "maxSites": 3,
    "categories": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/website-tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": [
        "pypi.org",
        "crates.io",
        "www.wikipedia.org",
    ],
    "maxSites": 3,
    "categories": [],
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/website-tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "pypi.org",
    "crates.io",
    "www.wikipedia.org"
  ],
  "maxSites": 3,
  "categories": []
}' |
apify call conserving_celerytop/website-tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/website-tech-stack-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BzFRXv35wcHzAm2Ky/builds/HgUguuqQYzOepcLNm/openapi.json
