# Tech Stack Detector: CMS, Ecommerce, Analytics, Payments (`pistachio_implementation/tech-stack-detector`) Actor

Find out what any website is built with. Give a list of domains and get the CMS, ecommerce platform, analytics and tag managers, payment providers, chat widgets, frameworks, CDN and hosting for each, with versions where visible. Fast, no browser, fair price per domain.

- **URL**: https://apify.com/pistachio\_implementation/tech-stack-detector.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** Lead generation, Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Tech Stack Detector: CMS, Ecommerce, Analytics, Payments

Find out what any website is built with. Give it a list of domains and get back, for each one, the **CMS, ecommerce platform, analytics and tag managers, payment providers, live chat widgets, JavaScript and web frameworks, CDN, hosting and web server**, with versions where the site exposes them.

Built for sales teams and lead generation agencies that qualify prospects by technology ("all Shopify stores using Klaviyo", "WordPress sites without a chat widget"), and for anyone auditing a list of sites.

### What it does

- Fetches each domain's **public home page and response headers** once, plus its robots.txt, and matches them against about **3,600 technology fingerprints** (headers, cookie names, meta tags, script sources, inline scripts, HTML and page structure).
- Returns handy summary columns (cms, ecommerce, analytics, payments, liveChat and more) **and** the full list of detections with version, confidence and category.
- Fast and cheap: no browser is started, so thousands of domains run in minutes.
- Polite: at most two requests per domain, honest user agent, robots.txt respected by default, retries with backoff on rate limits.
- No login, no cookies sent, no personal data collected.

### Input example

```json
{
    "domains": ["allbirds.com", "wordpress.org", "https://www.squarespace.com/"],
    "respectRobotsTxt": true,
    "minConfidence": 50,
    "includeAllTechnologies": true,
    "maxConcurrency": 5,
    "requestTimeoutSecs": 20
}
```

| Field | What it means |
|---|---|
| domains | Domains or URLs, one per line. Only the home page is checked. Duplicates are removed. |
| respectRobotsTxt | Skip a domain whose robots.txt forbids crawlers from the home page. Skipped domains are not charged. |
| minConfidence | Only report detections at or above this confidence (0 to 100). |
| includeAllTechnologies | Include the full detection list. Turn off for a compact table. |
| maxConcurrency | Domains checked in parallel (1 to 20). |
| requestTimeoutSecs | Give up on a slow site after this many seconds. |

### Output example

One row per domain (shortened here).

```json
{
    "domain": "allbirds.com",
    "finalUrl": "https://www.allbirds.com/",
    "statusCode": 200,
    "title": "Allbirds: Comfortable, Sustainable Shoes & Apparel",
    "technologyCount": 9,
    "cms": [],
    "ecommerce": ["Shopify"],
    "analytics": [],
    "tagManagers": ["Google Tag Manager"],
    "payments": ["PayPal", "Apple Pay"],
    "liveChat": [],
    "cdn": ["Cloudflare"],
    "security": ["HSTS"],
    "technologies": [
        { "name": "Shopify", "version": null, "confidence": 100, "categories": ["Ecommerce"], "website": "http://shopify.com" },
        { "name": "PayPal", "version": null, "confidence": 100, "categories": ["Payment processors"], "website": "https://paypal.com" }
    ],
    "possiblyBlocked": false,
    "error": null,
    "checkedAt": "2026-09-27T05:45:12.000Z"
}
```

Summary columns in every row: cms, ecommerce, analytics, tagManagers, payments, liveChat, marketingAutomation, crm, advertising, jsFrameworks, webFrameworks, uiFrameworks, jsLibraries, cdn, hosting, webServers, programmingLanguages, security, cookieCompliance, abTesting, reviews.

### Pricing

Pay per result: **$2 per 1,000 domains analyzed** ($0.002 each). You are charged only when the site actually served its home page. Domains that do not resolve, time out, block the request, or are skipped because of robots.txt are free.

### Limits

- **Static analysis only.** Tools that are injected later by a tag manager (for example Google Analytics loaded through Google Tag Manager, or a chat widget loaded on scroll) do not appear in the page source and may be missed. The tag manager itself is detected.
- Only the home page is checked. A shop that lives on a subdomain (shop.example.com) should be entered separately.
- Some large sites put bot walls in front of every automated visitor. Those rows come back with the HTTP status, `possiblyBlocked: true`, whatever could be read from the headers, and no charge.
- The fingerprint set is a fixed snapshot of the open source Wappalyzer database (version 6.10.54, MIT license). Technologies launched since early 2023 may be missing or named differently.

### FAQ

**How is this different from BuiltWith or Wappalyzer?** Same idea, pay as you go, no subscription, and bulk friendly. You pay a fraction of a cent per domain.

**Can I check thousands of domains?** Yes. Raise maxConcurrency to 10 or 20 for large lists.

**Why is a technology I know the site uses missing?** Most often it is loaded by JavaScript after the page opens. See Limits.

**Does it store anything about people?** No. It reads public pages and headers and returns technology names only.

**Something looks wrong.** Open an issue on the Issues tab with the domain and what you expected. Issues are answered quickly.

### Credits and license

Technology fingerprints and the matching engine come from Wappalyzer 6.10.54, the last release published under the MIT license (Copyright Elbert Alias and Wappalyzer contributors). The license text ships with the actor in `vendor/wappalyzer/LICENSE.txt`. This actor is not affiliated with Wappalyzer, BuiltWith or any vendor it detects. All product names are trademarks of their owners.

# Actor input Schema

## `domains` (type: `array`):

One domain or URL per line, for example shopify.com or https://www.example.com/. Only the home page of each domain is checked. Duplicates are removed.

## `respectRobotsTxt` (type: `boolean`):

When on, a domain whose robots.txt forbids crawlers from the home page is skipped and not charged.

## `minConfidence` (type: `integer`):

Only report technologies detected with at least this confidence (0 to 100).

## `includeAllTechnologies` (type: `boolean`):

Adds a technologies array with every detection, its version, confidence and categories. Turn off for a compact output with only the summary columns.

## `maxConcurrency` (type: `integer`):

Each domain gets at most two requests (robots.txt and the home page).

## `requestTimeoutSecs` (type: `integer`):

Give up on a slow site after this many seconds.

## Actor input object example

```json
{
  "domains": [
    "allbirds.com",
    "wordpress.org",
    "stripe.com"
  ],
  "respectRobotsTxt": true,
  "minConfidence": 50,
  "includeAllTechnologies": true,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "allbirds.com",
        "wordpress.org",
        "stripe.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "allbirds.com",
        "wordpress.org",
        "stripe.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "allbirds.com",
    "wordpress.org",
    "stripe.com"
  ]
}' |
apify call pistachio_implementation/tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/tech-stack-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/RbGCE5aUYJJWcOoT6/builds/hLiAaT0Ehsnb5j2Wo/openapi.json
