# Website Technology Detector (`ottomath/website-technology-detector`) Actor

Detect the CMS, e-commerce platform, analytics, CDN, frameworks and servers behind any list of websites. Plain HTTP, no browser. $5 per 1,000 sites.

- **URL**: https://apify.com/ottomath/website-technology-detector.md
- **Developed by:** [Ottomath](https://apify.com/ottomath) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 url analyseds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Technology Detector

Find out which technologies a website runs on: content management system, e-commerce platform, JavaScript frameworks, analytics, tag managers, CDN, hosting, web server, payment tools, chat widgets, cookie banners, fonts and more. Provide a list of URLs or domains and get one structured result per site. The Actor uses plain HTTP requests and no browser, so it is fast and inexpensive.

### What it does

For every URL or domain you provide, the Actor fetches the page once, follows up to 5 redirects, and reads everything a plain HTTP client can see: response headers, cookie names, the HTML, meta tags, script and stylesheet references, inline scripts, page text and the final URL. It matches these signals against a database of about 3,600 technology fingerprints in 108 categories. It also applies the database's relationships (WordPress implies PHP and MySQL, for example), reads the version when the page reveals it, and can look up the domain's DNS records to spot email, DNS and hosting providers.

The result is one dataset item per site with the technologies, their categories, versions and a confidence score. The Actor uses plain HTTP requests only. It does not run a browser, does not log in anywhere and does not disguise itself: every request carries the honest User-Agent `OttomathTechDetector/1.0 (+https://apify.com)`. It extracts technology names and versions only, never email addresses, phone numbers or names.

### Use cases

- **Prospecting and sales:** build a target list from a set of domains, for example every Shopify store or every WordPress site, and prioritise agencies, apps or services by the platform your prospects use.
- **Competitor and market research:** compare the analytics, tag manager, CDN and e-commerce stack of competitors, or measure how often a platform is used in a market segment.
- **Web agencies and migrations:** inventory the CMS, framework and server versions across the client sites you manage before an upgrade or a security review.
- **Data enrichment:** add a technology column to a CRM export or a dataset of company websites and re-run it on a schedule to track changes.

### Input

Provide the sites in `urls`. Each entry can be a full URL starting with `https://` or `http://`, or a bare domain such as `example.com` (https is assumed). Up to 10,000 entries per run; duplicates are analysed once. Local and private network addresses are rejected and reported in the `error` field.

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `urls` | list of strings | required | Websites to analyse, 1 to 10,000 per run. |
| `respectRobotsTxt` | boolean | true | Skip URLs that the site's robots.txt disallows for this Actor's User-Agent or for `*`. The path of the URL is checked, not only `/`. |
| `includeVersions` | boolean | true | Report versions when the page reveals them. When false, `version` is always null. |
| `minConfidence` | integer | 0 | Only report technologies with at least this confidence (0-100). Use 50 to hide weak single-signal guesses. |
| `detectDns` | boolean | true | Also read MX, TXT, NS, SOA and CNAME records to detect email, DNS and hosting providers. |
| `maxConcurrency` | integer | 10 | URLs processed in parallel (1-30). At most 2 requests per domain run at the same time. |
| `requestTimeoutSecs` | integer | 20 | Time limit for a single HTTP request (1-120). |

Example input:

```json
{
  "urls": ["https://example.com", "wordpress.org", "https://www.shopify.com"],
  "respectRobotsTxt": true,
  "includeVersions": true,
  "minConfidence": 0
}
```

### Output example

Each result is one item in the dataset, in the same order of fields. Dates are ISO 8601 in UTC. You can download the dataset as JSON, CSV, Excel, XML or HTML, or read it through the Apify API.

```json
{
  "url": "blog.example.com",
  "finalUrl": "https://blog.example.com/",
  "domain": "blog.example.com",
  "statusCode": 200,
  "technologies": [
    { "name": "WordPress", "slug": "wordpress", "categories": ["CMS", "Blogs"], "version": "6.4.2", "confidence": 100, "website": "https://wordpress.org" },
    { "name": "PHP", "slug": "php", "categories": ["Programming languages"], "version": "8.2.12", "confidence": 100, "website": "http://php.net" },
    { "name": "Nginx", "slug": "nginx", "categories": ["Web servers", "Reverse proxies"], "version": "1.24.0", "confidence": 100, "website": "http://nginx.org/en" }
  ],
  "categories": {
    "CMS": ["WordPress"],
    "Programming languages": ["PHP"],
    "Web servers": ["Nginx"]
  },
  "detectedAt": "2026-10-08T10:30:00.000Z",
  "error": null
}
```

`url` is exactly what you provided and `finalUrl` is where the redirects ended. `domain` is the host of the final URL without a leading `www.`. `confidence` is a number from 1 to 100: 100 means at least one strong signal, lower values mean weaker or indirect evidence. Technologies are sorted by category importance, so the CMS and e-commerce platform come first. `error` is null when the page was fetched and analysed. Otherwise it explains why, for example `Blocked by robots.txt`, `Timeout after 20 s`, `Domain not found (DNS lookup failed)` or `HTTP 404 Not Found`, and the other fields hold whatever could be determined.

### Pricing

This Actor uses Apify's pay-per-event model. You pay for what it delivers, not for how long it runs.

| Event | When it is charged | Price |
| --- | --- | --- |
| Actor start | Once when a run starts. Apify charges it automatically. | $0.005 |
| URL analysed (`url-analyzed`) | Once for each URL whose page was fetched with an HTTP status below 400 and analysed. | $0.005 |

URLs that fail, time out, are blocked by robots.txt, are invalid or return an error status (4xx or 5xx) are reported in the dataset and are free. Duplicates are analysed and charged once. Example: 1,000 working URLs cost $0.005 + 1,000 x $0.005 = $5.005.

You stay in control of the cost. Set the maximum charge for the run in the run options: the Actor never analyses more URLs than your limit pays for. When the limit is reached it stops by itself, keeps everything already delivered and ends the run with a clear status message.

### FAQ

**How do I cap my spending?**
Use the maximum total charge of the run. The Actor reads it, stops cleanly when the next URL would exceed it, and tells you how many URLs were not processed.

**Does it respect robots.txt?**
Yes, by default. The Actor reads robots.txt once per site and follows the rules for its own User-Agent, or for `*` when none are specific. If robots.txt cannot be read because of a server error, the URL is skipped. You are responsible for complying with the terms of the sites you analyse and with applicable law.

**Why is a technology I know about missing?**
Without a browser the Actor cannot see JavaScript variables or background requests, so a site that reveals its stack only that way may show fewer technologies. The fingerprint database is also a snapshot, see Limitations.

**What happens when a site blocks the request?**
You get an item with an error such as `HTTP 403 Forbidden`. The Actor does not rotate proxies, solve CAPTCHAs or change its identity. Technologies visible in the error response, such as a CDN header, are still listed, and the URL is free.

**How long does a run take?**
Mostly the time sites need to answer. A list of 100 typical sites finishes in a few minutes. Very large lists of 10,000 URLs take longer, so split them into batches if you need results quickly.

**Can I schedule it or call it from other tools?**
Yes. Like any Actor it can run on a schedule, through the Apify API, or from integrations such as webhooks.

### Performance

Analysing a page is CPU-bound and runs in worker threads. In a single thread on a development machine, five well-known home pages (45 KB to 1 MB of HTML) took 0.06 to 0.3 seconds of CPU time each, and up to 0.5 seconds for the first page of a run. A page of 2 MB, the maximum, took 0.6 to 1.1 seconds. A run with 1,024 MB of memory gets about a quarter of a CPU core on Apify, so allow roughly four times that as elapsed time per page.

Two measures keep this predictable. A fingerprint pattern is skipped when a keyword it requires is not on the page. A leading wildcard (`.+` or `.*`) is removed from patterns, which does not change what they find but avoids a very slow search in long single-line scripts; when a page holds several matches of such a pattern, the version comes from the first one. A page that still needs more than 20 seconds, for example deliberately malformed markup, is abandoned, reported with an error and not charged.

### Limitations

- Plain HTTP only: pages are not rendered, so JavaScript globals and background requests are not evaluated. Roughly 300 technologies in the database can only be recognised that way and will not be reported.
- One page per URL (usually the home page), plus robots.txt and redirects. The first 2 MB of the page are analysed.
- A page that takes more than 20 seconds to analyse (extremely large or deliberately malformed markup) is abandoned and reported with an error. It is not charged.
- The fingerprint database dates from January 2023. Newer tools may be missing; a small set of extra fingerprints for modern frameworks such as Next.js, Nuxt, Astro and SvelteKit is included.
- No login, no CAPTCHA solving and no proxy rotation. Sites that block automated requests return an error.
- DNS-based detection depends on the DNS records the domain publishes.

### Data source and licence

Fingerprints come from the technology data of the open-source Wappalyzer project, release 6.10.54 of its npm package, published under the MIT licence. The data is downloaded when the Actor is built and checked against a fixed SHA-256 hash. This Actor is not affiliated with or endorsed by that project. The licence terms and attributions are listed in `THIRD_PARTY_NOTICES.md`.

### Changelog

#### 1.0.0 (2026-10-08)

- Initial release.

# Actor input Schema

## `urls` (type: `array`):

One URL or domain per line, for example `https://example.com` or `example.com` (https is assumed). Up to 10,000 per run. Duplicates are analysed once.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs that the site's robots.txt disallows for this Actor's User-Agent (or for `*`). Skipped URLs are reported with an error message and are free.

## `includeVersions` (type: `boolean`):

Report the version of a technology when the page reveals it (for example WordPress 6.4.2).

## `minConfidence` (type: `integer`):

Only report technologies detected with at least this confidence (0-100). 0 reports everything, 50 hides weak single-signal guesses.

## `detectDns` (type: `boolean`):

Also read the domain's MX, TXT, NS, SOA and CNAME records to detect email, DNS and hosting providers. Uses normal DNS lookups only.

## `maxConcurrency` (type: `integer`):

How many URLs are processed in parallel. At most 2 requests per domain run at the same time.

## `requestTimeoutSecs` (type: `integer`):

Give up on a single HTTP request after this many seconds. The URL is then reported with a Timeout error.

## Actor input object example

```json
{
  "urls": [
    "https://example.com",
    "https://wordpress.org"
  ],
  "respectRobotsTxt": true,
  "includeVersions": true,
  "minConfidence": 0,
  "detectDns": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

One item per URL with the detected technologies (and free error rows for URLs that could not be analysed), as JSON. Use the dataset tabs in the Console for tables, CSV and Excel exports.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com",
        "https://wordpress.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ottomath/website-technology-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://example.com",
        "https://wordpress.org",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("ottomath/website-technology-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com",
    "https://wordpress.org"
  ]
}' |
apify call ottomath/website-technology-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ottomath/website-technology-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5FelLY2FDyV8g5rfI/builds/MZFqfs6T7JLgBKYKE/openapi.json
