# Website Tech Stack Detector - Evidence & Change Tracking (`maydit/website-tech-stack-detector`) Actor

Detect supported CMS, ecommerce, analytics and framework signals from public HTML and headers. Evidence per finding, CRM IDs and complete-scope repeat comparisons. No browser or API key.

- **URL**: https://apify.com/maydit/website-tech-stack-detector.md
- **Developed by:** [Brandt May](https://apify.com/maydit) (community)
- **Categories:** Developer tools, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 website observations

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Tech Stack Detector - Evidence & Change Tracking

Give this Actor a list of public websites. It reads their HTML and response headers, then returns supported technology signals with the exact asset, header or HTML attribute behind each finding. Use the results to segment supplied company lists, review a website portfolio, or compare the same pages on a schedule.

It supports CMS and ecommerce systems, frameworks, analytics, advertising, customer support, payments and hosting signals. Your own `externalId` travels with each result so you can join the output to existing records. No external API key is required.

### Start with two public examples

```json
{
  "websites": [
    { "url": "https://wordpress.org/", "externalId": "wordpress-demo" },
    { "url": "https://nextjs.org/", "externalId": "nextjs-demo" }
  ]
}
```

An omitted or empty `websites` array runs this two-site demo. No automatic site crawl occurs. You can supply up to 100 websites and at most two additional pages on each website's original origin:

```json
{
  "websites": [
    {
      "url": "https://wordpress.org/",
      "externalId": "site-123",
      "additionalUrls": ["https://wordpress.org/about/", "https://wordpress.org/news/"]
    }
  ],
  "maxRunSeconds": 240
}
```

Reading supplied inner pages helps find tools not visible on the homepage. Every supplied page must be inspected successfully for that website to enter the result dataset. A failed page is recorded in `SUMMARY`; the incomplete website is not billed as a result or used to update history.

### Output

Each successful website check produces one record. A simplified illustrative result:

```json
{
  "url": "https://nextjs.org/",
  "externalId": "nextjs-demo",
  "finalUrl": "https://nextjs.org/",
  "status": "ok",
  "httpStatus": 200,
  "pageCount": 1,
  "technologyCount": 2,
  "technologyNames": ["Next.js", "Vercel"],
  "technologiesByCategory": {
    "framework": ["Next.js"],
    "hosting": ["Vercel"]
  },
  "findings": [
    {
      "name": "Next.js",
      "category": "framework",
      "certainty": "direct",
      "evidence": [
        {
          "type": "header",
          "value": "x-powered-by: Next.js",
          "sourcePage": "https://nextjs.org/"
        }
      ]
    }
  ],
  "detectionMethod": "public_html_and_response_headers",
  "rulesVersion": "1.0.0",
  "changeStatus": "not_compared"
}
```

Actual records include `checkedAt`, `pageUrls`, `scopeKey`, `previousCheckedAt`, `addedTechnologies` and `noLongerObservedTechnologies`. Findings are alphabetized; evidence is capped at three examples per technology per page, up to nine per website. Asset evidence omits query strings and fragment identifiers. `certainty: direct` means an identifying artifact was observed, not that we verified a vendor subscription, executed its code, or measured how the site uses it.

`status` is `ok` when a supported signal was found, or `no_supported_signal` when the HTML was successfully inspected without a match. Both are completed, billable checks. A zero-signal result does not prove software is absent.

The default key-value store's `SUMMARY` record contains processed counts, skipped inputs, failures, per-page reasons and time-budget coverage. Failure classes include blocked pages, robots restrictions, HTTP errors, non-HTML responses, request errors and timeouts. If no website can be completely inspected, the run fails with a diagnostic instead of silently succeeding empty.

If your run charge limit prevents a result from being saved, the Actor stops with a clear warning and `SUMMARY.stoppedOnChargeLimit: true`. Unsaved results do not advance history. A charge limit reached before the first result is a normal empty budget stop, reported as `outcome: stopped_on_charge_limit`.

### Supported signal catalog

These are the explicitly implemented names; coverage is limited to their visible signals:

| Category | Technologies |
| --- | --- |
| CMS and site tools | WordPress, Drupal, Joomla, Ghost, TYPO3, Hugo, Webflow, Wix, Squarespace, Docusaurus |
| Ecommerce | WooCommerce, Shopify, PrestaShop, Magento |
| Frameworks and libraries | Next.js, Nuxt, Gatsby, Angular, SvelteKit, Express, jQuery, Alpine.js, Bootstrap |
| Analytics and tags | Google Tag Manager, Google tag, Google Analytics 4, Plausible Analytics, Microsoft Clarity, Hotjar, Segment, Mixpanel, PostHog, Fathom Analytics |
| Marketing and advertising | Google Ads, HubSpot, Marketo, Klaviyo, Mailchimp, Meta Pixel, TikTok Pixel, LinkedIn Insight Tag |
| Support | Intercom, Zendesk, Crisp, Tawk.to, Drift, Freshchat |
| Payments | Stripe.js, PayPal SDK, Braintree |
| Monitoring and consent | Sentry, New Relic, Datadog RUM, OneTrust, Cookiebot, Usercentrics |
| Infrastructure | Cloudflare, Vercel, Netlify, nginx, Apache, Microsoft IIS |

Signatures inspect script and stylesheet references, selected preload assets, generator tags, explicit HTML attributes, narrow inline JavaScript assignments and response headers. A vendor mentioned in ordinary page prose or a normal hyperlink does not count. Shopify CDN assets alone do not label a site as a Shopify store. A generic Google tag is distinguished from a GA4 tag identifier. No inferred backend or framework dependencies are added.

### Compare repeat runs

```json
{
  "websites": ["https://wordpress.org/", "https://nextjs.org/"],
  "comparePrevious": true,
  "snapshotName": "my-website-portfolio"
}
```

The **first successful run returns all successful checks and establishes a baseline**. Reuse the same store name and exact supplied page set on later runs. Results still include every successful website, with `changeStatus` of `baseline`, `baseline_reset`, `unchanged` or `changed`. Unchanged results are billable because the site was inspected again.

Snapshots contain URLs and technology names and are saved in the named key-value store in your Apify account. A changed supplied page scope creates a separate baseline. A rules version or final redirect destination change resets the baseline, avoiding a misleading comparison. Incomplete or failed website checks never overwrite a baseline. Avoid concurrent runs that share a snapshot store: the last completing run saves the latest snapshot.

`noLongerObservedTechnologies` means the current complete check did not observe those signals. A site may have changed consent settings, hidden assets or delivered different markup; this is not proof that a tool was uninstalled. There is no webhook or notification sending in this Actor. You can use the exported result with your own automation.

### Inputs and limits

| Input | Default | Meaning |
| --- | --- | --- |
| `websites` | Two-site demo in code | Up to 100 strings or `{url, externalId, additionalUrls}` objects |
| `maxWebsites` | 100 | Process this many input records after exact duplicate removal |
| `comparePrevious` | false | Save and compare complete observations |
| `snapshotName` | `website-tech-monitor` | Private named store; use a distinct name per independent monitor |
| `maxRunSeconds` | 240 | Wall-clock budget; minimum 30, maximum 3600 |

At most three supplied pages per website, five redirects per page, 2 MB HTML per page and 20 seconds per request are allowed within the overall budget. Respectful public HTTP requests check robots.txt, including redirect destinations. Requests blocked by access controls are reported. Sites requiring login or JavaScript execution cannot be fully inspected by this Actor. HTML consent gates and geography can change visible signals. This is an observed technology inventory, not a security or compliance audit, traffic estimate, revenue estimate, or complete Wappalyzer/BuiltWith replacement.

Exact repeated page scopes with the same `externalId` are deduplicated. Different IDs or different supplied page scopes produce separate checks. The Actor preserves the input record's first URL as `url`; the fetched final location is `finalUrl`.

### Pricing

The launch price is **$0.005 per successful website check ($5 per 1,000)** on Free/Bronze. Each check includes its supplied scope of up to three pages. A successful zero-signal check and an unchanged comparison are charged. Failed, robots-disallowed, incomplete or unprocessed websites are not charged as results. An Actor Start event of $0.00005 also applies. Silver is 20% lower and Gold/Platinum/Diamond 40% lower. The live Pricing tab is authoritative.

Normal Apify account credits and limits apply. No feature requires a separate API subscription. Do not run the same monitored scope concurrently if you need an ordered comparison history.

### Examples

The `tasks/` directory contains three proposed example inputs: WordPress CMS/header signals, Next.js framework/hosting signals, and WordPress evidence across three supplied pages. They are local proposals until separately validated on the platform and published.

# Actor input Schema

## `websites` (type: `array`):

Up to 100 URL strings or objects with url, optional externalId, and optional additionalUrls (at most two URLs on the same origin). Three total supplied pages per website; no automatic crawling. Omitted or empty input uses wordpress.org and nextjs.org.

## `maxWebsites` (type: `integer`):

Maximum input website records to process after exact duplicate removal. Remaining inputs are counted in SUMMARY.

## `comparePrevious` (type: `boolean`):

Save private complete snapshots in the named key-value store and compare repeat runs. The first successful run establishes a baseline and returns every successful website. Failed or incomplete website checks never replace a baseline. A missing signal is not proof that software was uninstalled.

## `snapshotName` (type: `string`):

Use the same name and identical supplied page scope for repeat comparisons. A separate name isolates independent monitors. Letters, numbers and hyphens, 1–63 characters. Avoid overlapping runs with the same store.

## `maxRunSeconds` (type: `integer`):

Stop before this wall-clock budget or the platform timeout. All requests and redirects share the remaining budget. Partial website scopes are not written or compared.

## Actor input object example

```json
{
  "websites": [
    {
      "url": "https://wordpress.org/"
    },
    {
      "url": "https://nextjs.org/"
    }
  ],
  "maxWebsites": 100,
  "comparePrevious": false,
  "snapshotName": "website-tech-monitor",
  "maxRunSeconds": 240
}
```

# Actor output Schema

## `results` (type: `string`):

One record per fully inspected supplied website scope, including successful zero-signal checks.

## `summary` (type: `string`):

Coverage counts, time-budget stops, website failures and unbilled partial checks.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        {
            "url": "https://wordpress.org/"
        },
        {
            "url": "https://nextjs.org/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("maydit/website-tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        { "url": "https://wordpress.org/" },
        { "url": "https://nextjs.org/" },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("maydit/website-tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    {
      "url": "https://wordpress.org/"
    },
    {
      "url": "https://nextjs.org/"
    }
  ]
}' |
apify call maydit/website-tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,maydit/website-tech-stack-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4hudy7lbdb5kOVqRg/builds/3ebtdodMseSVt1Cd2/openapi.json
