# Website Technology Detector: Tech Stack & CMS Lookup (`humble-echidna/tech-stack-detector`) Actor

Detect any website's tech stack: CMS, e-commerce platform, frameworks, analytics, tag managers, CDN, hosting and JavaScript libraries, with version, confidence and evidence, from the MIT-licensed Wappalyzer fingerprints plus our own. Monitor a list: get only sites that added or dropped a technology.

- **URL**: https://apify.com/humble-echidna/tech-stack-detector.md
- **Developed by:** [Michael Costa](https://apify.com/humble-echidna) (community)
- **Categories:** Lead generation, Developer tools, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 website analyseds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Website Technology Detector do?

**Website Technology Detector** is a bulk **tech stack detector**: it finds out **what a website is built with**.
Give it a list of sites; for each you get its CMS, e-commerce platform, frameworks, analytics, CDN and hosting, with
the **version**, a **confidence** score and the **evidence** behind every technology.

It fetches one page per URL like a normal visitor, without a browser, and matches it against 3,500+ technology
fingerprints (the MIT-licensed Wappalyzer data plus our own). That makes it fast and cheap, and it means technologies
that only show up once the page's JavaScript runs can be missed (see "How accurate is it?" below).

**Try it in one click:** the input comes pre-filled with two websites. That's 2 results, about $0.004 (2 × $0.002,
plus $0.00005 for the run start). **Then replace them with the sites you actually want to check.**

### Monitor sites: get only the ones whose tech stack changed, in Slack, email or a webhook

With **Only report sites whose stack changed** on, each run compares every site with the last run of the same list
and returns only the sites that **added or dropped a technology**, with what was added and removed (name and
categories): a site that moved from WooCommerce to Shopify, started using HubSpot, or dropped its analytics. For
sales and agency teams that's the moment to get in touch. Unchanged sites are rechecked but not returned, at $0.40
per 1,000 instead of $2.00 (a quiet week for 1,000 sites: $0.40).

1. Put your sites in **Website URLs**, turn on **Only report sites whose stack changed** (`"onlyStackChanges":
   true`), and click **Start**. This first run returns every site and is the baseline the next runs compare with.
2. Click **Save as a new task** (top right of the actor page). The comparison belongs to that exact list of sites,
   so the task keeps comparing with its own last run. Changing the list starts a new baseline.
3. In Apify Console, open **Schedules**, click **Create new**, set how often in **Schedule setup** (for example
   weekly, Monday at 06:00), then **Add** your task.
4. On the task, open the **Integrations** tab and pick where the changes go:
   - **Slack**: click **Configure**, sign in, pick the workspace and channel, and the "run succeeded" event. A
     useful message: `{{resource.statusMessage}}` and a link to the results,
     `<https://console.apify.com/storage/datasets/{{resource.defaultDatasetId}}|stack changes>`. Each result's
     `changeSummary` reads like "added: Shopify (Ecommerce); removed: WooCommerce (Ecommerce)".
   - **Gmail**: click **Connect with Google**, set the subject and body, and attach the dataset (for example as
     CSV). It sends after each successful run.
   - **HTTP webhook**: event `ACTOR.RUN.SUCCEEDED`, your URL. Apify POSTs `{"eventType": ..., "resource": {...}}`;
     `resource.defaultDatasetId` is the run's dataset, and
     `GET https://api.apify.com/v2/datasets/<defaultDatasetId>/items?format=json` (with your API token) returns the
     changed sites.

Apify's integrations fire after every successful run, including quiet ones: a quiet run's dataset is empty, and its
status message says how many sites were unchanged. The pre-filled two sites, run twice with the option on (local
runs, 2026-09-25): the first returned both ($0.004); the second returned 0 and said "2 websites analysed out of 2;
since the run of 2026-09-25 23:54 UTC: 0 changed, 0 new, 2 unchanged (not returned); 0 returned".

**A site that fails is never reported as having dropped its technologies.** When a site is down, blocked by
robots.txt or a bot check, refuses the request or isn't an HTML page, it isn't compared at all: it's listed as failed
in the log and `RUN_STATS`, isn't charged, and its last known stack is kept for the next run. A page that suddenly
shows no technologies at all (usually a maintenance, parking or error page) is treated the same way. See
[Stack changes since the last run](#stack-changes-since-the-last-run) for exactly what's compared.

### What data does Website Technology Detector return?

| Field | Example | Notes |
|---|---|---|
| `input`, `url`, `finalUrl` | `https://wordpress.org/` | What you typed, what was fetched, where it ended. |
| `statusCode`, `pageTitle` | `200`, `Blog Tool, Publishing Platform, and CMS – WordPress.org` | |
| `technologyNames` | `["Google Tag Manager", "Nginx", "PHP", "WordPress 7.2", ...]` | Name plus version, when known. |
| `technologies[].name`, `.categories` | `WordPress`, `["CMS", "Blogs"]` | |
| `technologies[].version` | `7.2` | `null` when the page doesn't reveal it. |
| `technologies[].confidence` | `100` | 0-100, summed over independent signals. |
| `technologies[].evidence` | `["meta generator: WordPress 7.2-alpha-63914"]` | The header, cookie name, meta tag, script URL or HTML that matched. |
| `categories` | `{"CMS": ["WordPress"], "Web servers": ["Nginx"]}` | Technologies grouped by category. |
| `technologyCount` | `11` | |
| `analysisComplete` | `true` | `false` when a time limit cut the matching short. |
| `changeType`, `technologiesAdded`, `technologiesRemoved`, `changeSummary` | `"changed"`, `[{"name": "Shopify", "categories": ["Ecommerce"], ...}]`, `[...]`, `"added: Shopify (Ecommerce)"` | With **Only report sites whose stack changed**, after the first run; `null` otherwise. |
| `sourceTitle`, `sourcePlaceId`, `sourceIndex` | `"Rosie's Bakery"`, `"ChIJ..."`, `12` | From the dataset item the website came from (see [Enrich Google Maps leads](#enrich-google-maps-leads-who-runs-shopify-wordpress-or-wix)); `null` for URLs you type in, and `sourcePlaceId` when the item has no `placeId`. |

One result per website. The full list is under [Output](#output).

### How much does it cost to detect a website's tech stack?

You pay per website analysed: **$2.00 per 1,000 websites**, plus $0.00005 each time a run starts. A URL that fails
(site down, blocked by robots.txt or a bot check, not an HTML page, wrong domain) returns no result and isn't
charged. With **Only report sites whose stack changed** on, a site that's rechecked and found unchanged isn't
returned and costs **$0.40 per 1,000** ($0.0004 a site).

**It's cheaper on paid Apify plans:** $1.80 per 1,000 websites on Starter, $1.60 on Scale and $1.40 on Business. The prices on this page are the Free-plan price, so on a paid plan you pay less than the examples show.

- **The example below:** 2 websites × $0.002 = $0.004, plus the start fee.
- **A month, for example:** an agency checking 1,000 new lead domains a week in one run each: 4 × 1,000 × $0.002 =
  $8.00, plus 4 starts ($0.0002): **about $8.00**.
- **Monitoring the same list, for example** 1,000 prospect sites weekly with **Only report sites whose stack
  changed**: week 1 (the baseline) 1,000 × $0.002 = $2.00; each later week where 20 sites changed, 20 × $0.002 +
  980 × $0.0004 = $0.43. **About $3.30 a month**, against $8.00 without the option.
- **Caps:** **Max results per run** in the input, and **Maximum cost per run** in the run options. The run stops
  cleanly at whichever comes first; it stops fetching as soon as the limit is covered, so a capped run is also a fast
  one.

### How to detect what a website is built with

1. Open Website Technology Detector and click **Try for free** (or **Start** if you're signed in).
2. Put the sites in **Website URLs**, one per line: a domain (`example.com`) or a full page address.
3. Click **Start**, then open the **Output** tab and export as JSON, CSV or Excel.

### Enrich Google Maps leads: who runs Shopify, WordPress or Wix?

Got a list of businesses from a Google Maps scraper? Website Technology Detector reads its dataset directly and
tells you what each business's website runs on, so you can pick out the Shopify shops, the WordPress sites or the
restaurants still on Wix.

1. Run a Google Maps scraper for your search (for example "bakeries in Denver") and let it finish.
2. Open Website Technology Detector and pick that run's dataset in **Or: websites from a dataset** (`datasetId`;
   the dataset id also works). **Website URLs** is ignored then.
3. Leave **Field with the website** empty: the `website` field is found automatically. The place's own Google Maps
   link (a Maps scraper's `url`) is never used. If your scraper puts the website somewhere else, type the field's
   name, for example `contact.website`.
4. Click **Start**. In the **Output** tab, the **Leads (from a dataset)** view shows each business's name
   (`sourceTitle`), its technologies, its Google `sourcePlaceId` and its row in your dataset (`sourceIndex`, 0 = the
   first). Export as CSV and join it to your leads on `sourcePlaceId` or `sourceIndex`.

What happens to each place:

- **No website, or not a website** (empty, a phone number, a Google Maps link): skipped, not fetched, not charged,
  and counted in the run's `RUN_STATS` record (`dataset.withoutWebsite`, `dataset.notAWebsite`,
  `dataset.googleMapsLinks`).
- **Branches sharing a website** (the same domain, with or without `www.`): analysed and charged once, with the
  first place's name and ids. Tracking parameters such as `utm_source` are dropped from the address.
- **Everything else** is analysed like a URL you typed in, at the same price, $2.00 per 1,000 websites; a site that
  fails isn't charged. 400 places with a website cost at most 400 × $0.002 = $0.80, plus the start fee. The Google
  Maps scraper's own run is billed separately by that actor.
- **Limits:** the first 20,000 items and 10,000 distinct websites of a dataset per run; **Max results per run** and
  **Maximum cost per run** stop the run as usual. The dataset is read with your own Apify account's access,
  read-only.

### Example: two websites

The pre-filled input:

```json
{"urls": ["https://wordpress.org/", "https://www.python.org/"]}
```

It returned 2 results, 11 technologies each. The wordpress.org one (real output from a local run on 2026-09-25;
4 of its 11 technologies shown, and some evidence lines left out):

```json
{
  "id": "31368c20a9917549c78ccac6",
  "input": "https://wordpress.org/",
  "url": "https://wordpress.org/",
  "finalUrl": "https://wordpress.org/",
  "statusCode": 200,
  "pageTitle": "Blog Tool, Publishing Platform, and CMS – WordPress.org",
  "technologyCount": 11,
  "analysisComplete": true,
  "technologyNames": ["Google Font API", "Google Tag Manager", "Gutenberg 24.0.0", "HSTS", "MySQL", "Nginx",
                      "Open Graph", "PHP", "Priority Hints", "RSS", "WordPress 7.2"],
  "technologies": [
    {"name": "WordPress", "categories": ["CMS", "Blogs"], "version": "7.2", "confidence": 100,
     "website": "https://wordpress.org",
     "evidence": ["header link: <https://wordpress.org/wp-json/>; rel=\"https://api.w.org/\"",
                  "meta generator: WordPress 7.2-alpha-63914"]},
    {"name": "Gutenberg", "categories": ["WordPress plugins", "Editors"], "version": "24.0.0", "confidence": 100,
     "website": "https://github.com/WordPress/gutenberg",
     "evidence": ["element link[href*='/wp-content/plugins/gutenberg/'] href: https://wordpress.org/wp-content/plugins/gutenberg/build/styles/block-library/navigation/style.min.css?ver=24.0.0"]},
    {"name": "Google Tag Manager", "categories": ["Tag managers"], "version": null, "confidence": 100,
     "website": "http://www.google.com/tagmanager",
     "evidence": ["inline script: googletagmanager.com/gtm.js", "inline script: 'dataLayer','GTM-P24PF4B'"]},
    {"name": "MySQL", "categories": ["Databases"], "version": null, "confidence": 100,
     "website": "http://mysql.com", "evidence": ["implied by WordPress"]}
  ],
  "categories": {"CMS": ["WordPress"], "Tag managers": ["Google Tag Manager"], "Web servers": ["Nginx"],
                 "Programming languages": ["PHP"], "Databases": ["MySQL"]},
  "scrapedAt": "2026-09-25T20:25:16.703568Z"
}
```

(`categories` shortened too.) The python.org result found Nginx, Varnish, Fastly and jQuery 1.8.2, among others.

### How accurate is it?

Measured on 181 public home pages (September 2026: 60 online shops on Shopify, WooCommerce, Magento, BigCommerce and
headless stacks, 48 SaaS marketing sites, 36 small businesses and organisations on Squarespace, Wix, WordPress and
Drupal, 13 docs sites, 12 news sites, 12 big brands) against a hand-checked answer for 28 technologies. Each "yes"
in the answer key is backed by something visible in that page's HTML, headers or cookie names (a script URL, a
generator tag, a response header), written down per site.

| Technology group | Found | Precision | Recall |
|---|---|---|---|
| CMS and site builders: WordPress, Drupal, Webflow, Squarespace, Wix | 48 of 51 | 100% | 94% |
| E-commerce: Shopify, WooCommerce, Magento, BigCommerce | 51 of 51 | 100% | 100% |
| Frameworks: Next.js, Nuxt.js, Gatsby | 41 of 41 | 100% | 100% |
| Analytics and tags: Google Analytics, Google Tag Manager, Segment, Plausible, Hotjar, Facebook Pixel | 124 of 126 | 100% | 98% |
| CDN: Cloudflare, Fastly, Amazon CloudFront | 135 of 135 | 100% | 100% |
| Marketing and consent: HubSpot, Marketo, Optimizely, OneTrust | 46 of 50 | 100% | 92% |
| **All 28 technologies** | **447 of 456** | **100%** | **98%** |

No false positives. On the 65 of those sites that were checked only after the fingerprints were final, recall was
96%. Before version 1.1 (the 2023 fingerprints alone) recall was 76%. What these numbers don't cover, honestly:

- **Tools a tag manager or script adds later, in a browser**, are invisible here, and they were not counted either
  (the answer key only has what the page itself shows). A browser-based tool will list more of those.
- **What it still misses:** WordPress used headless behind another front end (only media URLs show it), Google Tag
  Manager served from the site's own domain, and Optimizely set up in code rather than with its standard snippet.
- **Too few sites to measure:** chat widgets (Intercom, Drift) and payments (Stripe) are fingerprinted, but almost
  never appear in a home page's HTML: 1, 0 and 1 of 181 sites. Don't rely on this actor for those.
- Precision was measured for these 28 technologies; the other 3,500 are reported with their evidence so you can
  check them.
- About 1 site in 11 refuses automated visitors (HTTP 403, or a robots.txt that can't be read); those return no
  result and aren't charged. So does a site that answers with a bot check, block page or waiting room even with
  HTTP 200 (1 more site in the sample): it's reported as blocked, not analysed as if it were the site.

### Ready-to-run examples

Each example opens with the input already filled in. Run it as it is, or change the input first.

- **[Detect the CMS and tech stack of websites](https://apify.com/humble-echidna/tech-stack-detector/examples/detect-website-tech-stack)**: Find out which CMS, frameworks, analytics, e-commerce and hosting tools a list of websites uses, with a confidence level and the evidence for each detection.
- **[Build a lead list from websites’ tech stacks](https://apify.com/humble-echidna/tech-stack-detector/examples/website-tech-leads)**: Check a list of company websites and get each one’s CMS, e-commerce platform, frameworks, analytics, CDN and hosting, with the evidence for each detection.

### Input

| Field | What it does |
|---|---|
| **Website URLs** | The pages to analyse, one per line: a full address (`https://www.example.com/shop`) or just a domain (`example.com` means its `https://` home page). |
| Only report sites whose stack changed | Monitoring: return only the sites that added or dropped a technology since the last run of this list (the first run returns every site). Up to 10,000 sites per list. |
| Or: websites from a dataset (`datasetId`) | One of your Apify datasets, e.g. a Google Maps scraper's results: each item's website is analysed and **Website URLs** is ignored. See [Enrich Google Maps leads](#enrich-google-maps-leads-who-runs-shopify-wordpress-or-wix). |
| Field with the website (`datasetUrlField`) | Only with a dataset. Empty (default): found automatically among `website`, `url`, `domain`, `site` and `homepage`, then one level down (`contact.website`). |
| Max results per run | Cap the total number of websites analysed (with the option above: checked, returned or not). |

```json
{
  "urls": ["https://wordpress.org/", "shopify.com", "https://www.example.com/shop"],
  "maxResults": 100
}
```

The same page typed twice (`example.com` and `https://example.com/`) is analysed and charged once.

### Output

One result per website:

```json
{
  "id": "5d1e0c0f3c2e4a8b9f6a7d21",
  "input": "shop.example.com",
  "url": "https://shop.example.com/",
  "finalUrl": "https://shop.example.com/home",
  "statusCode": 200,
  "pageTitle": "A small WordPress shop",
  "technologyCount": 9,
  "analysisComplete": true,
  "technologyNames": ["Nginx 1.25.3", "PHP 8.2.1", "WordPress 6.4.2", "Google Analytics", "jQuery", "MySQL"],
  "technologies": [
    {
      "name": "WordPress",
      "categories": ["CMS", "Blogs"],
      "version": "6.4.2",
      "confidence": 100,
      "website": "https://wordpress.org",
      "evidence": [
        "meta generator: WordPress 6.4.2",
        "script: https://shop.example.com/wp-includes/js/jquery/jquery.min.js?ver=3.7.1"
      ]
    },
    {
      "name": "MySQL",
      "categories": ["Databases"],
      "version": null,
      "confidence": 100,
      "website": "http://mysql.com",
      "evidence": ["implied by WordPress"]
    }
  ],
  "categories": {"CMS": ["WordPress"], "Databases": ["MySQL"], "Web servers": ["Nginx"]},
  "scrapedAt": "2026-09-24T12:00:00Z"
}
```

(Shortened: the real result lists every technology found; `changeType`, `technologiesAdded`,
`technologiesRemoved`, `changeSummary` and `previousCheckedAt` are there too, `null`, and so are `sourceTitle`,
`sourcePlaceId` and `sourceIndex`, which are only filled for websites read from a dataset.) `id` is stable across runs, so
you can use it to deduplicate. `version` is `null` when the page doesn't reveal it. A page where nothing is
recognised is still a result, with an empty `technologies` list.

#### Stack changes since the last run

With **Only report sites whose stack changed** on, the run remembers each site's technologies (names and categories,
not the pages) in a key-value store in your own account (`tech-stack-detector-memory`, one record per list of sites)
and compares with it next time. After the first run (the baseline, returned as usual with these fields `null`), each
returned site has:

- `changeType`: `changed` (technologies added or removed) or `new` (no earlier result to compare with, e.g. the site
  failed on every earlier run).
- `technologiesAdded`: each technology found now and not last time, with its `categories` and `version`.
- `technologiesRemoved`: each technology found last time and not now, with its `categories`.
- `changeSummary`: the same in one line, for a Slack message or a spreadsheet column.
- `previousCheckedAt`: when the stack it's compared with was seen.

A changed site, for example (the shape of a real result; the site and values are illustrative):

```json
{
  "url": "https://shop.example.com/",
  "technologyCount": 12,
  "technologyNames": ["Cloudflare", "Klaviyo", "Shopify", "..."],
  "changeType": "changed",
  "technologiesAdded": [{"name": "Shopify", "categories": ["Ecommerce"], "version": null},
                        {"name": "Klaviyo", "categories": ["Marketing automation"], "version": null}],
  "technologiesRemoved": [{"name": "WooCommerce", "categories": ["Ecommerce", "WordPress plugins"]}],
  "changeSummary": "added: Klaviyo (Marketing automation), Shopify (Ecommerce); removed: WooCommerce (Ecommerce, WordPress plugins)",
  "previousCheckedAt": "2026-09-21T06:00:04Z",
  "scrapedAt": "2026-09-28T06:00:03.912Z"
}
```

What it deliberately doesn't report:

- **Removals from a failed fetch.** A site that is down, blocked or refuses isn't compared; its last known stack is
  kept, so when it answers again it's compared with that, not reported as new or re-added.
- **"Everything removed" from an empty page.** A site where nothing at all is found this time, after something was
  found before, is listed in `RUN_STATS` as `sitesUnsure`, not returned and not charged, and keeps its old stack.
- **Removals when the analysis was cut short.** When a time limit stopped the matching (`analysisComplete: false`),
  technologies it didn't find this time aren't reported as removed; additions still are.
- **Version changes.** Only technologies coming and going count: a version is only known when a page happens to
  show it, so comparing versions would report noise.

`RUN_STATS` has `baseline`, `previousRun` (when the list was last checked), `sitesNew`, `sitesChanged`,
`sitesUnchanged`, `sitesUnsure` and `unchangedSitesCharged`. A site cut by your max results or maximum cost per run
isn't remembered, so its change is reported on the next run instead.

### Run it on a schedule, or from your own code

1. Save your input as a **task** (**Save as a new task**, top right of the actor page) and add it to a
   **schedule** (Console → Schedules → **Create new**): for example weekly or monthly for a fixed list of client or
   competitor sites. To hear only about sites whose stack changed, turn on **Only report sites whose stack
   changed**; the steps are in
   [Monitor sites](#monitor-sites-get-only-the-ones-whose-tech-stack-changed-in-slack-email-or-a-webhook).
2. Collect results: download the dataset as JSON, CSV or Excel; fetch the latest run's results from the API
   (`GET https://api.apify.com/v2/actor-tasks/<task id>/runs/last/dataset/items?status=SUCCEEDED&format=csv`, with
   your API token); let a webhook tell your system when a run succeeds; or connect it to Make, Zapier or n8n
   through Apify's integrations.

#### Can I use Website Technology Detector from an AI agent (MCP)?

Yes, through Apify's MCP server: add `https://mcp.apify.com?tools=humble-echidna/tech-stack-detector` to your MCP
client (or let the agent find it with the server's actor search). The agent passes the sites, e.g.
`{"urls": ["example.com"]}`, and reads `technologyNames`, or `evidence` when it needs to check an answer. To enrich
another actor's results, it passes that run's dataset instead: `{"datasetId": "<defaultDatasetId>"}`.

### Who it's for

Agencies and freelancers qualifying leads by what a site runs on (every WordPress or Shopify shop in a list, sites
still on an old jQuery), SaaS teams researching which tools competitors and prospects use, and sales ops teams adding
a "built with" column to a lead list before it goes into the CRM. The recurring job: run each new batch of lead or
prospect domains through it, or re-check a fixed list every week or month. It sees what a site's home page (or
whichever page you give it) shows every visitor, so read "How accurate is it?" before relying on a given
technology.

### Why this one?

- **One flat price, no usage bill.** You pay per result at the price above and nothing else: no compute, proxy or storage charges on top, apart from the $0.00005 run-start fee, so you know what a run costs before it starts. Results you don't get aren't charged.
- **Evidence for every answer.** Each technology says why it was detected, so a surprising result can be checked in
  seconds instead of trusted blindly.
- **Versions and confidence.** Versions come from generator tags, headers and script URLs; confidence adds up across
  independent signals (0-100).
- **Stack changes as a monitor.** Re-run the same list on a schedule and get only the sites that added or dropped a
  technology, with both lists and their categories; unchanged sites cost a fifth of a result. A failed or blocked
  fetch is never reported as a removal.
- **You only pay for websites analysed.** A URL that fails (site down, blocked by robots.txt or a bot check, not an
  HTML page, wrong domain) returns no result and isn't charged.
- **Polite and safe.** It identifies itself honestly (User-Agent `HumbleEchidnaApify`), follows each site's
  robots.txt (read once per site per run) and Crawl-delay, makes at most 2 requests per site at once, and only
  requests public web addresses on the standard ports (80 and 443).
- **Reliable.** One failing URL never affects the others in your run, and no page can stall it (every page has a
  size and time limit). The run log and the `RUN_STATS` record say exactly which URL had a problem and why.

### Limits

- One page per URL, without running its JavaScript: tools a tag manager adds later in a browser are missed (see
  "How accurate is it?").
- Pages over 3 MB are cut at 3 MB, and matching has time limits per page (see `analysisComplete` in the FAQ).
- Only public web addresses on the standard ports (80 and 443); no login, no proxy. Sites that refuse automated
  visitors or answer with a bot check are reported and not charged.
- Fingerprints for products launched since 2023 may be missing unless we've added them.

### FAQ

#### Can I get only the sites whose technology stack changed since my last run?

Yes: turn on **Only report sites whose stack changed** and run the same list on a schedule. The first run returns
every site; after that each run returns only the sites that added or dropped a technology, with `technologiesAdded`,
`technologiesRemoved` (names and categories) and a one-line `changeSummary`. Unchanged sites are rechecked but not
returned ($0.40 per 1,000). See [Monitor sites](#monitor-sites-get-only-the-ones-whose-tech-stack-changed-in-slack-email-or-a-webhook).

#### Why are unchanged sites charged at all?

Because finding out that a site didn't change takes the same work as analysing it: the page is downloaded and
matched against every fingerprint again. Unchanged sites cost a fifth of a returned site, and never when the site
failed, was blocked or showed nothing.

#### A site went down or blocked the run. Will it show up as "all technologies removed"?

No. A site that fails (down, robots.txt, a bot check, a 403, not HTML) isn't compared and keeps its last known stack;
a page where nothing at all is found, after something was before, is treated the same way (`sitesUnsure` in
`RUN_STATS`). Neither is returned or charged. When the site answers normally again, it's compared with the stack it
had before.

#### Why wasn't a technology I know is there detected?

This actor reads one page's HTML, inline scripts, response
headers and cookie names, without running the page's JavaScript. Technologies that only show up once scripts run in
a browser (tools loaded later by a tag manager, single-page apps that build everything in the browser) and anything
only visible in DNS or TLS certificates are missed. Tools used only on other pages (a checkout, a blog) are found
only if you give that page's URL too. The fingerprints are the last openly (MIT) licensed version of the Wappalyzer
technology data, from January 2023, plus our own current fingerprints for the most-used tools (see "How accurate is
it?"), so products launched since 2023 that we haven't added may be unknown.

#### What does `confidence` mean?

Each fingerprint pattern carries a weight (100 unless the data says otherwise);
a technology's confidence is the sum of its matched patterns, capped at 100. A technology that is only implied by
another (PHP by WordPress) is never more confident than what implied it.

#### What does `analysisComplete: false` mean?

Every page has limits so that one unusual page can't stall a run:
pages over 3 MB are cut at 3 MB, the page source is matched on its first and last 1,500 lines (2,000 characters
each), CSS selectors look at the first 10,000 elements, each pattern gets 0.1 s per value and each page 15 s of
matching in total. When one of the time limits was hit, the result says so; everything listed was still really found.

#### Why did a URL come back as "blocked by robots.txt"?

The site's owner has asked crawlers not to fetch that page. This
actor respects that, and the URL is not charged.

#### Why does it refuse `localhost`, `10.x.x.x`, an internal hostname or a URL with a port?

It only fetches public web pages,
on the standard web ports (80 for `http://`, 443 for `https://`); a URL with any other port, such as `:8080`, is
refused. Every hostname (and every redirect) is resolved first, and a page is refused if any address it resolves to
is private, loopback, link-local (cloud metadata) or otherwise not on the public internet; the connection then goes
to the address that was checked. These refusals are reported as "not a public web address" and aren't charged. A
domain that doesn't exist is reported as such.

#### What happens when a site answers "403" or shows a bot check?

Some sites refuse automated visitors. This actor doesn't try to get
around that: the URL is reported as refused and not charged. The same goes for a challenge, block or waiting-room
page served with an ordinary "200" (Cloudflare "Just a moment...", Akamai "Access Denied" and failover pages,
PerimeterX / HUMAN, DataDome, Imperva / Incapsula, Kasada, Queue-it, or a near-empty "verify you are human" page):
the log says "blocked: the site answered with a bot check" and names it. A normal page that merely loads Cloudflare's
or another vendor's scripts is analysed as usual.

#### Something that used to work now fails. Why?

Sites change without notice. The run log names the URL and what went
wrong, and every other URL in the run is unaffected. Please open an issue with the input you used.

#### Is it legal to detect a website's technology stack?

It fetches only the pages you give it, once each, as a normal logged-out visitor, and follows
each site's `robots.txt`. It doesn't log in, get around any protection, or collect personal data: it reports which
products a site uses, from what the site sends to every visitor (cookie values are never output, only cookie
names). The Wappalyzer fingerprint data is used under its MIT licence, which ships with the actor; the supplementary
fingerprints are our own. What you do with the results
(and any site's own terms) is your responsibility.

### Related actors

| Actor | Use it when |
|---|---|
| [SEO Audit Crawler](https://apify.com/humble-echidna/seo-audit) | You want a full on-page audit (titles, H1s, canonicals, broken links per page) of the same sites, for a fuller picture of a prospect's website. |
| [Bulk URL & Broken Link Checker](https://apify.com/humble-echidna/url-checker) | You want to clean a lead list of dead or redirected domains before analysing it. |
| [Sitemap URL Extractor](https://apify.com/humble-echidna/sitemap-urls) | You want to analyse pages other than the home page: it lists every URL a site's sitemap has. |

### Feedback and support

Found a bug, a wrong detection or a missing technology? Open an issue on the **Issues** tab with the URL and what
you expected.

### Versions

Current version: **1.3**. See the Changelog tab for what changed in each version.

# Changelog

This Actor's version history is a separate document: https://apify.com/humble-echidna/tech-stack-detector/changelog.md

# Actor input Schema

## `urls` (type: `array`):

The pages to analyse, one per line: a full address (https://www.example.com/shop) or just a domain (example.com, which means its https:// home page). Each URL is fetched once and returns one result listing the technologies found on it. Ignored when datasetId is set.

## `datasetId` (type: `string`):

One of your Apify datasets, picked here or given by id, e.g. a Google Maps scraper's results: each item's website is analysed (urls is ignored). Items without a website, values that aren't a web address and Google Maps links are skipped; a website shared by several items is analysed once. Each result carries sourceTitle, sourcePlaceId and sourceIndex from its item. At most the first 20,000 items and 10,000 websites are read, with your own account's access, read-only.

## `datasetUrlField` (type: `string`):

Only with datasetId: the item field that holds the website, e.g. `website`, or a dotted path such as `contact.website`. Leave empty (the default) to find it automatically: the first of website, url, domain, site and homepage that has a website in the first 100 items, then the same names one level down (contact.website).

## `onlyStackChanges` (type: `boolean`):

For scheduled monitoring of the same list. Each run compares every site with the last run of this exact list and returns only the sites that added or dropped a technology, with what was added and removed (name and categories). The first run returns every site and is the baseline. Unchanged sites are rechecked but not returned. A site that fails or is blocked is never reported as having dropped anything. Changing the list starts a new baseline.

## `maxResults` (type: `integer`):

Stop after this many websites in total (with Only report sites whose stack changed: websites checked, returned or not). Leave empty for no limit. The run also stops cleanly at the maximum cost per run you set in the run options, whichever comes first.

## Actor input object example

```json
{
  "urls": [
    "https://wordpress.org/",
    "https://www.python.org/"
  ],
  "onlyStackChanges": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `runStats` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://wordpress.org/",
        "https://www.python.org/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("humble-echidna/tech-stack-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://wordpress.org/",
        "https://www.python.org/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("humble-echidna/tech-stack-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://wordpress.org/",
    "https://www.python.org/"
  ]
}' |
apify call humble-echidna/tech-stack-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,humble-echidna/tech-stack-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/IwXnuxGAHrWS9RHFY/builds/RR9ifTafqf91qtTCF/openapi.json
