# Changelog of Website Technology Detector: Tech Stack & CMS Lookup (`humble-echidna/tech-stack-detector`) Actor

- **URL**: https://apify.com/humble-echidna/tech-stack-detector/changelog.md
- **Full Actor documentation**: https://apify.com/humble-echidna/tech-stack-detector.md

## Changelog

Versions follow MAJOR.MINOR.PATCH (`src/version.py`); Apify shows MAJOR.MINOR from `.actor/actor.json`.
Every run logs its version and records it in the `RUN_STATS` key-value record.

### 1.3.2 (2026-09-28)

- 2 ready-to-run examples published on Apify Store and linked from the README.

### 1.3.1 (2026-09-28)

- Cheaper on paid Apify plans: $1.80 per 1,000 websites on Starter, $1.60 on Scale and $1.40 on Business ($2.00 on the Free plan).

### 1.3.0 (2026-09-27)

Enrich another actor's results, e.g. Google Maps leads: who runs Shopify, WordPress or Wix.

- New inputs **Or: websites from a dataset** (`datasetId`, a dataset picker with read access) and **Field with the
  website** (`datasetUrlField`). With a dataset, each item's website is analysed and `urls` is ignored. The field is
  found automatically when left empty: the first of `website`, `url`, `domain`, `site`, `homepage` with a website in
  the first 100 items, then the same names one level down (`contact.website`); a dotted path names a nested one.
- Websites sharing a host (with or without `www.`, e.g. a chain's branches) are analysed and charged once, with the
  first item's source fields.
- Items without a website, values that aren't a public web address (a phone number, `N/A`, an email) and Google
  Maps links (the place's own `url` in a Maps scraper's output) are skipped, never fetched or charged, and counted in
  the new `RUN_STATS.dataset` record (`itemsRead`, `websites`, `withoutWebsite`, `notAWebsite`, `googleMapsLinks`,
  `duplicates`, the field used and whether it was found automatically). Tracking parameters (`utm_*`, `gclid`,
  `fbclid`, ...) are dropped from the address. A dataset with no website at all, or one that can't be read, fails
  the run with the reason.
- New nullable fields on every result: `sourceTitle` (the item's `title`, else `name`), `sourcePlaceId` (its
  `placeId`) and `sourceIndex` (its position in the dataset, 0 = the first), to join results back to the leads;
  `null` for URLs typed in. A new "Leads (from a dataset)" dataset view shows them next to the key columns.
- Read with the run's own token (the user's access, read-only), only the fields needed (whole items only
  for the first 100 when the field is looked for one level down); at most the first 20,000
  items and 10,000 distinct websites per run. Max results / maximum cost per run apply as before. Outside Apify, a
  dataset in local storage with that name is read.
- Pricing unchanged: no new events; a website from a dataset costs the same as one typed in.
- Internal: the reader is shared (`mms_common.dataset_input`), as are its test stand-ins (`mms_testing`).
- Store SEO description now mentions Google Maps lead lists (`.actor/store.json`; synced by push.sh).

### 1.2.1 (2026-09-26)

- Internal: the detector and its fingerprint data moved into a shared package (`mms_stack`) that Hiring Signals
  also uses, so a fingerprint fix reaches both. Results are unchanged (the same tests pass).

### 1.2.0 (2026-09-25)

Stack changes: a monitoring mode that reports only the sites whose technologies changed since the last run.

- New input **Only report sites whose stack changed** (`onlyStackChanges`, off by default). Each run compares every
  site with the last run of the same list (one record per list in the `tech-stack-detector-memory` key-value store
  in the user's account, keyed by a hash of the sorted canonical URLs, so two lists never hide each other's
  changes) and returns only the sites that added or dropped a technology. New fields on every result (`null` when the
  mode is off and on its first, baseline run): `changeType` (`changed` or `new`), `technologiesAdded` (name,
  categories, version), `technologiesRemoved` (name, categories), `changeSummary` (one line) and `previousCheckedAt`.
  A "Stack changes" dataset view shows them.
- Never reported as removed: technologies of a site whose fetch failed or was blocked (not compared, last stack
  kept), of a page where nothing at all was found after something was before (`sitesUnsure`, not charged, last
  stack kept), or technologies missing from an analysis a time limit cut short (kept until a complete one). Version
  changes alone don't count.
- Charging in the mode: returned sites are charged as before (`apify-default-dataset-item`); a site rechecked and
  found unchanged is charged a new event, `unchanged-site`, only once that event is priced (proposed $0.40 per
  1,000; until then it's free); failed, blocked and "unsure" sites are free. The maximum cost per run limits the
  sites read, and a changed site the maximum cost can't pay for isn't returned and is reported on the next run.
  "Max results per run" caps the sites checked in the mode.
- Store description now mentions the monitor (actor.json; synced to the Store by push.sh).
- Up to 10,000 sites per list in the mode; a longer list fails at the start with a clear message.
- `RUN_STATS` gains `onlyStackChanges`, `baseline`, `previousRun`, `sitesNew`, `sitesChanged`, `sitesUnchanged`,
  `sitesUnsure` and `unchangedSitesCharged` (same keys on every exit path).
- Faster and cheaper for every run: pages are still downloaded 5 at a time, but analysed one at a time. Parallel
  analyses only competed for Python's interpreter lock: on the same 30 live sites, 30.2 CPU-seconds and 31 s before,
  10.6 CPU-seconds and 20 s after, with the same technology count on every site.
- Internal: the only-changed spending logic (`Spend`) moved from seo-audit to the shared library
  (`mms_common.spend`), now used by both actors.

### 1.1.1 (2026-09-25)

- A bot check, block page or waiting room served with HTTP 200 is no longer analysed and charged as if it were the
  site: it's reported as "blocked: the site answered with a bot check (...)" and not charged, like a 403. Recognised:
  Cloudflare challenge and block pages (and `cf-mitigated: challenge`), Akamai "Access Denied" and failover pages,
  PerimeterX / HUMAN, DataDome, Imperva / Incapsula, Kasada, Queue-it waiting rooms, and near-empty "verify you are
  human" / "enable JavaScript and cookies" pages. Normal pages that load those vendors' scripts are still analysed
  (the benchmark's 181 real pages: none refused; the one bot page in it, Patagonia's Akamai failover page, caught).
  The check lives in the shared library (`mms_common.bot_challenge`) for the other actors.

### 1.1.0 (2026-09-25)

Measured and improved accuracy. On 181 public home pages with hand-checked answers for 28 technologies, recall went
from 76% to 98% at 100% precision (0 false positives, down from 3); on the 65 sites checked only after the new
fingerprints were final, 78% to 96%. The benchmark, its answer key with the evidence per site, and the script that
re-runs it are in `benchmark/` (not shipped with the actor).

- Supplementary fingerprints of our own (`src/fingerprints/supplementary.json`), written from the evidence in those
  pages and each vendor's standard install snippet: Fastly (X-Served-By / X-Timer headers), Next.js (`/_next/static/`
  URLs, App Router pages, `x-nextjs-*` headers), and the inline install snippets of Google Tag Manager, Google
  Analytics, Segment, Hotjar, Facebook Pixel, Plausible (current `pa-*.js` script), Marketo (Munchkin and forms),
  OneTrust, Intercom, HubSpot and Stripe, plus headless BigCommerce image URLs. They add to the 2023 data.
- Inline scripts are now read: the data's `scripts` patterns (script content) are matched on each page's inline
  `<script>` blocks (first 100,000 characters of each, 1,000,000 in all; JSON blocks skipped). Before, snippets on a
  long one-line page were cut off by the 2,000-characters-per-line limit of the page source.
- Scripts named by `<link rel="preload" as="script">` or `<link rel="modulepreload">` count as script URLs.
- Four wrong rules in the 2023 data corrected: cdnjs no longer implies the site uses Cloudflare, RankMath no longer
  implies Google Analytics, Module Federation is no longer reported from the `data-webpack` attribute every webpack
  runtime writes, and RxJS no longer matches hashed bundle names ending in "rx.js".
- README: a "How accurate is it?" section with the measured numbers and what it can't detect.

### 1.0.0 (2026-09-24)

First release.

- For each website URL the user lists, fetches that one page (following redirects) and reports the technologies
  found on it: CMS, e-commerce platform, web frameworks, analytics, tag managers, advertising, CDN, hosting/PaaS, web
  server, programming language, JavaScript libraries and more. Per technology: name, categories, version (when the
  page reveals it), confidence (0-100) and up to 5 pieces of evidence (the header, cookie name, meta tag, script URL,
  HTML excerpt or element that matched). One result per URL, also with the final URL, HTTP status, page title, a
  flat "technologyNames" list and the technologies grouped by category.
- Evidence read: response headers, cookie names set by the page, `<meta>` tags, `<script src>` URLs, the page
  source, the visible text, the final URL, and CSS selectors made of one compound part (`tag#id.class[attr*=value]`).
  Not read: anything that needs a browser to run the page's JavaScript (window variables, XHR calls), DNS records or
  TLS certificates; technologies only visible that way are missed.
- Fingerprints: the last MIT-licensed release of the Wappalyzer technology data (commit `aff18efa`, 2023-01-23;
  3,588 technologies, 108 categories), vendored unmodified with its licence in `src/fingerprints/`. Newer versions of
  that data are GPL-3.0 and are not used, so technologies launched since early 2023 may be missing.
- Cookie values never appear in the output: a cookie is shown by its name only, also for the fingerprints that match
  the whole Set-Cookie header (ASP.NET Core, Empretienda, 1C-Bitrix), and a version is never read from a cookie's
  value.
- Implies/excludes/requires rules from the data applied (e.g. WordPress implies PHP and MySQL, with at most the
  implying technology's confidence).
- Bounded work per page, so no page can stall a run: pages over 3 MB are cut at 3 MB; the page source is matched on
  3,000 lines of 2,000 characters at most (the start and end of longer pages), selectors look at the first 10,000
  elements, every unbounded regex repeat is capped at 250, each pattern has 0.1 s per value, and each page has a 15 s
  matching budget. When a limit cuts matching short the result says `analysisComplete: false`.
- Honest identity: User-Agent `HumbleEchidnaApify/1.0 (+https://apify.com/humble-echidna)`. robots.txt checked
  before every request, redirect hops included (read once per site per run); Crawl-delay and Retry-After honoured;
  at most 2 requests per site at once and 5 overall; only public addresses on ports 80 and 443.
- Charged per website analysed through Apify's standard `apify-default-dataset-item` event. URLs that fail (robots.txt
  block, not found, not HTML, private or unresolvable address, server errors) return no result and aren't charged;
  a failed URL gives its result slot back. "Max results per run" input, and the maximum cost per run is honoured:
  no page is fetched that the limit won't pay for (skipped URLs are listed as `queriesSkipped` in `RUN_STATS`).
- The same page typed twice (`example.com` and `https://example.com/`) is analysed and charged once.
- Failure isolation: a failing URL only affects itself; the log and `RUN_STATS` say which URL failed and why.
