Versions follow MAJOR.MINOR.PATCH (src/version.py); Apify shows MAJOR.MINOR from .actor/actor.json.
Every run logs its version and records it in the RUN_STATS key-value record.
- 2 ready-to-run examples published on Apify Store and linked from the README.
- Cheaper on paid Apify plans: $1.80 per 1,000 websites on Starter, $1.60 on Scale and $1.40 on Business ($2.00 on the Free plan).
Enrich another actor's results, e.g. Google Maps leads: who runs Shopify, WordPress or Wix.
- New inputs Or: websites from a dataset (
datasetId, a dataset picker with read access) and Field with the
website (datasetUrlField). With a dataset, each item's website is analysed and urls is ignored. The field is
found automatically when left empty: the first of website, url, domain, site, homepage with a website in
the first 100 items, then the same names one level down (contact.website); a dotted path names a nested one.
- Websites sharing a host (with or without
www., e.g. a chain's branches) are analysed and charged once, with the
first item's source fields.
- Items without a website, values that aren't a public web address (a phone number,
N/A, an email) and Google
Maps links (the place's own url in a Maps scraper's output) are skipped, never fetched or charged, and counted in
the new RUN_STATS.dataset record (itemsRead, websites, withoutWebsite, notAWebsite, googleMapsLinks,
duplicates, the field used and whether it was found automatically). Tracking parameters (utm_*, gclid,
fbclid, ...) are dropped from the address. A dataset with no website at all, or one that can't be read, fails
the run with the reason.
- New nullable fields on every result:
sourceTitle (the item's title, else name), sourcePlaceId (its
placeId) and sourceIndex (its position in the dataset, 0 = the first), to join results back to the leads;
null for URLs typed in. A new "Leads (from a dataset)" dataset view shows them next to the key columns.
- Read with the run's own token (the user's access, read-only), only the fields needed (whole items only
for the first 100 when the field is looked for one level down); at most the first 20,000
items and 10,000 distinct websites per run. Max results / maximum cost per run apply as before. Outside Apify, a
dataset in local storage with that name is read.
- Pricing unchanged: no new events; a website from a dataset costs the same as one typed in.
- Internal: the reader is shared (
mms_common.dataset_input), as are its test stand-ins (mms_testing).
- Store SEO description now mentions Google Maps lead lists (
.actor/store.json; synced by push.sh).
- Internal: the detector and its fingerprint data moved into a shared package (
mms_stack) that Hiring Signals
also uses, so a fingerprint fix reaches both. Results are unchanged (the same tests pass).
Stack changes: a monitoring mode that reports only the sites whose technologies changed since the last run.
- New input Only report sites whose stack changed (
onlyStackChanges, off by default). Each run compares every
site with the last run of the same list (one record per list in the tech-stack-detector-memory key-value store
in the user's account, keyed by a hash of the sorted canonical URLs, so two lists never hide each other's
changes) and returns only the sites that added or dropped a technology. New fields on every result (null when the
mode is off and on its first, baseline run): changeType (changed or new), technologiesAdded (name,
categories, version), technologiesRemoved (name, categories), changeSummary (one line) and previousCheckedAt.
A "Stack changes" dataset view shows them.
- Never reported as removed: technologies of a site whose fetch failed or was blocked (not compared, last stack
kept), of a page where nothing at all was found after something was before (
sitesUnsure, not charged, last
stack kept), or technologies missing from an analysis a time limit cut short (kept until a complete one). Version
changes alone don't count.
- Charging in the mode: returned sites are charged as before (
apify-default-dataset-item); a site rechecked and
found unchanged is charged a new event, unchanged-site, only once that event is priced (proposed $0.40 per
1,000; until then it's free); failed, blocked and "unsure" sites are free. The maximum cost per run limits the
sites read, and a changed site the maximum cost can't pay for isn't returned and is reported on the next run.
"Max results per run" caps the sites checked in the mode.
- Store description now mentions the monitor (actor.json; synced to the Store by push.sh).
- Up to 10,000 sites per list in the mode; a longer list fails at the start with a clear message.
RUN_STATS gains onlyStackChanges, baseline, previousRun, sitesNew, sitesChanged, sitesUnchanged,
sitesUnsure and unchangedSitesCharged (same keys on every exit path).
- Faster and cheaper for every run: pages are still downloaded 5 at a time, but analysed one at a time. Parallel
analyses only competed for Python's interpreter lock: on the same 30 live sites, 30.2 CPU-seconds and 31 s before,
10.6 CPU-seconds and 20 s after, with the same technology count on every site.
- Internal: the only-changed spending logic (
Spend) moved from seo-audit to the shared library
(mms_common.spend), now used by both actors.
- A bot check, block page or waiting room served with HTTP 200 is no longer analysed and charged as if it were the
site: it's reported as "blocked: the site answered with a bot check (...)" and not charged, like a 403. Recognised:
Cloudflare challenge and block pages (and
cf-mitigated: challenge), Akamai "Access Denied" and failover pages,
PerimeterX / HUMAN, DataDome, Imperva / Incapsula, Kasada, Queue-it waiting rooms, and near-empty "verify you are
human" / "enable JavaScript and cookies" pages. Normal pages that load those vendors' scripts are still analysed
(the benchmark's 181 real pages: none refused; the one bot page in it, Patagonia's Akamai failover page, caught).
The check lives in the shared library (mms_common.bot_challenge) for the other actors.
Measured and improved accuracy. On 181 public home pages with hand-checked answers for 28 technologies, recall went
from 76% to 98% at 100% precision (0 false positives, down from 3); on the 65 sites checked only after the new
fingerprints were final, 78% to 96%. The benchmark, its answer key with the evidence per site, and the script that
re-runs it are in benchmark/ (not shipped with the actor).
- Supplementary fingerprints of our own (
src/fingerprints/supplementary.json), written from the evidence in those
pages and each vendor's standard install snippet: Fastly (X-Served-By / X-Timer headers), Next.js (/_next/static/
URLs, App Router pages, x-nextjs-* headers), and the inline install snippets of Google Tag Manager, Google
Analytics, Segment, Hotjar, Facebook Pixel, Plausible (current pa-*.js script), Marketo (Munchkin and forms),
OneTrust, Intercom, HubSpot and Stripe, plus headless BigCommerce image URLs. They add to the 2023 data.
- Inline scripts are now read: the data's
scripts patterns (script content) are matched on each page's inline
<script> blocks (first 100,000 characters of each, 1,000,000 in all; JSON blocks skipped). Before, snippets on a
long one-line page were cut off by the 2,000-characters-per-line limit of the page source.
- Scripts named by
<link rel="preload" as="script"> or <link rel="modulepreload"> count as script URLs.
- Four wrong rules in the 2023 data corrected: cdnjs no longer implies the site uses Cloudflare, RankMath no longer
implies Google Analytics, Module Federation is no longer reported from the
data-webpack attribute every webpack
runtime writes, and RxJS no longer matches hashed bundle names ending in "rx.js".
- README: a "How accurate is it?" section with the measured numbers and what it can't detect.
First release.
- For each website URL the user lists, fetches that one page (following redirects) and reports the technologies
found on it: CMS, e-commerce platform, web frameworks, analytics, tag managers, advertising, CDN, hosting/PaaS, web
server, programming language, JavaScript libraries and more. Per technology: name, categories, version (when the
page reveals it), confidence (0-100) and up to 5 pieces of evidence (the header, cookie name, meta tag, script URL,
HTML excerpt or element that matched). One result per URL, also with the final URL, HTTP status, page title, a
flat "technologyNames" list and the technologies grouped by category.
- Evidence read: response headers, cookie names set by the page,
<meta> tags, <script src> URLs, the page
source, the visible text, the final URL, and CSS selectors made of one compound part (tag#id.class[attr*=value]).
Not read: anything that needs a browser to run the page's JavaScript (window variables, XHR calls), DNS records or
TLS certificates; technologies only visible that way are missed.
- Fingerprints: the last MIT-licensed release of the Wappalyzer technology data (commit
aff18efa, 2023-01-23;
3,588 technologies, 108 categories), vendored unmodified with its licence in src/fingerprints/. Newer versions of
that data are GPL-3.0 and are not used, so technologies launched since early 2023 may be missing.
- Cookie values never appear in the output: a cookie is shown by its name only, also for the fingerprints that match
the whole Set-Cookie header (ASP.NET Core, Empretienda, 1C-Bitrix), and a version is never read from a cookie's
value.
- Implies/excludes/requires rules from the data applied (e.g. WordPress implies PHP and MySQL, with at most the
implying technology's confidence).
- Bounded work per page, so no page can stall a run: pages over 3 MB are cut at 3 MB; the page source is matched on
3,000 lines of 2,000 characters at most (the start and end of longer pages), selectors look at the first 10,000
elements, every unbounded regex repeat is capped at 250, each pattern has 0.1 s per value, and each page has a 15 s
matching budget. When a limit cuts matching short the result says
analysisComplete: false.
- Honest identity: User-Agent
HumbleEchidnaApify/1.0 (+https://apify.com/humble-echidna). robots.txt checked
before every request, redirect hops included (read once per site per run); Crawl-delay and Retry-After honoured;
at most 2 requests per site at once and 5 overall; only public addresses on ports 80 and 443.
- Charged per website analysed through Apify's standard
apify-default-dataset-item event. URLs that fail (robots.txt
block, not found, not HTML, private or unresolvable address, server errors) return no result and aren't charged;
a failed URL gives its result slot back. "Max results per run" input, and the maximum cost per run is honoured:
no page is fetched that the limit won't pay for (skipped URLs are listed as queriesSkipped in RUN_STATS).
- The same page typed twice (
example.com and https://example.com/) is analysed and charged once.
- Failure isolation: a failing URL only affects itself; the log and
RUN_STATS say which URL failed and why.