Website Technology Detector: Tech Stack & CMS Lookup
Pricing
from $1.40 / 1,000 website analyseds
Website Technology Detector: Tech Stack & CMS Lookup
Detect any website's tech stack: CMS, e-commerce platform, frameworks, analytics, tag managers, CDN, hosting and JavaScript libraries, with version, confidence and evidence, from the MIT-licensed Wappalyzer fingerprints plus our own. Monitor a list: get only sites that added or dropped a technology.
Pricing
from $1.40 / 1,000 website analyseds
Rating
0.0
(0)
Developer
Michael Costa
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
14 hours ago
Last modified
Categories
Share
What does Website Technology Detector do?
Website Technology Detector is a bulk tech stack detector: it finds out what a website is built with. Give it a list of sites; for each you get its CMS, e-commerce platform, frameworks, analytics, CDN and hosting, with the version, a confidence score and the evidence behind every technology.
It fetches one page per URL like a normal visitor, without a browser, and matches it against 3,500+ technology fingerprints (the MIT-licensed Wappalyzer data plus our own). That makes it fast and cheap, and it means technologies that only show up once the page's JavaScript runs can be missed (see "How accurate is it?" below).
Try it in one click: the input comes pre-filled with two websites. That's 2 results, about $0.004 (2 × $0.002, plus $0.00005 for the run start). Then replace them with the sites you actually want to check.
Monitor sites: get only the ones whose tech stack changed, in Slack, email or a webhook
With Only report sites whose stack changed on, each run compares every site with the last run of the same list and returns only the sites that added or dropped a technology, with what was added and removed (name and categories): a site that moved from WooCommerce to Shopify, started using HubSpot, or dropped its analytics. For sales and agency teams that's the moment to get in touch. Unchanged sites are rechecked but not returned, at $0.40 per 1,000 instead of $2.00 (a quiet week for 1,000 sites: $0.40).
- Put your sites in Website URLs, turn on Only report sites whose stack changed (), and click Start. This first run returns every site and is the baseline the next runs compare with."onlyStackChanges": true
- Click Save as a new task (top right of the actor page). The comparison belongs to that exact list of sites, so the task keeps comparing with its own last run. Changing the list starts a new baseline.
- In Apify Console, open Schedules, click Create new, set how often in Schedule setup (for example weekly, Monday at 06:00), then Add your task.
- On the task, open the Integrations tab and pick where the changes go:
- Slack: click Configure, sign in, pick the workspace and channel, and the "run succeeded" event. A
useful message:
{{resource.statusMessage}}and a link to the results,<https://console.apify.com/storage/datasets/{{resource.defaultDatasetId}}|stack changes>. Each result'schangeSummaryreads like "added: Shopify (Ecommerce); removed: WooCommerce (Ecommerce)". - Gmail: click Connect with Google, set the subject and body, and attach the dataset (for example as CSV). It sends after each successful run.
- HTTP webhook: event
ACTOR.RUN.SUCCEEDED, your URL. Apify POSTs{"eventType": ..., "resource": {...}};resource.defaultDatasetIdis the run's dataset, andGET https://api.apify.com/v2/datasets/<defaultDatasetId>/items?format=json(with your API token) returns the changed sites.
- Slack: click Configure, sign in, pick the workspace and channel, and the "run succeeded" event. A
useful message:
Apify's integrations fire after every successful run, including quiet ones: a quiet run's dataset is empty, and its status message says how many sites were unchanged. The pre-filled two sites, run twice with the option on (local runs, 2026-09-25): the first returned both ($0.004); the second returned 0 and said "2 websites analysed out of 2; since the run of 2026-09-25 23:54 UTC: 0 changed, 0 new, 2 unchanged (not returned); 0 returned".
A site that fails is never reported as having dropped its technologies. When a site is down, blocked by
robots.txt or a bot check, refuses the request or isn't an HTML page, it isn't compared at all: it's listed as failed
in the log and RUN_STATS, isn't charged, and its last known stack is kept for the next run. A page that suddenly
shows no technologies at all (usually a maintenance, parking or error page) is treated the same way. See
Stack changes since the last run for exactly what's compared.
What data does Website Technology Detector return?
| Field | Example | Notes |
|---|---|---|
input, url, finalUrl | https://wordpress.org/ | What you typed, what was fetched, where it ended. |
statusCode, pageTitle | 200, Blog Tool, Publishing Platform, and CMS – WordPress.org | |
technologyNames | ["Google Tag Manager", "Nginx", "PHP", "WordPress 7.2", ...] | Name plus version, when known. |
technologies[].name, .categories | WordPress, ["CMS", "Blogs"] | |
technologies[].version | 7.2 | null when the page doesn't reveal it. |
technologies[].confidence | 100 | 0-100, summed over independent signals. |
technologies[].evidence | ["meta generator: WordPress 7.2-alpha-63914"] | The header, cookie name, meta tag, script URL or HTML that matched. |
categories | {"CMS": ["WordPress"], "Web servers": ["Nginx"]} | Technologies grouped by category. |
technologyCount | 11 | |
analysisComplete | true | false when a time limit cut the matching short. |
changeType, technologiesAdded, technologiesRemoved, changeSummary | "changed", [{"name": "Shopify", "categories": ["Ecommerce"], ...}], [...], "added: Shopify (Ecommerce)" | With Only report sites whose stack changed, after the first run; null otherwise. |
sourceTitle, sourcePlaceId, sourceIndex | "Rosie's Bakery", "ChIJ...", 12 | From the dataset item the website came from (see Enrich Google Maps leads); null for URLs you type in, and sourcePlaceId when the item has no placeId. |
One result per website. The full list is under Output.
How much does it cost to detect a website's tech stack?
You pay per website analysed: $2.00 per 1,000 websites, plus $0.00005 each time a run starts. A URL that fails (site down, blocked by robots.txt or a bot check, not an HTML page, wrong domain) returns no result and isn't charged. With Only report sites whose stack changed on, a site that's rechecked and found unchanged isn't returned and costs $0.40 per 1,000 ($0.0004 a site).
It's cheaper on paid Apify plans: $1.80 per 1,000 websites on Starter, $1.60 on Scale and $1.40 on Business. The prices on this page are the Free-plan price, so on a paid plan you pay less than the examples show.
- The example below: 2 websites × $0.002 = $0.004, plus the start fee.
- A month, for example: an agency checking 1,000 new lead domains a week in one run each: 4 × 1,000 × $0.002 = $8.00, plus 4 starts ($0.0002): about $8.00.
- Monitoring the same list, for example 1,000 prospect sites weekly with Only report sites whose stack changed: week 1 (the baseline) 1,000 × $0.002 = $2.00; each later week where 20 sites changed, 20 × $0.002 + 980 × $0.0004 = $0.43. About $3.30 a month, against $8.00 without the option.
- Caps: Max results per run in the input, and Maximum cost per run in the run options. The run stops cleanly at whichever comes first; it stops fetching as soon as the limit is covered, so a capped run is also a fast one.
How to detect what a website is built with
- Open Website Technology Detector and click Try for free (or Start if you're signed in).
- Put the sites in Website URLs, one per line: a domain (
example.com) or a full page address. - Click Start, then open the Output tab and export as JSON, CSV or Excel.
Enrich Google Maps leads: who runs Shopify, WordPress or Wix?
Got a list of businesses from a Google Maps scraper? Website Technology Detector reads its dataset directly and tells you what each business's website runs on, so you can pick out the Shopify shops, the WordPress sites or the restaurants still on Wix.
- Run a Google Maps scraper for your search (for example "bakeries in Denver") and let it finish.
- Open Website Technology Detector and pick that run's dataset in Or: websites from a dataset (
datasetId; the dataset id also works). Website URLs is ignored then. - Leave Field with the website empty: the
websitefield is found automatically. The place's own Google Maps link (a Maps scraper'surl) is never used. If your scraper puts the website somewhere else, type the field's name, for examplecontact.website. - Click Start. In the Output tab, the Leads (from a dataset) view shows each business's name
(
sourceTitle), its technologies, its GooglesourcePlaceIdand its row in your dataset (sourceIndex, 0 = the first). Export as CSV and join it to your leads onsourcePlaceIdorsourceIndex.
What happens to each place:
- No website, or not a website (empty, a phone number, a Google Maps link): skipped, not fetched, not charged,
and counted in the run's
RUN_STATSrecord (dataset.withoutWebsite,dataset.notAWebsite,dataset.googleMapsLinks). - Branches sharing a website (the same domain, with or without
www.): analysed and charged once, with the first place's name and ids. Tracking parameters such asutm_sourceare dropped from the address. - Everything else is analysed like a URL you typed in, at the same price, $2.00 per 1,000 websites; a site that fails isn't charged. 400 places with a website cost at most 400 × $0.002 = $0.80, plus the start fee. The Google Maps scraper's own run is billed separately by that actor.
- Limits: the first 20,000 items and 10,000 distinct websites of a dataset per run; Max results per run and Maximum cost per run stop the run as usual. The dataset is read with your own Apify account's access, read-only.
Example: two websites
The pre-filled input:
{"urls": ["https://wordpress.org/", "https://www.python.org/"]}
It returned 2 results, 11 technologies each. The wordpress.org one (real output from a local run on 2026-09-25; 4 of its 11 technologies shown, and some evidence lines left out):
{"id": "31368c20a9917549c78ccac6","input": "https://wordpress.org/","url": "https://wordpress.org/","finalUrl": "https://wordpress.org/","statusCode": 200,"pageTitle": "Blog Tool, Publishing Platform, and CMS – WordPress.org","technologyCount": 11,"analysisComplete": true,"technologyNames": ["Google Font API", "Google Tag Manager", "Gutenberg 24.0.0", "HSTS", "MySQL", "Nginx","Open Graph", "PHP", "Priority Hints", "RSS", "WordPress 7.2"],"technologies": [{"name": "WordPress", "categories": ["CMS", "Blogs"], "version": "7.2", "confidence": 100,"website": "https://wordpress.org","evidence": ["header link: <https://wordpress.org/wp-json/>; rel=\"https://api.w.org/\"","meta generator: WordPress 7.2-alpha-63914"]},{"name": "Gutenberg", "categories": ["WordPress plugins", "Editors"], "version": "24.0.0", "confidence": 100,"website": "https://github.com/WordPress/gutenberg","evidence": ["element link[href*='/wp-content/plugins/gutenberg/'] href: https://wordpress.org/wp-content/plugins/gutenberg/build/styles/block-library/navigation/style.min.css?ver=24.0.0"]},{"name": "Google Tag Manager", "categories": ["Tag managers"], "version": null, "confidence": 100,"website": "http://www.google.com/tagmanager","evidence": ["inline script: googletagmanager.com/gtm.js", "inline script: 'dataLayer','GTM-P24PF4B'"]},{"name": "MySQL", "categories": ["Databases"], "version": null, "confidence": 100,"website": "http://mysql.com", "evidence": ["implied by WordPress"]}],"categories": {"CMS": ["WordPress"], "Tag managers": ["Google Tag Manager"], "Web servers": ["Nginx"],"Programming languages": ["PHP"], "Databases": ["MySQL"]},"scrapedAt": "2026-09-25T20:25:16.703568Z"}
(categories shortened too.) The python.org result found Nginx, Varnish, Fastly and jQuery 1.8.2, among others.
How accurate is it?
Measured on 181 public home pages (September 2026: 60 online shops on Shopify, WooCommerce, Magento, BigCommerce and headless stacks, 48 SaaS marketing sites, 36 small businesses and organisations on Squarespace, Wix, WordPress and Drupal, 13 docs sites, 12 news sites, 12 big brands) against a hand-checked answer for 28 technologies. Each "yes" in the answer key is backed by something visible in that page's HTML, headers or cookie names (a script URL, a generator tag, a response header), written down per site.
| Technology group | Found | Precision | Recall |
|---|---|---|---|
| CMS and site builders: WordPress, Drupal, Webflow, Squarespace, Wix | 48 of 51 | 100% | 94% |
| E-commerce: Shopify, WooCommerce, Magento, BigCommerce | 51 of 51 | 100% | 100% |
| Frameworks: Next.js, Nuxt.js, Gatsby | 41 of 41 | 100% | 100% |
| Analytics and tags: Google Analytics, Google Tag Manager, Segment, Plausible, Hotjar, Facebook Pixel | 124 of 126 | 100% | 98% |
| CDN: Cloudflare, Fastly, Amazon CloudFront | 135 of 135 | 100% | 100% |
| Marketing and consent: HubSpot, Marketo, Optimizely, OneTrust | 46 of 50 | 100% | 92% |
| All 28 technologies | 447 of 456 | 100% | 98% |
No false positives. On the 65 of those sites that were checked only after the fingerprints were final, recall was 96%. Before version 1.1 (the 2023 fingerprints alone) recall was 76%. What these numbers don't cover, honestly:
- Tools a tag manager or script adds later, in a browser, are invisible here, and they were not counted either (the answer key only has what the page itself shows). A browser-based tool will list more of those.
- What it still misses: WordPress used headless behind another front end (only media URLs show it), Google Tag Manager served from the site's own domain, and Optimizely set up in code rather than with its standard snippet.
- Too few sites to measure: chat widgets (Intercom, Drift) and payments (Stripe) are fingerprinted, but almost never appear in a home page's HTML: 1, 0 and 1 of 181 sites. Don't rely on this actor for those.
- Precision was measured for these 28 technologies; the other 3,500 are reported with their evidence so you can check them.
- About 1 site in 11 refuses automated visitors (HTTP 403, or a robots.txt that can't be read); those return no result and aren't charged. So does a site that answers with a bot check, block page or waiting room even with HTTP 200 (1 more site in the sample): it's reported as blocked, not analysed as if it were the site.
Ready-to-run examples
Each example opens with the input already filled in. Run it as it is, or change the input first.
- Detect the CMS and tech stack of websites: Find out which CMS, frameworks, analytics, e-commerce and hosting tools a list of websites uses, with a confidence level and the evidence for each detection.
- Build a lead list from websites’ tech stacks: Check a list of company websites and get each one’s CMS, e-commerce platform, frameworks, analytics, CDN and hosting, with the evidence for each detection.
Input
| Field | What it does |
|---|---|
| Website URLs | The pages to analyse, one per line: a full address (https://www.example.com/shop) or just a domain (example.com means its https:// home page). |
| Only report sites whose stack changed | Monitoring: return only the sites that added or dropped a technology since the last run of this list (the first run returns every site). Up to 10,000 sites per list. |
Or: websites from a dataset (datasetId) | One of your Apify datasets, e.g. a Google Maps scraper's results: each item's website is analysed and Website URLs is ignored. See Enrich Google Maps leads. |
Field with the website (datasetUrlField) | Only with a dataset. Empty (default): found automatically among website, url, domain, site and homepage, then one level down (contact.website). |
| Max results per run | Cap the total number of websites analysed (with the option above: checked, returned or not). |
{"urls": ["https://wordpress.org/", "shopify.com", "https://www.example.com/shop"],"maxResults": 100}
The same page typed twice (example.com and https://example.com/) is analysed and charged once.
Output
One result per website:
{"id": "5d1e0c0f3c2e4a8b9f6a7d21","input": "shop.example.com","url": "https://shop.example.com/","finalUrl": "https://shop.example.com/home","statusCode": 200,"pageTitle": "A small WordPress shop","technologyCount": 9,"analysisComplete": true,"technologyNames": ["Nginx 1.25.3", "PHP 8.2.1", "WordPress 6.4.2", "Google Analytics", "jQuery", "MySQL"],"technologies": [{"name": "WordPress","categories": ["CMS", "Blogs"],"version": "6.4.2","confidence": 100,"website": "https://wordpress.org","evidence": ["meta generator: WordPress 6.4.2","script: https://shop.example.com/wp-includes/js/jquery/jquery.min.js?ver=3.7.1"]},{"name": "MySQL","categories": ["Databases"],"version": null,"confidence": 100,"website": "http://mysql.com","evidence": ["implied by WordPress"]}],"categories": {"CMS": ["WordPress"], "Databases": ["MySQL"], "Web servers": ["Nginx"]},"scrapedAt": "2026-09-24T12:00:00Z"}
(Shortened: the real result lists every technology found; changeType, technologiesAdded,
technologiesRemoved, changeSummary and previousCheckedAt are there too, null, and so are sourceTitle,
sourcePlaceId and sourceIndex, which are only filled for websites read from a dataset.) id is stable across runs, so
you can use it to deduplicate. version is null when the page doesn't reveal it. A page where nothing is
recognised is still a result, with an empty technologies list.
Stack changes since the last run
With Only report sites whose stack changed on, the run remembers each site's technologies (names and categories,
not the pages) in a key-value store in your own account (tech-stack-detector-memory, one record per list of sites)
and compares with it next time. After the first run (the baseline, returned as usual with these fields null), each
returned site has:
changeType:changed(technologies added or removed) ornew(no earlier result to compare with, e.g. the site failed on every earlier run).technologiesAdded: each technology found now and not last time, with itscategoriesandversion.technologiesRemoved: each technology found last time and not now, with itscategories.changeSummary: the same in one line, for a Slack message or a spreadsheet column.previousCheckedAt: when the stack it's compared with was seen.
A changed site, for example (the shape of a real result; the site and values are illustrative):
{"url": "https://shop.example.com/","technologyCount": 12,"technologyNames": ["Cloudflare", "Klaviyo", "Shopify", "..."],"changeType": "changed","technologiesAdded": [{"name": "Shopify", "categories": ["Ecommerce"], "version": null},{"name": "Klaviyo", "categories": ["Marketing automation"], "version": null}],"technologiesRemoved": [{"name": "WooCommerce", "categories": ["Ecommerce", "WordPress plugins"]}],"changeSummary": "added: Klaviyo (Marketing automation), Shopify (Ecommerce); removed: WooCommerce (Ecommerce, WordPress plugins)","previousCheckedAt": "2026-09-21T06:00:04Z","scrapedAt": "2026-09-28T06:00:03.912Z"}
What it deliberately doesn't report:
- Removals from a failed fetch. A site that is down, blocked or refuses isn't compared; its last known stack is kept, so when it answers again it's compared with that, not reported as new or re-added.
- "Everything removed" from an empty page. A site where nothing at all is found this time, after something was
found before, is listed in
RUN_STATSassitesUnsure, not returned and not charged, and keeps its old stack. - Removals when the analysis was cut short. When a time limit stopped the matching (
analysisComplete: false), technologies it didn't find this time aren't reported as removed; additions still are. - Version changes. Only technologies coming and going count: a version is only known when a page happens to show it, so comparing versions would report noise.
RUN_STATS has baseline, previousRun (when the list was last checked), sitesNew, sitesChanged,
sitesUnchanged, sitesUnsure and unchangedSitesCharged. A site cut by your max results or maximum cost per run
isn't remembered, so its change is reported on the next run instead.
Run it on a schedule, or from your own code
- Save your input as a task (Save as a new task, top right of the actor page) and add it to a schedule (Console → Schedules → Create new): for example weekly or monthly for a fixed list of client or competitor sites. To hear only about sites whose stack changed, turn on Only report sites whose stack changed; the steps are in Monitor sites.
- Collect results: download the dataset as JSON, CSV or Excel; fetch the latest run's results from the API
(
GET https://api.apify.com/v2/actor-tasks/<task id>/runs/last/dataset/items?status=SUCCEEDED&format=csv, with your API token); let a webhook tell your system when a run succeeds; or connect it to Make, Zapier or n8n through Apify's integrations.
Can I use Website Technology Detector from an AI agent (MCP)?
Yes, through Apify's MCP server: add https://mcp.apify.com?tools=humble-echidna/tech-stack-detector to your MCP
client (or let the agent find it with the server's actor search). The agent passes the sites, e.g.
{"urls": ["example.com"]}, and reads technologyNames, or evidence when it needs to check an answer. To enrich
another actor's results, it passes that run's dataset instead: {"datasetId": "<defaultDatasetId>"}.
Who it's for
Agencies and freelancers qualifying leads by what a site runs on (every WordPress or Shopify shop in a list, sites still on an old jQuery), SaaS teams researching which tools competitors and prospects use, and sales ops teams adding a "built with" column to a lead list before it goes into the CRM. The recurring job: run each new batch of lead or prospect domains through it, or re-check a fixed list every week or month. It sees what a site's home page (or whichever page you give it) shows every visitor, so read "How accurate is it?" before relying on a given technology.
Why this one?
- One flat price, no usage bill. You pay per result at the price above and nothing else: no compute, proxy or storage charges on top, apart from the $0.00005 run-start fee, so you know what a run costs before it starts. Results you don't get aren't charged.
- Evidence for every answer. Each technology says why it was detected, so a surprising result can be checked in seconds instead of trusted blindly.
- Versions and confidence. Versions come from generator tags, headers and script URLs; confidence adds up across independent signals (0-100).
- Stack changes as a monitor. Re-run the same list on a schedule and get only the sites that added or dropped a technology, with both lists and their categories; unchanged sites cost a fifth of a result. A failed or blocked fetch is never reported as a removal.
- You only pay for websites analysed. A URL that fails (site down, blocked by robots.txt or a bot check, not an HTML page, wrong domain) returns no result and isn't charged.
- Polite and safe. It identifies itself honestly (User-Agent
HumbleEchidnaApify), follows each site's robots.txt (read once per site per run) and Crawl-delay, makes at most 2 requests per site at once, and only requests public web addresses on the standard ports (80 and 443). - Reliable. One failing URL never affects the others in your run, and no page can stall it (every page has a
size and time limit). The run log and the
RUN_STATSrecord say exactly which URL had a problem and why.
Limits
- One page per URL, without running its JavaScript: tools a tag manager adds later in a browser are missed (see "How accurate is it?").
- Pages over 3 MB are cut at 3 MB, and matching has time limits per page (see
analysisCompletein the FAQ). - Only public web addresses on the standard ports (80 and 443); no login, no proxy. Sites that refuse automated visitors or answer with a bot check are reported and not charged.
- Fingerprints for products launched since 2023 may be missing unless we've added them.
FAQ
Can I get only the sites whose technology stack changed since my last run?
Yes: turn on Only report sites whose stack changed and run the same list on a schedule. The first run returns
every site; after that each run returns only the sites that added or dropped a technology, with technologiesAdded,
technologiesRemoved (names and categories) and a one-line changeSummary. Unchanged sites are rechecked but not
returned ($0.40 per 1,000). See Monitor sites.
Why are unchanged sites charged at all?
Because finding out that a site didn't change takes the same work as analysing it: the page is downloaded and matched against every fingerprint again. Unchanged sites cost a fifth of a returned site, and never when the site failed, was blocked or showed nothing.
A site went down or blocked the run. Will it show up as "all technologies removed"?
No. A site that fails (down, robots.txt, a bot check, a 403, not HTML) isn't compared and keeps its last known stack;
a page where nothing at all is found, after something was before, is treated the same way (sitesUnsure in
RUN_STATS). Neither is returned or charged. When the site answers normally again, it's compared with the stack it
had before.
Why wasn't a technology I know is there detected?
This actor reads one page's HTML, inline scripts, response headers and cookie names, without running the page's JavaScript. Technologies that only show up once scripts run in a browser (tools loaded later by a tag manager, single-page apps that build everything in the browser) and anything only visible in DNS or TLS certificates are missed. Tools used only on other pages (a checkout, a blog) are found only if you give that page's URL too. The fingerprints are the last openly (MIT) licensed version of the Wappalyzer technology data, from January 2023, plus our own current fingerprints for the most-used tools (see "How accurate is it?"), so products launched since 2023 that we haven't added may be unknown.
What does confidence mean?
Each fingerprint pattern carries a weight (100 unless the data says otherwise); a technology's confidence is the sum of its matched patterns, capped at 100. A technology that is only implied by another (PHP by WordPress) is never more confident than what implied it.
What does analysisComplete: false mean?
Every page has limits so that one unusual page can't stall a run: pages over 3 MB are cut at 3 MB, the page source is matched on its first and last 1,500 lines (2,000 characters each), CSS selectors look at the first 10,000 elements, each pattern gets 0.1 s per value and each page 15 s of matching in total. When one of the time limits was hit, the result says so; everything listed was still really found.
Why did a URL come back as "blocked by robots.txt"?
The site's owner has asked crawlers not to fetch that page. This actor respects that, and the URL is not charged.
Why does it refuse localhost, 10.x.x.x, an internal hostname or a URL with a port?
It only fetches public web pages,
on the standard web ports (80 for http://, 443 for https://); a URL with any other port, such as :8080, is
refused. Every hostname (and every redirect) is resolved first, and a page is refused if any address it resolves to
is private, loopback, link-local (cloud metadata) or otherwise not on the public internet; the connection then goes
to the address that was checked. These refusals are reported as "not a public web address" and aren't charged. A
domain that doesn't exist is reported as such.
What happens when a site answers "403" or shows a bot check?
Some sites refuse automated visitors. This actor doesn't try to get around that: the URL is reported as refused and not charged. The same goes for a challenge, block or waiting-room page served with an ordinary "200" (Cloudflare "Just a moment...", Akamai "Access Denied" and failover pages, PerimeterX / HUMAN, DataDome, Imperva / Incapsula, Kasada, Queue-it, or a near-empty "verify you are human" page): the log says "blocked: the site answered with a bot check" and names it. A normal page that merely loads Cloudflare's or another vendor's scripts is analysed as usual.
Something that used to work now fails. Why?
Sites change without notice. The run log names the URL and what went wrong, and every other URL in the run is unaffected. Please open an issue with the input you used.
Is it legal to detect a website's technology stack?
It fetches only the pages you give it, once each, as a normal logged-out visitor, and follows
each site's robots.txt. It doesn't log in, get around any protection, or collect personal data: it reports which
products a site uses, from what the site sends to every visitor (cookie values are never output, only cookie
names). The Wappalyzer fingerprint data is used under its MIT licence, which ships with the actor; the supplementary
fingerprints are our own. What you do with the results
(and any site's own terms) is your responsibility.
Related actors
| Actor | Use it when |
|---|---|
| SEO Audit Crawler | You want a full on-page audit (titles, H1s, canonicals, broken links per page) of the same sites, for a fuller picture of a prospect's website. |
| Bulk URL & Broken Link Checker | You want to clean a lead list of dead or redirected domains before analysing it. |
| Sitemap URL Extractor | You want to analyse pages other than the home page: it lists every URL a site's sitemap has. |
Feedback and support
Found a bug, a wrong detection or a missing technology? Open an issue on the Issues tab with the URL and what you expected.
Versions
Current version: 1.3. See the Changelog tab for what changed in each version.