Website Technology Detector
Pricing
from $5.00 / 1,000 url analyseds
Website Technology Detector
Detect the CMS, e-commerce platform, analytics, CDN, frameworks and servers behind any list of websites. Plain HTTP, no browser. $5 per 1,000 sites.
Pricing
from $5.00 / 1,000 url analyseds
Rating
0.0
(0)
Developer
Ottomath
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Find out which technologies a website runs on: content management system, e-commerce platform, JavaScript frameworks, analytics, tag managers, CDN, hosting, web server, payment tools, chat widgets, cookie banners, fonts and more. Provide a list of URLs or domains and get one structured result per site. The Actor uses plain HTTP requests and no browser, so it is fast and inexpensive.
What it does
For every URL or domain you provide, the Actor fetches the page once, follows up to 5 redirects, and reads everything a plain HTTP client can see: response headers, cookie names, the HTML, meta tags, script and stylesheet references, inline scripts, page text and the final URL. It matches these signals against a database of about 3,600 technology fingerprints in 108 categories. It also applies the database's relationships (WordPress implies PHP and MySQL, for example), reads the version when the page reveals it, and can look up the domain's DNS records to spot email, DNS and hosting providers.
The result is one dataset item per site with the technologies, their categories, versions and a confidence score. The Actor uses plain HTTP requests only. It does not run a browser, does not log in anywhere and does not disguise itself: every request carries the honest User-Agent OttomathTechDetector/1.0 (+https://apify.com). It extracts technology names and versions only, never email addresses, phone numbers or names.
Use cases
- Prospecting and sales: build a target list from a set of domains, for example every Shopify store or every WordPress site, and prioritise agencies, apps or services by the platform your prospects use.
- Competitor and market research: compare the analytics, tag manager, CDN and e-commerce stack of competitors, or measure how often a platform is used in a market segment.
- Web agencies and migrations: inventory the CMS, framework and server versions across the client sites you manage before an upgrade or a security review.
- Data enrichment: add a technology column to a CRM export or a dataset of company websites and re-run it on a schedule to track changes.
Input
Provide the sites in urls. Each entry can be a full URL starting with https:// or http://, or a bare domain such as example.com (https is assumed). Up to 10,000 entries per run; duplicates are analysed once. Local and private network addresses are rejected and reported in the error field.
| Field | Type | Default | Description |
|---|---|---|---|
urls | list of strings | required | Websites to analyse, 1 to 10,000 per run. |
respectRobotsTxt | boolean | true | Skip URLs that the site's robots.txt disallows for this Actor's User-Agent or for *. The path of the URL is checked, not only /. |
includeVersions | boolean | true | Report versions when the page reveals them. When false, version is always null. |
minConfidence | integer | 0 | Only report technologies with at least this confidence (0-100). Use 50 to hide weak single-signal guesses. |
detectDns | boolean | true | Also read MX, TXT, NS, SOA and CNAME records to detect email, DNS and hosting providers. |
maxConcurrency | integer | 10 | URLs processed in parallel (1-30). At most 2 requests per domain run at the same time. |
requestTimeoutSecs | integer | 20 | Time limit for a single HTTP request (1-120). |
Example input:
{"urls": ["https://example.com", "wordpress.org", "https://www.shopify.com"],"respectRobotsTxt": true,"includeVersions": true,"minConfidence": 0}
Output example
Each result is one item in the dataset, in the same order of fields. Dates are ISO 8601 in UTC. You can download the dataset as JSON, CSV, Excel, XML or HTML, or read it through the Apify API.
{"url": "blog.example.com","finalUrl": "https://blog.example.com/","domain": "blog.example.com","statusCode": 200,"technologies": [{ "name": "WordPress", "slug": "wordpress", "categories": ["CMS", "Blogs"], "version": "6.4.2", "confidence": 100, "website": "https://wordpress.org" },{ "name": "PHP", "slug": "php", "categories": ["Programming languages"], "version": "8.2.12", "confidence": 100, "website": "http://php.net" },{ "name": "Nginx", "slug": "nginx", "categories": ["Web servers", "Reverse proxies"], "version": "1.24.0", "confidence": 100, "website": "http://nginx.org/en" }],"categories": {"CMS": ["WordPress"],"Programming languages": ["PHP"],"Web servers": ["Nginx"]},"detectedAt": "2026-10-08T10:30:00.000Z","error": null}
url is exactly what you provided and finalUrl is where the redirects ended. domain is the host of the final URL without a leading www.. confidence is a number from 1 to 100: 100 means at least one strong signal, lower values mean weaker or indirect evidence. Technologies are sorted by category importance, so the CMS and e-commerce platform come first. error is null when the page was fetched and analysed. Otherwise it explains why, for example Blocked by robots.txt, Timeout after 20 s, Domain not found (DNS lookup failed) or HTTP 404 Not Found, and the other fields hold whatever could be determined.
Pricing
This Actor uses Apify's pay-per-event model. You pay for what it delivers, not for how long it runs.
| Event | When it is charged | Price |
|---|---|---|
| Actor start | Once when a run starts. Apify charges it automatically. | $0.005 |
URL analysed (url-analyzed) | Once for each URL whose page was fetched with an HTTP status below 400 and analysed. | $0.005 |
URLs that fail, time out, are blocked by robots.txt, are invalid or return an error status (4xx or 5xx) are reported in the dataset and are free. Duplicates are analysed and charged once. Example: 1,000 working URLs cost $0.005 + 1,000 x $0.005 = $5.005.
You stay in control of the cost. Set the maximum charge for the run in the run options: the Actor never analyses more URLs than your limit pays for. When the limit is reached it stops by itself, keeps everything already delivered and ends the run with a clear status message.
FAQ
How do I cap my spending? Use the maximum total charge of the run. The Actor reads it, stops cleanly when the next URL would exceed it, and tells you how many URLs were not processed.
Does it respect robots.txt?
Yes, by default. The Actor reads robots.txt once per site and follows the rules for its own User-Agent, or for * when none are specific. If robots.txt cannot be read because of a server error, the URL is skipped. You are responsible for complying with the terms of the sites you analyse and with applicable law.
Why is a technology I know about missing? Without a browser the Actor cannot see JavaScript variables or background requests, so a site that reveals its stack only that way may show fewer technologies. The fingerprint database is also a snapshot, see Limitations.
What happens when a site blocks the request?
You get an item with an error such as HTTP 403 Forbidden. The Actor does not rotate proxies, solve CAPTCHAs or change its identity. Technologies visible in the error response, such as a CDN header, are still listed, and the URL is free.
How long does a run take? Mostly the time sites need to answer. A list of 100 typical sites finishes in a few minutes. Very large lists of 10,000 URLs take longer, so split them into batches if you need results quickly.
Can I schedule it or call it from other tools? Yes. Like any Actor it can run on a schedule, through the Apify API, or from integrations such as webhooks.
Performance
Analysing a page is CPU-bound and runs in worker threads. In a single thread on a development machine, five well-known home pages (45 KB to 1 MB of HTML) took 0.06 to 0.3 seconds of CPU time each, and up to 0.5 seconds for the first page of a run. A page of 2 MB, the maximum, took 0.6 to 1.1 seconds. A run with 1,024 MB of memory gets about a quarter of a CPU core on Apify, so allow roughly four times that as elapsed time per page.
Two measures keep this predictable. A fingerprint pattern is skipped when a keyword it requires is not on the page. A leading wildcard (.+ or .*) is removed from patterns, which does not change what they find but avoids a very slow search in long single-line scripts; when a page holds several matches of such a pattern, the version comes from the first one. A page that still needs more than 20 seconds, for example deliberately malformed markup, is abandoned, reported with an error and not charged.
Limitations
- Plain HTTP only: pages are not rendered, so JavaScript globals and background requests are not evaluated. Roughly 300 technologies in the database can only be recognised that way and will not be reported.
- One page per URL (usually the home page), plus robots.txt and redirects. The first 2 MB of the page are analysed.
- A page that takes more than 20 seconds to analyse (extremely large or deliberately malformed markup) is abandoned and reported with an error. It is not charged.
- The fingerprint database dates from January 2023. Newer tools may be missing; a small set of extra fingerprints for modern frameworks such as Next.js, Nuxt, Astro and SvelteKit is included.
- No login, no CAPTCHA solving and no proxy rotation. Sites that block automated requests return an error.
- DNS-based detection depends on the DNS records the domain publishes.
Data source and licence
Fingerprints come from the technology data of the open-source Wappalyzer project, release 6.10.54 of its npm package, published under the MIT licence. The data is downloaded when the Actor is built and checked against a fixed SHA-256 hash. This Actor is not affiliated with or endorsed by that project. The licence terms and attributions are listed in THIRD_PARTY_NOTICES.md.
Changelog
1.0.0 (2026-10-08)
- Initial release.