# Website Technology Detector – CMS, Analytics, Hosting (`tinlark/website-technology-detector`) Actor

Paste a list of websites and get the technology stack of each: CMS, shop platform, framework, analytics, CDN, hosting, payment and chat widgets, with version, confidence and evidence. Reads each home page once, honours robots.txt. Free during launch; from November 2026: $0.02 per site.

- **URL**: https://apify.com/tinlark/website-technology-detector.md
- **Developed by:** [Tinlark](https://apify.com/tinlark) (community)
- **Categories:** Developer tools, Lead generation, SEO tools
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Technology Detector – CMS, Analytics, Hosting

Paste a list of websites and get the technology stack of each one: **CMS, shop platform, JavaScript framework, analytics and tag managers, CDN and hosting, payment and chat widgets, cookie consent tools, fonts, maps and video players**. Every technology comes with its category, the version when the site shows one, a confidence level and the evidence found (a header, cookie, meta tag, script URL or page marker).

The Actor reads each site's home page once, checks the site's robots.txt first, identifies itself with a `TinlarkBot` User-Agent, and needs no API key. It recognises 500+ technologies with fingerprints Tinlark wrote itself.

### Use cases

- **Bulk tech stack lookup.** Which CMS, shop platform, framework, analytics tool, CDN or host does each site in a list use? One row per site, in one schema.
- **Find sites that use a technology.** Run a prospect or competitor list and keep the rows whose `technologyNames` contain Shopify, WooCommerce, HubSpot, Klaviyo or any other name in the list.
- **Sales and agency prospecting.** Add the stack to an account list: shop platform, payment provider, chat widget, consent tool, A/B testing tool.
- **Migration and upgrade targets.** Sites on an old CMS or library version, where the site publishes the version (`version` field).
- **Market and competitor research.** Count how often a technology appears in a list of companies, and which categories they cover.
- **Quick checks from an API or a schedule.** Run it from the Apify API, a Schedule or an integration; the input and output schemas are defined.

### Who it is for

Sales and RevOps teams, agencies and freelancers who qualify prospects by technology, analysts who need a technology footprint for many sites at once, and developers who want a plain API for "what is this site built with". If you also need contacts, social links, SEO meta data or DNS for the same domains, see Related Tinlark Actors below.

### Input

| Field | What it does |
|---|---|
| `urls` (required) | One website per entry: a domain such as `stripe.com` or a full URL. Up to 5,000 per run. A bare domain is tried as `https://domain`, then `https://www.domain`, then `http://domain`. A full URL is fetched as written. |
| `categories` | Keep only technologies in these categories (for example `CMS`, `E-commerce`, `Analytics`, `CDN`, `Hosting`, `Payments`). Empty means all. The 35 categories are listed in the input form. |
| `includeEvidence` | Add the evidence behind each technology (default on). |
| `includeVersions` | Add the version when the site shows one (default on). |
| `followRedirects` | Follow up to 5 redirects (default on). robots.txt is checked again for every host the redirects lead to. When off, the redirect answer itself is analysed. |
| `timeoutSecs` | Per-request timeout (default 20). |
| `maxConcurrency` | Sites analysed in parallel (default 10, maximum 20). Requests to one host are always at least one second apart. |

robots.txt is always honoured; there is no switch for it.

#### Example input

```json
{
  "urls": ["https://stripe.com", "https://shopify.com", "https://wordpress.org"],
  "categories": [],
  "includeEvidence": true,
  "includeVersions": true
}
```

### Output

One row per site in the default dataset (JSON, CSV, Excel, XML or API). The *Overview*, *Technologies* (one row per site and technology), *Categories* and *Not analysed* views show the main fields.

#### Example row (real output, shortened to four of nine technologies)

```json
{
  "url": "wordpress.org",
  "finalUrl": "https://wordpress.org/",
  "domain": "wordpress.org",
  "status": "ok",
  "httpStatus": 200,
  "technologyCount": 9,
  "technologies": [
    { "name": "WordPress", "category": "CMS", "version": "7.2", "confidence": "high", "evidence": ["header:link", "meta:generator", "link:/wp-content/"] },
    { "name": "Google Tag Manager", "category": "Tag managers", "version": null, "confidence": "high", "evidence": ["iframe:googletagmanager.com/ns.html", "html:GTM-P24PF4B"] },
    { "name": "Nginx", "category": "Web servers", "version": null, "confidence": "high", "evidence": ["header:server"] },
    { "name": "PHP", "category": "Backend", "version": null, "confidence": "low", "evidence": ["implied by WordPress"] }
  ],
  "technologyNames": ["Jetpack Stats", "Jetpack Site Accelerator", "WordPress", "Google Fonts", "Google Tag Manager", "Nginx", "Jetpack", "Gutenberg blocks", "PHP"],
  "categories": {
    "Analytics": ["Jetpack Stats"], "Backend": ["PHP"], "CDN": ["Jetpack Site Accelerator"], "CMS": ["WordPress"], "Fonts": ["Google Fonts"],
    "Tag managers": ["Google Tag Manager"], "Web servers": ["Nginx"], "WordPress": ["Jetpack", "Gutenberg blocks"]
  },
  "server": "nginx",
  "poweredBy": null,
  "generator": "WordPress 7.2-alpha-64071",
  "checkedAt": "2026-10-03T14:52:21Z",
  "error": null
}
```

#### Field notes

- `status` is `ok`, `blocked-by-robots` (the site's robots.txt does not allow TinlarkBot), `unreachable` (no answer, an HTTP error such as 403 or 404, or an invalid entry) or `not-html` (the URL serves an image, PDF or JSON). The `error` field says why. Only `ok` rows are ever charged.
- `technologies[].confidence`: `high` when a response header, cookie, meta generator tag or a script from the vendor's own host shows it, or when two signals agree. `medium` for a single page-level signal (a file name such as `jquery.min.js`, a loader snippet in an inline script or a marker in the page source). `low` for a technology implied by another one (PHP when WordPress is found).
- `technologies[].evidence` starts with where the signal was found: `header:`, `cookie:`, `meta:`, `script:`, `link:` (stylesheet, preload or preconnect), `img:`, `iframe:`, `inline:` (a loader snippet in an inline script), `html:` (a marker in the page source) or `implied by`. Query strings are cut off.
- `categories` groups the names by category. `server`, `poweredBy` and `generator` repeat the Server header, the X-Powered-By header and the generator meta tag as the site sends them.
- The `SUMMARY` record of the run's key-value store has counts per status, billable sites, request totals and the most common technologies.

### Pricing

**Free during launch (until 31 October 2026).** You pay only Apify's own platform usage for your runs.

Measured platform usage: about $0.04 per 1,000 sites (a mixed run of 100 well-known sites cost $0.004 at 1,024 MB and took 58 seconds). On the Apify free plan that usage is covered by your monthly credit.

**From 1 November 2026: pay per event.** $0.02 per site on the Bronze plan. Planned prices, shown here so you can budget:

| Apify plan | Site analysed | Per 1,000 sites |
|---|---|---|
| Free | $0.030 | $30 |
| Bronze | $0.020 | $20 |
| Silver | $0.016 | $16 |
| Gold | $0.012 | $12 |

The event is charged once per site whose home page was read and analysed (status `ok`). Not charged: sites that are unreachable, blocked by robots.txt, answer an HTTP error, serve no HTML, or are invalid input. Example at the Bronze price: 100 entries of which 97 are analysed cost $1.94. Set *Maximum cost per run* in the run options to cap spending once pricing is active; the Actor stops cleanly at the cap.

### Limits and honest notes

- **No JavaScript is executed.** The Actor reads the HTML, headers and cookies the server sends. A site that builds its page in the browser shows only what is in the first HTML, and tools that a tag manager loads after the page starts are not visible. The shorter the technology list of a single-page app, the more likely this is the cause.
- **One page per site.** Only the URL you give is fetched (a bare domain means its home page). Other pages of the site may use other tools.
- **Back-end and hosting facts are inferred from what the server reveals.** Many sites hide their server software, language and framework. A missing technology means "not visible", not "not used".
- **Versions are rare.** Only a few technologies publish their version (a generator tag, a Server header, a file name with a version). `version` is `null` otherwise.
- **Some sites will not answer.** In a run of 100 well-known sites, 97 were analysed, 1 disallowed TinlarkBot in robots.txt and 2 were unreachable. Sites behind bot protection often answer 403; the Actor reports that and does not use proxies or other means to get past it. A robots.txt that answers 401, 403 or a server error counts as "do not crawl".
- **Fingerprints are Tinlark's own.** The list covers 535 technologies, written from vendors' public documentation and from reading real pages; no third-party fingerprint list was copied. A run of 100 well-known sites found 177 different technologies. Fingerprints for rarer tools follow the vendors' documented embed code but may not yet have been seen on a live page, so a rare tool can be missed or, rarely, named wrongly. Report such cases in the Issues tab.
- **Pace.** Requests to one host are at least one second apart. A mixed list of 100 sites took 58 seconds at 1,024 MB (the default) and 212 seconds at 256 MB, which has a smaller share of a CPU core. The default run *Timeout* is 2 hours; raise it for very long lists of slow sites.

### Data sources, terms and your responsibility

- **The listed sites themselves:** their public home pages, response headers and robots.txt, read only where robots.txt allows.
- **Tinlark's fingerprint list:** written by Tinlark, shipped with the Actor. The Actor calls no third-party API and sends the sites you list to no one else.

You are responsible for how you use the output and for following the terms of the sites you list. This Actor is not affiliated with any of the technologies or companies it detects.

### FAQ

**Does it use a browser or a proxy?** No. It makes plain HTTP requests, which keeps runs cheap. Sites that need JavaScript to show their content are read as the server sends them.

**What does "unreachable" mean?** The host did not answer, or the home page answered an HTTP error (403, 404, 5xx) on every address tried (https, https with www, http). The `httpStatus` and `error` fields say which. These rows are free.

**Why is my site `blocked-by-robots`?** Its robots.txt disallows TinlarkBot (or all robots) for the page, or the robots.txt request itself was refused with 401, 403 or a server error. The Actor does not fetch the page then, and the row is free.

**Can I check a page other than the home page?** Yes: give the full URL. It is fetched as written, once.

**How do I see only the shop platform and analytics?** Set `categories` to `["E-commerce", "Analytics"]`. `technologyCount` then counts only those.

**Is this a Wappalyzer or BuiltWith alternative?** It does the same core job as their lookups: name the technologies a website runs. It works differently: it checks the sites you give it at run time and keeps no database of past scans or lists of sites per technology, it reads the HTML the server sends without running JavaScript, its fingerprints are Tinlark's own, and it runs on Apify with the output in a dataset you can export or call from the API. Tinlark is not affiliated with Wappalyzer or BuiltWith.

**Can I run it on a schedule or from code?** Yes: Apify Schedules, the API, client libraries or the Apify MCP server.

**Disclaimers and legality: is this allowed?** The Actor reads one public page per site that the site owner serves to any visitor, only where robots.txt allows it, at a low request rate and under a declared `TinlarkBot` User-Agent. It does not log in, solve captchas, run scripts or get around blocks. Whether you may use a given output depends on your jurisdiction and on the terms of the sites you list; you are responsible for that. The data is technical (which software a site runs), not personal data, but check your own obligations if you combine it with other data. Tinlark is not affiliated with the sites or technologies it detects, and this is not legal advice.

### Related Tinlark Actors

- [Domain Tech Stack, Contacts & SEO Profiler](https://apify.com/tinlark/domain-intelligence-profiler): the same technology detection, plus published business contacts, social links, SEO meta, sitemap size, security headers and DNS for each domain. Use it when you need more than the stack.

### Support

Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the site) and what you expected. A technology that is missing or wrongly named can be added to the fingerprint list.

# Actor input Schema

## `urls` (type: `array`):

One website per entry: a domain such as stripe.com or a full URL such as https://www.shopify.com. Up to 5,000 entries per run. A bare domain is tried as https://domain, then https://www.domain, then http://domain.

## `categories` (type: `array`):

Keep only technologies in these categories, for example CMS, E-commerce, Analytics, CDN, Hosting or Payments. Leave empty for all categories.

## `includeEvidence` (type: `boolean`):

Add the evidence behind each technology: the header, cookie, meta tag, script URL or page marker that showed it.

## `includeVersions` (type: `boolean`):

Add the version number when the site shows one (null otherwise). Switch off to leave the field out.

## `followRedirects` (type: `boolean`):

Follow up to 5 redirects to the final page. robots.txt is checked again for every host the redirects lead to. When off, the redirect answer itself is analysed.

## `timeoutSecs` (type: `integer`):

Give up on a single request after this many seconds.

## `maxConcurrency` (type: `integer`):

Sites analysed at the same time. Requests to one host are always at least one second apart.

## Actor input object example

```json
{
  "urls": [
    "stripe.com",
    "https://www.shopify.com"
  ],
  "categories": [],
  "includeEvidence": true,
  "includeVersions": true,
  "followRedirects": true,
  "timeoutSecs": 20,
  "maxConcurrency": 10
}
```

# Actor output Schema

## `sites` (type: `string`):

One row per website with the technologies found.

## `technologies` (type: `string`):

One row per website and technology.

## `summary` (type: `string`):

Counts per status, billable sites, request totals and the most common technologies.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://stripe.com",
        "https://shopify.com",
        "https://wordpress.org"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tinlark/website-technology-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://stripe.com",
        "https://shopify.com",
        "https://wordpress.org",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("tinlark/website-technology-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://stripe.com",
    "https://shopify.com",
    "https://wordpress.org"
  ]
}' |
apify call tinlark/website-technology-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tinlark/website-technology-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gtzfGwsS8p1ZL1Zow/builds/xj0PBFeC3Ke3CkQID/openapi.json
