# Website Tech Stack Lookup - Technology Detector (`datagrit/website-tech-stack-lookup`) Actor

Detect the technology stack of any list of domains: CMS, ecommerce, analytics, CDN, frameworks with versions, plus mail provider, SPF and DMARC.

- **URL**: https://apify.com/datagrit/website-tech-stack-lookup.md
- **Developed by:** [datagrit](https://apify.com/datagrit) (community)
- **Categories:** Lead generation, SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Website Tech Stack Lookup do?

Website Tech Stack Lookup tells you which technologies a list of websites runs on: CMS, ecommerce platform, analytics, tag manager, CDN, hosting, web server, JavaScript framework, marketing and chat tools, payment providers and hundreds more, with the version when the site exposes it, a confidence score and the evidence behind every detection. It also reads the public DNS of each domain to report the **mail provider (MX)**, the **SPF record** with the email services it authorises, the **DMARC policy** and the DNS provider. Paste domains or URLs, run, and export JSON, CSV or Excel, or call it from the Apify API, n8n, Make or an AI agent through MCP.

### Why use a website technology detector?

- **Lead generation and list building** – keep only the Shopify, WooCommerce, Magento or HubSpot sites from a list of thousands (Only sites using any of these technologies) and get the mail provider for outreach planning.
- **Sales qualification** – see whether a prospect uses a competitor, which analytics and marketing automation it pays for, and whether its email is on Google Workspace or Microsoft 365.
- **Competitor and market research** – measure how often a platform, framework or CDN appears across a market segment.
- **Email deliverability audits** – check SPF and DMARC for a whole portfolio of domains in one run.
- **Agency and M\&A due diligence** – get the server, hosting, CMS version and plugin footprint of a site before a pitch or a deal.

### Example output

| input | cms | ecommerce | cdn | hosting | mailProvider | dmarcPolicy | technologyCount |
|---|---|---|---|---|---|---|---|
| porterandyork.com | WordPress | WooCommerce | Cloudflare | WP Engine | Google Workspace | none | 11 |
| coxandcox.co.uk | | Magento 2 | Fastly | Adobe Commerce Cloud, Platform.sh | Google Workspace | quarantine | 16 |
| jquery.com | WordPress 7.1 | | Cloudflare | | Forward Email | reject | 6 |

```json
{
  "input": "porterandyork.com",
  "domain": "porterandyork.com",
  "finalUrl": "https://porterandyork.com/",
  "statusCode": 200,
  "title": "Buy Premium Meat Online | Porter & York Butcher Shop",
  "language": "en-US",
  "technologyCount": 11,
  "technologyNames": "Cloudflare, WordPress, Cloudflare DNS, Google Workspace, WP Engine, Google Search Console, WooCommerce, Google Tag Manager, Cloudflare Bot Management, PHP, Rank Math",
  "technologies": [
    { "name": "WooCommerce", "category": "Ecommerce", "version": null, "confidence": 99, "evidence": ["html /wp-content/plugins/woocommerce/", "html wc_add_to_cart_params"] },
    { "name": "Google Workspace", "category": "Email provider", "version": null, "confidence": 100, "evidence": ["dns mx aspmx.l.google.com", "dns spf include:_spf.google.com"] }
  ],
  "cms": "WordPress",
  "ecommerce": "WooCommerce",
  "cdn": "Cloudflare",
  "hosting": "WP Engine",
  "mailProvider": "Google Workspace",
  "hasSpf": true,
  "spfRecord": "v=spf1 include:_spf.google.com ~all",
  "hasDmarc": true,
  "dmarcPolicy": "none",
  "dnsChecked": true,
  "found": true
}
```

The **Technologies with evidence** view unwinds the list into one row per technology.

### How much does it cost?

You pay per website analyzed. Pricing depends on your Apify plan: a small fee when a run starts, then a price per analyzed website that is lower on paid plans. The Apify free plan includes monthly credit you can use to try it. Sites that could not be analyzed (domain does not exist, timeout, bot challenge page, HTTP error) get a row with found: false and the reason, and are never charged; sites skipped by the technology filter are not charged either. Set a maximum spend on the run and the Actor stops when it is reached. Each site costs one page request plus a few DNS lookups, so runs are fast and light on platform resources.

### Input

- **Domains or URLs** – bare domains (shopify.com), hostnames (www.bbc.co.uk) or full URLs. A bare domain is tried at https://domain/, then https://www.domain/, then http://domain/. Duplicates are merged and invalid entries are listed in the run status. Leave empty for a free example run.
- **Technology categories to list** – limit the technologies list to categories such as CMS, Ecommerce or Analytics; the summary columns stay complete.
- **Only sites using any of these technologies** – lead filter by technology name, for example Shopify, WooCommerce, Magento, HubSpot or Klaviyo.
- **Follow redirects**, **Check DNS**, **Parallel sites**, **Request timeout**, **Proxy configuration**.

### What is detected and how?

The Actor loads the home page (or the URL you give) over plain HTTP and matches 847 technology fingerprints written for this Actor against the response headers, cookie names, meta tags, script URLs and page HTML. DNS adds MX, SPF, DMARC, NS, CNAME, A and PTR signals: mail providers and security gateways, sending services authorised in SPF, DNS hosts, hosting and CDN networks, and services the domain was verified with. Versions come from generator tags, headers such as Server and X-Powered-By, and versioned CDN paths. Technologies implied by another one (PHP by WordPress, React by Next.js) are marked "implied by" and carry a lower confidence.

### FAQ

#### Is it legal to use this data?

The Actor reads only what any browser or DNS client receives from a public website and public DNS: it does not log in, does not use your cookies and does not solve captchas. How you use the results is your responsibility, including data protection and anti-spam rules for outreach. This description is not legal advice.

#### How accurate is it?

On 1 October 2026 the page and header detections on 23 well-known sites (shopify.com, wordpress.org, vercel.com, hubspot.com, stripe.com, bbc.co.uk, several Shopify, WooCommerce and Magento stores and others) were checked by hand against their page source: 175 detections, 2 false positives found in the first pass (WordPress.com hosting inferred from a Jetpack statistics script, and a Vue.js signal that also appears in Inertia/React pages), both fixed; after the fix all 175 were correct. DNS detections on the same sites (199) are read directly from public records; a TXT verification record shows that the domain was verified with a service at some point, not that it is still in use. The same sample was used to tune the fingerprints, so treat this as a sanity check rather than a benchmark. Technologies loaded only by JavaScript after the page renders, or only on inner pages, are not visible to a single HTTP request.

#### Why is a site reported as blocked-by-bot-protection?

Some sites answer requests from data center addresses with a challenge page (Cloudflare, Akamai, Fastly, DataDome and similar). The Actor recognises these pages, returns found: false with the reason and does not charge. A residential proxy in Proxy configuration often gets through.

#### What are the limits?

One page per site (after redirects, up to 8 hops and one meta refresh), the first 2 MB of HTML, 20 seconds per request by default. There is no limit on the number of domains. The Actor analyzes 10 sites at a time by default (up to 50), and never more than 2 URLs of the same host at once; in a local test on 1 October 2026, 12 well-known sites took 5.9 seconds in total including DNS. Slow or unreachable sites take longer, up to the request timeout for each address tried.

#### How often should I run it?

Tech stacks change slowly: monthly runs through an Apify schedule are enough for most lead lists; weekly for competitive monitoring.

### Related Actors

Other datagrit Actors cover job boards with salary data (Greenhouse, Ashby) and company registers, useful next to a tech stack when you build company profiles.

# Changelog

This Actor's version history is a separate document: https://apify.com/datagrit/website-tech-stack-lookup/changelog.md

# Actor input Schema

## `domains` (type: `array`):

Domains (shopify.com), hostnames (www.bbc.co.uk) or full URLs (https://example.com/pricing), one per line. A bare domain is fetched at https://domain/, then https://www.domain/ and http://domain/ if the first address does not answer; a full URL is fetched exactly as given. Duplicates (with or without www, http or https) are merged, and entries that are not domains or http(s) URLs are listed in the run status. Leave empty to get a free example for two well-known sites.

## `categories` (type: `array`):

Optional. Limit the technologies list (technologies, technologyNames, technologyCategories, technologyCount) to these categories. The summary columns (cms, ecommerce, analytics, mailProvider and the rest) are always filled from every detected technology. Categories: A/B testing, Accessibility, Advertising, Affiliate, Analytics, Blog, Build tool, Business software, CDN, CMS, CRM, Captcha, Comments, Cookie consent, Customer data platform, Customer support, DNS provider, Documentation, Ecommerce, Email marketing, Email provider, Email service, Feature flags, Font service, Forms, Forum, Headless CMS, Hosting, Image CDN, JavaScript CDN, JavaScript framework, JavaScript library, LMS, Live chat, Loyalty & referrals, Maps, Marketing automation, Monitoring, Operating system, PaaS, Payment, Personalization, Popups, Product adoption, Product search, Programming language, Push notifications, Reviews, SEO, Scheduling, Security, Session replay, Social, Static site generator, Subscriptions, Tag manager, Translation, UI framework, Video, Web framework, Web server, Website builder.

## `requireTechnologies` (type: `array`):

Optional lead filter. Return (and charge for) a site only when it uses at least one of these technologies, for example Shopify, WooCommerce or HubSpot. Names are matched case-insensitively against the technology names in the output; unknown names are listed in the run status. Sites that do not match are skipped without charge. Unreachable sites still get a free row with found: false.

## `followRedirects` (type: `boolean`):

Follow HTTP redirects (and one meta refresh) to the final page and analyze that page. Turn off to analyze only the first response; a redirect is then returned with its status code and redirectLocation.

## `includeDns` (type: `boolean`):

Look up MX, TXT (SPF and verification records), DMARC, NS, CNAME and A/PTR records of each domain to report the mail provider, SPF and DMARC, the DNS provider and services that verified the domain. Turn off for faster runs when you need only the website stack.

## `maxConcurrency` (type: `integer`):

How many sites are analyzed at the same time. Every site is a different server, so a higher value mostly speeds up long lists.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for each page before trying the next address or reporting the site as unreachable (reason timeout).

## `proxyConfiguration` (type: `object`):

Optional proxy for the page requests. Most sites answer without one; sites behind strict bot protection may block data center addresses and are then reported with found: false and reason blocked-by-bot-protection. DNS lookups do not use the proxy.

## Actor input object example

```json
{
  "domains": [
    "shopify.com",
    "wordpress.org",
    "hubspot.com"
  ],
  "categories": [],
  "requireTechnologies": [],
  "followRedirects": true,
  "includeDns": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All extracted records as a dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "shopify.com",
        "wordpress.org",
        "hubspot.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagrit/website-tech-stack-lookup").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "shopify.com",
        "wordpress.org",
        "hubspot.com",
    ],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("datagrit/website-tech-stack-lookup").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "shopify.com",
    "wordpress.org",
    "hubspot.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call datagrit/website-tech-stack-lookup --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagrit/website-tech-stack-lookup"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fjJHVotQRkTPLfpcT/builds/bj062eTlLTb1BAHhn/openapi.json
