# Website Tech Stack Scraper — CMS, Shop, Analytics, Payments (`chorelet/website-tech-stack-scraper`) Actor

Detect what a website runs on: CMS, e-commerce platform, frontend framework, hosting and CDN, analytics, ad pixels, marketing and support tools, payment providers and bot protection — with the evidence behind every detection. No login or API key.

- **URL**: https://apify.com/chorelet/website-tech-stack-scraper.md
- **Developed by:** [Chorelet](https://apify.com/chorelet) (community)
- **Categories:** Developer tools, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Tech Stack Scraper — CMS, Shop, Analytics, Payments

Find out what a website runs on. Paste domains and get, for each one, the **CMS, e-commerce platform, frontend framework, hosting and CDN, analytics, ad pixels, marketing and support tools, payment providers and bot protection** — with the exact evidence behind every detection. JSON, CSV or Excel, or through the API.

No login, no API key, no browser: the Actor reads the homepage, its response headers and `robots.txt`, the same things a visitor's browser receives first.

### Why this Actor

- **Evidence, not guesses.** Every technology comes with the header line, script URL or HTML snippet that proved it, so you can check a detection instead of trusting it.
- **First-party matching.** A `wp-content` image linked from someone else's domain will not label a site WordPress — a mistake most detectors make.
- **robots.txt too.** Platforms usually name themselves there, which catches shops and CMSes whose homepage is a JavaScript shell.
- **Built for lists.** Flat `cms` / `ecommerce` / `analytics` / `payments` columns, so a CSV of a thousand domains is usable without unpacking anything.

### Sample output

One item of the dataset (long values shortened):

```json
{
  "website": "https://nytimes.com/",
  "title": "The New York Times - Breaking News, US News, World News and Videos",
  "cms": "WordPress",
  "ecommerce": null,
  "frameworks": [
    "Svelte",
    "SvelteKit"
  ],
  "analytics": [
    "Google Tag Manager"
  ],
  "payments": [],
  "technologyCount": 6
}
```

### What you get

- 120 technologies across 19 categories, each with the header line, script URL or HTML snippet that proved it
- Versions where the page states them (WordPress, jQuery, nginx, Angular…)
- Flat columns — `cms`, `ecommerce`, `frameworks`, `analytics`, `payments`, `hosting` — so a CSV is usable as-is
- `robots.txt` fingerprints, which often name the platform even when the homepage is a JavaScript shell
- First-party matching: a `wp-content` image borrowed from another domain does not make a site "WordPress"
- Monitored daily

### Input

- **Websites** — domains or URLs, one per line.
- **Include evidence** — keep the snippet that proved each detection (on by default).
- **Also read robots.txt** — one extra small request per site, not charged separately.
- **Timeout per site**, **Websites in parallel**.

### Limits and notes

- This is static detection: everything a browser loads *later* — tags injected by Google Tag Manager, widgets added by JavaScript — is invisible here. Expect a handful of solid detections per site rather than the long list a browser extension shows after the page finishes loading.
- What it does catch reliably: platforms (Shopify, WordPress, Webflow, Wix, Magento…), frameworks (Next.js, Nuxt, Astro, Remix…), hosting and CDN, and any script the HTML references directly.
- Sites behind an aggressive bot wall (some marketplaces and airlines) answer 403 to any plain HTTP client; those rows come back with the status and an error instead of a guess.
- Every site is charged once, whether or not anything was detected.

### Input example

```json
{
  "websites": [
    "gymshark.com",
    "vercel.com",
    "plausible.io"
  ],
  "includeEvidence": true,
  "checkRobotsTxt": true,
  "requestTimeoutSecs": 20,
  "concurrency": 5
}
```

### How much does it cost?

Pay per website — no subscription, no minimum, no charge for platform usage.

| Volume | Price |
|---|---|
| 1,000 websites | $2.00 |
| 10,000 websites | $20.00 |
| 100,000 websites | $200.00 |

The Apify **free plan includes $5 of usage every month** — about 2,500 websites with this Actor, no card needed. Nothing else is charged: platform usage is included in the price, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off these prices.

### Use it from code, n8n, Make, Zapier or an AI agent

Run the Actor and download the dataset in one call (JSON by default; add `&format=csv` or `xlsx`):

```bash
curl -X POST "https://api.apify.com/v2/acts/chorelet~website-tech-stack-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"websites": ["gymshark.com", "vercel.com", "plausible.io"], "includeEvidence": true, "checkRobotsTxt": true, "requestTimeoutSecs": 20, "concurrency": 5}'
```

Python:

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("chorelet/website-tech-stack-scraper").call(run_input={"websites": ["gymshark.com", "vercel.com", "plausible.io"], "includeEvidence": true, "checkRobotsTxt": true, "requestTimeoutSecs": 20, "concurrency": 5})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

- **n8n, Make, Zapier** — use the Apify node/module: run the Actor, then "get dataset items".
- **Google Sheets, Slack, webhooks** — add an integration on the run's *Integrations* tab.
- **AI agents** — the Actor is available as a tool through the Apify MCP server; the dataset schema describes every field for the model.
- **Schedules** — run it hourly, daily or weekly from the *Schedules* tab.

### FAQ

**How does this compare with a browser extension?**

An extension sees the page after JavaScript has run, so it also catches tags injected by Google Tag Manager. This Actor reads the HTML, headers and robots.txt — fewer detections per site, but hundreds of sites a minute and no browser to pay for.

**Which technologies are covered?**

About 120: e-commerce platforms, CMSes and site builders, frontend and backend frameworks, hosting, CDNs and servers, analytics, ad pixels, marketing, email and support tools, payment providers, monitoring, auth, search, captchas, fonts and media.

**Why did a site come back with only two technologies?**

Because that is what its HTML actually reveals. Sites that load everything through a tag manager expose very little before JavaScript runs — the fields you do get (platform, framework, hosting) are the reliable ones.

**Can I find every site that uses Shopify?**

Not with this Actor — it checks the domains you give it. Feed it a list from your CRM, a directory or another Actor, then filter on `ecommerce`.

**What does a 403 mean?**

The site blocks plain HTTP clients. The row keeps the status and the error so you can tell "blocked" apart from "nothing detected".

### Support

Questions, missing fields or a source that changed? Open an issue on the *Issues* tab or write to support@chorelet.app — problems are usually fixed within a day, and the Actor is checked every morning by an automated test run. If the Actor saved you time, a short review on its Store page helps other people find it.

# Actor input Schema

## `websites` (type: `array`):

Domains or URLs, one per line: `gymshark.com`, `https://example.com`. The Actor follows redirects and tries the www/non-www twin if the first address does not answer.

## `includeEvidence` (type: `boolean`):

Keep the snippet that proved each detection (a header line, a script URL or a piece of HTML). Turn it off for a smaller dataset.

## `checkRobotsTxt` (type: `boolean`):

One extra small request per site. robots.txt often names the platform (Shopify, WordPress, Magento) even when the homepage is a JavaScript shell. Not charged separately.

## `requestTimeoutSecs` (type: `integer`):

Slow sites are given up on after this many seconds.

## `concurrency` (type: `integer`):

Higher is faster; lower is gentler on small sites.

## Actor input object example

```json
{
  "websites": [
    "gymshark.com",
    "vercel.com",
    "plausible.io"
  ],
  "includeEvidence": true,
  "checkRobotsTxt": true,
  "requestTimeoutSecs": 20,
  "concurrency": 5
}
```

# Actor output Schema

## `stacks` (type: `string`):

All websites — items of the default dataset. Use ?format=csv or xlsx on this URL for spreadsheets.

## `summary` (type: `string`):

Sites scanned, detections found, sites that could not be read.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "gymshark.com",
        "vercel.com",
        "plausible.io"
    ],
    "includeEvidence": true,
    "checkRobotsTxt": true,
    "requestTimeoutSecs": 20,
    "concurrency": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("chorelet/website-tech-stack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": [
        "gymshark.com",
        "vercel.com",
        "plausible.io",
    ],
    "includeEvidence": True,
    "checkRobotsTxt": True,
    "requestTimeoutSecs": 20,
    "concurrency": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("chorelet/website-tech-stack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "gymshark.com",
    "vercel.com",
    "plausible.io"
  ],
  "includeEvidence": true,
  "checkRobotsTxt": true,
  "requestTimeoutSecs": 20,
  "concurrency": 5
}' |
apify call chorelet/website-tech-stack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,chorelet/website-tech-stack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/NK30HKs0N0zV7wxSA/builds/IRWWgtHTyr03BLDbj/openapi.json
