# Wappalyzer Scraper (`lergassy/wappalyzer-scraper`) Actor

\[$0.01/site] Wappalyzer-style technology lookup for any list of websites: CMS, e-commerce, framework, analytics, ad pixels, payments, hosting, CRM, chat — 209 signatures in 32 categories, one row per site, plus e-mails and social profiles from the same page. No run fee, no browser.

- **URL**: https://apify.com/lergassy/wappalyzer-scraper.md
- **Developed by:** [Matvey](https://apify.com/lergassy) (community)
- **Categories:** Lead generation, Developer tools, Agents
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $7.00 / 1,000 websites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Wappalyzer Scraper** looks up the technology stack of any list of websites the way Wappalyzer does — CMS, e-commerce platform, front-end framework, analytics, advertising pixels, payments, hosting and CDN, CRM and e-mail tools, live chat, cookie consent, reviews, A/B testing — **209 signatures across 32 categories** — and gives it back as one flat row per website. **$0.01 per website**, no start fee, no API key, no browser extension.

Because the home page is already fetched, the same row also carries the **e-mails, phone numbers and social profiles** found on it. One run gives you the stack *and* the way to contact whoever runs it.

Sites that refuse to answer come back as an error row and are **never billed**.

### What is Wappalyzer Scraper?

It is a **Wappalyzer lookup you can run in bulk** — from the Apify console, the API, a spreadsheet, n8n, Make or an AI agent — without the browser extension, the monthly plan or the per-lookup credits. Paste one domain or ten thousand; per website you get what it runs, in which category, whether the evidence was in the page or in a response header, and how to reach the owner.

Typical uses: prospect lists ("Shopify stores that do not use Klaviyo yet"), qualifying inbound leads before the call, competitive research on a market segment, auditing a portfolio of client sites, and enriching a CRM with technographic fields.

### What it detects

| Category | Examples |
|---|---|
| E-commerce | Shopify, WooCommerce, Magento, BigCommerce, PrestaShop, Ecwid |
| CMS and builders | WordPress, Wix, Squarespace, Webflow, Drupal, Joomla, Tilda, HubSpot CMS |
| Front-end | Next.js, Nuxt, React, Vue, Angular, Svelte, Gatsby, Astro, jQuery, Bootstrap, Tailwind |
| Analytics | GA4, Universal Analytics, Matomo, Plausible, Hotjar, Clarity, Mixpanel, Amplitude, Heap |
| Ad pixels | Meta, Google Ads, TikTok, LinkedIn, X, Pinterest, Snapchat, Reddit, Criteo, Taboola |
| Marketing and CRM | HubSpot, Marketo, Klaviyo, Mailchimp, ActiveCampaign, Omnisend, Segment, Tealium, Braze |
| Payments | Stripe, PayPal, Square, Klarna, Afterpay, Zip, Apple Pay |
| Hosting, CDN, server | Cloudflare, Vercel, Netlify, Akamai, Fastly, CloudFront, nginx, Apache, LiteSpeed, IIS |
| Consent and security | OneTrust, Cookiebot, CookieYes, Usercentrics, reCAPTCHA, hCaptcha, Turnstile |
| Chat, booking, reviews | Intercom, Zendesk, Tawk, Crisp, Calendly, Acuity, Mindbody, Yotpo, Judge.me, Trustpilot |
| A/B, search, video, monitoring | Optimizely, VWO, Adobe Target, Algolia, Vimeo, Wistia, Sentry, New Relic, Datadog |

### What the numbers actually are

Measured against 97 real websites — small businesses taken from social-profile links and well-known brands across retail, media, SaaS and local services, in several countries:

| | |
|---|---|
| Websites that answered | **88%** directly, **92%** when blocked ones are retried through a proxy |
| Technologies found per website | **4 on average**, median 4, best 12 |
| Websites with at least one technology | **95%** of those that answered |
| Social profiles found | **78%** of websites |
| E-mail on the home page | **28%** of websites |
| Median time per website | about **1.5 seconds** |

The 8% that refuse everything are heavy retail and media sites behind enterprise bot protection. They come back as an error row with the reason, and they cost nothing.

### How much does it cost?

**$0.01 per website** — one row, one charge, no start fee. A thousand domains is **$10**.

- Sites that refuse every attempt: **$0**.
- Sites removed by your own filters: **$0**.
- Contacts, categories, titles and headers: included, not an add-on.

#### Bulk export: what 50,000 rows actually cost

This Actor is built for bulk jobs — put thousands of domains into one run, or call it from the API on
a schedule. There is **no fee per run, no fee per page and no proxy charge**: you pay for the rows
you keep, and error rows are free. That is what decides the bill once you scan a whole market
rather than a handful of sites.

| Job | This Actor | Most-used Actor in this category |
|---|---|---|
| 50,000 websites across 100 runs | **$500** | $5,000 |

Checked on the Apify Store on 23 September 2026 against the Actor with the most monthly users in
this category, which charges $0.10 per website. Some Actors here ask less per row — this table
compares against the one buyers actually use most.

### How to use it in three steps

1. Paste domains into **🌐 Websites** — `gymshark.com`, `allbirds.com`, or a list of ten thousand from a sheet.
2. If you are building a prospect list, fill in **✅ Keep only sites that use** and **🚫 Drop sites that use**. Filtered-out sites are not billed.
3. Press **Start**, then download JSON, CSV or Excel — or call the run from the API, n8n, Make, Zapier or an AI agent.

### ⬇️ Input

```json
{
  "domains": ["gymshark.com", "allbirds.com", "brooklinen.com"],
  "mustUse": ["shopify"],
  "mustNotUse": ["klaviyo"],
  "includeContacts": true
}
```

That input answers a real sales question: *Shopify stores that have not bought Klaviyo yet* — with the e-mail address to write to, in the same row.

#### Filters

**✅ Keep only sites that use** requires every named technology; **🚫 Drop sites that use** removes a site if any of them is present. Names are matched loosely, so `shopify` matches both Shopify and Shopify Plus. **🗂️ Only these categories** narrows what appears in the row when you only care about, say, payments.

### ⬆️ Output

```json
{
  "type": "website",
  "domain": "gymshark.com",
  "stack": "Shopify · Google Tag Manager · Klaviyo · Cloudflare · Meta Pixel",
  "technologies": ["Shopify", "Google Tag Manager", "Klaviyo", "Cloudflare", "Meta Pixel"],
  "technologyCount": 5,
  "categories": ["ads_pixel", "cdn", "ecommerce", "email_marketing", "tag_manager"],
  "technologiesByCategory": { "ecommerce": ["Shopify"], "cdn": ["Cloudflare"], "email_marketing": ["Klaviyo"] },
  "ecommerce": "Shopify",
  "hosting": ["Cloudflare"],
  "marketing": ["Klaviyo"],
  "adsPixels": ["Meta Pixel"],
  "emails": ["press@gymshark.com"],
  "instagram": "https://instagram.com/gymshark",
  "title": "Gymshark",
  "readVia": "direct",
  "scrapedAt": "2026-09-23T12:04:11.004Z"
}
```

`technologiesDetailed` adds, per technology, its category and whether the evidence was in the page or in a response header. Two table views come with the dataset: **Tech stack** and **Contacts**.

### Use cases

#### Prospect lists for anyone selling to online stores

Filter by platform, exclude the competitor's tool, keep the e-mail. That is the whole workflow, in one run.

#### Qualifying inbound leads

Before the call, know whether they are on Shopify or Magento, whether they already run a consent platform, and which analytics they trust.

#### Competitive and market research

Scan a market segment and count platforms, pixels and payment providers. `technologiesByCategory` makes the aggregation a one-liner.

#### Agencies auditing client sites

Check a portfolio for missing pixels, missing consent banners or an outdated framework, on a schedule.

#### AI agents

Connect through the Apify MCP server and ask *"what does gymshark.com run on?"* — the agent gets rows, not a screenshot.

### Integrations

Every run is available through the Apify API, and the Actor works out of the box with n8n, Make, Zapier, Google Sheets, LangChain and the Apify MCP server. Schedule it, webhook it, or export straight to CSV.

### FAQ

**Is this the official Wappalyzer?** No. It is an independent Actor with its own signature set, built to give the same kind of answer — the technologies a site runs — in bulk and at $0.01 per site.

**Why is a site missing a technology I know it uses?** Detection reads the home page and its response headers. A tool loaded only on inner pages or after login is not visible from the home page. Use `technologiesDetailed` to see the evidence for what was found.

**Do I need a proxy?** Usually not. 88% of sites answer a direct request; the rest are retried through residential addresses automatically when **Retry blocked sites through a proxy** is on.

**Is this legal?** The Actor reads publicly available home pages and response headers, the same thing a browser does. Contact data is limited to what a site publishes about itself; use it in line with GDPR, CCPA and the site's terms.

**Is there a run fee, a monthly plan or a minimum?** No. One charge per delivered website, and nothing else.

# Actor input Schema

## `domains` (type: `array`):

Domains or full URLs to look up, the way you would type them into Wappalyzer — <b>shopify.com</b>, <b>https://gymshark.com</b>. One row per website, $0.01 each; sites that do not answer are free.

## `startUrls` (type: `array`):

The same thing from a file or a Google Sheet, for long lists.

## `categoriesFilter` (type: `array`):

Leave empty for everything the site uses. Pick a few to keep the rows narrow — for example only e-commerce and payments.

## `mustUse` (type: `array`):

Technology names, matched loosely: <b>shopify</b>, <b>klaviyo</b>. A site must use all of them to be delivered — and a site that is filtered out is not charged.

## `mustNotUse` (type: `array`):

The opposite: skip anything already running these. <b>hubspot</b>, <b>intercom</b> — the classic way to find prospects who have not bought your competitor yet.

## `includeContacts` (type: `boolean`):

E-mails, tel: numbers and social profiles found on the home page, in the same row as the technologies. Free — it is the page we already fetched.

## `includeSignals` (type: `boolean`):

Page title, meta description, generator tag, server and x-powered-by headers.

## `useProxyWhenBlocked` (type: `boolean`):

Most sites answer a direct request. The ones that refuse are retried from residential addresses, which lifted coverage from 88% to 92% in testing. Turn it off for the cheapest, fastest run.

## `concurrency` (type: `integer`):

How many websites to read at the same time.

## `maxItems` (type: `integer`):

A hard ceiling on what one run delivers and bills.

## `proxyConfiguration` (type: `object`):

Used only for sites that refuse a direct request.

## Actor input object example

```json
{
  "domains": [
    "gymshark.com",
    "allbirds.com",
    "notion.so"
  ],
  "startUrls": [],
  "categoriesFilter": [],
  "mustUse": [],
  "mustNotUse": [],
  "includeContacts": true,
  "includeSignals": true,
  "useProxyWhenBlocked": true,
  "concurrency": 10,
  "maxItems": 10000,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `websites` (type: `string`):

One row per website: every technology detected with its category, Wappalyzer-style, plus contacts from the same page.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "gymshark.com",
        "allbirds.com",
        "notion.so"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("lergassy/wappalyzer-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "gymshark.com",
        "allbirds.com",
        "notion.so",
    ],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("lergassy/wappalyzer-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "gymshark.com",
    "allbirds.com",
    "notion.so"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call lergassy/wappalyzer-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lergassy/wappalyzer-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7gTzgXG6qhLZGM1KQ/builds/PfYOHQjVxYHztNL6w/openapi.json
