# Website Technographic & Marketing Stack Scraper (`jungle_synthesizer/website-technographic-marketing-stack-scraper`) Actor

Detects the martech, ad-pixel, CMP, commerce-app and network stack behind any domain — not just the CMS. Returns pixel/analytics account ids, consent-platform vendor, payment and loyalty apps, plus DNS/TLS network enrichment (email provider, ASN org, DNS provider).

- **URL**: https://apify.com/jungle\_synthesizer/website-technographic-marketing-stack-scraper.md
- **Developed by:** [BowTiedRaccoon](https://apify.com/jungle_synthesizer) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.80 / 1,000 record scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Technographic & Marketing Stack Scraper

Detects the martech, ad-pixel, and commerce-app stack behind any domain — not just the CMS. Point it at a list of domains and get back pixel and analytics account IDs, the consent-management vendor, payment and loyalty apps, plus DNS/TLS network facts like email provider and hosting ASN, for as many domains as you feed it.

***

### Website Technographic Scraper Features

- Extracts ad-pixel vendors **with the account ID attached** — `meta_pixel(1234567890123456)`, not just "has a Meta Pixel"
- Identifies the consent-management platform (OneTrust, Cookiebot, Usercentrics, and six others) — the field most technographic tools skip entirely
- Covers the full commerce-app layer: personalisation, review widgets, subscription/loyalty apps, payment processors, shipping/fulfilment tools
- Resolves network facts no HTML-only tool can reach — mail provider from MX, DNS operator from NS, hosting org from ASN, TLS certificate issuer and expiry
- Detects ecommerce platform, CMS, and JS framework in the same pass
- Attaches a matched-signature and confidence score to every detected field, so you can see exactly why a value was returned
- Optional headless-render retry for domains whose first pass turns up almost nothing, billed only when it actually runs

***

### Who Uses Technographic Data?

- **Martech and Shopify-app vendors** — build a prospect list filtered to sites NOT already running a competing app
- **Ad agencies** — qualify inbound leads by checking which pixels and ad platforms a prospect already runs
- **RevOps teams** — feed technographic fields into Clay or HubSpot for ABM segment building
- **Privacy and consent-platform vendors** — target sites with no CMP detected, or a competitor's CMP
- **MSPs and email-security resellers** — filter by email provider (Google Workspace vs. Microsoft 365) as a standing ICP signal

***

### How the Technographic Scraper Works

1. Feed it a list of domains — bare hostnames like `gymshark.com`, no scheme needed.
2. Each domain's homepage is fetched and checked against a signature table covering 15+ technology categories, from ad pixels to shipping apps.
3. In parallel, the domain's MX, NS, and A records resolve to an email provider, DNS operator, and hosting ASN, and its TLS certificate is read for issuer and expiry.
4. A domain that comes back with almost no signal — usually a heavier client-rendered app shell — can optionally get one retry through a full browser render, billed as an add-on only when it fires.

***

### Input

```json
{
  "domains": ["gymshark.com", "stripe.com"],
  "enableRenderFallback": false,
  "maxItems": 100
}
```

| Field                  | Type    | Default | Description |
|------------------------|---------|---------|-------------|
| `domains`              | array   | —       | Domains to fingerprint, e.g. `gymshark.com`. One record is returned per domain. |
| `enableRenderFallback` | boolean | `false` | Retry a domain through a headless browser when the plain fetch finds fewer than 3 signals. Billed as an additional "Record scraped - premium" event on top of the base charge for that domain. |
| `maxItems`             | integer | —       | Maximum number of domain records to return. |

***

### Website Technographic Scraper Output Fields

Real output from a run against `allbirds.com` (`third_party_script_domains` trimmed for length):

```json
{
  "domain": "allbirds.com",
  "final_url": "https://www.allbirds.com/",
  "http_status": 200,
  "ecommerce_platform": "shopify",
  "cms": null,
  "js_framework": [],
  "analytics": ["gtm(GTM-TH8KRSBJ)"],
  "martech": ["attentive"],
  "cmp_vendor": null,
  "cdn": ["cloudflare"],
  "waf_vendor": "cloudflare",
  "hosting_provider": "CLOUDFLARENET - Cloudflare, Inc., US",
  "asn": "AS13335",
  "asn_org": "CLOUDFLARENET - Cloudflare, Inc., US",
  "email_provider": "microsoft_365",
  "dns_provider": "ns3.markmonitor.com",
  "ssl_issuer": "Let's Encrypt",
  "ssl_expires_at": "2026-11-09T00:48:32.000Z",
  "security_headers": { "hsts": true, "csp": true, "x_frame_options": true },
  "server_header": "cloudflare",
  "tag_manager_containers": ["GTM-TH8KRSBJ"],
  "third_party_script_domains": ["cdn.shopify.com", "www.googletagmanager.com", "www.facebook.com", "www.tiktok.com", "..."],
  "render_used": false
}
```

| Field                        | Type     | Description |
|------------------------------|----------|-------------|
| `domain`                     | string   | The normalized input domain (bare host, no scheme/www). |
| `final_url`                  | string   | The URL the homepage actually resolved to, after any redirect (catches geo/locale redirects). |
| `http_status`                | number   | HTTP status code of the final response. |
| `ecommerce_platform`         | string   | shopify | woocommerce | magento | bigcommerce | salesforce\_cc | shopware | prestashop | wix\_stores | squarespace\_commerce | null |
| `cms`                        | string   | wordpress | drupal | contentful | sanity | webflow | hubspot\_cms | adobe\_aem | typo3 | craft | ghost | null |
| `cms_version`                | string   | CMS version, where the generator meta or a versioned asset path exposes it. |
| `js_framework`               | array    | Detected JS frameworks: react | next | nuxt | vue | svelte | angular | remix | astro |
| `ad_pixels`                  | array    | Ad pixel vendors with account ID where extractable, e.g. `meta_pixel(123456789012345)`. |
| `analytics`                  | array    | Analytics vendors with measurement/container ID where extractable, e.g. `ga4(G-ABC1234567)`. |
| `session_replay`             | array    | hotjar(site\_id) | fullstory | clarity | contentsquare | quantum\_metric | logrocket |
| `martech`                    | array    | Marketing-automation vendors with account ID where extractable, e.g. `klaviyo(company_id)`. |
| `crm_signals`                | array    | hubspot\_forms | salesforce\_web2lead | pipedrive | zoho | intercom | drift | zendesk\_widget | freshchat |
| `cmp_vendor`                 | string   | onetrust | cookiebot | usercentrics | trustarc | quantcast | didomi | osano | klaro | null |
| `personalisation`            | array    | dynamic\_yield | optimizely | vwo | ab\_tasty | monetate | bloomreach | algolia | constructor\_io |
| `review_platform`            | array    | trustpilot | yotpo | okendo | judge\_me | bazaarvoice | reviews\_io | stamped | loox |
| `subscription_loyalty`       | array    | recharge | bold | smile\_io | loyaltylion | yotpo\_loyalty | stay\_ai |
| `payment_processors`         | array    | stripe | braintree | adyen | paypal | klarna | afterpay | affirm | checkout\_com | square |
| `shipping_fulfilment`        | array    | shippo | shipbob | narvar | aftership | route | shipstation |
| `cdn`                        | array    | The content delivery network(s) fronting the site, detected from response headers. |
| `hosting_provider`           | string   | Hosting org resolved from the homepage's A record ASN. |
| `asn`                        | string   | Autonomous system number of the homepage's A record, e.g. `AS13335`. |
| `asn_org`                    | string   | Organization name registered against that ASN. |
| `email_provider`             | string   | google\_workspace | microsoft\_365 | zoho | proofpoint | mimecast | self\_hosted | null — from MX records |
| `dns_provider`               | string   | Nameserver operator, classified from the domain's NS records. |
| `ssl_issuer`                 | string   | Certificate authority organization from the live TLS certificate. |
| `ssl_expires_at`             | date     | TLS certificate expiry, ISO 8601. |
| `security_headers`           | object   | `{ hsts, csp, x_frame_options }` — boolean presence of each header. |
| `server_header`              | string   | Raw Server response header. |
| `powered_by_header`          | string   | Raw X-Powered-By response header. |
| `waf_vendor`                 | string   | The bot-management / firewall vendor protecting the site, when detectable from response headers. |
| `tag_manager_containers`     | array    | GTM container IDs extracted from the analytics signature matches. |
| `third_party_script_domains` | array    | Every distinct external script/asset host referenced by the page — the raw evidence trail. |
| `detection_evidence`         | object   | Per detected field: which signature matched and where (html | header). |
| `confidence`                 | object   | Per detected field/value, a 0-1 confidence score (1.0 when an account ID was extracted). |
| `render_used`                | boolean  | True when the record required the render fallback pass. |
| `scraped_at`                 | datetime | Timestamp the record was emitted. |

***

### FAQ

#### How do I detect what technology a website uses?

Feed a domain list into this actor's `domains` input. Each domain returns one record covering its ecommerce platform, CMS, ad pixels, analytics, consent platform, and commerce apps in a single pass.

#### Does it return ad pixel account IDs, or just which pixel is installed?

Both, when the ID is extractable from the page — `meta_pixel(1234567890123456)`, `tiktok_pixel(...)`, `ga4(G-ABC1234567)`. The account ID is what turns a technology flag into a usable ad-spend prospect record.

#### Can it check thousands of domains in one run?

Yes — `domains` accepts an arbitrary list and `maxItems` caps how many records come back. Very large lists (tens of thousands+) are best split across a few runs.

#### Does it detect the CMP (cookie consent) vendor?

Yes — `cmp_vendor` covers OneTrust, Cookiebot, Usercentrics, TrustArc, Quantcast, Didomi, Osano, and Klaro.

#### What if a site is a JavaScript app shell with almost nothing in the raw HTML?

Set `enableRenderFallback: true`. A domain whose first pass finds fewer than 3 signals gets one retry through a full page render, billed as an add-on event only on the domains where it actually runs.

***

### Need More Features?

Open an issue on the actor's Apify Store page and we'll take a look.

### Why Use This Technographic Scraper?

- **Goes past the CMS guess** — ad-pixel IDs, consent platform, and the full commerce-app layer (personalisation, reviews, loyalty, payments, shipping) in one output row
- **Network-side facts an HTML-only tool can't reach** — email provider, DNS operator, and hosting org, resolved from the domain's own DNS and TLS records
- **Every detection is auditable** — `detection_evidence` and `confidence` show exactly what matched and how sure the actor is, instead of a bare boolean

# Actor input Schema

## `sp_intended_usage` (type: `string`):

What will this data feed? E.g. lead lists, KYB checks, price tracking.

## `sp_improvement_suggestions` (type: `string`):

Provide any feedback or suggestions for improvements.

## `sp_contact` (type: `string`):

We'll personally help with your use case. No spam.

## `domains` (type: `array`):

Domains to fingerprint, e.g. gymshark.com. One record is returned per domain.

## `enableRenderFallback` (type: `boolean`):

When the plain HTML fetch finds fewer than 3 technology signals (typical of a client-rendered app shell), retry the domain through a headless browser render pass. Billed as an additional "Record scraped - premium" event on top of the base charge for that domain, since the render pass carries real added compute cost.

## `maxItems` (type: `integer`):

Maximum number of domain records to return

## Actor input object example

```json
{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "domains": [
    "gymshark.com",
    "stripe.com"
  ],
  "enableRenderFallback": false,
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "domains": [
        "gymshark.com",
        "stripe.com"
    ],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("jungle_synthesizer/website-technographic-marketing-stack-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "domains": [
        "gymshark.com",
        "stripe.com",
    ],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("jungle_synthesizer/website-technographic-marketing-stack-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "domains": [
    "gymshark.com",
    "stripe.com"
  ],
  "maxItems": 10
}' |
apify call jungle_synthesizer/website-technographic-marketing-stack-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jungle_synthesizer/website-technographic-marketing-stack-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Kuoo7r0Nn5nqWY0cg/builds/ep1i30fSkOCM7YHXQ/openapi.json
