# Website stack · Apps, ATS, colors & llms.txt by URL (`corent1robert/website-signals-scraper`) Actor

Paste company URLs — get CMS, HubSpot, Klaviyo, Shopify apps, CMP, demo link, role emails, llms.txt, brand colors, hiring ATS and careers tool mentions. One row per site, CSV-ready. No login. No API key.

- **URL**: https://apify.com/corent1robert/website-signals-scraper.md
- **Developed by:** [Corentin Robert](https://apify.com/corent1robert) (community)
- **Categories:** Lead generation, E-commerce, Marketing
- **Stats:** 2 total users, 2 monthly users, 65.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 homepage stacks

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Website stack · Apps, ATS, colors & llms.txt by URL

Paste **company URLs**. Get **CMS**, marketing pixels (HubSpot, GTM, Klaviyo), **Shopify apps** (Gorgias, Recharge…), **CMP**, **locales**, **llms.txt**, **brand colors**, on-page facts, newsletter presence, and **careers ATS + tool mentions** — one row per site, CSV-ready.

**No login. No API key. No browser.**

### Who is this for?

| You are… | Typical goal | Suggested setup |
|----------|--------------|-----------------|
| **Shopify / HubSpot agency** | Find Shopify without Klaviyo, Gorgias, Recharge, reviews or loyalty | Default run, filter `shopify_without_*` |
| **Outbound SDR / BDR** | Personalize by stack, demo link, generic inbox, colors, hiring | Colors ON, careers ON |
| **Email / ESP vendor** | See who already uses Klaviyo, Mailchimp, Brevo, Substack | Default run — we never subscribe |
| **AI tooling vendor** | See who publishes `llms.txt`, or if jobs pages mention Cursor / Copilot | Default run + careers ON |
| **RevOps** | Enrich a CRM domain list before sequencing | Paste domains |
| **Designer / mockup** | Pull 3–5 brand hex codes for a one-pager | Colors ON, careers off |

What you get by default: CMS and shop platform, HubSpot/GTM/Klaviyo flags, Shopify apps and CMP from script/CDN hosts (not competitor blog links), locales, public `/llms.txt` if it is a real markdown file, Calendly/Cal.com **event URLs** or an on-site `/demo` link when present in the HTML, a contact page plus **role mailboxes only** (`hello@`, `contact@` — never `jean.dupont@`), newsletter form URL (no signup), 3–5 hex colors, title/H1/schema, Google vs Microsoft 365 from MX, careers URL, ATS (Ashby, Lever, Teamtailor…), tool names **only** from jobs pages.

When to turn options off: skip **careers** if you only need homepage stack — you then pay the homepage event only (no hiring-page event). Skip **colors** if you only need pixels; skip **email MX** if you already ran a DNS stack check. Colors and MX stay inside the homepage event.

### Quick start

1. Open the Actor in [Apify Console](https://console.apify.com/actors/RYCFeU4YU8i9LiMJo).
2. Paste domains or URLs, one per line (`enky.com`, `payfit.com`).
3. Leave the three include toggles **on** for the full row.
4. Click **Start** — rows appear in the Dataset tab as each site finishes.
5. Export CSV / JSON / Excel.

### Ready-made examples

Saved inputs live in `published-tasks/inputs/` (publish from Console for Store landing pages):

| Example | Best for |
|---------|----------|
| Shopify stores without Klaviyo | Agency outbound |
| Companies already on HubSpot | HubSpot ISV / agency |
| Hiring pages and careers ATS | Talent / AI tooling |
| Brand colors for outreach mockups | SDR personalization |

### What it extracts

| Category | Fields |
|----------|--------|
| Identity | `domain`, `final_url`, `page_title`, `h1`, `lang`, `locales` |
| CMS / shop | `cms`, `ecommerce`, `shopify_without_klaviyo`, `shopify_without_gorgias`, `shopify_without_recharge`, `shopify_without_reviews`, `shopify_without_loyalty` |
| Marketing stack | `hubspot`, `gtm`, `ga4`, `klaviyo`, `intercom`, `crisp`, `calendly`, `stripe`, `meta_pixel`, `hotjar`, `linkedin_insight`, `typeform` |
| Shopify apps | `gorgias`, `judgeme`, `loox`, `yotpo`, `smile`, `recharge`, `skio`, `loyaltylion` |
| Consent | `consent_manager` (axeptio, didomi, cookiebot, onetrust, cookieyes, tarteaucitron) |
| AI crawl file | `has_llms_txt`, `llms_txt_url` (origin `/llms.txt` only; body not stored) |
| Outreach | `demo_url`, `meeting_tool`, `contact_url`, `generic_emails` (homepage HTML only; no personal inboxes) |
| Newsletter | `has_newsletter`, `newsletter_provider`, `newsletter_form_url` (never submitted) |
| Brand | `colors`, `theme_color` |
| On-page | `meta_description`, `schema_types`, `has_blog`, `has_pricing` |
| Email | `email_platform`, `mx_host` |
| Careers | `careers_url`, `hiring_now`, `careers_ats`, `careers_stack_mentions` |

`cms` values include `shopify`, `wordpress`, `webflow`, `wix`, `squarespace`, `framer`, `ghost`, `drupal`, `joomla`, `hubspot_cms`, `bigcommerce`, `magento`, `prestashop`, `bubble`, or `custom`. WordPress shops are `cms: wordpress` + `ecommerce: woocommerce`.

### Typical fill rates

From a 30-site mixed probe (SaaS + ecommerce + CMS vendors, HTTP homepage):

| Signal | Indicative rate |
|--------|-----------------|
| Homepage loaded | ~100% on public sites |
| GTM | ~50% |
| Hiring page found | ~70% when careers ON |
| Newsletter form / ESP | ~10–15% |
| Intercom / Calendly **widget** on the homepage HTML | often 0% (loads in JS) |
| Calendly / Cal.com **event URL** in an `href` / `iframe` | when they published a bookable link |

Treat pixels as “present in homepage HTML”, not “the company never uses this tool”.

### How much does it cost to scrape website stack signals?

Pay-per-event. **You pay for what landed in the row**, not for empty 404s. HTTP-only — compute stays low. There is **no** actor-start fee.

| Event | Price | When it fires | What you get in that charge |
|-------|-------|----------------|-----------------------------|
| **Homepage stack** (`website-homepage`) | **$0.003** | Homepage HTML loaded | CMS, shop flag, pixels, Shopify apps, CMP, locales, `/llms.txt` check, demo / contact / role emails, newsletter URL (never submitted), on-page title/H1/schema, brand colors, Google vs Microsoft 365 MX |
| **Hiring page found** (`website-careers`) | **$0.002** | A public jobs page or ATS was found (`hiring_now` or `careers_ats`) | Careers URL, ATS (Ashby, Lever, Teamtailor, Welcome to the Jungle…), tool names on that page |

**Not billed**

- Homepage that failed to load (`fetched: false`)
- Careers scan that only hit 404s / empty shells (no jobs page, no ATS)
- Extra CSS files fetched for colors
- Extra GET for `/llms.txt` (404s and HTML soft-404s stay `has_llms_txt: false`)
- `includeCareers` **On** with no hiring page — homepage event only
- Turning colors / MX **Off** does not change the homepage price (same event; fewer fields)

**Typical mix** (same 30-site probe as fill rates): ~70% of sites with careers **On** produce a hiring page → most full runs are **$0.003 + $0.002 = $0.005** per site, not $0.005 on every URL.

| Scenario | Events | Approx. cost |
|----------|--------|----------------|
| 25 sites, homepage only (careers Off) | 25 × $0.003 | **$0.08** |
| 25 sites, careers On, ~70% hiring pages | 25 × $0.003 + ~18 × $0.002 | **~$0.11** |
| 25 sites, every site has a jobs page | 25 × $0.003 + 25 × $0.002 | **$0.13** |
| 1,000 sites, homepage only | 1,000 × $0.003 | **$3.00** |
| 1,000 sites, careers On, ~70% hiring | 1,000 × $0.003 + ~700 × $0.002 | **~$4.40** |
| 1,000 sites, every site hiring | 1,000 × $0.005 | **$5.00** |
| 50,000 sites, homepage only | 50,000 × $0.003 | **$150** |
| 50,000 sites, ~70% hiring | 50,000 × $0.003 + ~35,000 × $0.002 | **~$220** |

Dataset rows are still written when the homepage fails (with `error`) — those rows are **not** charged.

### Is scraping public website signals free?

A 25-site try is about **$0.08–$0.13** (pay-per-event). It is not an unlimited free export. Compute is HTTP-only and small next to the events above.

### Is it legal to scrape public website signals?

This Actor only reads **public HTML, CSS, DNS MX, and `/llms.txt`** that the company already publishes. It does **not** log in, submit newsletter forms, crawl inboxes, or collect personal email addresses. As with any dataset of company identifiers, ensure your use complies with GDPR and applicable regulations.

### Input

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `urls` | string\[] | — | Company domains or URLs (required) |
| `includeCareers` | boolean | `true` | Scan jobs pages. Extra **$0.002** only when a hiring page / ATS is found |
| `includeColors` | boolean | `true` | Extract 3–5 brand hex colors |
| `includeEmailStack` | boolean | `true` | Detect Google Workspace vs Microsoft 365 from MX |

API-only: `maxUrls` (integer, default 0 = no cap after dedupe), `verboseLogs` (boolean) — technical skip reasons in the run log. `proxyConfiguration` — omitted from the Console form; cloud runs use Apify datacenter proxy by default. Power users can pass it in JSON/API input.

```json
{
  "urls": ["enky.com", "payfit.com", "alan.com"],
  "includeCareers": true,
  "includeColors": true,
  "includeEmailStack": true
}
```

### Output example

```json
{
  "domain": "enky.com",
  "cms": "shopify",
  "ecommerce": "shopify",
  "shopify_without_klaviyo": false,
  "shopify_without_gorgias": true,
  "shopify_without_recharge": true,
  "shopify_without_reviews": true,
  "shopify_without_loyalty": true,
  "gorgias": false,
  "recharge": false,
  "consent_manager": null,
  "locales": ["en"],
  "has_llms_txt": false,
  "llms_txt_url": null,
  "demo_url": null,
  "meeting_tool": null,
  "contact_url": "https://enky.com/pages/contact",
  "generic_emails": ["hello@enky.com"],
  "hubspot": false,
  "gtm": true,
  "klaviyo": true,
  "has_newsletter": true,
  "newsletter_provider": "klaviyo",
  "newsletter_form_url": "https://enky.com/",
  "colors": ["#f0977a", "#0e7ea5", "#6a8691"],
  "page_title": "Enky",
  "email_platform": "google_workspace",
  "careers_url": "https://www.welcometothejungle.com/fr/companies/enky/jobs",
  "careers_ats": "welcometothejungle",
  "hiring_now": true,
  "careers_stack_mentions": [],
  "fetched": true,
  "error": null
}
```

Klaviyo on the homepage is a pixel/form signal — it can vary by theme and cookie banner.

### How it works

1. **Normalize** pasted URLs to one origin per company (www deduped).
2. **Fetch the homepage** (HTTP + HTML). Detect CMS, pixels, Shopify apps (CDN hosts only), CMP, locales, newsletter, demo/meeting links, contact page and role mailboxes, on-page facts, colors from CSS. GET `/llms.txt` at the origin — keep the URL only if it looks like markdown, not a soft-404 HTML page.
3. **Optional MX** — Google Workspace vs Microsoft 365.
4. **Optional careers** — follow homepage job links in any language, then locale-aware fallbacks (`/karriere`, `/empleo`, `/vagas`…). Read ATS hosts and tool names **only** on those pages.

### Local development

```bash
cd website-signals-scraper
npm install
npm test
apify run
```

The CLI validates `storage/key_value_stores/default/INPUT.json` against the input schema. Use `.actor/INPUT.json` as the Console prefill, or:

```bash
apify run --input-file=./.actor/INPUT.json
node scripts/local-probe.mjs
```

When running locally, keys in the simulated KV input **override** a root `input.json` if you add one.

### Limitations

- Homepage HTML plus a few stylesheets and careers URLs. No Lighthouse, no 50-page crawl, no Playwright in v1.
- A Cursor mention on a jobs page is not proof they use it every day.
- Heavy JS-only sites may look `custom` with empty pixels.
- We never subscribe to newsletters.
- We never collect personal inboxes (`jean.dupont@`). Role mailboxes on the homepage only.
- A Calendly **widget** with no public event `href` leaves `demo_url` empty.

### Also available

**[Welcome to the Jungle · Hiring Companies & Profiles](https://apify.com/corent1robert/wttj-hiring-signal-scraper)** — enrich a WTTJ company profile or discover companies hiring by sector.

**[Tech Stack · M365 & Google Workspace](https://apify.com/corent1robert/collaboration-platform-detector)** — bulk DNS check for Microsoft 365 vs Google Workspace (thousands of domains, no HTML).

### Support

Contact <corentin@outreacher.fr> if you need a custom scraper or tailored automation.

### FAQ

**Do I need Shopify / HubSpot admin?** No. Public homepage HTML only.

**Will you subscribe to their newsletter?** No. We only record that a form exists and which ESP it uses.

**Why is Intercom false on Intercom’s own site?** Many widgets inject after JavaScript. v1 reads the first HTML response, not a full browser session.

**Do you scrape employee emails?** No. Only role mailboxes already in homepage `mailto:` links (`hello@`, `contact@`, `sales@`).

**Why is `demo_url` empty if they use Calendly?** We keep a URL only when the HTML already has a bookable link (`calendly.com/team/…`, Cal.com, HubSpot Meetings) or an on-site `/demo` path. The JS widget alone is not enough.

**What is `llms.txt`?** A public markdown file at `https://example.com/llms.txt` ([llmstxt.org](https://llmstxt.org/)) so AI tools can read a site summary. We record whether it exists and the URL — not the file body.

**Do HubSpot / Klaviyo / Shopify / Gorgias flags cost extra?** No. They are part of the **$0.003** homepage event.

**Do I pay if careers is On but there is no jobs page?** No second event. Only the homepage charge if the site loaded.

**Can I pay homepage-only?** Yes. Turn **Scan careers** Off. Still **$0.003** per loaded site.

# Actor input Schema

## `urls` (type: `array`):

Domains or full URLs, one per line (enky.com). Each site becomes one row.

## `includeCareers` (type: `boolean`):

Find hiring pages in any language and ATS / tool names. Extra charge only when a jobs page is found. Off = homepage stack only.

## `includeColors` (type: `boolean`):

Pull 3–5 hex colors for outreach mockups.

## `includeEmailStack` (type: `boolean`):

See Microsoft 365 vs Google Workspace from public MX. No extra DNS actor to run.

## `maxUrls` (type: `integer`):

API-only. Cap after dedupe. 0 = no cap (the run still stops at the time limit). Console volume is the URL list.

## Actor input object example

```json
{
  "urls": [
    "enky.com",
    "payfit.com"
  ],
  "includeCareers": true,
  "includeColors": true,
  "includeEmailStack": true,
  "maxUrls": 0
}
```

# Actor output Schema

## `dataset` (type: `string`):

Full export — one analyzed website per row.

## `overview` (type: `string`):

CMS, pixels, ATS and hiring at a glance.

## `outreach` (type: `string`):

Shopify gaps, demo URL, role emails, CMP, llms.txt, colors, newsletter URL, careers link.

## `runLog` (type: `string`):

Text log mirrored from the run console.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "enky.com",
        "payfit.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("corent1robert/website-signals-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "enky.com",
        "payfit.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("corent1robert/website-signals-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "enky.com",
    "payfit.com"
  ]
}' |
apify call corent1robert/website-signals-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,corent1robert/website-signals-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/RYCFeU4YU8i9LiMJo/builds/iUT3LRLRD3mRnqT0W/openapi.json
