Website stack · Apps, ATS, colors & llms.txt by URL
Pricing
from $3.00 / 1,000 homepage stacks
Website stack · Apps, ATS, colors & llms.txt by URL
Paste company URLs — get CMS, HubSpot, Klaviyo, Shopify apps, CMP, demo link, role emails, llms.txt, brand colors, hiring ATS and careers tool mentions. One row per site, CSV-ready. No login. No API key.
Pricing
from $3.00 / 1,000 homepage stacks
Rating
0.0
(0)
Developer
Corentin Robert
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
2
Monthly active users
4 days ago
Last modified
Categories
Share
Paste company URLs. Get CMS, marketing pixels (HubSpot, GTM, Klaviyo), Shopify apps (Gorgias, Recharge…), CMP, locales, llms.txt, brand colors, on-page facts, newsletter presence, and careers ATS + tool mentions — one row per site, CSV-ready.
No login. No API key. No browser.
Who is this for?
| You are… | Typical goal | Suggested setup |
|---|---|---|
| Shopify / HubSpot agency | Find Shopify without Klaviyo, Gorgias, Recharge, reviews or loyalty | Default run, filter shopify_without_* |
| Outbound SDR / BDR | Personalize by stack, demo link, generic inbox, colors, hiring | Colors ON, careers ON |
| Email / ESP vendor | See who already uses Klaviyo, Mailchimp, Brevo, Substack | Default run — we never subscribe |
| AI tooling vendor | See who publishes llms.txt, or if jobs pages mention Cursor / Copilot | Default run + careers ON |
| RevOps | Enrich a CRM domain list before sequencing | Paste domains |
| Designer / mockup | Pull 3–5 brand hex codes for a one-pager | Colors ON, careers off |
What you get by default: CMS and shop platform, HubSpot/GTM/Klaviyo flags, Shopify apps and CMP from script/CDN hosts (not competitor blog links), locales, public /llms.txt if it is a real markdown file, Calendly/Cal.com event URLs or an on-site /demo link when present in the HTML, a contact page plus role mailboxes only (hello@, contact@ — never jean.dupont@), newsletter form URL (no signup), 3–5 hex colors, title/H1/schema, Google vs Microsoft 365 from MX, careers URL, ATS (Ashby, Lever, Teamtailor…), tool names only from jobs pages.
When to turn options off: skip careers if you only need homepage stack — you then pay the homepage event only (no hiring-page event). Skip colors if you only need pixels; skip email MX if you already ran a DNS stack check. Colors and MX stay inside the homepage event.
Quick start
- Open the Actor in Apify Console.
- Paste domains or URLs, one per line (
enky.com,payfit.com). - Leave the three include toggles on for the full row.
- Click Start — rows appear in the Dataset tab as each site finishes.
- Export CSV / JSON / Excel.
Ready-made examples
Saved inputs live in published-tasks/inputs/ (publish from Console for Store landing pages):
| Example | Best for |
|---|---|
| Shopify stores without Klaviyo | Agency outbound |
| Companies already on HubSpot | HubSpot ISV / agency |
| Hiring pages and careers ATS | Talent / AI tooling |
| Brand colors for outreach mockups | SDR personalization |
What it extracts
| Category | Fields |
|---|---|
| Identity | domain, final_url, page_title, h1, lang, locales |
| CMS / shop | cms, ecommerce, shopify_without_klaviyo, shopify_without_gorgias, shopify_without_recharge, shopify_without_reviews, shopify_without_loyalty |
| Marketing stack | hubspot, gtm, ga4, klaviyo, intercom, crisp, calendly, stripe, meta_pixel, hotjar, linkedin_insight, typeform |
| Shopify apps | gorgias, judgeme, loox, yotpo, smile, recharge, skio, loyaltylion |
| Consent | consent_manager (axeptio, didomi, cookiebot, onetrust, cookieyes, tarteaucitron) |
| AI crawl file | has_llms_txt, llms_txt_url (origin /llms.txt only; body not stored) |
| Outreach | demo_url, meeting_tool, contact_url, generic_emails (homepage HTML only; no personal inboxes) |
| Newsletter | has_newsletter, newsletter_provider, newsletter_form_url (never submitted) |
| Brand | colors, theme_color |
| On-page | meta_description, schema_types, has_blog, has_pricing |
email_platform, mx_host | |
| Careers | careers_url, hiring_now, careers_ats, careers_stack_mentions |
cms values include shopify, wordpress, webflow, wix, squarespace, framer, ghost, drupal, joomla, hubspot_cms, bigcommerce, magento, prestashop, bubble, or custom. WordPress shops are cms: wordpress + ecommerce: woocommerce.
Typical fill rates
From a 30-site mixed probe (SaaS + ecommerce + CMS vendors, HTTP homepage):
| Signal | Indicative rate |
|---|---|
| Homepage loaded | ~100% on public sites |
| GTM | ~50% |
| Hiring page found | ~70% when careers ON |
| Newsletter form / ESP | ~10–15% |
| Intercom / Calendly widget on the homepage HTML | often 0% (loads in JS) |
Calendly / Cal.com event URL in an href / iframe | when they published a bookable link |
Treat pixels as “present in homepage HTML”, not “the company never uses this tool”.
How much does it cost to scrape website stack signals?
Pay-per-event. You pay for what landed in the row, not for empty 404s. HTTP-only — compute stays low. There is no actor-start fee.
| Event | Price | When it fires | What you get in that charge |
|---|---|---|---|
Homepage stack (website-homepage) | $0.003 | Homepage HTML loaded | CMS, shop flag, pixels, Shopify apps, CMP, locales, /llms.txt check, demo / contact / role emails, newsletter URL (never submitted), on-page title/H1/schema, brand colors, Google vs Microsoft 365 MX |
Hiring page found (website-careers) | $0.002 | A public jobs page or ATS was found (hiring_now or careers_ats) | Careers URL, ATS (Ashby, Lever, Teamtailor, Welcome to the Jungle…), tool names on that page |
Not billed
- Homepage that failed to load (
fetched: false) - Careers scan that only hit 404s / empty shells (no jobs page, no ATS)
- Extra CSS files fetched for colors
- Extra GET for
/llms.txt(404s and HTML soft-404s stayhas_llms_txt: false) includeCareersOn with no hiring page — homepage event only- Turning colors / MX Off does not change the homepage price (same event; fewer fields)
Typical mix (same 30-site probe as fill rates): ~70% of sites with careers On produce a hiring page → most full runs are $0.003 + $0.002 = $0.005 per site, not $0.005 on every URL.
| Scenario | Events | Approx. cost |
|---|---|---|
| 25 sites, homepage only (careers Off) | 25 × $0.003 | $0.08 |
| 25 sites, careers On, ~70% hiring pages | 25 × $0.003 + ~18 × $0.002 | ~$0.11 |
| 25 sites, every site has a jobs page | 25 × $0.003 + 25 × $0.002 | $0.13 |
| 1,000 sites, homepage only | 1,000 × $0.003 | $3.00 |
| 1,000 sites, careers On, ~70% hiring | 1,000 × $0.003 + ~700 × $0.002 | ~$4.40 |
| 1,000 sites, every site hiring | 1,000 × $0.005 | $5.00 |
| 50,000 sites, homepage only | 50,000 × $0.003 | $150 |
| 50,000 sites, ~70% hiring | 50,000 × $0.003 + ~35,000 × $0.002 | ~$220 |
Dataset rows are still written when the homepage fails (with error) — those rows are not charged.
Is scraping public website signals free?
A 25-site try is about $0.08–$0.13 (pay-per-event). It is not an unlimited free export. Compute is HTTP-only and small next to the events above.
Is it legal to scrape public website signals?
This Actor only reads public HTML, CSS, DNS MX, and /llms.txt that the company already publishes. It does not log in, submit newsletter forms, crawl inboxes, or collect personal email addresses. As with any dataset of company identifiers, ensure your use complies with GDPR and applicable regulations.
Input
| Parameter | Type | Default | Description |
|---|---|---|---|
urls | string[] | — | Company domains or URLs (required) |
includeCareers | boolean | true | Scan jobs pages. Extra $0.002 only when a hiring page / ATS is found |
includeColors | boolean | true | Extract 3–5 brand hex colors |
includeEmailStack | boolean | true | Detect Google Workspace vs Microsoft 365 from MX |
API-only: maxUrls (integer, default 0 = no cap after dedupe), verboseLogs (boolean) — technical skip reasons in the run log. proxyConfiguration — omitted from the Console form; cloud runs use Apify datacenter proxy by default. Power users can pass it in JSON/API input.
{"urls": ["enky.com", "payfit.com", "alan.com"],"includeCareers": true,"includeColors": true,"includeEmailStack": true}
Output example
{"domain": "enky.com","cms": "shopify","ecommerce": "shopify","shopify_without_klaviyo": false,"shopify_without_gorgias": true,"shopify_without_recharge": true,"shopify_without_reviews": true,"shopify_without_loyalty": true,"gorgias": false,"recharge": false,"consent_manager": null,"locales": ["en"],"has_llms_txt": false,"llms_txt_url": null,"demo_url": null,"meeting_tool": null,"contact_url": "https://enky.com/pages/contact","generic_emails": ["hello@enky.com"],"hubspot": false,"gtm": true,"klaviyo": true,"has_newsletter": true,"newsletter_provider": "klaviyo","newsletter_form_url": "https://enky.com/","colors": ["#f0977a", "#0e7ea5", "#6a8691"],"page_title": "Enky","email_platform": "google_workspace","careers_url": "https://www.welcometothejungle.com/fr/companies/enky/jobs","careers_ats": "welcometothejungle","hiring_now": true,"careers_stack_mentions": [],"fetched": true,"error": null}
Klaviyo on the homepage is a pixel/form signal — it can vary by theme and cookie banner.
How it works
- Normalize pasted URLs to one origin per company (www deduped).
- Fetch the homepage (HTTP + HTML). Detect CMS, pixels, Shopify apps (CDN hosts only), CMP, locales, newsletter, demo/meeting links, contact page and role mailboxes, on-page facts, colors from CSS. GET
/llms.txtat the origin — keep the URL only if it looks like markdown, not a soft-404 HTML page. - Optional MX — Google Workspace vs Microsoft 365.
- Optional careers — follow homepage job links in any language, then locale-aware fallbacks (
/karriere,/empleo,/vagas…). Read ATS hosts and tool names only on those pages.
Local development
cd website-signals-scrapernpm installnpm testapify run
The CLI validates storage/key_value_stores/default/INPUT.json against the input schema. Use .actor/INPUT.json as the Console prefill, or:
apify run --input-file=./.actor/INPUT.jsonnode scripts/local-probe.mjs
When running locally, keys in the simulated KV input override a root input.json if you add one.
Limitations
- Homepage HTML plus a few stylesheets and careers URLs. No Lighthouse, no 50-page crawl, no Playwright in v1.
- A Cursor mention on a jobs page is not proof they use it every day.
- Heavy JS-only sites may look
customwith empty pixels. - We never subscribe to newsletters.
- We never collect personal inboxes (
jean.dupont@). Role mailboxes on the homepage only. - A Calendly widget with no public event
hrefleavesdemo_urlempty.
Also available
Welcome to the Jungle · Hiring Companies & Profiles — enrich a WTTJ company profile or discover companies hiring by sector.
Tech Stack · M365 & Google Workspace — bulk DNS check for Microsoft 365 vs Google Workspace (thousands of domains, no HTML).
Support
Contact corentin@outreacher.fr if you need a custom scraper or tailored automation.
FAQ
Do I need Shopify / HubSpot admin? No. Public homepage HTML only.
Will you subscribe to their newsletter? No. We only record that a form exists and which ESP it uses.
Why is Intercom false on Intercom’s own site? Many widgets inject after JavaScript. v1 reads the first HTML response, not a full browser session.
Do you scrape employee emails? No. Only role mailboxes already in homepage mailto: links (hello@, contact@, sales@).
Why is demo_url empty if they use Calendly? We keep a URL only when the HTML already has a bookable link (calendly.com/team/…, Cal.com, HubSpot Meetings) or an on-site /demo path. The JS widget alone is not enough.
What is llms.txt? A public markdown file at https://example.com/llms.txt (llmstxt.org) so AI tools can read a site summary. We record whether it exists and the URL — not the file body.
Do HubSpot / Klaviyo / Shopify / Gorgias flags cost extra? No. They are part of the $0.003 homepage event.
Do I pay if careers is On but there is no jobs page? No second event. Only the homepage charge if the site loaded.
Can I pay homepage-only? Yes. Turn Scan careers Off. Still $0.003 per loaded site.