Website stack · Apps, ATS, colors & llms.txt by URL avatar

Website stack · Apps, ATS, colors & llms.txt by URL

Pricing

from $3.00 / 1,000 homepage stacks

Go to Apify Store
Website stack · Apps, ATS, colors & llms.txt by URL

Website stack · Apps, ATS, colors & llms.txt by URL

Paste company URLs — get CMS, HubSpot, Klaviyo, Shopify apps, CMP, demo link, role emails, llms.txt, brand colors, hiring ATS and careers tool mentions. One row per site, CSV-ready. No login. No API key.

Pricing

from $3.00 / 1,000 homepage stacks

Rating

0.0

(0)

Developer

Corentin Robert

Corentin Robert

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

2

Monthly active users

4 days ago

Last modified

Share

Paste company URLs. Get CMS, marketing pixels (HubSpot, GTM, Klaviyo), Shopify apps (Gorgias, Recharge…), CMP, locales, llms.txt, brand colors, on-page facts, newsletter presence, and careers ATS + tool mentions — one row per site, CSV-ready.

No login. No API key. No browser.

Who is this for?

You are…Typical goalSuggested setup
Shopify / HubSpot agencyFind Shopify without Klaviyo, Gorgias, Recharge, reviews or loyaltyDefault run, filter shopify_without_*
Outbound SDR / BDRPersonalize by stack, demo link, generic inbox, colors, hiringColors ON, careers ON
Email / ESP vendorSee who already uses Klaviyo, Mailchimp, Brevo, SubstackDefault run — we never subscribe
AI tooling vendorSee who publishes llms.txt, or if jobs pages mention Cursor / CopilotDefault run + careers ON
RevOpsEnrich a CRM domain list before sequencingPaste domains
Designer / mockupPull 3–5 brand hex codes for a one-pagerColors ON, careers off

What you get by default: CMS and shop platform, HubSpot/GTM/Klaviyo flags, Shopify apps and CMP from script/CDN hosts (not competitor blog links), locales, public /llms.txt if it is a real markdown file, Calendly/Cal.com event URLs or an on-site /demo link when present in the HTML, a contact page plus role mailboxes only (hello@, contact@ — never jean.dupont@), newsletter form URL (no signup), 3–5 hex colors, title/H1/schema, Google vs Microsoft 365 from MX, careers URL, ATS (Ashby, Lever, Teamtailor…), tool names only from jobs pages.

When to turn options off: skip careers if you only need homepage stack — you then pay the homepage event only (no hiring-page event). Skip colors if you only need pixels; skip email MX if you already ran a DNS stack check. Colors and MX stay inside the homepage event.

Quick start

  1. Open the Actor in Apify Console.
  2. Paste domains or URLs, one per line (enky.com, payfit.com).
  3. Leave the three include toggles on for the full row.
  4. Click Start — rows appear in the Dataset tab as each site finishes.
  5. Export CSV / JSON / Excel.

Ready-made examples

Saved inputs live in published-tasks/inputs/ (publish from Console for Store landing pages):

ExampleBest for
Shopify stores without KlaviyoAgency outbound
Companies already on HubSpotHubSpot ISV / agency
Hiring pages and careers ATSTalent / AI tooling
Brand colors for outreach mockupsSDR personalization

What it extracts

CategoryFields
Identitydomain, final_url, page_title, h1, lang, locales
CMS / shopcms, ecommerce, shopify_without_klaviyo, shopify_without_gorgias, shopify_without_recharge, shopify_without_reviews, shopify_without_loyalty
Marketing stackhubspot, gtm, ga4, klaviyo, intercom, crisp, calendly, stripe, meta_pixel, hotjar, linkedin_insight, typeform
Shopify appsgorgias, judgeme, loox, yotpo, smile, recharge, skio, loyaltylion
Consentconsent_manager (axeptio, didomi, cookiebot, onetrust, cookieyes, tarteaucitron)
AI crawl filehas_llms_txt, llms_txt_url (origin /llms.txt only; body not stored)
Outreachdemo_url, meeting_tool, contact_url, generic_emails (homepage HTML only; no personal inboxes)
Newsletterhas_newsletter, newsletter_provider, newsletter_form_url (never submitted)
Brandcolors, theme_color
On-pagemeta_description, schema_types, has_blog, has_pricing
Emailemail_platform, mx_host
Careerscareers_url, hiring_now, careers_ats, careers_stack_mentions

cms values include shopify, wordpress, webflow, wix, squarespace, framer, ghost, drupal, joomla, hubspot_cms, bigcommerce, magento, prestashop, bubble, or custom. WordPress shops are cms: wordpress + ecommerce: woocommerce.

Typical fill rates

From a 30-site mixed probe (SaaS + ecommerce + CMS vendors, HTTP homepage):

SignalIndicative rate
Homepage loaded~100% on public sites
GTM~50%
Hiring page found~70% when careers ON
Newsletter form / ESP~10–15%
Intercom / Calendly widget on the homepage HTMLoften 0% (loads in JS)
Calendly / Cal.com event URL in an href / iframewhen they published a bookable link

Treat pixels as “present in homepage HTML”, not “the company never uses this tool”.

How much does it cost to scrape website stack signals?

Pay-per-event. You pay for what landed in the row, not for empty 404s. HTTP-only — compute stays low. There is no actor-start fee.

EventPriceWhen it firesWhat you get in that charge
Homepage stack (website-homepage)$0.003Homepage HTML loadedCMS, shop flag, pixels, Shopify apps, CMP, locales, /llms.txt check, demo / contact / role emails, newsletter URL (never submitted), on-page title/H1/schema, brand colors, Google vs Microsoft 365 MX
Hiring page found (website-careers)$0.002A public jobs page or ATS was found (hiring_now or careers_ats)Careers URL, ATS (Ashby, Lever, Teamtailor, Welcome to the Jungle…), tool names on that page

Not billed

  • Homepage that failed to load (fetched: false)
  • Careers scan that only hit 404s / empty shells (no jobs page, no ATS)
  • Extra CSS files fetched for colors
  • Extra GET for /llms.txt (404s and HTML soft-404s stay has_llms_txt: false)
  • includeCareers On with no hiring page — homepage event only
  • Turning colors / MX Off does not change the homepage price (same event; fewer fields)

Typical mix (same 30-site probe as fill rates): ~70% of sites with careers On produce a hiring page → most full runs are $0.003 + $0.002 = $0.005 per site, not $0.005 on every URL.

ScenarioEventsApprox. cost
25 sites, homepage only (careers Off)25 × $0.003$0.08
25 sites, careers On, ~70% hiring pages25 × $0.003 + ~18 × $0.002~$0.11
25 sites, every site has a jobs page25 × $0.003 + 25 × $0.002$0.13
1,000 sites, homepage only1,000 × $0.003$3.00
1,000 sites, careers On, ~70% hiring1,000 × $0.003 + ~700 × $0.002~$4.40
1,000 sites, every site hiring1,000 × $0.005$5.00
50,000 sites, homepage only50,000 × $0.003$150
50,000 sites, ~70% hiring50,000 × $0.003 + ~35,000 × $0.002~$220

Dataset rows are still written when the homepage fails (with error) — those rows are not charged.

Is scraping public website signals free?

A 25-site try is about $0.08–$0.13 (pay-per-event). It is not an unlimited free export. Compute is HTTP-only and small next to the events above.

This Actor only reads public HTML, CSS, DNS MX, and /llms.txt that the company already publishes. It does not log in, submit newsletter forms, crawl inboxes, or collect personal email addresses. As with any dataset of company identifiers, ensure your use complies with GDPR and applicable regulations.

Input

ParameterTypeDefaultDescription
urlsstring[]Company domains or URLs (required)
includeCareersbooleantrueScan jobs pages. Extra $0.002 only when a hiring page / ATS is found
includeColorsbooleantrueExtract 3–5 brand hex colors
includeEmailStackbooleantrueDetect Google Workspace vs Microsoft 365 from MX

API-only: maxUrls (integer, default 0 = no cap after dedupe), verboseLogs (boolean) — technical skip reasons in the run log. proxyConfiguration — omitted from the Console form; cloud runs use Apify datacenter proxy by default. Power users can pass it in JSON/API input.

{
"urls": ["enky.com", "payfit.com", "alan.com"],
"includeCareers": true,
"includeColors": true,
"includeEmailStack": true
}

Output example

{
"domain": "enky.com",
"cms": "shopify",
"ecommerce": "shopify",
"shopify_without_klaviyo": false,
"shopify_without_gorgias": true,
"shopify_without_recharge": true,
"shopify_without_reviews": true,
"shopify_without_loyalty": true,
"gorgias": false,
"recharge": false,
"consent_manager": null,
"locales": ["en"],
"has_llms_txt": false,
"llms_txt_url": null,
"demo_url": null,
"meeting_tool": null,
"contact_url": "https://enky.com/pages/contact",
"generic_emails": ["hello@enky.com"],
"hubspot": false,
"gtm": true,
"klaviyo": true,
"has_newsletter": true,
"newsletter_provider": "klaviyo",
"newsletter_form_url": "https://enky.com/",
"colors": ["#f0977a", "#0e7ea5", "#6a8691"],
"page_title": "Enky",
"email_platform": "google_workspace",
"careers_url": "https://www.welcometothejungle.com/fr/companies/enky/jobs",
"careers_ats": "welcometothejungle",
"hiring_now": true,
"careers_stack_mentions": [],
"fetched": true,
"error": null
}

Klaviyo on the homepage is a pixel/form signal — it can vary by theme and cookie banner.

How it works

  1. Normalize pasted URLs to one origin per company (www deduped).
  2. Fetch the homepage (HTTP + HTML). Detect CMS, pixels, Shopify apps (CDN hosts only), CMP, locales, newsletter, demo/meeting links, contact page and role mailboxes, on-page facts, colors from CSS. GET /llms.txt at the origin — keep the URL only if it looks like markdown, not a soft-404 HTML page.
  3. Optional MX — Google Workspace vs Microsoft 365.
  4. Optional careers — follow homepage job links in any language, then locale-aware fallbacks (/karriere, /empleo, /vagas…). Read ATS hosts and tool names only on those pages.

Local development

cd website-signals-scraper
npm install
npm test
apify run

The CLI validates storage/key_value_stores/default/INPUT.json against the input schema. Use .actor/INPUT.json as the Console prefill, or:

apify run --input-file=./.actor/INPUT.json
node scripts/local-probe.mjs

When running locally, keys in the simulated KV input override a root input.json if you add one.

Limitations

  • Homepage HTML plus a few stylesheets and careers URLs. No Lighthouse, no 50-page crawl, no Playwright in v1.
  • A Cursor mention on a jobs page is not proof they use it every day.
  • Heavy JS-only sites may look custom with empty pixels.
  • We never subscribe to newsletters.
  • We never collect personal inboxes (jean.dupont@). Role mailboxes on the homepage only.
  • A Calendly widget with no public event href leaves demo_url empty.

Also available

Welcome to the Jungle · Hiring Companies & Profiles — enrich a WTTJ company profile or discover companies hiring by sector.

Tech Stack · M365 & Google Workspace — bulk DNS check for Microsoft 365 vs Google Workspace (thousands of domains, no HTML).

Support

Contact corentin@outreacher.fr if you need a custom scraper or tailored automation.

FAQ

Do I need Shopify / HubSpot admin? No. Public homepage HTML only.

Will you subscribe to their newsletter? No. We only record that a form exists and which ESP it uses.

Why is Intercom false on Intercom’s own site? Many widgets inject after JavaScript. v1 reads the first HTML response, not a full browser session.

Do you scrape employee emails? No. Only role mailboxes already in homepage mailto: links (hello@, contact@, sales@).

Why is demo_url empty if they use Calendly? We keep a URL only when the HTML already has a bookable link (calendly.com/team/…, Cal.com, HubSpot Meetings) or an on-site /demo path. The JS widget alone is not enough.

What is llms.txt? A public markdown file at https://example.com/llms.txt (llmstxt.org) so AI tools can read a site summary. We record whether it exists and the URL — not the file body.

Do HubSpot / Klaviyo / Shopify / Gorgias flags cost extra? No. They are part of the $0.003 homepage event.

Do I pay if careers is On but there is no jobs page? No second event. Only the homepage charge if the site loaded.

Can I pay homepage-only? Yes. Turn Scan careers Off. Still $0.003 per loaded site.