B2B Tech Stack Enricher: Filter Your Lead List by Technology avatar

B2B Tech Stack Enricher: Filter Your Lead List by Technology

Pricing

from $3.00 / 1,000 website analyzeds

Go to Apify Store
B2B Tech Stack Enricher: Filter Your Lead List by Technology

B2B Tech Stack Enricher: Filter Your Lead List by Technology

Give it company websites or a lead dataset. It returns the technologies visibly in use (CMS, e-commerce, analytics, CRM, payments, email provider, CDN) with evidence and confidence, keeps your original columns, and can filter 'uses Shopify but not Klaviyo'.

Pricing

from $3.00 / 1,000 website analyzeds

Rating

0.0

(0)

Developer

MST MORIUM AKTHER MAYA

MST MORIUM AKTHER MAYA

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Give it a list of company websites, or an existing lead dataset. For each site it reports which platforms and tools are visibly in use (CMS, e-commerce platform, analytics, CRM and marketing tools, chat, payments, consent, CDN/hosting and the email provider) together with the evidence behind each finding, a confidence level, an honest fetch status, and your original columns untouched. You can also filter: "uses Shopify or WooCommerce, has a CRM, but no live chat", and every row gets a matched flag.

It reads each site's public homepage over plain HTTP (optionally a few same-site pages). No login, no browser, no proxies.

Who it is for

  • Sales, agency and partnership teams that want to qualify or segment a lead list by the technology a company visibly runs ("has a CRM", "runs Shopify or WooCommerce", "no live chat yet").
  • Data and automation builders who need a flat, predictable row per website to feed a spreadsheet, CRM, n8n/Make flow or an AI agent.

What it is, and what it is not. It is a focused, evidence-first enrichment step for B2B lists: every finding shows what it was based on, and a site it could not read is reported as unknown, never as "uses nothing". It is not a browser-based crawler, not a database of thousands of fingerprints, and it does not promise to find every technology on every site.

What you get

  • Works on your list. Pass websites, or pick a dataset (for example the output of a lead scraper). Every row comes back with all original columns first and the analysis next to them. Duplicate domains are analysed once; the repeats are free.
  • Filter layer. Keep or flag rows by technology (requireTechnologies = all of, requireAnyTechnologies = at least one of, excludeTechnologies = none of) and by category (requireCategories, excludeCategories). Each row gets matched: true / false / null. null means "cannot tell" (for example the site blocked us), so a blocked site is never reported as "does not use X".
  • Evidence and confidence on every detection. Up to three short pieces of evidence per technology (for example script src: https://js.stripe.com/v3/) and a high / medium / low confidence. Versions are reported only when the site states them explicitly.
  • Honest about failures. fetchStatus says what happened (ok, blocked, robots_disallowed, dns_failure, ...). A site we could not read is never shown as "no technology found". You are charged only for websites that were fetched and analysed.
  • Email and DNS signals. Mail provider (Google Workspace, Microsoft 365, Zoho, ...), mail-security gateway, sending services seen in SPF, DMARC policy, name servers. These also work when the website itself blocks us.
  • Flat sales-ready columns. cms, ecommercePlatform, crm, marketingAutomation, emailMarketing, supportChat, emailProvider, plus lists analyticsTools, advertisingPixels, paymentProviders, salesIntelligence, so you can sort and segment in a spreadsheet without unpacking JSON.
  • A curated catalog, not a giant one. 234 technologies in 37 categories that matter for prospecting, each with a fingerprint we can explain. A technology that is not in the catalog is simply not reported.

Quick start

  1. Open Input and paste websites, one per line. example.com, www.example.com and https://example.com/about are the same site.
  2. (Optional) Fill a filter, for example Must use at least one of: Shopify, WooCommerce.
  3. Click Start. Open the Overview view for one line per website, or Technologies with evidence for details.

Tip: set Maximum rows to process to 5 for a first test run.

Example input

{
"websites": ["allbirds.com", "https://www.example-agency.com/about"],
"requireAnyTechnologies": ["Shopify", "WooCommerce"],
"excludeCategories": ["Live chat / support"],
"minConfidence": "medium"
}

Bulk and dataset use

Pick a dataset under Dataset with websites. If its website column is not called website, url or domain, type the column name under Column that holds the website. Datasets are read page by page, so large inputs do not have to fit in memory. Every original column is kept; if one of our field names would collide with yours (you already have a domain column), yours is kept and ours is written as enrichment_domain.

  • Large lists: run with 1 GB memory or more. The Actor processes many sites in parallel, keeps memory flat, and pushes results continuously, so a long run can be stopped at any time and everything finished so far is already in the dataset.
  • Set a maximum cost on the run. The Actor never starts more work than that budget allows, and rows it could not start are returned as free skipped_budget_limit rows, so nothing silently disappears. skipped_time_limit works the same way if the run timeout is reached.
  • Re-runs: if a run is resumed after a restart, rows already written are not processed or charged again.

Output

Example output (abridged, from a test page)

{
"website": "shop.example",
"inputIndex": 0,
"domain": "shop.example",
"fetchStatus": "ok",
"detection": "full",
"technologyCount": 3,
"technologies": [
{"name": "Shopify", "category": "Ecommerce", "confidence": "high", "version": null,
"evidence": ["script src: https://cdn.shopify.com/s/files/1/0001/theme.js", "cookie name: _shopify_y"]},
{"name": "Stripe", "category": "Payments", "confidence": "high", "version": null,
"evidence": ["script src: https://js.stripe.com/v3/"]},
{"name": "nginx", "category": "Web server", "confidence": "high", "version": "1.25.3",
"evidence": ["server: nginx/1.25.3"]}
],
"ecommercePlatform": "Shopify",
"crm": null,
"paymentProviders": ["Stripe"],
"emailProvider": "Google Workspace",
"companyName": "Acme Shop Inc",
"matched": true,
"duplicateOf": null
}

Output fields

FieldMeaning
your original columnsfirst, unchanged
inputIndexposition of the row in your input (0-based). Rows of one batch can be stored a few positions out of order; sort by inputIndex to restore the exact order
domain, finalUrl, redirectedthe site analysed, where it ended up, whether a redirect was followed (notes flags a redirect to a different domain)
fetchStatus, fetchReason, httpStatusoutcome of reading the page (see Fetch status)
detectionfull = page read and analysed · headers_only = a response came back but the page was unusable (blocked, error, not HTML), only headers/cookies were examined · none = nothing was examined on the website
technologies[]name, category, confidence, version (or null), evidence[]
technologyCountnumber of entries in technologies. 0 with detection: full means "page read, nothing in the catalog matched". 0 with any other detection means "not analysed", not "uses nothing"
cms, ecommercePlatform, crm, marketingAutomation, emailMarketing, supportChat, emailProviderconvenience columns: the top technology of that category at medium confidence or better, else null
analyticsTools[], advertisingPixels[], paymentProviders[], salesIntelligence[]names of detected technologies of those categories (medium confidence or better)
companyName, pageTitle, pageDescription, language, canonicalUrl, socialProfilesobserved on the page only. Company social-profile links only; no e-mail addresses, phone numbers or personal profiles are collected
dns, dnsStatusapex domain, MX hosts, name servers, SPF record, DMARC policy, whether the domain accepts mail
matched, matchReasonresult of your filter: true, false or null (cannot tell) and why
duplicateOfinputIndex of the first row with the same site; the result is copied and not charged again
blockedBywhen blocked, the protection product we recognised (for example cloudflare)
pagesAnalyzed, truncated, noteswhich pages were read, whether the page was cut at 1.5 MB, and short flags such as fell_back_to_www_variant
siteKey, checkedAtinternal de-duplication key and UTC time

A run summary (counts by status, stop reason, billing, number of HTTP requests and megabytes received) is saved as the SUMMARY record in the key-value store.

Fetch status

fetchStatusMeaningCharged
okpage fetched and analysedyes
invalid_inputnot a usable web address (empty, IP address, credentials, unsupported port, ...)no
unsafe_targetresolves to a private/internal addressno
dns_failure, connection_failed, tls_error, timeout, too_many_redirectssite unreachableno
robots_disallowedthe site's robots.txt forbids our bot (or robots.txt returned a server error)no
blockedthe site answered 403/429 or showed a bot-check pageno
http_error, server_error, non_html4xx/5xx or not an HTML pageno
skipped_time_limit, skipped_budget_limitnot started because the run ran out of time or reached your maximum cost; re-run these rowsno
analysis_errorunexpected internal error for that rowno

Evidence and confidence

  • high: at least one essentially unique fingerprint (vendor-only script host, cookie, header, generator tag, platform CNAME) or two medium ones of different types.
  • medium: one specific but shareable signal (for example a script path that other sites can copy, or an SPF include:).
  • low: a single weak hint. Low detections are listed, but only count in the filter if you set minConfidence: low.

Only structural evidence counts: script/link/iframe/image/form URLs, meta tags, cookies, response headers, inline script code, DNS records. Visible text and ordinary links never count, so a blog post that mentions Shopify does not make a site a Shopify site. For path-style fingerprints (for example /wp-content/), a file hot-linked from another site counts one level lower than the same file served by the site itself.

The levels are rule-based, not statistical probabilities: "high" does not mean "99% likely". An SPF include: shows a service is authorised to send mail for the domain; it can be stale, so it is medium. An MX record (where mail is delivered now) is strong.

Supported technologies and categories

  • CMS: Adobe Experience Manager, Craft CMS, Drupal, Framer, Ghost, HubSpot CMS, Joomla, Sitecore, Squarespace, TYPO3, Webflow, Weebly, Wix, WordPress
  • Ecommerce: BigCommerce, Ecwid, Magento, OpenCart, PrestaShop, Salesforce Commerce Cloud, Shopify, WooCommerce
  • Ecommerce app: Judge.me, Loox, Recharge, Yotpo
  • Page builder: Elementor
  • Frontend framework: Angular, AngularJS, Astro, Gatsby, Next.js, Nuxt, React, Remix, SvelteKit, VitePress, Vue.js
  • Web framework: ASP.NET, Django, Express, Laravel, Ruby on Rails
  • Programming language: Node.js, PHP, Python
  • JavaScript library: jQuery
  • UI framework: Bootstrap, Tailwind CSS
  • Analytics: Adobe Analytics, Amplitude, Crazy Egg, FullStory, Google Analytics, Heap, Hotjar, LogRocket, Lucky Orange, Matomo, Microsoft Clarity, Mixpanel, Mouseflow, Pendo, Plausible, PostHog, Segment
  • Tag management: Adobe Experience Platform Launch, Google Tag Manager, Tealium
  • A/B testing: AB Tasty, Optimizely, VWO
  • Advertising: Criteo, Google Ads, Google AdSense, LinkedIn Insight Tag, Meta Pixel, Microsoft Advertising (UET), Pinterest Tag, Reddit Pixel, Snap Pixel, Taboola, TikTok Pixel, X Pixel
  • Marketing automation: ActiveCampaign, Attentive, HubSpot, Marketo, Oracle Eloqua, Pardot, Postscript
  • Email marketing: Braze, Brevo, Constant Contact, Customer.io, Drip, Kit (ConvertKit), Klaviyo, Mailchimp, Omnisend
  • CRM: Pipedrive, Salesforce, Zoho CRM
  • Sales intelligence: 6sense, Clearbit, Demandbase, Leadfeeder, ZoomInfo
  • Form builder: Contact Form 7, Gravity Forms, Jotform, Typeform, WPForms
  • Scheduling: Acuity Scheduling, Cal.com, Calendly, Chili Piper, HubSpot Meetings
  • Live chat / support: Crisp, Drift, Freshchat, Freshdesk, Gorgias, Help Scout, Intercom, LiveChat, Olark, Podium, Qualified, Tawk.to, Tidio, Zendesk, Zoho SalesIQ
  • Payments: 2Checkout / Verifone, Adyen, Afterpay, Authorize.Net, Braintree, Klarna, Mollie, Paddle, PayPal, Razorpay, Shop Pay / Shopify Payments, Square, Stripe
  • Consent management: Complianz, Cookiebot, CookieYes, Didomi, iubenda, OneTrust, Osano, Quantcast Choice, Termly, Usercentrics
  • CAPTCHA: Cloudflare Turnstile, Google reCAPTCHA, hCaptcha
  • Security / WAF: DataDome, HUMAN (PerimeterX), Imperva, Sucuri
  • CDN: Akamai, Amazon CloudFront, Bunny CDN, Cloudflare, Fastly, KeyCDN, Varnish
  • Hosting: Amazon S3, AWS Elastic Load Balancing, Kinsta, Pantheon, WP Engine
  • PaaS: Cloudflare Pages, Fly.io, GitHub Pages, Heroku, Netlify, Render, Vercel
  • Cloud provider: Amazon Web Services, Google Cloud, Microsoft Azure
  • Web server: Apache HTTP Server, Caddy, LiteSpeed, Microsoft IIS, nginx, OpenResty
  • DNS provider: Akamai Edge DNS, Amazon Route 53, Azure DNS, Cloudflare DNS, DNSimple, GoDaddy DNS, Google Cloud DNS, Namecheap DNS, NS1, Oracle Dyn DNS, UltraDNS
  • Email provider: Fastmail, GoDaddy Email, Google Workspace, Hostinger Email, IONOS Mail, Microsoft 365, Namecheap Private Email, Proton Mail, Rackspace Email, Titan Email, Zoho Mail
  • Email security gateway: Barracuda Email Security, Cisco Secure Email, Mimecast, Proofpoint
  • Transactional email: Amazon SES, Mailgun, Mailjet, Mandrill, Postmark, SendGrid, SparkPost
  • Font service: Adobe Fonts, Font Awesome, Google Fonts
  • Maps: Google Maps, Mapbox
  • Video: Vidyard, Vimeo embed, Wistia, YouTube embed
  • SEO: Yoast SEO

Technology names and logos belong to their respective owners. This Actor is independent and not affiliated with or endorsed by any of them. The catalog is written from scratch from each technology's own observable behaviour; it is not copied from Wappalyzer or similar rule sets.

Options

OptionDefaultNotes
Websites / Dataset-use either or both (websites first, then the dataset)
Must use (all of) / Must use at least one of / Must NOT use-technology names or aliases, case-insensitive. A typo fails the run immediately with a suggestion, before anything is charged
Must use a technology in these categories / Must NOT use any in these categories-category names from the list above, for example CRM, Live chat / support, Payments
Minimum confidence for the filtermediumsee Evidence and confidence
What to analyse per inputdomaindomain: one analysis per site homepage. exact: analyse exactly the URL given
Include DNS signalsonadds email provider, SPF/DMARC, hosting hints
Pages per website12-3 also reads same-site shop, pricing, contact or demo pages. In our own test on 45 real sites, reading 3 pages instead of 1 found one extra technology on 2 of them and took about 3 seconds longer per site, so the default is 1. Charged per website, not per page
Maximum rows to processallcheap way to test, e.g. 5

Pricing

Pay per event: one website-analyzed event for each website that was successfully fetched and analysed. Failures, blocked sites, robots-disallowed sites, duplicates and skipped rows are free, so a dead or protected site in your list never costs you anything. Reading more pages of the same website does not cost more. The current price is shown at the top of this page and on the Pricing tab. Set a maximum cost on the run to cap spend.

Limits (please read)

  • HTTP only. The Actor reads the HTML the server sends. Technology that only appears after JavaScript runs in a browser (some single-page apps, tags that a tag manager injects later) may be invisible. A tool that a site loads only through Google Tag Manager is typically not visible here; the tag manager itself is. Scripts that a cookie banner holds back until consent are read when they are present in the HTML. A site showing no technologies has not proven to use none.
  • Not every site can be read. Large, well-known sites block automated requests more often than small business sites. In our own test on the 1,000 most popular domains (a demanding list, not a typical lead list), 54% of rows were readable and therefore chargeable; the rest were unreachable, blocked, disallowed by robots.txt or not web pages. Expect different numbers on your own list, and use fetchStatus to see exactly what happened to each row.
  • Bot protection. The Actor identifies itself honestly (B2BTechStackEnricherBot) and does not try to bypass CAPTCHAs, challenges or blocks. Protected sites come back as blocked.
  • Homepage by default. Some tools only load on checkout or login pages.
  • Certificate problems. Sites whose certificate is expired, self-signed, for another name or served with an incomplete chain end as tls_error; fetchReason says which (for example unknown_or_incomplete_certificate_chain). Certificates are never ignored.
  • Versions are reported only when the site states them; most sites do not.
  • Accuracy. We do not publish an accuracy percentage. There is no public labelled benchmark that matches this catalog, and a number without one would not be honest. Every detection carries its evidence so you can judge it.
  • Apex-domain rule for DNS. The domain's zone is found by walking up the name until a name server answers. A sub-domain delegated to its own DNS zone is treated as its own domain.

Responsible use

  • robots.txt is always respected, for our bot name, again for every redirect and for every extra page; this cannot be switched off. A Crawl-delay is honoured between extra pages of the same site.
  • A handful of requests per site at most: one homepage fetch (plus a robots.txt fetch and a www/http retry only when the first attempt fails to connect), and up to two extra pages if you ask for them.
  • Requests to private, loopback, link-local and other internal addresses are refused, including when a public name or a redirect points at one.
  • Do not use the results to violate a site's terms, to harass anyone, or to process personal data without a legal basis. The Actor does not collect personal data.

API and code examples

Replace YOUR_USERNAME with the account that publishes the Actor and YOUR_API_TOKEN with your Apify API token (keep it secret).

cURL

curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~b2b-tech-stack-enricher/run-sync-get-dataset-items?token=YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{"websites": ["allbirds.com", "example.com"], "requireAnyTechnologies": ["Shopify", "WooCommerce"]}'

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("YOUR_USERNAME/b2b-tech-stack-enricher").call(run_input={
"datasetId": "YOUR_LEADS_DATASET_ID",
"urlField": "website",
"requireCategories": ["CRM"],
"excludeTechnologies": ["HubSpot"],
})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row["matched"] is True:
print(row["domain"], row["crm"], row["emailProvider"])

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('YOUR_USERNAME/b2b-tech-stack-enricher').call({
websites: ['allbirds.com', 'example.com'],
requireAnyTechnologies: ['Shopify'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.filter((r) => r.matched === true).map((r) => r.domain));

n8n

Use the Apify node (or an HTTP Request node) to run the Actor with the JSON input above, wait for the run to finish, then read the dataset items. A typical flow: Google Sheets (list of companies) -> Apify (Run Actor) -> Apify (Get dataset items) -> IF matched is true -> CRM or email tool. Treat rows with matched: null as "re-check later", not as "no".

Using it from an AI agent

The input is plain JSON and the output is one flat row per website with explicit statuses, which suits tool-calling agents. Two rules for agents: pass maxWebsites to keep a run small, and read matched/fetchStatus literally (null and non-ok statuses mean "unknown", never "absent"). Whether this Actor can be called through the Apify MCP server depends on your account's Apify settings; check the Actor's page in Apify Console.

FAQ

Why does a site I know uses X show nothing? Either the technology is not in the catalog, it is only added by JavaScript in the browser, or the site blocked us. Check detection and fetchStatus first; technologyCount: 0 only means "nothing found" when detection is full.

What does matched: null mean? We cannot tell. The site was blocked, unreachable, not fully analysed, or the evidence was below your minimum confidence.

Am I charged for blocked or dead sites? No. Only websites that returned a readable page and were analysed.

Are duplicates charged? No. A domain that appears several times is analysed once; repeats are free and carry duplicateOf.

Does it use a browser or proxies? No. Plain HTTP requests from Apify's network, so some protected sites will block it.

Can I add my own technology? Not at the moment. Tell us which technology is missing; the catalog is extended from each vendor's own documented install snippet.

Are versions exact? Only when the site states one (a generator meta tag, a Server header). Otherwise version is null.

Changelog

  • 1.0.1: scripts held back by cookie-banner tools (for example OneTrust or Cookiebot) and app tags listed in JSON loaders (for example on Shopify stores) are now read; a vendor's own CDN hint no longer makes a site look like it runs that vendor's platform; VitePress added.
  • 1.0.0: first release.