B2B Tech Stack Enricher: Filter Your Lead List by Technology
Pricing
from $3.00 / 1,000 website analyzeds
B2B Tech Stack Enricher: Filter Your Lead List by Technology
Give it company websites or a lead dataset. It returns the technologies visibly in use (CMS, e-commerce, analytics, CRM, payments, email provider, CDN) with evidence and confidence, keeps your original columns, and can filter 'uses Shopify but not Klaviyo'.
Pricing
from $3.00 / 1,000 website analyzeds
Rating
0.0
(0)
Developer
MST MORIUM AKTHER MAYA
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Give it a list of company websites, or an existing lead dataset. For each site it reports which platforms and tools are visibly in use (CMS, e-commerce platform, analytics, CRM and marketing tools, chat, payments, consent, CDN/hosting and the email provider) together with the evidence behind each finding, a confidence level, an honest fetch status, and your original columns untouched. You can also filter: "uses Shopify or WooCommerce, has a CRM, but no live chat", and every row gets a matched flag.
It reads each site's public homepage over plain HTTP (optionally a few same-site pages). No login, no browser, no proxies.
Who it is for
- Sales, agency and partnership teams that want to qualify or segment a lead list by the technology a company visibly runs ("has a CRM", "runs Shopify or WooCommerce", "no live chat yet").
- Data and automation builders who need a flat, predictable row per website to feed a spreadsheet, CRM, n8n/Make flow or an AI agent.
What it is, and what it is not. It is a focused, evidence-first enrichment step for B2B lists: every finding shows what it was based on, and a site it could not read is reported as unknown, never as "uses nothing". It is not a browser-based crawler, not a database of thousands of fingerprints, and it does not promise to find every technology on every site.
What you get
- Works on your list. Pass websites, or pick a dataset (for example the output of a lead scraper). Every row comes back with all original columns first and the analysis next to them. Duplicate domains are analysed once; the repeats are free.
- Filter layer. Keep or flag rows by technology (
requireTechnologies= all of,requireAnyTechnologies= at least one of,excludeTechnologies= none of) and by category (requireCategories,excludeCategories). Each row getsmatched: true / false / null.nullmeans "cannot tell" (for example the site blocked us), so a blocked site is never reported as "does not use X". - Evidence and confidence on every detection. Up to three short pieces of evidence per technology (for example
script src: https://js.stripe.com/v3/) and ahigh / medium / lowconfidence. Versions are reported only when the site states them explicitly. - Honest about failures.
fetchStatussays what happened (ok,blocked,robots_disallowed,dns_failure, ...). A site we could not read is never shown as "no technology found". You are charged only for websites that were fetched and analysed. - Email and DNS signals. Mail provider (Google Workspace, Microsoft 365, Zoho, ...), mail-security gateway, sending services seen in SPF, DMARC policy, name servers. These also work when the website itself blocks us.
- Flat sales-ready columns.
cms,ecommercePlatform,crm,marketingAutomation,emailMarketing,supportChat,emailProvider, plus listsanalyticsTools,advertisingPixels,paymentProviders,salesIntelligence, so you can sort and segment in a spreadsheet without unpacking JSON. - A curated catalog, not a giant one. 234 technologies in 37 categories that matter for prospecting, each with a fingerprint we can explain. A technology that is not in the catalog is simply not reported.
Quick start
- Open Input and paste websites, one per line.
example.com,www.example.comandhttps://example.com/aboutare the same site. - (Optional) Fill a filter, for example Must use at least one of: Shopify, WooCommerce.
- Click Start. Open the Overview view for one line per website, or Technologies with evidence for details.
Tip: set Maximum rows to process to 5 for a first test run.
Example input
{"websites": ["allbirds.com", "https://www.example-agency.com/about"],"requireAnyTechnologies": ["Shopify", "WooCommerce"],"excludeCategories": ["Live chat / support"],"minConfidence": "medium"}
Bulk and dataset use
Pick a dataset under Dataset with websites. If its website column is not called website, url or domain, type the column name under Column that holds the website. Datasets are read page by page, so large inputs do not have to fit in memory. Every original column is kept; if one of our field names would collide with yours (you already have a domain column), yours is kept and ours is written as enrichment_domain.
- Large lists: run with 1 GB memory or more. The Actor processes many sites in parallel, keeps memory flat, and pushes results continuously, so a long run can be stopped at any time and everything finished so far is already in the dataset.
- Set a maximum cost on the run. The Actor never starts more work than that budget allows, and rows it could not start are returned as free
skipped_budget_limitrows, so nothing silently disappears.skipped_time_limitworks the same way if the run timeout is reached. - Re-runs: if a run is resumed after a restart, rows already written are not processed or charged again.
Output
Example output (abridged, from a test page)
{"website": "shop.example","inputIndex": 0,"domain": "shop.example","fetchStatus": "ok","detection": "full","technologyCount": 3,"technologies": [{"name": "Shopify", "category": "Ecommerce", "confidence": "high", "version": null,"evidence": ["script src: https://cdn.shopify.com/s/files/1/0001/theme.js", "cookie name: _shopify_y"]},{"name": "Stripe", "category": "Payments", "confidence": "high", "version": null,"evidence": ["script src: https://js.stripe.com/v3/"]},{"name": "nginx", "category": "Web server", "confidence": "high", "version": "1.25.3","evidence": ["server: nginx/1.25.3"]}],"ecommercePlatform": "Shopify","crm": null,"paymentProviders": ["Stripe"],"emailProvider": "Google Workspace","companyName": "Acme Shop Inc","matched": true,"duplicateOf": null}
Output fields
| Field | Meaning |
|---|---|
| your original columns | first, unchanged |
inputIndex | position of the row in your input (0-based). Rows of one batch can be stored a few positions out of order; sort by inputIndex to restore the exact order |
domain, finalUrl, redirected | the site analysed, where it ended up, whether a redirect was followed (notes flags a redirect to a different domain) |
fetchStatus, fetchReason, httpStatus | outcome of reading the page (see Fetch status) |
detection | full = page read and analysed · headers_only = a response came back but the page was unusable (blocked, error, not HTML), only headers/cookies were examined · none = nothing was examined on the website |
technologies[] | name, category, confidence, version (or null), evidence[] |
technologyCount | number of entries in technologies. 0 with detection: full means "page read, nothing in the catalog matched". 0 with any other detection means "not analysed", not "uses nothing" |
cms, ecommercePlatform, crm, marketingAutomation, emailMarketing, supportChat, emailProvider | convenience columns: the top technology of that category at medium confidence or better, else null |
analyticsTools[], advertisingPixels[], paymentProviders[], salesIntelligence[] | names of detected technologies of those categories (medium confidence or better) |
companyName, pageTitle, pageDescription, language, canonicalUrl, socialProfiles | observed on the page only. Company social-profile links only; no e-mail addresses, phone numbers or personal profiles are collected |
dns, dnsStatus | apex domain, MX hosts, name servers, SPF record, DMARC policy, whether the domain accepts mail |
matched, matchReason | result of your filter: true, false or null (cannot tell) and why |
duplicateOf | inputIndex of the first row with the same site; the result is copied and not charged again |
blockedBy | when blocked, the protection product we recognised (for example cloudflare) |
pagesAnalyzed, truncated, notes | which pages were read, whether the page was cut at 1.5 MB, and short flags such as fell_back_to_www_variant |
siteKey, checkedAt | internal de-duplication key and UTC time |
A run summary (counts by status, stop reason, billing, number of HTTP requests and megabytes received) is saved as the SUMMARY record in the key-value store.
Fetch status
fetchStatus | Meaning | Charged |
|---|---|---|
ok | page fetched and analysed | yes |
invalid_input | not a usable web address (empty, IP address, credentials, unsupported port, ...) | no |
unsafe_target | resolves to a private/internal address | no |
dns_failure, connection_failed, tls_error, timeout, too_many_redirects | site unreachable | no |
robots_disallowed | the site's robots.txt forbids our bot (or robots.txt returned a server error) | no |
blocked | the site answered 403/429 or showed a bot-check page | no |
http_error, server_error, non_html | 4xx/5xx or not an HTML page | no |
skipped_time_limit, skipped_budget_limit | not started because the run ran out of time or reached your maximum cost; re-run these rows | no |
analysis_error | unexpected internal error for that row | no |
Evidence and confidence
- high: at least one essentially unique fingerprint (vendor-only script host, cookie, header, generator tag, platform CNAME) or two medium ones of different types.
- medium: one specific but shareable signal (for example a script path that other sites can copy, or an SPF
include:). - low: a single weak hint. Low detections are listed, but only count in the filter if you set
minConfidence: low.
Only structural evidence counts: script/link/iframe/image/form URLs, meta tags, cookies, response headers, inline script code, DNS records. Visible text and ordinary links never count, so a blog post that mentions Shopify does not make a site a Shopify site. For path-style fingerprints (for example /wp-content/), a file hot-linked from another site counts one level lower than the same file served by the site itself.
The levels are rule-based, not statistical probabilities: "high" does not mean "99% likely". An SPF include: shows a service is authorised to send mail for the domain; it can be stale, so it is medium. An MX record (where mail is delivered now) is strong.
Supported technologies and categories
- CMS: Adobe Experience Manager, Craft CMS, Drupal, Framer, Ghost, HubSpot CMS, Joomla, Sitecore, Squarespace, TYPO3, Webflow, Weebly, Wix, WordPress
- Ecommerce: BigCommerce, Ecwid, Magento, OpenCart, PrestaShop, Salesforce Commerce Cloud, Shopify, WooCommerce
- Ecommerce app: Judge.me, Loox, Recharge, Yotpo
- Page builder: Elementor
- Frontend framework: Angular, AngularJS, Astro, Gatsby, Next.js, Nuxt, React, Remix, SvelteKit, VitePress, Vue.js
- Web framework: ASP.NET, Django, Express, Laravel, Ruby on Rails
- Programming language: Node.js, PHP, Python
- JavaScript library: jQuery
- UI framework: Bootstrap, Tailwind CSS
- Analytics: Adobe Analytics, Amplitude, Crazy Egg, FullStory, Google Analytics, Heap, Hotjar, LogRocket, Lucky Orange, Matomo, Microsoft Clarity, Mixpanel, Mouseflow, Pendo, Plausible, PostHog, Segment
- Tag management: Adobe Experience Platform Launch, Google Tag Manager, Tealium
- A/B testing: AB Tasty, Optimizely, VWO
- Advertising: Criteo, Google Ads, Google AdSense, LinkedIn Insight Tag, Meta Pixel, Microsoft Advertising (UET), Pinterest Tag, Reddit Pixel, Snap Pixel, Taboola, TikTok Pixel, X Pixel
- Marketing automation: ActiveCampaign, Attentive, HubSpot, Marketo, Oracle Eloqua, Pardot, Postscript
- Email marketing: Braze, Brevo, Constant Contact, Customer.io, Drip, Kit (ConvertKit), Klaviyo, Mailchimp, Omnisend
- CRM: Pipedrive, Salesforce, Zoho CRM
- Sales intelligence: 6sense, Clearbit, Demandbase, Leadfeeder, ZoomInfo
- Form builder: Contact Form 7, Gravity Forms, Jotform, Typeform, WPForms
- Scheduling: Acuity Scheduling, Cal.com, Calendly, Chili Piper, HubSpot Meetings
- Live chat / support: Crisp, Drift, Freshchat, Freshdesk, Gorgias, Help Scout, Intercom, LiveChat, Olark, Podium, Qualified, Tawk.to, Tidio, Zendesk, Zoho SalesIQ
- Payments: 2Checkout / Verifone, Adyen, Afterpay, Authorize.Net, Braintree, Klarna, Mollie, Paddle, PayPal, Razorpay, Shop Pay / Shopify Payments, Square, Stripe
- Consent management: Complianz, Cookiebot, CookieYes, Didomi, iubenda, OneTrust, Osano, Quantcast Choice, Termly, Usercentrics
- CAPTCHA: Cloudflare Turnstile, Google reCAPTCHA, hCaptcha
- Security / WAF: DataDome, HUMAN (PerimeterX), Imperva, Sucuri
- CDN: Akamai, Amazon CloudFront, Bunny CDN, Cloudflare, Fastly, KeyCDN, Varnish
- Hosting: Amazon S3, AWS Elastic Load Balancing, Kinsta, Pantheon, WP Engine
- PaaS: Cloudflare Pages, Fly.io, GitHub Pages, Heroku, Netlify, Render, Vercel
- Cloud provider: Amazon Web Services, Google Cloud, Microsoft Azure
- Web server: Apache HTTP Server, Caddy, LiteSpeed, Microsoft IIS, nginx, OpenResty
- DNS provider: Akamai Edge DNS, Amazon Route 53, Azure DNS, Cloudflare DNS, DNSimple, GoDaddy DNS, Google Cloud DNS, Namecheap DNS, NS1, Oracle Dyn DNS, UltraDNS
- Email provider: Fastmail, GoDaddy Email, Google Workspace, Hostinger Email, IONOS Mail, Microsoft 365, Namecheap Private Email, Proton Mail, Rackspace Email, Titan Email, Zoho Mail
- Email security gateway: Barracuda Email Security, Cisco Secure Email, Mimecast, Proofpoint
- Transactional email: Amazon SES, Mailgun, Mailjet, Mandrill, Postmark, SendGrid, SparkPost
- Font service: Adobe Fonts, Font Awesome, Google Fonts
- Maps: Google Maps, Mapbox
- Video: Vidyard, Vimeo embed, Wistia, YouTube embed
- SEO: Yoast SEO
Technology names and logos belong to their respective owners. This Actor is independent and not affiliated with or endorsed by any of them. The catalog is written from scratch from each technology's own observable behaviour; it is not copied from Wappalyzer or similar rule sets.
Options
| Option | Default | Notes |
|---|---|---|
| Websites / Dataset | - | use either or both (websites first, then the dataset) |
| Must use (all of) / Must use at least one of / Must NOT use | - | technology names or aliases, case-insensitive. A typo fails the run immediately with a suggestion, before anything is charged |
| Must use a technology in these categories / Must NOT use any in these categories | - | category names from the list above, for example CRM, Live chat / support, Payments |
| Minimum confidence for the filter | medium | see Evidence and confidence |
| What to analyse per input | domain | domain: one analysis per site homepage. exact: analyse exactly the URL given |
| Include DNS signals | on | adds email provider, SPF/DMARC, hosting hints |
| Pages per website | 1 | 2-3 also reads same-site shop, pricing, contact or demo pages. In our own test on 45 real sites, reading 3 pages instead of 1 found one extra technology on 2 of them and took about 3 seconds longer per site, so the default is 1. Charged per website, not per page |
| Maximum rows to process | all | cheap way to test, e.g. 5 |
Pricing
Pay per event: one website-analyzed event for each website that was successfully fetched and analysed. Failures, blocked sites, robots-disallowed sites, duplicates and skipped rows are free, so a dead or protected site in your list never costs you anything. Reading more pages of the same website does not cost more. The current price is shown at the top of this page and on the Pricing tab. Set a maximum cost on the run to cap spend.
Limits (please read)
- HTTP only. The Actor reads the HTML the server sends. Technology that only appears after JavaScript runs in a browser (some single-page apps, tags that a tag manager injects later) may be invisible. A tool that a site loads only through Google Tag Manager is typically not visible here; the tag manager itself is. Scripts that a cookie banner holds back until consent are read when they are present in the HTML. A site showing no technologies has not proven to use none.
- Not every site can be read. Large, well-known sites block automated requests more often than small business sites. In our own test on the 1,000 most popular domains (a demanding list, not a typical lead list), 54% of rows were readable and therefore chargeable; the rest were unreachable, blocked, disallowed by robots.txt or not web pages. Expect different numbers on your own list, and use
fetchStatusto see exactly what happened to each row. - Bot protection. The Actor identifies itself honestly (
B2BTechStackEnricherBot) and does not try to bypass CAPTCHAs, challenges or blocks. Protected sites come back asblocked. - Homepage by default. Some tools only load on checkout or login pages.
- Certificate problems. Sites whose certificate is expired, self-signed, for another name or served with an incomplete chain end as
tls_error;fetchReasonsays which (for exampleunknown_or_incomplete_certificate_chain). Certificates are never ignored. - Versions are reported only when the site states them; most sites do not.
- Accuracy. We do not publish an accuracy percentage. There is no public labelled benchmark that matches this catalog, and a number without one would not be honest. Every detection carries its evidence so you can judge it.
- Apex-domain rule for DNS. The domain's zone is found by walking up the name until a name server answers. A sub-domain delegated to its own DNS zone is treated as its own domain.
Responsible use
- robots.txt is always respected, for our bot name, again for every redirect and for every extra page; this cannot be switched off. A
Crawl-delayis honoured between extra pages of the same site. - A handful of requests per site at most: one homepage fetch (plus a
robots.txtfetch and awww/httpretry only when the first attempt fails to connect), and up to two extra pages if you ask for them. - Requests to private, loopback, link-local and other internal addresses are refused, including when a public name or a redirect points at one.
- Do not use the results to violate a site's terms, to harass anyone, or to process personal data without a legal basis. The Actor does not collect personal data.
API and code examples
Replace YOUR_USERNAME with the account that publishes the Actor and YOUR_API_TOKEN with your Apify API token (keep it secret).
cURL
curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~b2b-tech-stack-enricher/run-sync-get-dataset-items?token=YOUR_API_TOKEN" \-H "Content-Type: application/json" \-d '{"websites": ["allbirds.com", "example.com"], "requireAnyTechnologies": ["Shopify", "WooCommerce"]}'
Python
from apify_client import ApifyClientclient = ApifyClient("YOUR_API_TOKEN")run = client.actor("YOUR_USERNAME/b2b-tech-stack-enricher").call(run_input={"datasetId": "YOUR_LEADS_DATASET_ID","urlField": "website","requireCategories": ["CRM"],"excludeTechnologies": ["HubSpot"],})for row in client.dataset(run["defaultDatasetId"]).iterate_items():if row["matched"] is True:print(row["domain"], row["crm"], row["emailProvider"])
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });const run = await client.actor('YOUR_USERNAME/b2b-tech-stack-enricher').call({websites: ['allbirds.com', 'example.com'],requireAnyTechnologies: ['Shopify'],});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items.filter((r) => r.matched === true).map((r) => r.domain));
n8n
Use the Apify node (or an HTTP Request node) to run the Actor with the JSON input above, wait for the run to finish, then read the dataset items. A typical flow: Google Sheets (list of companies) -> Apify (Run Actor) -> Apify (Get dataset items) -> IF matched is true -> CRM or email tool. Treat rows with matched: null as "re-check later", not as "no".
Using it from an AI agent
The input is plain JSON and the output is one flat row per website with explicit statuses, which suits tool-calling agents. Two rules for agents: pass maxWebsites to keep a run small, and read matched/fetchStatus literally (null and non-ok statuses mean "unknown", never "absent"). Whether this Actor can be called through the Apify MCP server depends on your account's Apify settings; check the Actor's page in Apify Console.
FAQ
Why does a site I know uses X show nothing? Either the technology is not in the catalog, it is only added by JavaScript in the browser, or the site blocked us. Check detection and fetchStatus first; technologyCount: 0 only means "nothing found" when detection is full.
What does matched: null mean? We cannot tell. The site was blocked, unreachable, not fully analysed, or the evidence was below your minimum confidence.
Am I charged for blocked or dead sites? No. Only websites that returned a readable page and were analysed.
Are duplicates charged? No. A domain that appears several times is analysed once; repeats are free and carry duplicateOf.
Does it use a browser or proxies? No. Plain HTTP requests from Apify's network, so some protected sites will block it.
Can I add my own technology? Not at the moment. Tell us which technology is missing; the catalog is extended from each vendor's own documented install snippet.
Are versions exact? Only when the site states one (a generator meta tag, a Server header). Otherwise version is null.
Changelog
- 1.0.1: scripts held back by cookie-banner tools (for example OneTrust or Cookiebot) and app tags listed in JSON loaders (for example on Shopify stores) are now read; a vendor's own CDN hint no longer makes a site look like it runs that vendor's platform; VitePress added.
- 1.0.0: first release.