Website Tech Stack Scraper — CMS, Shop, Analytics, Payments avatar

Website Tech Stack Scraper — CMS, Shop, Analytics, Payments

Pricing

from $1.40 / 1,000 websites

Go to Apify Store
Website Tech Stack Scraper — CMS, Shop, Analytics, Payments

Website Tech Stack Scraper — CMS, Shop, Analytics, Payments

Detect what a website runs on: CMS, e-commerce platform, frontend framework, hosting and CDN, analytics, ad pixels, marketing and support tools, payment providers and bot protection — with the evidence behind every detection. No login or API key.

Pricing

from $1.40 / 1,000 websites

Rating

0.0

(0)

Developer

Chorelet

Chorelet

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Find out what a website runs on. Paste domains and get, for each one, the CMS, e-commerce platform, frontend framework, hosting and CDN, analytics, ad pixels, marketing and support tools, payment providers and bot protection — with the exact evidence behind every detection. JSON, CSV or Excel, or through the API.

No login, no API key, no browser: the Actor reads the homepage, its response headers and robots.txt, the same things a visitor's browser receives first.

Why this Actor

  • Evidence, not guesses. Every technology comes with the header line, script URL or HTML snippet that proved it, so you can check a detection instead of trusting it.
  • First-party matching. A wp-content image linked from someone else's domain will not label a site WordPress — a mistake most detectors make.
  • robots.txt too. Platforms usually name themselves there, which catches shops and CMSes whose homepage is a JavaScript shell.
  • Built for lists. Flat cms / ecommerce / analytics / payments columns, so a CSV of a thousand domains is usable without unpacking anything.

Sample output

One item of the dataset (long values shortened):

{
"website": "https://nytimes.com/",
"title": "The New York Times - Breaking News, US News, World News and Videos",
"cms": "WordPress",
"ecommerce": null,
"frameworks": [
"Svelte",
"SvelteKit"
],
"analytics": [
"Google Tag Manager"
],
"payments": [],
"technologyCount": 6
}

What you get

  • 120 technologies across 19 categories, each with the header line, script URL or HTML snippet that proved it
  • Versions where the page states them (WordPress, jQuery, nginx, Angular…)
  • Flat columns — cms, ecommerce, frameworks, analytics, payments, hosting — so a CSV is usable as-is
  • robots.txt fingerprints, which often name the platform even when the homepage is a JavaScript shell
  • First-party matching: a wp-content image borrowed from another domain does not make a site "WordPress"
  • Monitored daily

Input

  • Websites — domains or URLs, one per line.
  • Include evidence — keep the snippet that proved each detection (on by default).
  • Also read robots.txt — one extra small request per site, not charged separately.
  • Timeout per site, Websites in parallel.

Limits and notes

  • This is static detection: everything a browser loads later — tags injected by Google Tag Manager, widgets added by JavaScript — is invisible here. Expect a handful of solid detections per site rather than the long list a browser extension shows after the page finishes loading.
  • What it does catch reliably: platforms (Shopify, WordPress, Webflow, Wix, Magento…), frameworks (Next.js, Nuxt, Astro, Remix…), hosting and CDN, and any script the HTML references directly.
  • Sites behind an aggressive bot wall (some marketplaces and airlines) answer 403 to any plain HTTP client; those rows come back with the status and an error instead of a guess.
  • Every site is charged once, whether or not anything was detected.

Input example

{
"websites": [
"gymshark.com",
"vercel.com",
"plausible.io"
],
"includeEvidence": true,
"checkRobotsTxt": true,
"requestTimeoutSecs": 20,
"concurrency": 5
}

How much does it cost?

Pay per website — no subscription, no minimum, no charge for platform usage.

VolumePrice
1,000 websites$2.00
10,000 websites$20.00
100,000 websites$200.00

The Apify free plan includes $5 of usage every month — about 2,500 websites with this Actor, no card needed. Nothing else is charged: platform usage is included in the price, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off these prices.

Use it from code, n8n, Make, Zapier or an AI agent

Run the Actor and download the dataset in one call (JSON by default; add &format=csv or xlsx):

curl -X POST "https://api.apify.com/v2/acts/chorelet~website-tech-stack-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"websites": ["gymshark.com", "vercel.com", "plausible.io"], "includeEvidence": true, "checkRobotsTxt": true, "requestTimeoutSecs": 20, "concurrency": 5}'

Python:

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("chorelet/website-tech-stack-scraper").call(run_input={"websites": ["gymshark.com", "vercel.com", "plausible.io"], "includeEvidence": true, "checkRobotsTxt": true, "requestTimeoutSecs": 20, "concurrency": 5})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)
  • n8n, Make, Zapier — use the Apify node/module: run the Actor, then "get dataset items".
  • Google Sheets, Slack, webhooks — add an integration on the run's Integrations tab.
  • AI agents — the Actor is available as a tool through the Apify MCP server; the dataset schema describes every field for the model.
  • Schedules — run it hourly, daily or weekly from the Schedules tab.

FAQ

How does this compare with a browser extension?

An extension sees the page after JavaScript has run, so it also catches tags injected by Google Tag Manager. This Actor reads the HTML, headers and robots.txt — fewer detections per site, but hundreds of sites a minute and no browser to pay for.

Which technologies are covered?

About 120: e-commerce platforms, CMSes and site builders, frontend and backend frameworks, hosting, CDNs and servers, analytics, ad pixels, marketing, email and support tools, payment providers, monitoring, auth, search, captchas, fonts and media.

Why did a site come back with only two technologies?

Because that is what its HTML actually reveals. Sites that load everything through a tag manager expose very little before JavaScript runs — the fields you do get (platform, framework, hosting) are the reliable ones.

Can I find every site that uses Shopify?

Not with this Actor — it checks the domains you give it. Feed it a list from your CRM, a directory or another Actor, then filter on ecommerce.

What does a 403 mean?

The site blocks plain HTTP clients. The row keeps the status and the error so you can tell "blocked" apart from "nothing detected".

Support

Questions, missing fields or a source that changed? Open an issue on the Issues tab or write to support@chorelet.app — problems are usually fixed within a day, and the Actor is checked every morning by an automated test run. If the Actor saved you time, a short review on its Store page helps other people find it.