# Changelog of Website Tech Stack Detector (Wappalyzer API Scraper) (`rl1987/wappalyzer-tech-lookup`) Actor

- **URL**: https://apify.com/rl1987/wappalyzer-tech-lookup/changelog.md
- **Full Actor documentation**: https://apify.com/rl1987/wappalyzer-tech-lookup.md

## Changelog

All notable changes to the Wappalyzer Technology Lookup Actor are documented here.

### \[1.0.26] - 2026-09-29

#### Fixed

- **Rows that deliver no technology are no longer charged.** A domain where nothing was recognised, and a domain whose lookup failed, each produced a row that cost an `output-row` event ($0.0005) despite carrying no detection — measured on a 40-URL run that charged 2,092 events for 2,092 rows, three of which were bare `{"domain": ...}` placeholders. Those rows are still written, because they are how you see which domains were processed, but they now go out in a separate uncharged `push_data` call. Verified on a 6-domain run: 191 rows written, 188 charged.

#### Added

- **Three fields the Wappalyzer API returns and the Actor was discarding**: `technologySlug` (Wappalyzer's stable key for the technology, on every row), `versions` (the detected version(s), on roughly one row in five — the README already promised this), and `cpe`, the identifier CVE databases key on, on the rows where a version was detected. `cpe` turns the output from "what does this site run" into "which of these run something with a known CVE".

### \[1.0.17] - 2026-09-21

#### Fixed

- **Lookups now go through residential proxies instead of datacenter proxies.** The Wappalyzer API rate-limits per IP. The shared Apify datacenter pool is small and already rate-limited by Wappalyzer, so routing through it returned `HTTP 429` for nearly every request and the Actor then retried each one five times — producing empty results and runs that dragged until they timed out. Measured on the same 81-URL input: datacenter returned 81 empty rows after 405 retries; residential returns the full 3,974 rows with zero 429s, in 32s instead of 515s. `proxyConfiguration` is now an input, so you can choose a different pool or disable the proxy.
- **Concurrency now fits the run's memory**, and is configurable via `maxConcurrency`. Default run memory is raised to 512 MB; 20 parallel lookups was being killed by the 128 MB limit.
- **Runs no longer time out on large inputs, and a run that does run out of time still delivers its results.** Rows were collected in memory until every lookup had finished, then written one at a time with a separate charge call per row - two API round trips each, measured at 9.6 rows/sec on build 1.0.11. A 1000-URL input spent ~26 minutes on output alone, a 3000-URL input could not finish inside the one-hour platform timeout at all, and because nothing was written until the last lookup returned, a timed-out run produced **zero** rows despite every lookup having succeeded. Results are now flattened and pushed in batches of 100 as each lookup completes, with one `push_data` call per batch that both writes and charges.
- Reduced the per-request timeout from 60s to 30s. A single rate-limited URL could previously consume 5 x 60s plus backoff = 315s on its own, which at 5 concurrent lookups meant ~105 minutes for 100 URLs.
- Added a run deadline: the Actor now stops starting new work shortly before the platform would kill the run and flushes what it has collected, instead of being killed with rows still buffered.
- Corrected a comment that described the charge as $0.50 per row; it is $0.0005.

### \[1.0] - 2026-06-18

#### Added

- Initial release.
- Bulk technology-stack lookup: accepts a list of URLs or bare domains and detects the technologies behind each site via the Wappalyzer API.
- **Flat, CSV-ready output** — one row per detected technology with `domain`, `technologyName`, `trafficRank`, `confirmedAt`, `iconUrl`, and comma-joined `categories`.
- Direct PNG `iconUrl` for every technology.
- Input normalization: bare domains default to `https://`, invalid entries are skipped, and duplicate URLs are looked up only once.
- Requests routed through Apify's datacenter proxy with a rotating IP per attempt to avoid the shared API key's per-IP rate limit; failed lookups retry on a fresh IP with backoff and are recorded with an `error` field instead of failing the run.
- Pay-per-event pricing: **$0.50 per 1,000 output rows** ($0.0005 each), charged once per row and capped by the user's per-run spending limit.
- Minimal footprint: runs in 128 MB of memory.
