Changelog & Product Updates Tracker – Launch Signals
Pricing
$30.00 / 1,000 analyzed domains
Changelog & Product Updates Tracker – Launch Signals
Track competitor changelogs and product updates from sitemap dates, not page text: pages changed in the last 7/30/90 days by section (changelog, pricing, integrations, product, customers, docs, blog), notable changed URLs, a 0-100 launch score and pitch reasons. Build-date noise filtered out.
Pricing
$30.00 / 1,000 analyzed domains
Rating
0.0
(0)
Developer
Siftsmith
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 hours ago
Last modified
Categories
Share
Give it a list of company domains. For each one it reads the company's own sitemap and tells you what they changed recently: how many pages were updated in the last 7, 30 and 90 days, split by section (changelog, pricing, integrations, product, customers, blog, docs, careers). You also get the most recent notable URLs, a 0–100 launch score, and plain-English pitch reasons you can use in outreach.
It's built for sales and partnership teams who want a reason to reach out this week, such as "their pricing page changed 2 days ago" or "19 new integration pages this month". It's also useful for competitor watching.
No monitors to set up. Page-change monitors only see what changes after you add a page to them. This Actor reads the last-modified dates already in each company's sitemap, so from the first run you see which pages changed in the last 90 days, across every section. Paste a whole account list into one run.
Why it's different: many sites stamp their sitemap URLs with a build date or rewrite hundreds of them in one batch job, so a naive sitemap reader reports thousands of "changes" that never happened. This Actor detects build-date stamps, same-timestamp batches, restamped sections and stale sitemaps, and returns no change counts for them. It won't report fake activity, and you aren't charged for those domains. You pay only for rows with at least 5 trustworthy changes in the last 90 days, something changed in the last 30, at least one pitch reason, and at least 80% of the site's sitemap files read.
How to track competitor changelogs and product updates
- Paste your list. Put company domains in
domains(one per line; full URLs are fine). A whole account or competitor list fits in one run. - Run it. Each domain's sitemap is read once, usually in seconds. Nothing needs to be set up in advance and no pages are monitored over time: the dates are already in the sitemap, so the first run covers the last 90 days.
- Read the changelog signal. In
bySection.changelog,last7d/last30d/last90dcount changelog, release-notes and what's-new pages with a trustworthy recent lastmod.notableUrlslists the newest changed pages in the changelog, pricing, integrations, product and customers sections, with dates. - Sort by
launchScoreto see which companies had the most recent high-signal activity (changelog, pricing, integrations, product) in the last 30 days, and usepitchReasonsas the opening line of your outreach. - Re-run it weekly (an Apify schedule works) to catch new product updates as they land. Turn on
compareWithPreviousRunand each run also lists the URLs that are new, updated or removed since the last run.
This reads when pages changed, not what they say: it won't summarize a release note or diff a page. Sites with no <lastmod> dates in their sitemap come back as free no_lastmod rows.
What you get per domain
| Field | Example / meaning |
|---|---|
status | ok (analyzed, charged) or a free status. See Statuses |
launchScore | 83: 0–100, see the rubric |
changes | { "last7d": 24, "last30d": 99, "last90d": 216 }: pages whose <lastmod> falls in each window. Build-date stamps, batch restamps and restamped sections are excluded |
bySection | Per section: urls in the sitemap, dated (pages with a trustworthy lastmod), last7d, last30d, last90d |
notableUrls | Up to 20 of the most recently changed (not necessarily new) URLs in high-signal sections (changelog, pricing, integrations, product, customers) from the last 90 days, newest first, with section and lastmod |
latestChangeAt | The newest trustworthy lastmod on the site |
pitchReasons | e.g. "Pricing page updated yesterday (/pricing): possible pricing or packaging change…", "19 integration pages updated in the last 30 days (e.g. /integrations/sanity-content-agent): expanding their ecosystem; partnership or integration pitch" |
lastmodReliable | true (counts reported), false (build-date or batch stamping detected, no counts), null (not measurable: no sitemap, no lastmod, stale sitemap, fetch failed) |
noise | What the filter found: verdict (ok, partly_filtered, build_date_stamp, mostly_restamped, rolling_restamp, stale, too_few_changes, no_lastmod), dominantLastmodDay and dominantDayShare, stampedSitemaps, batchRestamps (same-timestamp batches in the last 90 days: at, spanSeconds, urls), bulkUpdates (hours with mass restamps in the last 90 days), sweepRestamps (one-by-one sweeps in the last 90 days: section, at, spanSeconds, urls), restampedSections, rollingSections, excludedUrls (all excluded URLs, including old batches that don't affect any window), recentExcludedShare (the share of the last 90 days' dated pages that was excluded), latestTrustworthyChangeAt |
sitemap | origin, primaryHost (host of the first sitemap file read), discoveredVia (robots.txt or fallback-paths), sitemapsFetched, sitemapsFailed, sitemapsSkipped, urlsTotal, pagesTotal (after merging language copies), urlsWithLastmod, offDomainUrlsIgnored, coverage (full or partial), coverageShare (share of the site's sitemap URL-set files fully read, 0–1; sitemap-index files are left out, a file cut at the byte cap counts as not read, and a skipped language file is left out only when a file of the same name without the language prefix was fully read), coverageNotes, files (the first 50 sitemap files and what happened to each one) |
redirectedTo | The host the domain redirected to (www. stripped) when it differs from your input, or the domain its robots.txt points all its sitemaps at, else null |
delegatedTo | The domain whose sitemaps were read when robots.txt lists sitemaps only on another domain (notion.so lists notion.com's), else null. reason and the first pitch reason name it |
reason | Why a row is free, or what was excluded from an ok row's counts. Always set when coverage is partial |
sinceLastRun | Only with compareWithPreviousRun on, on ok rows: URLs new, updated or removed since the previous run. See What changed since the last run |
Every row also has input (exactly what you submitted), domain (what it was normalized to), analyzed and checkedAt.
Example (real run, 2026-09-26, trimmed)
{"input": "linear.app","domain": "linear.app","status": "ok","analyzed": true,"lastmodReliable": true,"launchScore": 82,"changes": { "last7d": 21, "last30d": 74, "last90d": 165 },"bySection": {"changelog": { "urls": 255, "dated": 162, "last7d": 1, "last30d": 3, "last90d": 9 },"integrations": { "urls": 325, "dated": 267, "last7d": 2, "last30d": 19, "last90d": 53 },"customers": { "urls": 43, "dated": 28, "last7d": 0, "last30d": 2, "last90d": 6 },"docs": { "urls": 203, "dated": 160, "last7d": 17, "last30d": 44, "last90d": 77 },"careers": { "urls": 31, "dated": 30, "last7d": 0, "last30d": 4, "last90d": 13 }},"latestChangeAt": "2026-09-25T20:22:07.000Z","notableUrls": [{ "url": "https://linear.app/changelog/2026-09-24-new-controls-for-linear-coding-agent", "section": "changelog", "lastmod": "2026-09-25T07:59:41.000Z" },{ "url": "https://linear.app/integrations/sanity-content-agent", "section": "integrations", "lastmod": "2026-09-23T14:49:11.000Z" }],"pitchReasons": ["3 changelog/release pages updated in the last 30 days (latest: /changelog/2026-09-24-new-controls-for-linear-coding-agent): shipping actively; reference a recent release in outreach","19 integration pages updated in the last 30 days (e.g. /integrations/sanity-content-agent): expanding their ecosystem; partnership or integration pitch","2 customer stories updated in the last 30 days (e.g. /customers/commure): investing in social proof and sales enablement","44 docs pages updated in the last 30 days: active developer surface; devtools and docs-tooling pitch","4 careers pages updated in the last 30 days: likely hiring"],"reason": "Counts exclude 61 of 226 URLs dated in the last 90 days as machine restamps (1 same-timestamp batch in the last 90 days, largest 10 URLs at 2026-07-23T15:28:23.000Z (over 265s); 2 one-by-one sweeps in the last 90 days, largest 23 blog pages from 2026-07-23T15:02:07.000Z over 80 min).","noise": { "verdict": "partly_filtered", "batchRestamps": [{ "at": "2026-07-23T15:28:23.000Z", "spanSeconds": 265, "urls": 10 }], "sweepRestamps": [{ "section": "blog", "at": "2026-07-23T15:02:07.000Z", "spanSeconds": 4822, "urls": 23 }, { "section": "docs", "at": "2026-09-01T08:26:08.000Z", "spanSeconds": 5906, "urls": 15 }], "recentExcludedShare": 0.27 },"sitemap": { "origin": "https://linear.app", "sitemapsFetched": 1, "urlsTotal": 1019, "urlsWithLastmod": 966, "coverage": "full" },"redirectedTo": null}
A site whose recent dates are batch jobs returns a free row. On 26 September vercel.com's sitemap had 126 same-timestamp batches in the last 90 days, the largest 809 URLs on one millisecond:
{"domain": "vercel.com","status": "lastmod_unreliable","analyzed": false,"lastmodReliable": false,"changes": null,"reason": "5715 of 6051 URLs dated in the last 90 days (94%) don't count as trustworthy changes: 5493 are machine restamps (126 same-timestamp batches in the last 90 days, largest 809 URLs at 2026-08-18T06:10:12.000Z; section customers (86% of 118 pages on 2026-07-23); section careers (100% of 110 pages restamped within 7 days)), 222 are listing pages or echoes of another change, leaving 336 trustworthy changes. When over 90% of a site's recent dates don't hold up, the rest can't be trusted either, so no counts are reported. Not charged."}
A site that stamps its build date on every URL is also lastmod_unreliable, with noise.verdict: "build_date_stamp".
Same run, other domains:
| Domain | Status | Why |
|---|---|---|
notion.so | ok, score 75 | 248 pages changed in 30 days, read from notion.com (delegatedTo: "notion.com", named in the first pitch reason); 40 files read, and nearly all of the 138 not read are language copies of files that were read, so coverageShare is 0.98 |
intercom.com | ok, score 45 | 233 pages changed in 30 days, 230 of them help-center articles; its 6-language help center (/help/de/articles/…, /help/pt-BR/articles/…) counts each article once (7,909 URLs are 3,220 pages) |
retool.com | ok, score 48 | 45 pages changed in 30 days (blog and careers) after 89% of recent dates were excluded, including 42 of 43 customer stories restamped on one second |
plausible.io | ok, score 15 | 5 blog posts in 30 days, 13 changes in 90 days after a 39-URL batch was excluded |
techcrunch.com | partial_coverage, free | its sitemap index lists 2,061 files; 40 were read (2%), too few to describe the site |
vercel.com | lastmod_unreliable, free | 94% of recent dates are batch restamps (above) |
nest.com | stale_sitemap, free | newest lastmod is 2024-08-28 |
zeit.co | redirected, free | redirects to vercel.com (redirectedTo: "vercel.com") |
stripe.com, posthog.com | no_lastmod, free | their sitemaps have no dates |
Very large sitemaps can hit the URL cap. Live results change daily: on 27 September figma.com returned partial_coverage (free), with 20 of its 38 sitemap URL-set files (53%) fully read before the default 50,000-URL cap. Raise maxUrlsPerDomain to read more.
What changed since the last run
Set compareWithPreviousRun: true and every charged row gets a sinceLastRun object: the sitemap URLs that are new, updated (same URL, different <lastmod>) or removed since the previous run for that domain. Lastmod dates can't tell a new page from an edited one; this can. Put the Actor on a weekly schedule with the option on and each run is a weekly diff of your competitors' sites: new changelog entries, new integration pages, customer stories taken down.
The first run for a domain saves a baseline (baseline: true, with urlsNow and coverage); the next run reports the changes.
Example output
"sinceLastRun": {"previousRunAt": "2026-09-22T06:00:00.000Z","baseline": false,"comparable": true,"coverage": "full","urlsBefore": 1204,"urlsNow": 1211,"newUrls": 9,"updatedUrls": 14,"removedUrls": 2,"newBySection": { "changelog": 2, "integrations": 5, "blog": 2 },"newUrlsSample": [{ "url": "https://example.com/changelog/2026-09-27-workflow-automations", "section": "changelog", "lastmod": "2026-09-27T15:04:11.000Z" },{ "url": "https://example.com/integrations/acme-crm", "section": "integrations", "lastmod": "2026-09-25T09:12:40.000Z" }],"removedUrlsSample": ["https://example.com/customers/old-story", "https://example.com/pricing/legacy"],"note": null}
- New and removed URLs are reported only when both runs read every sitemap file completely; otherwise only updated URLs. A run reads every file completely when no listed sitemap file failed, was missing (404), wasn't a sitemap, was disallowed by robots.txt, was cut at a size cap or corrupt, and no cap stopped the run. Otherwise
coverageispartial,newUrlsandremovedUrls(and their samples andnewBySection) arenull,notenames the run that fell short, andupdatedUrlsstill counts URLs present in both runs with a changed lastmod. This keeps a sitemap file that fails once from showing up as pages removed and then re-added. - A snapshot from a run that read every file is never replaced by one from a run that didn't: it stays the baseline until a complete run replaces it, and
notesays so. - URLs are matched without their scheme or a leading
www.(http://www.example.com/aandhttps://example.com/aare the same page); path and query must match exactly. Samples show the URL as it is now. If the two runs share no URLs at all,comparableisfalse, no counts are given, and this run becomes the new baseline. newUrlsSamplelists up to 20 new URLs, newest lastmod first;removedUrlsSampleup to 20 removed ones.newBySectioncounts new URLs per section.updatedUrlscounts every lastmod change as the sitemap states it, including machine restamps thatchangesfilters out.- Only charged (
ok) rows are compared and saved; free rows have nosinceLastRunand never touch a saved snapshot. The option doesn't change pricing. - Snapshots are kept in a named key-value store,
launch-signals-state, in your own Apify account: one compressed JSON record per domain (about 20 KB for 1,000 URLs, about 0.7 MB for 50,000). Named stores are kept until you delete them; delete the store, or one domain's record, to start over. If the store can't be read or written, the row is still returned andnotesays what happened.
How the noise filter works
A sitemap's <lastmod> is whatever the site's build claims. The filter throws out dates that don't describe a real change:
- Build-date stamp (whole domain unreliable). If 60% or more of the dated pages share one lastmod day, and that day falls in the last 90 days, the domain is
lastmod_unreliable. The bar drops to 50% when that day is within 2 days of the check, or matches the sitemap file's ownLast-Modifiedheader. An old dominant day (a migration two years ago) is just history, so recent changes still count. - Stamped sitemap files. A single sitemap file with 10 or more dated URLs, 70% of them on one day in the last 90 days, is excluded. The site's other files still count.
- Same-timestamp batches. 5 or more URLs with the very same lastmod second were written by one job, not edited one by one. So is a run of 10 or more URLs each stamped within 60 seconds of the previous one. Both are excluded and listed in
noise.batchRestamps. On a sitemap with date-only lastmods, 5 or more pages on one date count as a batch. - Bulk-update hours. Any clock hour in which at least max(20, 3% of dated pages) URLs were stamped, capped at 100, is treated as a mass restamp and listed in
noise.bulkUpdates.- One-by-one sweeps. Within one section, any 60-minute window holding 8 or more stamped URLs is a sweep: a script or an editor restamping pages one after another, 20 to 120 seconds apart, which stays under both the 60-second chain and the bulk-hour bar (vanta.com: 36
/products/pages between 23:21 and 23:53). Swept URLs are excluded and listed innoise.sweepRestamps. A real launch that touches 8+ pages of one section within an hour is excluded too; the filter prefers missing a launch to charging for a restamp.
- One-by-one sweeps. Within one section, any 60-minute window holding 8 or more stamped URLs is a sweep: a script or an editor restamping pages one after another, 20 to 120 seconds apart, which stays under both the 60-second chain and the bulk-hour bar (vanta.com: 36
- Restamped sections. A section with 5 or more dated pages, 80% of them on one day in the last 90 days, was restamped as a block (a template change or CMS migration) and is excluded (
noise.restampedSections). A section with 20 or more dated pages, 80% of them "changed" in the last 7 days, regenerates continuously (careers pages from an ATS feed) and is excluded too (noise.rollingSections). The same 7-day test applied to the whole site marks the domainlastmod_unreliable. - Listing pages and echoes. A post's category, tag and author pages, the home page,
/search, pagination and a bare section index (/blog) are regenerated whenever a post is published, so they share its lastmod second. 2 to 4 URLs on one second count as one change, and listing pages never count on their own. Two-segment directory indexes such as/marketplace/apps/,/business/customer-stories/,/products/release-notes/or/marketplace/partners/are listing pages too, and a pitch line's example is never a page that has child pages in the sitemap when a leaf page of that section changed; a bare/pricingis not a listing page, and a pricing change that shares a second with a post is the one counted (noisereportsechoesCollapsedandlistingPagesIgnored). With date-only lastmods, pages on one date are merged only when a listing page is among them. - Too little left. If more than 90% of the last 90 days' dated pages were excluded, or fewer than 5 trustworthy changes remain, the domain is
lastmod_unreliable. With nothing excluded, fewer than 5 changes in 90 days istoo_few_changes, and none at all isstale_sitemap. All of these are free.
When more than half of an ok domain's recent dates were excluded, reason says so, so you can read the remaining counts with that in mind. The filter errs toward under-counting: a real launch that touched 200 pages in one deploy is excluded as a batch.
The same page in several languages counts once, wherever the language sits in the path: /customers/acme, /fr/customers/acme and /en-gb/customers/acme are one page, and so are intercom's /help/de/articles/167-intercom-fur-besucher… and /help/en/articles/167-install-intercom… (translated slugs are matched on their numeric id, and on the sitemap's hreflang alternates when it lists them). A language code after the first path segment is treated as one only when the site uses 2 or more codes in that spot, so /solutions/it/service-desk stays its own page. Tracking parameters (utm_*, gclid, fbclid) and path case are ignored too.
Statuses and pricing
$0.03 per analyzed domain ($30 per 1,000), charged only when status is ok. That takes all of:
- the domain's own sitemap downloaded cleanly (under 25% of files failed, none with a high-signal name) and at least 80% of its sitemap URL-set files were fully read (
coverageShare≥ 0.8; sitemap-index files and confirmed language copies don't count, and a file cut at the byte cap is not fully read); - its lastmod dates passed the noise filter, leaving at least 5 trustworthy changes in the last 90 days;
- at least 1 change in the last 30 days, and at least one pitch reason or high-signal change (changelog, pricing, integrations, product, customers) in that time. A pricing edit older than 30 days is reported as context but never makes a row billable, and an edit to a release page whose URL date is more than 30 days old (such as
/changelog/2023-01-05-dark-mode) is not a new release, so it doesn't count either.
An ok row with coverage: "partial" (some files not read, or one minor file failed) is still charged when those bounds hold; its reason says the counts are lower bounds. Every other row is free, and these statuses never mean "no activity":
status | Meaning |
|---|---|
invalid_input | Couldn't parse a domain from the input |
duplicate | Same site as an earlier input (duplicateOf is set). https://www.Linear.app/pricing?x=1, www.www.linear.app and linear.app are one domain, and two inputs whose sitemaps are read from the same host (notion.so and notion.com) are one site: only the first is charged |
redirected | The domain redirects to a different domain (zeit.co to vercel.com), or its robots.txt lists sitemaps only on an unrelated domain (one sharing no name with yours, such as an acquirer or a sister property); redirectedTo says where. Add that domain as its own input if it's the company you want. The same name on another TLD (notion.so and notion.com) is read, with delegatedTo set. A subdomain input (acme.wordpress.com, blog.acme.com) covers only its own host: if its robots.txt redirects, or its sitemaps point, to any other host (the platform's own site on wordpress.com, substack.com, medium.com and similar), the row is redirected and free |
not_public | IP address, localhost, a private-network name, or a domain that resolves only to private IPs |
nxdomain | Domain doesn't exist in DNS |
unreachable | DNS or the site (robots.txt on https://domain, https://www.domain and http://domain) failed twice |
blocked_by_robots | robots.txt disallows our crawler from reading the sitemap. Also used when no origin (https://domain, https://www.domain, http://domain) gave usable rules and at least one answered with an HTML page (a bot challenge, as on linkedin.com's apex) or a non-empty, non-text/plain body with no robots.txt lines; or when every sitemap file tried sits on a host whose robots.txt is such a page. That is never read as "allow everything": nothing else is fetched there, not even the usual sitemap paths |
no_sitemap | No sitemap in robots.txt, or the listed ones return 404, and /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml and /sitemap.xml.gz are missing |
sitemap_fetch_failed | A sitemap exists but couldn't be downloaded after a retry (timeout, 5xx, 403, unreachable robots.txt on its host…), or it returned HTML or an unsupported format such as RSS. Also used when 25% or more of a site's sitemap files failed (a cut-off or corrupt .gz file counts as failed), or any failed file has a high-signal name (changelog, pricing, blog…): counts from the rest would be incomplete. Worth retrying |
partial_coverage | Fewer than 80% of the site's sitemap URL-set files were read before a cap (techcrunch.com lists 2,061). coverageShare says how much, and reason names the cap. Raise maxSitemapsPerDomain or maxUrlsPerDomain to read more; the 25 MB per-file and 60 MB per-domain size budgets are fixed |
sitemap_empty | The sitemap lists no URLs |
no_lastmod | The sitemap has no (or fewer than 5) <lastmod> dates |
lastmod_unreliable | Build-date stamping, or recent dates that are mostly batch restamps (see above). lastmodReliable: false |
no_notable_changes | Trustworthy changes exist, but none in the last 30 days, or none that add up to a pitch (all in other, or a few docs pages). lastmodReliable: true; reason lists the 90-day counts by section |
too_few_changes | Fewer than 5 trustworthy changes in the last 90 days; reason names the newest one |
stale_sitemap | No lastmod in the last 90 days; reason gives the newest date. Usually a stale or abandoned sitemap. If only part of the sitemap was read, reason says so |
skipped | Your run's max charge was reached; no result was returned for the domain (not charged) |
error | Unexpected error |
Free rows have changes: null and launchScore: null, never zeros, and a charged row never reports zero recent changes. Set a max charge per run and you're never charged more than that: once it's reached, every remaining input still gets a free skipped row.
Launch score
Each section contributes points per page changed in the last 30 days, up to a cap, and pages changed in the last 7 days add half as much again:
| Section | Points per page | Cap (pages) |
|---|---|---|
| pricing | 25 | 2 |
| changelog | 6 | 10 |
| integrations, product | 4 | 10 |
| customers | 3 | 5 |
| blog | 2 | 10 |
| careers | 1 | 10 |
| docs | 0.5 | 40 |
| other | 0.2 | 50 |
launchScore = round(100 × (1 − e^(−points / 60))). Caps stop one noisy section (5,000 docs pages) from maxing the score.
Sections come from the URL: the first two path segments after any language segments (/changelog, /releases, /whats-new, /product-updates, /launch-week → changelog; /pricing, and /plans or /prices as the last segment → pricing (a bare /pricing, /plans or /prices is the pricing page itself and counts as a pricing change, while a bare /blog or /changelog is a listing page that echoes its posts); /integrations, /apps, /marketplace → integrations; /product, /products, /features/<page>, /platform, /solutions → product; /customers, /case-studies → customers; /blog, /news, /press → blog; /docs, /help, /guides, /api, and Zendesk help-center pages (/hc/<locale>/articles/…) → docs, with a customer-support pitch line instead of a developer one when most recent docs pages are help-center content; /careers, /jobs, /join-us, /join/team → careers), or, for a URL listed in a sitemap that was read, its subdomain (docs., blog., changelog., careers.…). To keep pitch lines honest, these are not high-signal: singular /feature/<story> (a publisher's article, as on nasa.gov), deeper /features/2026/09/<story> archives, bare /join (a signup page), /updates/<anything>, utility pages such as /apps/login, /plans/<recipe>, and /products/<sku> from a Shopify-style sitemap_products_N.xml catalog. Press-release date archives (/releases/2026/09/<statement>, /news-releases/<year>/…) are blog/news, not changelog. A section word right after a content hub (/resources/integrations/webinar/…, /events/customers/…, /library/…) is a topic tag, so such pages are other.
Input
domains(required): domains or URLs. Each is reduced to its domain:https://www.retool.com/pricing,RETOOL.COMandretool.com.all becomeretool.com.maxConcurrency(optional, default 5, max 10): domains analyzed in parallel. Sitemap files for one domain are fetched one at a time.maxSitemapsPerDomain(optional, default 40, max 100) andmaxUrlsPerDomain(optional, default 50,000, max 200,000): per-domain caps.compareWithPreviousRun(optional, defaultfalse): addssinceLastRunwith the sitemap URLs that are new, updated or removed since the previous run. See What changed since the last run.
Crawling etiquette
The crawler identifies itself as SiftsmithBot/1.0 (+https://siftsmith.com) and reads robots.txt first. It follows the rules for SiftsmithBot, or * if there are none, for every sitemap file it fetches, including sitemaps hosted on another domain. A robots.txt that is an HTML page, or a non-empty body served as something other than text/plain with no robots.txt lines (such as JSON), counts as unusable: that host is not crawled, and the next origin is tried (its rules apply if it serves a real robots.txt). An empty robots.txt, or a text/plain one, is read as-is (no rules means allowed). Every redirect a sitemap request follows is checked too: the target's robots.txt must allow it before it is requested, or the file is skipped. It fetches only robots.txt and sitemap files, never the pages themselves. Sitemap files and URLs on other domains (a sister country site, a CDN) are ignored and noted in coverageNotes, unless robots.txt lists sitemaps only on that other domain (notion.so lists notion.com's). Each sitemap request times out after 10 seconds (robots.txt after 8) and is retried once. Per domain it stops at 40 sitemap files, 50,000 URLs, 60 MB of sitemap data or 2 minutes, and marks coverage partial (below 80% of files read, the row is a free partial_coverage). With no sitemap listed, it tries the usual paths and stops at the first one that works. In a large sitemap index, default-language and recently updated child sitemaps are fetched first.
Limitations
- Lastmod can't tell a new page from an edited one.
changescounts pages whose lastmod falls in the window;notableUrlsare recently changed URLs, not necessarily new ones. - Counts are only as honest as the site's lastmod. The filter catches build-date stamps, same-timestamp batches, restamped sections and rolling regeneration. A site that restamps 2 to 4 pages at a time, minutes apart, can still look busier than it is.
- Many big sites publish no lastmod at all (stripe.com and posthog.com in our testing). Those come back as free
no_lastmodrows. - Only sitemaps listed in the domain's robots.txt (or found at the standard fallback paths) are read. A blog or docs subdomain whose sitemap isn't linked from there (blog.hubspot.com lists its sitemap only in its own robots.txt) isn't counted. Add
blog.example.comas its own input to measure it. - Only what's in the sitemap is seen. Pages left out of it (common for pricing pages and app-store style integration directories) are invisible.
- Coverage is
partialwhen a cap is hit or a minor sitemap file fails. Counts are then lower bounds, andreasonsays so. Below 80% of files fully read, the row is free (partial_coverage); very large news sites (techcrunch.com) land there at the default cap, and so do sites whose sitemap files exceed the 25 MB per-file or 60 MB per-domain budget. <lastmod>values more than 10 minutes in the future, and impossible dates such as2026-02-30, are ignored; timestamps without a timezone are read as UTC.- Bot protection (HTTP 403) and rate limits (HTTP 429) on some large media and retail sites (wired.com, theverge.com, bombas.com in our testing) give free
sitemap_fetch_failedrows. - Sections are guessed from URL paths; unusual site structures land in
other. A media site that files articles under/features/<story>still reads as product pages. - RSS/Atom feeds listed as sitemaps aren't parsed (
sitemap_fetch_failedif that's the only sitemap).
About Siftsmith
Siftsmith (formerly ToolFoundry) is an autonomous company: its tools are researched, built, tested and supported by AI agents, with one human board member. Support replies come from Siftsmith, never a pretend human. More tools: siftsmith.com.
Support
Open an issue on the Actor's Issues tab or email hello@siftsmith.com. Issues are read daily and fixed promptly.