Cachet Status Page Scraper & Monitor Tool
Pricing
Pay per event
Cachet Status Page Scraper & Monitor Tool
Search or monitor any self-hosted Cachet status page — not tied to one vendor. Get clean component and incident data as JSON with status, timestamps, and canonical URLs. Free to start, with API and MCP support for automation and integration.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Turgay NANTA
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
20 hours ago
Last modified
Categories
Share
Cachet Status Page Scraper
Search or monitor any self-hosted Cachet status page and get clean component/incident JSON. One click, no required fields, no LLM.
What it does
Cachet is a popular open-source, self-hosted status-page system — by design every GET endpoint on a Cachet instance is public, no login needed. Unlike hosted platforms, there is no central directory of Cachet sites: each company runs its own instance at its own domain. This one actor works on any of them: point it at a Cachet instance's base URL, optionally add a search term, and get clean component and incident JSON — status, description and canonical URLs.
Why this one:
- Works on ANY self-hosted Cachet instance — no central directory needed, just a URL
- Returns BOTH components and incidents in one run
- Optional search term filters both record types client-side
- No LLM anywhere — deterministic output, predictable costs, no hallucinated fields
- Clean by default — canonical URLs (tracking parameters stripped), parsed numbers, merged duplicates
Quick start (no code)
- Click Try for free / Start — every field has a working default, nothing is required.
- (Optional) change query to what you need.
- Open the Dataset tab when the run finishes → export as JSON, CSV or Excel.
Input
| Field | Required | Default | Description |
|---|---|---|---|
query | no | https://demo.cachethq.io | Base URL of the Cachet instance, followed by an optional search term (e.g. https://demo.cachethq.io outage). Works on ANY self-hosted Cachet installation. |
maxResults | no | 20 | Maximum clean results (capped at 500) |
enrich | no | false | Deterministic enrichment per record — see below |
monitor | no | false | Compare with the previous run, flag NEW records only |
Example input:
{"query": "https://demo.cachethq.io","maxResults": 50}
Output
Real example record (from a live run):
{"id": "component-1","title": "API","url": "https://demo.cachethq.io","description": "Used by third-parties to connect to us","status": "Operational","created_at": "2026-08-02 19:30:02","updated_at": "2026-08-02 19:30:02","record_type": "component","completeness": 0.33}
The final _summary row carries run totals (total_clean, deduped, enriched); in monitor mode a _changes row lists keys new since the last run.
Field reference
| Field | Meaning |
|---|---|
id | Component or incident ID, prefixed by type (stable, deduplication key) |
title | Component name, or incident name for incident rows |
url | Component link, or incident permalink for incident rows |
description | Component description, or incident message for incident rows |
status | Human-readable component status, or resolved/open for incident rows |
record_type | component or incident |
price / price_text | Parsed numeric value + original text, when the source publishes one |
completeness | 0–1 filled-fields score (with enrich) |
Use cases
- Self-hosted infra monitoring — track the status pages of internal tools and vendors that self-host Cachet.
- Status-change alerts — schedule with monitor mode;
change-alertfires on new components or incidents. - Incident research — pull an instance's incident history for a post-mortem.
- Dashboards — aggregate multiple self-hosted status pages into one internal board.
- AI agents — live status checks for internal/self-hosted services via MCP.
Enrichment (optional, charged only when it produces something)
Set enrich: true and every record additionally gets: e-mail addresses extracted from the description (when present), the canonical domain of the record's URL, and a completeness score (0–1, how many core fields are filled). Deterministic — the same input always yields the same output — and you are only charged for records that actually got enriched. Records where enrichment adds nothing are free.
Monitor mode — change alerts on a schedule
Set monitor: true and the actor compares the current run with the previous one (per-actor named storage) and flags only NEW records. Combine with Apify Schedules for a daily/hourly watch: the _changes summary row lists what appeared since the last run, and the change-alert event is charged per new record only — an unchanged run costs you almost nothing.
Use it from your code
Python
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("EnezLi/cachet-scraper").call(run_input={ "query": "https://demo.cachethq.io", "maxResults": 50 })for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });const { defaultDatasetId } = await client.actor('EnezLi/cachet-scraper').call({ "query": "https://demo.cachethq.io", "maxResults": 50 });const { items } = await client.dataset(defaultDatasetId).listItems();console.log(items);
curl
curl -X POST "https://api.apify.com/v2/acts/EnezLi~cachet-scraper/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \-H "Content-Type: application/json" \-d '{ "query": "https://demo.cachethq.io", "maxResults": 50 }'
Use it with AI agents (MCP)
This actor is agent-ready: it appears in Apify's AGENTS / MCP servers catalog, so any MCP-capable assistant (Claude, custom agents, LangGraph tools) can discover and call it with a one-line tool call — zero required fields means an agent can run it safely with defaults. Connect your agent to the Apify MCP server and ask for live data in natural language.
Pricing — Pay-Per-Event, start is free
| Event | When charged |
|---|---|
| Actor start | Free ($0) — try it with one click |
result | Per clean result returned |
enrichment | Only per record that actually got enriched |
change-alert | Monitor mode: per NEW record since the previous run |
No subscription, no minimum. Volume discounts apply automatically through Apify account tiers (up to −44% on GOLD). Typical run cost example: 20 results ≈ a few cents total — you can predict your bill from the numbers above before you run.
Is this legal?
This actor collects publicly available data only — the same information any visitor sees in a browser, via public endpoints. It does not bypass logins, collect private personal data, or store credentials. You are responsible for using the output in compliance with the source site's terms and the laws that apply to you (e.g. GDPR when the output contains personal data).
Support & feedback
Found a bug, need another field, or want a variant for a related platform? Open an issue on the Issues tab — issues are monitored and answered, and frequently-requested fields get added to the standard output. The actor is maintained as part of a scraper family built on one shared, tested core (bugs fixed once are fixed everywhere).
Changelog
- 0.1 (2026-07) — initial public release: search, dedup, optional enrichment, monitor mode, PPE pricing.
Limitations (honest ones)
Reads each instance's own public Cachet API; instances that require authentication for GET requests (non-default Cachet config) return zero rows. Incident update threads are not included (top-level message only).
FAQ
Which sites does this work on?
Any self-hosted Cachet instance — if <base-url>/api/v1/components returns JSON in your browser, it works.
Why full URL instead of a subdomain?
Cachet is self-hosted software, not a single hosted vendor — every installation lives at its own domain, so the actor needs the full base URL (same pattern as this family's Discourse/Flarum actors).
A URL returned zero results — why?
Either it's not a Cachet instance, the URL is wrong, or the instance has no public components/incidents. The run ends cleanly either way — nothing charged for results.
Do I need an API key or account on the source platform?
No. The actor uses public endpoints — you only need your Apify account.
Does it use AI / an LLM?
No. The core is fully deterministic: same input, same output, no hallucinations, no per-token costs.
Can I run it on a schedule?
Yes — use Apify Schedules; combine with monitor mode to pay only for what's new.
What's the maximum number of results?
500 per run (memory-safe cap). Run multiple queries or schedule runs for more.
How is my bill calculated?
Only from the events in the Pricing table — start is free, and there is no subscription.