Automation Health Watchdog avatar

Automation Health Watchdog

Pricing

from $1.00 / 1,000 resolved workflow checks

Go to Apify Store
Automation Health Watchdog

Automation Health Watchdog

Detect missed runs, explicit failures, empty output, and dead-letter growth in workflow observations you supply. Returns auditable verdicts and explicit unknowns instead of false green status.

Pricing

from $1.00 / 1,000 resolved workflow checks

Rating

0.0

(0)

Developer

Mehdi Badawi

Mehdi Badawi

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

7 days ago

Last modified

Share

Supply workflow definitions, run observations, and source attestations to evaluate automation health. This Actor does not connect to n8n, Make, Zapier, or Apify accounts by itself. No input runs a clearly labeled synthetic demo.

The input and output below are the canonical contract surface (contracts/*.schema.json, contractVersion 1.0.0). Optional connectors are a documented extension point; the current Actor works on manually supplied observations and source attestations, and connector credentials never enter the input, output, or state payloads.

What it does

Answers, once per scheduled run: did each client workflow run when expected, fail explicitly, report success with missing or invalid output, or accumulate dead-letter and incomplete work — or can none of that be determined honestly?

You declare the workflows you maintain for clients and feed normalized run observations plus source fetch attestations — from a heartbeat, a small script, or a read-only connector. The Actor applies deterministic rules with hysteresis and writes one status row per workflow plus a run envelope, ready for Apify schedules and platform webhooks.

Status vocabulary

StatusMeaning
okPositive trusted healthy evidence: a healthy run, an in-bounds in-progress run, a verified-idle declared trigger, or fully excused-and-attested windows — and every required check verifiable and clear.
staleExpected run provably absent or never completed: a trusted source's coverage attests the due window is empty, the platform attests the schedule cannot fire, or a run stayed running past maxRunSeconds with no finish. Unverifiable absence is unknown, never stale.
erroredProven violation: explicit run failure, invalid declared success signal, step-level error under a green run, output records missing required fields, executions past a declared bound, or a reported dead trigger.
empty-outputClaimed success whose output is empty or below the declared minimum.
dlq-growthDead-letter or incomplete-execution depth grew beyond the declared bound.
unknownEvidence missing, malformed, unauthorized, rate-limited, temporally ambiguous, contradictory, or truncated by caps. Never a healthy state.

Determinism and honesty guarantees

  • Reordering or duplicating observations does not change the verdicts.
  • All comparisons use explicit ISO-8601 instants in UTC; pin evaluationTime to replay a run bit-for-bit.
  • Known upstream failure, malformed input, or missing evidence never produces ok.
  • A failed fetch proves nothing — an empty result after an auth or parse error is unknown, never a healthy empty list.
  • Configurable hysteresis (config.hysteresis, overridable per workflow) prevents notification flapping across scheduled runs.

Input

The input is the canonical contract input (contracts/input.schema.json):

  • workflows (required): platform, id, label, schedule (deadline/interval/adhoc with optional trigger), optional sourceRef, activatedAt, pauses, outputContract (minOutputItems, requiredFields, successSignal), backlog, maxRunSeconds, maxRunsPerWindow, hysteresis, tags.
  • observations: normalized records (id, runId, workflow, observedAt, finishedAt, claimed, signals, output, queue, steps, error, provenance).
  • sources: fetch receipts (fetch.status, httpStatus, retries), coverage attestations (coverage.completeThrough — what licenses stale), and optional claims about upstream state (schedule enablement, trigger subscriptions, identity hints, per-run claims).
  • config: lookbackSeconds, defaultGraceSeconds, maxFutureSkewSeconds, hysteresis, limits — all resource caps.
  • evaluationTime: explicit "now" instant; pinned in tests for bit-identical replays, filled from the run start on schedules.
  • priorState: hysteresis state from the previous run's nextState (the Actor supplies it from the named state store).
  • Times: every instant field (observedAt, finishedAt, fetchedAt, evaluationTime, expectedAt, anchorAt, activatedAt, completeThrough, verifiedAt) is an ISO-8601 timestamp with seconds and an explicit UTC offset, e.g. 2026-09-12T09:05:00Z.

Output

  • Dataset — one canonical status row per workflow (contracts/status.schema.json): key, workflow, status, detection, times, reasons, evidence, hysteresis, provenance, contractVersion, ruleVersion.
  • OUTPUT record — the run-output envelope (contracts/run-output.schema.json): runStatus, counts, transitions, warnings, caps, nextState.
  • STATE record — the nextState snapshot persisted to the named store, consumed as priorState by the next scheduled run.

Sample rows

Status row (dataset item):

{
"contractVersion": "1.0.0",
"ruleVersion": "1.0.0",
"key": "n8n/acme-shopify-order-sync",
"workflow": { "platform": "n8n", "id": "acme-shopify-order-sync", "label": "Acme — Shopify order sync" },
"status": "empty-output",
"detection": "empty-output",
"times": {
"evaluatedAt": "2026-09-12T09:05:00Z",
"expectedAt": "2026-09-12T09:00:00Z",
"deadlineAt": "2026-09-12T09:15:00Z",
"observedAt": "2026-09-12T09:00:31Z"
},
"reasons": [
{ "rule": "output-empty", "status": "empty-output", "detail": "claimed success produced 0 items against minOutputItems 1" }
],
"evidence": {
"observations": [{ "id": "obs-1", "observedAt": "2026-09-12T09:00:31Z", "claimed": "success", "trust": "trusted", "output": { "itemCount": 0 } }],
"attestations": [{ "sourceRef": "src-n8n", "completeThrough": "2026-09-12T09:05:00Z", "trusted": true }],
"unverifiable": []
},
"hysteresis": { "stableStatus": "empty-output", "stableSince": "2026-09-12T09:05:00Z", "pendingStatus": null, "pendingCount": 0, "consecutiveDetections": 1 },
"provenance": { "generatedBy": "ahw-core", "contractVersion": "1.0.0", "ruleVersion": "1.0.0", "mechanisms": ["connector"], "connectors": [{ "name": "n8n-readonly", "version": "0.1.0" }] }
}

Run envelope (OUTPUT record):

{
"contractVersion": "1.0.0",
"ruleVersion": "1.0.0",
"runStatus": "ok",
"evaluatedAt": "2026-09-12T09:05:00Z",
"counts": { "ok": 9, "stale": 1, "errored": 1, "empty-output": 1, "dlq-growth": 1, "unknown": 1 },
"transitions": [
{ "key": "n8n/acme-shopify-order-sync", "from": "ok", "to": "empty-output", "kind": "failure-entered", "notify": true }
],
"warnings": [{ "code": "orphan-observation", "detail": "1 observation matched no declared workflow" }],
"caps": { "limits": { "maxObservations": 10000 }, "observed": { "workflows": 14, "observations": 96 }, "exceeded": [] },
"nextState": { "stateVersion": "1.0.0", "workflows": { "n8n/acme-shopify-order-sync": { "stableStatus": "empty-output", "stableSince": "2026-09-12T09:05:00Z", "pendingStatus": null, "pendingCount": 0, "consecutiveDetections": 1, "lastDetection": "empty-output", "lastEvaluatedAt": "2026-09-12T09:05:00Z" } } }
}

What this Actor does not do

  • It does not fix or rerun workflows; it reports their health.
  • Connectors are read-only when they ship. Nothing is posted to, changed in, or subscribed on your platforms, and credentials are held outside the payload — no token fields exist in this contract.
  • A platform without API coverage (or with an experimental, gated API) reports unknown for the affected workflows instead of guessing healthy.

Scheduling and webhooks

Attach an Apify schedule for a daily or hourly cycle. Platform run-webhooks POST a run-event payload (run ID, dataset and store IDs) to your n8n/Make/Zapier endpoint; the receiver fetches the dataset rows and OUTPUT envelope via the Apify API — rows are not inlined into the webhook body. The dataset rows and OUTPUT envelope are the stable contract for downstream automation.

Optionally, set the AHW_WEBHOOK_URL environment variable to have the Actor POST a small run summary (run status, per-status counts, capped transitions) to an HTTPS endpoint after each run. This is a bounded best-effort notification: targets are validated (public HTTPS hosts only), the payload is capped, delivery happens after all outputs are written, and a delivery failure never changes computed statuses. AHW_WEBHOOK_TOKEN sets an optional bearer token.

Supported producer quick start

The supported first-release producer is a customer-owned heartbeat or export job that submits one normalized observation per workflow run and one source receipt describing what time window was actually checked. Use a stable workflow.platform + workflow.id pair, a unique observation.id, and the same sourceRef on the workflow, observation, and source receipt. Only set fetch.status to ok and advance coverage.completeThrough after the producer has successfully read the complete declared window. Authentication, timeout, rate-limit, parse, or partial-export outcomes must use their matching failure status; the Actor will return unknown instead of inventing health.

Run the producer first, pass its normalized JSON to this Actor, then schedule both at the same cadence. Pin evaluationTime only for replay or testing. For production schedules, leave it unset so each Actor run uses its own start time. Keep the optional outbound webhook disabled during initial qualification.

Recovery and hysteresis example

Hysteresis is stateful, but uncertainty is fail-safe: a failed-source cycle demotes an inferred ok state to unknown immediately. With recoverAfter: 2, the first trustworthy healthy cycle then records pendingStatus: "ok" and pendingCount: 1 while the visible status remains unknown; the second consecutive healthy cycle returns the stable status to ok.

To recover from unreadable or incorrect state, stop overlapping writers, export the named store's STATE record, and replay the last trusted input with an explicit priorState. If the replay is accepted, write its nextState back or start a new uniquely named state store. Never share one stateStoreName between independent tasks or customers.

State, retention, and support

Use one unique stateStoreName and one writer per independently monitored fleet. Minimize uploaded workflow identifiers and observations. Export STATE before recovery work; delete the named key-value store when your retention period ends. Synthetic examples contain no customer data. Support owner: Mehdi Badawi through the Apify Store support channel, with an initial-response target of two business days.