Automation Health Watchdog
Pricing
from $1.00 / 1,000 resolved workflow checks
Automation Health Watchdog
Detect missed runs, explicit failures, empty output, and dead-letter growth in workflow observations you supply. Returns auditable verdicts and explicit unknowns instead of false green status.
Pricing
from $1.00 / 1,000 resolved workflow checks
Rating
0.0
(0)
Developer
Mehdi Badawi
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 days ago
Last modified
Categories
Share
Supply workflow definitions, run observations, and source attestations to evaluate automation health. This Actor does not connect to n8n, Make, Zapier, or Apify accounts by itself. No input runs a clearly labeled synthetic demo.
The input and output below are the canonical contract surface (contracts/*.schema.json, contractVersion 1.0.0). Optional connectors are a documented extension point; the current Actor works on manually supplied observations and source attestations, and connector credentials never enter the input, output, or state payloads.
What it does
Answers, once per scheduled run: did each client workflow run when expected, fail explicitly, report success with missing or invalid output, or accumulate dead-letter and incomplete work — or can none of that be determined honestly?
You declare the workflows you maintain for clients and feed normalized run observations plus source fetch attestations — from a heartbeat, a small script, or a read-only connector. The Actor applies deterministic rules with hysteresis and writes one status row per workflow plus a run envelope, ready for Apify schedules and platform webhooks.
Status vocabulary
| Status | Meaning |
|---|---|
ok | Positive trusted healthy evidence: a healthy run, an in-bounds in-progress run, a verified-idle declared trigger, or fully excused-and-attested windows — and every required check verifiable and clear. |
stale | Expected run provably absent or never completed: a trusted source's coverage attests the due window is empty, the platform attests the schedule cannot fire, or a run stayed running past maxRunSeconds with no finish. Unverifiable absence is unknown, never stale. |
errored | Proven violation: explicit run failure, invalid declared success signal, step-level error under a green run, output records missing required fields, executions past a declared bound, or a reported dead trigger. |
empty-output | Claimed success whose output is empty or below the declared minimum. |
dlq-growth | Dead-letter or incomplete-execution depth grew beyond the declared bound. |
unknown | Evidence missing, malformed, unauthorized, rate-limited, temporally ambiguous, contradictory, or truncated by caps. Never a healthy state. |
Determinism and honesty guarantees
- Reordering or duplicating observations does not change the verdicts.
- All comparisons use explicit ISO-8601 instants in UTC; pin
evaluationTimeto replay a run bit-for-bit. - Known upstream failure, malformed input, or missing evidence never produces
ok. - A failed fetch proves nothing — an empty result after an auth or parse error is
unknown, never a healthy empty list. - Configurable hysteresis (
config.hysteresis, overridable per workflow) prevents notification flapping across scheduled runs.
Input
The input is the canonical contract input (contracts/input.schema.json):
workflows(required):platform,id,label,schedule(deadline/interval/adhocwith optionaltrigger), optionalsourceRef,activatedAt,pauses,outputContract(minOutputItems,requiredFields,successSignal),backlog,maxRunSeconds,maxRunsPerWindow,hysteresis,tags.observations: normalized records (id,runId,workflow,observedAt,finishedAt,claimed,signals,output,queue,steps,error,provenance).sources: fetch receipts (fetch.status,httpStatus,retries), coverage attestations (coverage.completeThrough— what licensesstale), and optionalclaimsabout upstream state (schedule enablement, trigger subscriptions, identity hints, per-run claims).config:lookbackSeconds,defaultGraceSeconds,maxFutureSkewSeconds,hysteresis,limits— all resource caps.evaluationTime: explicit "now" instant; pinned in tests for bit-identical replays, filled from the run start on schedules.priorState: hysteresis state from the previous run'snextState(the Actor supplies it from the named state store).- Times: every instant field (
observedAt,finishedAt,fetchedAt,evaluationTime,expectedAt,anchorAt,activatedAt,completeThrough,verifiedAt) is an ISO-8601 timestamp with seconds and an explicit UTC offset, e.g.2026-09-12T09:05:00Z.
Output
- Dataset — one canonical status row per workflow (
contracts/status.schema.json):key,workflow,status,detection,times,reasons,evidence,hysteresis,provenance,contractVersion,ruleVersion. - OUTPUT record — the run-output envelope (
contracts/run-output.schema.json):runStatus,counts,transitions,warnings,caps,nextState. - STATE record — the
nextStatesnapshot persisted to the named store, consumed aspriorStateby the next scheduled run.
Sample rows
Status row (dataset item):
{"contractVersion": "1.0.0","ruleVersion": "1.0.0","key": "n8n/acme-shopify-order-sync","workflow": { "platform": "n8n", "id": "acme-shopify-order-sync", "label": "Acme — Shopify order sync" },"status": "empty-output","detection": "empty-output","times": {"evaluatedAt": "2026-09-12T09:05:00Z","expectedAt": "2026-09-12T09:00:00Z","deadlineAt": "2026-09-12T09:15:00Z","observedAt": "2026-09-12T09:00:31Z"},"reasons": [{ "rule": "output-empty", "status": "empty-output", "detail": "claimed success produced 0 items against minOutputItems 1" }],"evidence": {"observations": [{ "id": "obs-1", "observedAt": "2026-09-12T09:00:31Z", "claimed": "success", "trust": "trusted", "output": { "itemCount": 0 } }],"attestations": [{ "sourceRef": "src-n8n", "completeThrough": "2026-09-12T09:05:00Z", "trusted": true }],"unverifiable": []},"hysteresis": { "stableStatus": "empty-output", "stableSince": "2026-09-12T09:05:00Z", "pendingStatus": null, "pendingCount": 0, "consecutiveDetections": 1 },"provenance": { "generatedBy": "ahw-core", "contractVersion": "1.0.0", "ruleVersion": "1.0.0", "mechanisms": ["connector"], "connectors": [{ "name": "n8n-readonly", "version": "0.1.0" }] }}
Run envelope (OUTPUT record):
{"contractVersion": "1.0.0","ruleVersion": "1.0.0","runStatus": "ok","evaluatedAt": "2026-09-12T09:05:00Z","counts": { "ok": 9, "stale": 1, "errored": 1, "empty-output": 1, "dlq-growth": 1, "unknown": 1 },"transitions": [{ "key": "n8n/acme-shopify-order-sync", "from": "ok", "to": "empty-output", "kind": "failure-entered", "notify": true }],"warnings": [{ "code": "orphan-observation", "detail": "1 observation matched no declared workflow" }],"caps": { "limits": { "maxObservations": 10000 }, "observed": { "workflows": 14, "observations": 96 }, "exceeded": [] },"nextState": { "stateVersion": "1.0.0", "workflows": { "n8n/acme-shopify-order-sync": { "stableStatus": "empty-output", "stableSince": "2026-09-12T09:05:00Z", "pendingStatus": null, "pendingCount": 0, "consecutiveDetections": 1, "lastDetection": "empty-output", "lastEvaluatedAt": "2026-09-12T09:05:00Z" } } }}
What this Actor does not do
- It does not fix or rerun workflows; it reports their health.
- Connectors are read-only when they ship. Nothing is posted to, changed in, or subscribed on your platforms, and credentials are held outside the payload — no token fields exist in this contract.
- A platform without API coverage (or with an experimental, gated API) reports
unknownfor the affected workflows instead of guessing healthy.
Scheduling and webhooks
Attach an Apify schedule for a daily or hourly cycle. Platform run-webhooks POST a run-event payload (run ID, dataset and store IDs) to your n8n/Make/Zapier endpoint; the receiver fetches the dataset rows and OUTPUT envelope via the Apify API — rows are not inlined into the webhook body. The dataset rows and OUTPUT envelope are the stable contract for downstream automation.
Optionally, set the AHW_WEBHOOK_URL environment variable to have the Actor POST a small run summary (run status, per-status counts, capped transitions) to an HTTPS endpoint after each run. This is a bounded best-effort notification: targets are validated (public HTTPS hosts only), the payload is capped, delivery happens after all outputs are written, and a delivery failure never changes computed statuses. AHW_WEBHOOK_TOKEN sets an optional bearer token.
Supported producer quick start
The supported first-release producer is a customer-owned heartbeat or export
job that submits one normalized observation per workflow run and one source
receipt describing what time window was actually checked. Use a stable
workflow.platform + workflow.id pair, a unique observation.id, and the
same sourceRef on the workflow, observation, and source receipt. Only set
fetch.status to ok and advance coverage.completeThrough after the producer
has successfully read the complete declared window. Authentication, timeout,
rate-limit, parse, or partial-export outcomes must use their matching failure
status; the Actor will return unknown instead of inventing health.
Run the producer first, pass its normalized JSON to this Actor, then schedule
both at the same cadence. Pin evaluationTime only for replay or testing. For
production schedules, leave it unset so each Actor run uses its own start time.
Keep the optional outbound webhook disabled during initial qualification.
Recovery and hysteresis example
Hysteresis is stateful, but uncertainty is fail-safe: a failed-source cycle
demotes an inferred ok state to unknown immediately. With recoverAfter: 2,
the first trustworthy healthy cycle then records pendingStatus: "ok" and
pendingCount: 1 while the visible status remains unknown; the second
consecutive healthy cycle returns the stable status to ok.
To recover from unreadable or incorrect state, stop overlapping writers,
export the named store's STATE record, and replay the last trusted input with
an explicit priorState. If the replay is accepted, write its nextState back
or start a new uniquely named state store. Never share one stateStoreName
between independent tasks or customers.
State, retention, and support
Use one unique stateStoreName and one writer per independently monitored
fleet. Minimize uploaded workflow identifiers and observations. Export STATE
before recovery work; delete the named key-value store when your retention
period ends. Synthetic examples contain no customer data. Support owner: Mehdi
Badawi through the Apify Store support channel, with an initial-response target
of two business days.