# Automation Health Watchdog (`mehdi_badawi/automation-health-watchdog`) Actor

Detect missed runs, explicit failures, empty output, and dead-letter growth in workflow observations you supply. Returns auditable verdicts and explicit unknowns instead of false green status.

- **URL**: https://apify.com/mehdi\_badawi/automation-health-watchdog.md
- **Developed by:** [Mehdi Badawi](https://apify.com/mehdi_badawi) (community)
- **Categories:** Automation, Developer tools, Integrations
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 resolved workflow checks

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Automation Health Watchdog

Supply workflow definitions, run observations, and source attestations to
evaluate automation health. This Actor does not connect to n8n, Make, Zapier,
or Apify accounts by itself. No input runs a clearly labeled synthetic demo.

The input and output below are the canonical contract surface (`contracts/*.schema.json`, `contractVersion` 1.0.0). Optional connectors are a documented extension point; the current Actor works on manually supplied observations and source attestations, and connector credentials never enter the input, output, or state payloads.

### What it does

Answers, once per scheduled run: did each client workflow run when expected, fail explicitly, report success with missing or invalid output, or accumulate dead-letter and incomplete work — or can none of that be determined honestly?

You declare the workflows you maintain for clients and feed normalized run observations plus source fetch attestations — from a heartbeat, a small script, or a read-only connector. The Actor applies deterministic rules with hysteresis and writes one status row per workflow plus a run envelope, ready for Apify schedules and platform webhooks.

### Status vocabulary

| Status | Meaning |
|---|---|
| `ok` | Positive trusted healthy evidence: a healthy run, an in-bounds in-progress run, a verified-idle declared trigger, or fully excused-and-attested windows — and every required check verifiable and clear. |
| `stale` | Expected run provably absent or never completed: a trusted source's coverage attests the due window is empty, the platform attests the schedule cannot fire, or a run stayed `running` past `maxRunSeconds` with no finish. Unverifiable absence is `unknown`, never `stale`. |
| `errored` | Proven violation: explicit run failure, invalid declared success signal, step-level error under a green run, output records missing required fields, executions past a declared bound, or a reported dead trigger. |
| `empty-output` | Claimed success whose output is empty or below the declared minimum. |
| `dlq-growth` | Dead-letter or incomplete-execution depth grew beyond the declared bound. |
| `unknown` | Evidence missing, malformed, unauthorized, rate-limited, temporally ambiguous, contradictory, or truncated by caps. Never a healthy state. |

### Determinism and honesty guarantees

- Reordering or duplicating observations does not change the verdicts.
- All comparisons use explicit ISO-8601 instants in UTC; pin `evaluationTime` to replay a run bit-for-bit.
- Known upstream failure, malformed input, or missing evidence never produces `ok`.
- A failed fetch proves nothing — an empty result after an auth or parse error is `unknown`, never a healthy empty list.
- Configurable hysteresis (`config.hysteresis`, overridable per workflow) prevents notification flapping across scheduled runs.

### Input

The input is the canonical contract input (`contracts/input.schema.json`):

- **`workflows`** (required): `platform`, `id`, `label`, `schedule` (`deadline`/`interval`/`adhoc` with optional `trigger`), optional `sourceRef`, `activatedAt`, `pauses`, `outputContract` (`minOutputItems`, `requiredFields`, `successSignal`), `backlog`, `maxRunSeconds`, `maxRunsPerWindow`, `hysteresis`, `tags`.
- **`observations`**: normalized records (`id`, `runId`, `workflow`, `observedAt`, `finishedAt`, `claimed`, `signals`, `output`, `queue`, `steps`, `error`, `provenance`).
- **`sources`**: fetch receipts (`fetch.status`, `httpStatus`, `retries`), coverage attestations (`coverage.completeThrough` — what licenses `stale`), and optional `claims` about upstream state (schedule enablement, trigger subscriptions, identity hints, per-run claims).
- **`config`**: `lookbackSeconds`, `defaultGraceSeconds`, `maxFutureSkewSeconds`, `hysteresis`, `limits` — all resource caps.
- **`evaluationTime`**: explicit "now" instant; pinned in tests for bit-identical replays, filled from the run start on schedules.
- **`priorState`**: hysteresis state from the previous run's `nextState` (the Actor supplies it from the named state store).
- **Times**: every instant field (`observedAt`, `finishedAt`, `fetchedAt`, `evaluationTime`, `expectedAt`, `anchorAt`, `activatedAt`, `completeThrough`, `verifiedAt`) is an ISO-8601 timestamp with seconds and an explicit UTC offset, e.g. `2026-09-12T09:05:00Z`.

### Output

- **Dataset** — one canonical status row per workflow (`contracts/status.schema.json`): `key`, `workflow`, `status`, `detection`, `times`, `reasons`, `evidence`, `hysteresis`, `provenance`, `contractVersion`, `ruleVersion`.
- **OUTPUT record** — the run-output envelope (`contracts/run-output.schema.json`): `runStatus`, `counts`, `transitions`, `warnings`, `caps`, `nextState`.
- **STATE record** — the `nextState` snapshot persisted to the named store, consumed as `priorState` by the next scheduled run.

### Sample rows

Status row (dataset item):

```json
{
    "contractVersion": "1.0.0",
    "ruleVersion": "1.0.0",
    "key": "n8n/acme-shopify-order-sync",
    "workflow": { "platform": "n8n", "id": "acme-shopify-order-sync", "label": "Acme — Shopify order sync" },
    "status": "empty-output",
    "detection": "empty-output",
    "times": {
        "evaluatedAt": "2026-09-12T09:05:00Z",
        "expectedAt": "2026-09-12T09:00:00Z",
        "deadlineAt": "2026-09-12T09:15:00Z",
        "observedAt": "2026-09-12T09:00:31Z"
    },
    "reasons": [
        { "rule": "output-empty", "status": "empty-output", "detail": "claimed success produced 0 items against minOutputItems 1" }
    ],
    "evidence": {
        "observations": [{ "id": "obs-1", "observedAt": "2026-09-12T09:00:31Z", "claimed": "success", "trust": "trusted", "output": { "itemCount": 0 } }],
        "attestations": [{ "sourceRef": "src-n8n", "completeThrough": "2026-09-12T09:05:00Z", "trusted": true }],
        "unverifiable": []
    },
    "hysteresis": { "stableStatus": "empty-output", "stableSince": "2026-09-12T09:05:00Z", "pendingStatus": null, "pendingCount": 0, "consecutiveDetections": 1 },
    "provenance": { "generatedBy": "ahw-core", "contractVersion": "1.0.0", "ruleVersion": "1.0.0", "mechanisms": ["connector"], "connectors": [{ "name": "n8n-readonly", "version": "0.1.0" }] }
}
```

Run envelope (OUTPUT record):

```json
{
    "contractVersion": "1.0.0",
    "ruleVersion": "1.0.0",
    "runStatus": "ok",
    "evaluatedAt": "2026-09-12T09:05:00Z",
    "counts": { "ok": 9, "stale": 1, "errored": 1, "empty-output": 1, "dlq-growth": 1, "unknown": 1 },
    "transitions": [
        { "key": "n8n/acme-shopify-order-sync", "from": "ok", "to": "empty-output", "kind": "failure-entered", "notify": true }
    ],
    "warnings": [{ "code": "orphan-observation", "detail": "1 observation matched no declared workflow" }],
    "caps": { "limits": { "maxObservations": 10000 }, "observed": { "workflows": 14, "observations": 96 }, "exceeded": [] },
    "nextState": { "stateVersion": "1.0.0", "workflows": { "n8n/acme-shopify-order-sync": { "stableStatus": "empty-output", "stableSince": "2026-09-12T09:05:00Z", "pendingStatus": null, "pendingCount": 0, "consecutiveDetections": 1, "lastDetection": "empty-output", "lastEvaluatedAt": "2026-09-12T09:05:00Z" } } }
}
```

### What this Actor does not do

- It does not fix or rerun workflows; it reports their health.
- Connectors are read-only when they ship. Nothing is posted to, changed in, or subscribed on your platforms, and credentials are held outside the payload — no token fields exist in this contract.
- A platform without API coverage (or with an experimental, gated API) reports `unknown` for the affected workflows instead of guessing healthy.

### Scheduling and webhooks

Attach an Apify schedule for a daily or hourly cycle. Platform run-webhooks POST a run-event payload (run ID, dataset and store IDs) to your n8n/Make/Zapier endpoint; the receiver fetches the dataset rows and OUTPUT envelope via the Apify API — rows are not inlined into the webhook body. The dataset rows and OUTPUT envelope are the stable contract for downstream automation.

Optionally, set the `AHW_WEBHOOK_URL` environment variable to have the Actor POST a small run summary (run status, per-status counts, capped transitions) to an HTTPS endpoint after each run. This is a bounded best-effort notification: targets are validated (public HTTPS hosts only), the payload is capped, delivery happens after all outputs are written, and a delivery failure never changes computed statuses. `AHW_WEBHOOK_TOKEN` sets an optional bearer token.

### Supported producer quick start

The supported first-release producer is a customer-owned heartbeat or export
job that submits one normalized `observation` per workflow run and one `source`
receipt describing what time window was actually checked. Use a stable
`workflow.platform` + `workflow.id` pair, a unique `observation.id`, and the
same `sourceRef` on the workflow, observation, and source receipt. Only set
`fetch.status` to `ok` and advance `coverage.completeThrough` after the producer
has successfully read the complete declared window. Authentication, timeout,
rate-limit, parse, or partial-export outcomes must use their matching failure
status; the Actor will return `unknown` instead of inventing health.

Run the producer first, pass its normalized JSON to this Actor, then schedule
both at the same cadence. Pin `evaluationTime` only for replay or testing. For
production schedules, leave it unset so each Actor run uses its own start time.
Keep the optional outbound webhook disabled during initial qualification.

### Recovery and hysteresis example

Hysteresis is stateful, but uncertainty is fail-safe: a failed-source cycle
demotes an inferred `ok` state to `unknown` immediately. With `recoverAfter: 2`,
the first trustworthy healthy cycle then records `pendingStatus: "ok"` and
`pendingCount: 1` while the visible status remains `unknown`; the second
consecutive healthy cycle returns the stable status to `ok`.

To recover from unreadable or incorrect state, stop overlapping writers,
export the named store's `STATE` record, and replay the last trusted input with
an explicit `priorState`. If the replay is accepted, write its `nextState` back
or start a new uniquely named state store. Never share one `stateStoreName`
between independent tasks or customers.

### State, retention, and support

Use one unique `stateStoreName` and one writer per independently monitored
fleet. Minimize uploaded workflow identifiers and observations. Export `STATE`
before recovery work; delete the named key-value store when your retention
period ends. Synthetic examples contain no customer data. Support owner: Mehdi
Badawi through the Apify Store support channel, with an initial-response target
of two business days.

# Actor input Schema

## `contractVersion` (type: `string`):

Must equal 1.0.0 when provided. The Actor stamps the current contract version when omitted.

## `evaluationTime` (type: `string`):

ISO-8601 instant treated as 'now' for all comparisons, with seconds and an explicit offset (e.g. 2026-09-12T09:05:00Z). The contract requires an explicit instant; leave null in saved schedule inputs so the Actor fills in each run's own start instant. Pin this in tests and replays for deterministic output.

## `workflows` (type: `array`):

Client workflows to evaluate, in the canonical contract shape. Each entry needs platform, id, label, and a schedule.

## `sources` (type: `array`):

Optional fetch receipts and coverage attestations in the canonical shape {id, platform, mechanism?, connector?, fetchedAt?, fetch{status, detail?, httpStatus?, retries?, uri?}, coverage?{completeThrough, coversAllWorkflows?, workflowKeys?, returnedCount?}, claims?{workflowStates, runClaims}}. fetch.status is ok|unauthorized|rate\_limited|timeout|malformed|unavailable; only status 'ok' attests absence. Claims attest upstream state (scheduleEnabled, trigger subscription, identity hints, per-run claimed status).

## `observations` (type: `array`):

Normalized run records in the canonical shape {id, runId?, workflow{platform,id}, sourceRef?, observedAt, finishedAt?, claimed, signals?, output?{itemCount, byteSize?, digest?, sample?}, queue?{dlqCount, incompleteCount}, steps?, error?, outputTruncated?, provenance?{mechanism, connector?, retrievedAt?, uri?, note?, requestId?}}. claimed is success|failure|running|unknown. Order does not matter; malformed entries are quarantined as evidence, never silently healthy.

## `config` (type: `object`):

Optional knobs {lookbackSeconds, defaultGraceSeconds, maxFutureSkewSeconds, hysteresis{enterAfter, recoverAfter}, limits{maxWorkflows, maxObservations, maxObservationsPerWorkflow, maxInputBytes, maxConnectorTimeoutMs, maxConnectorRetries, maxReasonsPerRow, maxStringLength, maxSignalsFields}}. Limits may only tighten the contract caps; out-of-range values fail the run.

## `priorState` (type: `object`):

Optional hysteresis state blob {stateVersion, workflows{<key>:{stableStatus, stableSince?, pendingStatus, pendingCount, consecutiveDetections, lastDetection, lastEvaluatedAt, lastTrustedDlqCount?, lastTrustedIncompleteCount?}}} as returned by a previous run's nextState. The Actor normally supplies this from the named state store; pass it explicitly to replay or seed state.

## `stateStoreName` (type: `string`):

Actor adapter setting (not part of the deterministic core input): name of the Apify key-value store used to persist nextState between scheduled runs. Set to null to disable persistence (each run then evaluates hysteresis from scratch, which can re-alert). Must be unique per task or schedule: two tasks sharing a store contaminate each other's hysteresis state.

## Actor input object example

```json
{
  "contractVersion": "1.0.0",
  "sources": [],
  "observations": [],
  "stateStoreName": "automation-health-watchdog-state"
}
```

# Actor output Schema

## `statusRows` (type: `string`):

One dataset item per declared workflow: canonical key, workflow identity, status, detection, times, reasons, evidence, hysteresis, provenance, rule/contract versions.

## `runOutput` (type: `string`):

The canonical run-output envelope: runStatus, per-status counts, transitions, warnings, caps audit, and nextState.

## `hysteresisState` (type: `string`):

Debug snapshot of nextState — the per-workflow hysteresis state persisted to the named state store and consumed as priorState by the next run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("mehdi_badawi/automation-health-watchdog").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("mehdi_badawi/automation-health-watchdog").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call mehdi_badawi/automation-health-watchdog --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mehdi_badawi/automation-health-watchdog"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XsOxvlVZesbV4tFNs/builds/xJ9MdYYs7P3bE2JWd/openapi.json
