Workflow Heartbeat Monitor - n8n, Make & Cron
Pricing
from $0.20 / 1,000 monitor auditeds
Workflow Heartbeat Monitor - n8n, Make & Cron
Detect workflows that never finish. Add one heartbeat at the final successful step, then audit n8n, Make, cron and Apify schedules for silence, failed completions and zero useful output.
Pricing
from $0.20 / 1,000 monitor auditeds
Rating
0.0
(0)
Developer
Vadim Bezrukov
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Your automation can fail without throwing an error.
Add one heartbeat at the final successful step and detect workflows that stop running, stop finishing, or start returning zero useful output. The check-in is a normal Actor run. A second scheduled run audits those heartbeats, and an Apify Monitoring alert on that audit emails you or posts to Slack when a workflow needs attention. This Actor does not receive public webhooks.
You can call it directly from MCP at https://mcp.apify.com?tools=automa-flow/workflow-heartbeat-monitor. An agent should ask it to record a final-step heartbeat for a named workflow, or to audit a known list and return workflows whose status is not HEALTHY. Execution uses your own Apify account. Anonymous MCP search can find the Actor; calling it requires an authenticated session.
What it detects
A workflow is unhealthy when the last stored heartbeat is missing, too old, reported ok: false, or produced less useful output than you required.
| Status | Reason code | Meaning |
|---|---|---|
HEALTHY | WITHIN_EXPECTED_WINDOW | The last heartbeat is recent enough and meets the count and success rules. |
LATE | HEARTBEAT_OVERDUE | The last heartbeat is older than maxSilenceMinutes. |
FAILED | LAST_CHECKIN_REPORTED_FAILURE | The last check-in said the final step did not succeed. |
ZERO_OUTPUT | LAST_CHECKIN_ZERO_OUTPUT | You required at least one useful record and the last count was 0. |
BELOW_MIN_COUNT | LAST_CHECKIN_BELOW_MIN_COUNT | The last count is below minCount, and this is not the zero-output case above. |
NEVER_SEEN | NO_HEARTBEAT_RECORDED | This key has no stored heartbeat. |
INVALID_STATE | STATE_MALFORMED, STATE_VERSION_UNSUPPORTED or HEARTBEAT_AFTER_AS_OF | The stored record cannot be judged. A missing count, when minCount is set, is STATE_MALFORMED. A heartbeat newer than a backdated asOf is HEARTBEAT_AFTER_AS_OF. |
NEVER_SEEN is a normal audit result. It is not a storage failure. A storage failure fails the run and never marks a monitor HEALTHY.
Why error monitoring is not enough
A workflow can accept a trigger, start without an exception, and still never complete the business task. Error alerts only see thrown errors. This Actor watches the final successful step: if that step does not check in, the next audit says so.
It is not an APM, a webhook gateway, or a guarantee that a workflow was delivered.
Quick start
- Check in. At the end of the workflow, after the useful work, run this Actor in
checkinmode. - Audit on a schedule. Save an
auditinput with"problemsOnly": trueas an Apify Task and schedule it. You do not need a second n8n or Make scenario for the audit. - Add the alert. On that Task, open Monitoring and create an alert: Metric
Number of results, Conditionis greater than0, turn on Email or Slack, then click Save & enable.
Steps 2 and 3 take about three minutes in Apify Console. Set up the scheduled audit walks through every click.
That is the whole setup. With problemsOnly on, a healthy audit writes no rows and succeeds quietly. When a workflow is late, failed or produced too little, the audit writes one row per problem, still succeeds, and the alert fires because the result count is above 0. The notification links to the run, where the Dataset lists each workflow, its status and the reason.
The audit run does not fail because a workflow is unhealthy. If you prefer a failed run as the alert, see Alerts.
The Console form prefills a demo check-in for daily-customer-import with ok set to true. That run succeeds and writes one Dataset row. An API call must send monitorKey and ok itself: leaving either out fails the run and stores nothing. Replace the demo key before you depend on it.
Check-in input:
{"mode": "checkin","monitorKey": "daily-customer-import","ok": true,"count": 142,"durationMs": 18750,"message": "Imported 142 customers"}
Audit input:
{"mode": "audit","monitors": [{"monitorKey": "daily-customer-import","maxSilenceMinutes": 90,"minCount": 1,"requireOk": true}],"problemsOnly": true}
A healthy audit row looks like this. You see it when problemsOnly is off; with it on, only rows whose status is not HEALTHY are written. previousSeenAt is null until a second check-in. The message text is not copied here.
{"recordType": "AUDIT","schemaVersion": 1,"monitorKey": "daily-customer-import","status": "HEALTHY","reasonCode": "WITHIN_EXPECTED_WINDOW","lastSeenAt": "2026-09-28T12:00:00.000Z","previousSeenAt": "2026-09-27T12:00:00.000Z","ageSeconds": 1250,"maxSilenceSeconds": 5400,"lastOk": true,"lastCount": 142,"lastDurationMs": 18750,"observedAt": "2026-09-28T12:20:50.000Z","source": "workflow-heartbeat","sourceId": "daily-customer-import","scrapedAt": "2026-09-28T12:20:50.000Z","fingerprint": "cd444a0fa72d0bf4e2d194040ba37d1c229644f070eaf26c5a8050d4081a62d9"}
The run summary is the key-value record RUN_SUMMARY. Read that before you page through the Dataset.
Set up the scheduled audit
You do this once, in Apify Console. It replaces a second n8n or Make scenario.
-
Fill in the audit. Open this Actor. Set Mode to Audit: evaluate saved heartbeats. Skip the Check-in section; an audit ignores it. In the Audit section, list your workflows in Monitors to audit. In the Alerts section, turn on Write rows only for workflows that need attention. Or switch the input to JSON and paste this, with your own keys:
{"mode": "audit","monitors": [{ "monitorKey": "customer-import-prod", "maxSilenceMinutes": 90, "minCount": 1 },{ "monitorKey": "nightly-export", "maxSilenceMinutes": 1500 }],"problemsOnly": true}Use exactly the
monitorKeyvalues your workflows send at check-in. -
Save it as a Task. Click Save as a new task at the top right of the page and give it a name such as
audit-workflow-heartbeats. -
Add the alert. On the new Task, open Monitoring and create an alert: Metric
Number of results, Conditionis greater than0, turn on Email or Slack, then click Save & enable. -
Test the alert once. Click Start on the Task before your workflows have checked in. Each workflow is
NEVER_SEEN, the run writes a row for each, and the alert should reach you within a few minutes. That proves the whole chain. It costs one start fee plus $0.0002 per workflow. -
Schedule it. Open Schedules, click Create new, and set how often to run, for example the cron expression
*/15 * * * *for every 15 minutes. Then click Add, choose a task, and pick the Task from step 2.
From then on the audit runs quietly while everything is healthy, and you hear about it only when a workflow needs attention. To change the list later, edit the Task input; the schedule and the alert stay as they are.
n8n example
On the last successful node, add an HTTP Request node. Store the Apify token in an n8n Header Auth credential, not in the workflow JSON or the URL: header name Authorization, value Bearer <your Apify token>.
POST https://api.apify.com/v2/acts/automa-flow~workflow-heartbeat-monitor/runs
{"mode": "checkin","monitorKey": "customer-import-prod","ok": true,"count": 142}
Then add customer-import-prod to the audit Task from Set up the scheduled audit. The audit should run more often than the silence window. With a 90 minute window, every 15 or 30 minutes is a practical choice: you only learn about a miss when the audit runs, so the schedule is the alert delay. A daily audit cannot notice a 90 minute miss until the next day.
If the workflow fails before this node, it never checks in. The audit then becomes LATE or stays NEVER_SEEN. That is the signal.
Make, cron and other tools
Every tool sends the same check-in: one POST to the run endpoint, with your token in the Authorization header.
- Make: add an HTTP module as the last step of the scenario, after the useful work. Method
POST, the URL above, headerAuthorization: Bearer <your Apify token>, JSON body withmode,monitorKey,okandcount. - Cron, a backend job or CI: run curl after the job's last successful command.
- Another Apify Actor: call the same endpoint from its last step, after it has saved its useful output, and send that output count as
count.
curl -X POST "https://api.apify.com/v2/acts/automa-flow~workflow-heartbeat-monitor/runs" \-H "Authorization: Bearer $APIFY_TOKEN" \-H "Content-Type: application/json" \-d "{\"mode\":\"checkin\",\"monitorKey\":\"nightly-export\",\"ok\":true,\"count\":18}"
Set APIFY_TOKEN from your own Apify account in the job's environment. Do not commit it or put it in the URL, where it can end up in logs.
The audit side is the same for every tool: add each monitorKey to the Task from Set up the scheduled audit.
Check-in mode
Required: monitorKey and ok. Neither has a default. The form prefills them for the demo, and a real workflow must send both. false means the final step reported failure. Omitting ok or monitorKey fails the run and does not store anything. mode is checkin when you leave it out and send no monitors.
Optional: count (useful records, integer from 0 to 1,000,000,000), durationMs (0 to 70 days), and message (at most 500 characters, no control characters).
monitorKey is trimmed. It may contain letters, numbers, spaces, and _ - . /. The raw key is not used as a storage key. The record id is heartbeat- plus a SHA-256 of the trimmed key.
The heartbeat time is the Actor clock. asOf is rejected on check-in. An explicit "mode": "checkin" that also includes monitors is rejected, so a wrong mode cannot silently skip an audit.
One Dataset row is written:
{"recordType": "CHECKIN","schemaVersion": 1,"monitorKey": "daily-customer-import","status": "RECORDED","reasonCode": null,"seenAt": "2026-09-28T12:00:00.000Z","ok": true,"count": 142,"durationMs": 18750,"totalCheckins": 1,"source": "workflow-heartbeat","sourceId": "daily-customer-import","scrapedAt": "2026-09-28T12:00:00.000Z","fingerprint": "cd444a0fa72d0bf4e2d194040ba37d1c229644f070eaf26c5a8050d4081a62d9"}
SUPERSEDED means a newer heartbeat was already stored, so this run did not replace it and did not bill. NOT_RECORDED means the stored record uses a schema version this build does not understand. It is left unchanged, nothing is billed, and the run fails after that row is saved so a green run is not mistaken for a new heartbeat.
A malformed record is replaced by the new check-in. When schemaVersion, lastSeenAt and lastOk can still be read, and the stored key matches, the new row keeps that time as previousSeenAt and increments totalCheckins. If those fields cannot be read, the replacement starts at 1. A stored time in the future is not chained, because it cannot be a previous observation.
Stored state:
{"schemaVersion": 1,"monitorKey": "daily-customer-import","lastSeenAt": "2026-09-28T12:00:00.000Z","previousSeenAt": "2026-09-27T12:00:00.000Z","lastOk": true,"lastCount": 142,"lastDurationMs": 18750,"lastMessage": "Imported 142 customers","totalCheckins": 41,"fingerprint": "cd444a0fa72d0bf4e2d194040ba37d1c229644f070eaf26c5a8050d4081a62d9"}
Two check-ins for the same key can race. The Actor keeps the newer timestamp and retries a short read-after-write so an older update does not cover a newer one. A lost race is not billed. The counter can skip a number if two writes overlap. Do not send two check-ins for the same key at the same moment if you need a perfect count.
Audit mode
Required: monitors (1 to 500). mode is audit whenever monitors is present, so you may leave it out. Each monitor needs monitorKey and maxSilenceMinutes (1 to 100800). Optional minCount and requireOk (default true).
problemsOnly defaults to false and only hides HEALTHY rows. Turn it on for a scheduled audit with an alert. Hidden healthy monitors are still evaluated and still billed. failRunOnProblem defaults to false; see Alerts.
asOf is an ISO-8601 time with a timezone. It only changes the age calculation against the current heartbeat. It does not replay history, and scrapedAt stays the real run time while observedAt is asOf. A value more than five minutes ahead of the run clock is rejected.
Rows are ordered by monitorKey. Audit never updates heartbeat timestamps.
Evaluation order for each monitor:
- No record:
NEVER_SEEN. - Malformed record, or a timestamp more than five minutes ahead of the run clock:
INVALID_STATE/STATE_MALFORMED. A valid heartbeat more than five minutes after a backdatedasOf:INVALID_STATE/HEARTBEAT_AFTER_AS_OF, because only the latest heartbeat is stored. - Unsupported
schemaVersion:INVALID_STATE/STATE_VERSION_UNSUPPORTED. requireOkandlastOkfalse:FAILED. This wins over lateness and over a low count.minCountis set andlastCountis missing:INVALID_STATE/STATE_MALFORMED.minCountis 1 andlastCountis 0:ZERO_OUTPUT.lastCountis belowminCount:BELOW_MIN_COUNT.- Age is strictly greater than the silence window:
LATE. - Otherwise:
HEALTHY.
Age equal to the window is still healthy. minCount 0 allows a stored count of 0. A future heartbeat of at most five minutes is treated as age 0.
Scheduling
Run the audit more often than maxSilenceMinutes. The silence window is how late the workflow may be. The schedule is how soon you hear about it.
A 15 minute workflow with a 20 minute window should be audited about every 5 minutes. An hourly workflow with a 90 minute window should be audited about every 15 to 30 minutes. Auditing faster than you need increases the number of monitor-audited events.
Check-in stays on the workflow. Audit stays on an Apify schedule.
Alerts
The Actor itself does not send email or Slack. Apify does, from the audit Task.
Recommended: a Monitoring alert on the result count.
- Save the audit as a Task with
"problemsOnly": true. - On the Task, open Monitoring and create an alert: Metric
Number of results, Conditionis greater than0. - Choose email or Slack.
| Audit result | Run status | Dataset rows | Alert |
|---|---|---|---|
| Every workflow healthy | Succeeded | none | no |
| One or more workflows need attention | Succeeded | one per problem | yes |
The status message says it in words, for example 2 of 5 workflows need attention. See the Dataset. or All 5 workflows are healthy. problemsOnly is on, so no rows were written. RUN_SUMMARY has the same count as problemCount.
Apify checks this kind of alert during the run and again when it finishes. If you forget problemsOnly, healthy rows count as results and the alert fires on every audit, so a missing setting is noisy, not silent.
Alternative: a failed run. Set "failRunOnProblem": true if you would rather alert on failed runs, for example with a failed-run notification or a webhook on ACTOR.RUN.FAILED. The audit saves its rows and RUN_SUMMARY, then fails. The default is off.
Either way, a charge limit, a storage failure, or an uncertain charge fails the run, because a partial audit must not look complete. A second Monitoring alert on the run status FAILED catches those.
Pricing
You pay per event. Platform usage is included, so there is no separate compute bill.
| Event | When it is charged | Price |
|---|---|---|
apify-actor-start | Once per run, check-in or audit | $0.0001 |
heartbeat-recorded | After the heartbeat is stored and the Dataset row is written | $0.001 |
monitor-audited | After that monitor is evaluated | $0.0002 |
Not charged: invalid input, a failed read or write, a superseded check-in, an unsupported stored version, and monitors that were not evaluated because the run charge limit was too low. NEVER_SEEN and INVALID_STATE are charged, because the audit did judge that monitor. problemsOnly does not reduce the audit charge. Dataset rows are not a second fee.
What one workflow costs, including the start fee of every run:
| Pattern | Runs per day | Cost per day | About per month |
|---|---|---|---|
| Daily check-in, audit every 6 hours | 1 check-in + 4 audits | $0.0023 | $0.07 |
| Hourly check-in, audit every 30 minutes | 24 check-ins + 48 audits | $0.0408 | $1.22 |
| 15 minute check-in, audit every 5 minutes | 96 check-ins + 288 audits | $0.192 | $5.76 |
One audit run can judge many workflows, so they share one start fee. Ten workflows in one audit every 30 minutes cost 48 x (10 x $0.0002 + $0.0001) = $0.1008 a day for the audits. Each extra workflow adds $0.0002 per run, not another start fee.
A missed daily import is worth more than about $1 a month of hourly monitoring. This is still more than a free ping service: Healthchecks.io, observed on 28 September 2026, offers 20 jobs on its free Hobbyist plan and 100 jobs at $20 a month on Business. Use this Actor when you want the success flag and the useful-output count judged inside Apify, next to the workflows you already run there.
One monitor-audited event is one workflow judged. One heartbeat-recorded event is one stored completion. The run summary reports billableEvents and chargedEvents.
For a local or non-PPE run, chargedEvents stays 0 and billableEvents still counts the work that would have been billed.
Set the run's maximum charge high enough for the whole audit: about 0.0002 times the number of monitors, plus the start event. If the cap is lower, the Actor evaluates a sorted prefix, writes those rows, and fails. notAudited counts the rest. They are not marked healthy.
Privacy
No website is contacted. The Actor stores only the monitor key, timestamps, the boolean result, the numeric count, the numeric duration, and the optional message. The message is kept on the heartbeat record and is not copied into audit rows or the run summary.
Do not put secrets, tokens, customer payloads, or personal data in message or monitorKey. There is no field for an arbitrary JSON payload, and unknown input fields are rejected.
Heartbeat records live in the named key-value store workflow-heartbeat-monitor until you delete that store. A later audit reads the same store. RUN_SUMMARY stays on the run's own default store, which is what the output link opens. The legal posture is low risk: this is your operational signal in your own storage, not a scrape of someone else's site.
Limits
- 1 to 500 monitors per audit.
- Monitor keys are ASCII, 1 to 100 characters after trimming.
- Silence window: 1 to 100800 minutes.
- Message: 500 characters, no control characters.
- Memory is 256 MB by default and cannot be set above 1024 MB, so the platform start event is charged once.
- The default run timeout is 900 seconds.
- There is no public webhook, standby server, dashboard, or external database.
- Simultaneous check-ins for one key can race. The newer timestamp is kept. The total can skip.
Failure semantics
no_results and source_failed are different outcomes. A missing heartbeat is NEVER_SEEN on that monitor and problems_found on the run. It is not source_failed.
| Situation | Run outcome | Billed |
|---|---|---|
| Check-in stored | ok | one heartbeat-recorded |
| Newer heartbeat already stored | ok, status SUPERSEDED | no |
| Unsupported stored version | not_recorded, run fails, status NOT_RECORDED | no |
| Audit, every judged monitor healthy | ok | one monitor-audited per monitor |
| Audit found unhealthy monitors (default) | problems_found, run succeeds | one event per judged monitor |
Same problems, failRunOnProblem true | problems_found, run fails after the rows are saved | one event per judged monitor |
| Key-value store or Dataset cannot be used | source_failed | only monitors whose rows were already saved |
| Charge limit stops the audit early | budget_exceeded | only the judged prefix |
| Charge result is unknown after a successful write | billing_uncertain | unknown, reconcile the run |
| Input is invalid | invalid_input | no |
Invalid mode, keys, counts, and silence windows that the input schema can express are rejected before a run starts. Identical monitor objects are rejected there too. A few rules can only be checked after start: audit requires monitors, the same monitorKey with different settings is rejected, check-in requires monitorKey and ok and rejects monitors and asOf, and asOf cannot be more than five minutes in the future. Those runs fail as invalid input and do not write a heartbeat.
FAQ
Is this a webhook monitor? No. Your workflow calls the Actor at the final step. Nothing is waiting on a public URL.
Does the audit change the heartbeat? No. Only checkin writes it.
A workflow is late, but the audit run succeeded. Is that right? Yes. The audit did its job: it found the problem and wrote a row for it. The alert comes from the Monitoring alert on the result count, not from a failed run. Turn on failRunOnProblem only if you want the run to fail instead.
How do I get an email or Slack message? Follow the three steps in Alerts. Without an alert, problems are only visible in the run's Dataset and status message.
What if the workflow errors before the last step? There is no new heartbeat. The audit becomes LATE, or NEVER_SEEN if it never succeeded once.
What if the final step runs and reports failure? Send ok: false. The audit returns FAILED, not LATE.
Can I monitor n8n, Make, cron, and another Actor the same way? Yes. Each one sends the same check-in JSON. See Make, cron and other tools.
For agents
Call this Actor to record that a workflow finished, or to judge a list of at most 500 monitor keys. Do not call it to deliver webhooks, page an on-call rotation, or inspect n8n internals.
The bill is one apify-actor-start ($0.0001) per run, plus one heartbeat-recorded ($0.001) for a stored check-in or one monitor-audited ($0.0002) per monitor actually judged. A check-in is not idempotent: retrying one that already succeeded stores and bills a second heartbeat. Read RUN_SUMMARY first. If problemCount is greater than zero, alert a person. If outcome is source_failed or budget_exceeded, do not treat the workflows as healthy and do not loop the audit. If notAudited is greater than zero, raise maxTotalChargeUsd and run the audit again.
Latest updates are in the CHANGELOG.md.