Monitoring failed workflows in Airflow, Prefect, Dagster, Temporal, Hatchet, Inngest and Trigger.dev
Every one of these orchestrators can tell you about a run that failed. None of them can tell you, from inside a run, about the run that never started: the worker was down, the schedule was paused, or a deploy dropped the job. Their failure hooks run as part of a run, and their built-in alerts depend on their own scheduler or control plane. An external dead-man's-switch keeps its own copy of the schedule and alerts when the expected check-in does not arrive, whatever the cause.
Your assistant can do this for you
Monitor my scheduled workflows with LastPing
On this page
At a glance
What each one reports natively, and the run it cannot report from inside.
| Orchestrator | Native failure signal | Cannot report from inside | Guide |
|---|---|---|---|
| Apache Airflow | on_failure_callback and on_success_callback (Dag or task); Deadline Alerts (3.1, experimental) | Paused Dag, scheduler down, Dag file removed | Airflow |
| Prefect | State hooks (on_failure, on_crashed); automations, including proactive triggers | A run that never starts (hooks); an outage of the Prefect server itself (automations) | Prefect |
| Dagster | Run status sensors and hooks; Dagster+ alert policies | dagster-daemon down, which also stops the sensors | Dagster |
| Temporal | Failed status in the UI; SDK and Cloud metrics for your own alerts | No worker polling the task queue; a paused Schedule | Temporal |
| Hatchet | On-failure task; tenant alerting to Slack and email | No worker running; a paused workflow dropping cron ticks | Hatchet |
| Inngest | onFailure handler, inngest/function.failed event; metrics and a Datadog integration | Paused function, function dropped at sync, app unreachable | Inngest |
| Trigger.dev | Run-failed alerts (email, Slack, webhook); onFailure hook | Task not in the current deployment; deactivated schedule | Trigger.dev |
Orchestrator by orchestrator
Native alerting, the blind spot, and the fix. Each heading links to the full setup.
Apache Airflow
Native: on_success_callback and on_failure_callback at Dag or task level, plus task-only callbacks for retries, execution start and skips. In Airflow 3, Dag-level callbacks run in the Dag processor. Deadline Alerts, new in Airflow 3.1 and marked experimental, fire when a Dag run exceeds a time threshold.
Blind spot: paused Dags are not scheduled by the scheduler, and a Dag file removed from the Dags folder is marked deactivated. Callbacks and deadlines belong to Dag runs, so when no Dag run is created there is nothing for them to fire on.
Fix: ping LastPing from on_success_callback, POST to /fail from on_failure_callback, and give the monitor the Dag's schedule.
Prefect
Native: state change hooks (on_completion, on_failure, on_crashed, on_cancellation, on_running) run client side, in the same process as the flow. Automations, in Prefect Cloud and in a Prefect server you run, can notify on state changes and can be proactive: they fire when an expected event does not happen, such as a flow staying Running for too long. The server marks a scheduled run Late when no worker picks it up, and a worker that misses three heartbeats is considered offline.
Blind spot: Prefect's own docs say hook execution cannot be guaranteed, and a hook cannot fire for a run that never starts. Automations live on the Prefect server, so they cannot report that server being down, and a paused or cleared deployment schedule creates no runs to become Late.
Fix: on_completion pings, on_failure and on_crashed ping /fail, and the monitor carries the deployment's schedule.
Dagster
Native: success and failure hooks on jobs, and run status sensors that can post to Slack or another service on a failed run. Dagster+ adds alert policies for run failures, schedule and sensor tick failures, code locations that fail to load, and Hybrid agents that stop heartbeating.
Blind spot: the dagster-daemon creates runs from schedules and sensors. If it stops, schedules stop creating runs, and the sensors that would send the alert are not evaluated either. The daemon's heartbeat shows on the Daemons tab, which someone has to look at.
Fix: a success_hook pings, a failure_hook pings /fail, and the monitor carries the schedule's cron.
Temporal
Native: failed Workflow Executions show in the Temporal UI. Temporal Cloud's notifications cover platform incidents, expiring credentials, billing and failovers; for workflow and worker health its docs point to metrics such as schedule-to-start latency and task queue backlog, alerted on in your own observability tool.
Blind spot: if no Worker polls the task queue, the Workflow Execution exists but makes no progress, and no workflow code runs to report it. A paused Schedule starts nothing.
Fix: run a small ping activity at the end of the workflow and a /fail activity on exception, with the monitor on the same schedule.
Hatchet
Native: an on-failure task runs as the last step of a workflow with a failed task, and the engine has tenant alerting settings for Slack and email. Tasks that wait in the queue past their schedule_timeout (five minutes by default) are cancelled.
Blind spot: the on-failure task is a task, so it needs a worker. With no worker running nothing executes, and a paused workflow set to drop cron ticks creates no run at all.
Fix: an on-success task pings, an on-failure task POSTs the errors to /fail, and the monitor carries the workflow's cron.
Inngest
Native: an onFailure handler runs after a function exhausts its retries, and an inngest/function.failed system event lets one function watch every failure. The dashboard shows failure rate and backlog per function, and the Datadog integration can alert on those metrics.
Blind spot: the failure handler is another function in your app, so it cannot run when Inngest cannot reach your app. A paused function or an archived app starts no new runs, and a function dropped from your code is gone after the next sync.
Fix: a final step.run pings, onFailure pings /fail, and the monitor carries the function's cron.
Trigger.dev
Native: alerts by email, Slack or webhook when a run fails or a deployment fails, and onSuccess and onFailure lifecycle hooks available on every task.
Blind spot: Trigger.dev's docs say a scheduled task in Staging or Production only triggers if it is in the current deployment, and in Dev only while the dev CLI runs. A deactivated schedule creates no runs.
Fix: onSuccess pings, onFailure POSTs the error to /fail, and the monitor carries the task's cron.
How the external check covers it
Same three parts on every orchestrator.
A monitor with the same schedule
Create a LastPing monitor with the job's cron and a grace window of its expected runtime plus a margin. LastPing now knows when a check-in is due, independently of the orchestrator.
A ping on success, a /fail on failure
From the orchestrator's own success and failure hooks, GET the ping URL on success and POST the error to /fail on failure. The failure pages immediately, with the error stored on the ping.
GET https://ping.lastping.dev/<your-monitor-id> # success
POST https://ping.lastping.dev/<your-monitor-id>/fail # failure, body = error
Silence is the alert
If no ping arrives inside the grace window, LastPing opens an incident. It does not need to know why: a dead worker, a paused schedule and a deploy that dropped the job all look the same from outside, and all of them alert. The next successful run closes the incident.
Questions people ask
Front-loaded answers: the most important fact first.
-
Which orchestrator has the best built-in failure alerting?
For runs that happened, most are adequate: Trigger.dev, Dagster+ and the Hatchet engine have alerting in the product, Prefect has automations, Airflow and Inngest give you failure callbacks and handlers to send your own, and Temporal points you to metrics in your own observability tool. The difference that matters is the run that never starts. Prefect's proactive automation triggers and Dagster+ agent and code location alerts cover some of it, but each still depends on the orchestrator's own control plane being up.
-
Why can't a failure hook report a run that never started?
A failure hook or callback is code that runs as part of a run, on a worker or in a process the orchestrator manages. If the worker is down, the schedule is paused, or the deployment no longer contains the job, there is no run, so there is nothing for the hook to be part of.
-
What does an external dead-man's-switch add?
It runs outside the orchestrator and keeps its own copy of the schedule. Each successful run sends a check-in; when one is missing after the grace window, it opens an incident, whatever the reason the run did not happen. Failures can also be sent to its
/failpath so they page immediately. -
Do I still need the orchestrator's own alerts?
They complement each other. The orchestrator knows which task failed and why; the external check knows a run was expected and did not arrive. The guides on this page send the failure detail to LastPing from the orchestrator's own failure hook, so one incident carries both.
-
Is this free?
Yes. LastPing is free for individuals. See vs Healthchecks.io and vs Cronitor.
The run that never started is the one nobody reports.
A matching schedule and two hooks, on any orchestrator. First alert in under a minute, and LastPing monitors itself the same way.