Skip to content
LastPing
GUIDE · HATCHET · WORKFLOWS · FREE

Monitor Hatchet workflows: alerts for failed, stalled and missed runs

Hatchet retries failed tasks, cancels tasks that wait too long in the queue, and can run an on-failure task when a workflow fails. All of that runs inside Hatchet, and the on-failure task needs a worker like any other task. If no worker is running, or the workflow is paused and its cron ticks are dropped, nothing runs and nothing reports it. Add an on-success and an on-failure task that ping LastPing, give the monitor the same cron, and a run that fails, hangs or never starts is reported from outside Hatchet.

Your assistant can do this for you

Monitor my Hatchet workflows with LastPing

On this page
  1. Setup: an on-success and an on-failure task
  2. What this catches
  3. Questions people ask

Setup: an on-success and an on-failure task

Ping on success, POST the errors to /fail on failure. A matching monitor schedule catches the run that never happened.

Create a heartbeat monitor in LastPing

In the console, create a monitor with the same cron as your Hatchet workflow. Hatchet's cron is the time it enqueues the run, not the time the run starts, so set the grace window to queue time plus the workflow's expected runtime plus a margin. Copy the ping URL.

https://ping.lastping.dev/<your-monitor-id>

Python: on_success_task and on_failure_task

An on-success task runs as the last step of a workflow whose tasks all succeeded; an on-failure task runs as the last step of a workflow with at least one failed task, and ctx.task_run_errors maps each failed task to its error message:

from datetime import timedelta

import httpx
from hatchet_sdk import Context, EmptyModel, Hatchet

hatchet = Hatchet()
PING = "https://ping.lastping.dev/<your-monitor-id>"

nightly = hatchet.workflow(name="nightly-etl", on_crons=["0 3 * * *"])

@nightly.task(execution_timeout=timedelta(minutes=30))
def extract(input: EmptyModel, ctx: Context) -> dict[str, str]:
    ...   # your work
    return {"status": "done"}

@nightly.on_success_task()
def ping_success(input: EmptyModel, ctx: Context) -> None:
    httpx.get(PING, timeout=10)

@nightly.on_failure_task()
def ping_failure(input: EmptyModel, ctx: Context) -> None:
    httpx.post(f"{PING}/fail", content=str(ctx.task_run_errors)[:10_000], timeout=10)

Register nightly on your worker as usual. The POSTed body (up to 10,000 bytes) is stored with the fail ping, so the incident carries the error.

TypeScript: onSuccess and onFailure

The TypeScript SDK has the same pair on a workflow declaration; ctx.errors() returns the upstream errors inside the failure task:

import { hatchet } from './hatchet-client';

const PING = 'https://ping.lastping.dev/<your-monitor-id>';

export const nightly = hatchet.workflow({
  name: 'nightly-etl',
  on: { cron: '0 3 * * *' },
});

nightly.task({
  name: 'extract',
  executionTimeout: '30m',
  fn: async () => {
    // your work
    return { status: 'done' };
  },
});

nightly.onSuccess({
  name: 'ping-success',
  fn: async () => {
    await fetch(PING);
  },
});

nightly.onFailure({
  name: 'ping-failure',
  fn: async (_input, ctx) => {
    await fetch(`${PING}/fail`, {
      method: 'POST',
      body: JSON.stringify(ctx.errors()).slice(0, 10000),
    });
  },
});

The run that never starts needs no code

With no worker running, the cron still enqueues the task, but nothing executes it, and Hatchet cancels it once it has waited longer than its schedule_timeout (five minutes by default). The on-failure task needs a worker too, so it never runs. A workflow paused with cron runs set to drop creates no run at all. In every one of these cases LastPing receives no ping, and the incident opens on its own when the grace window ends.

Bound long tasks with execution_timeout

Hatchet treats a task that exceeds its execution_timeout (60 seconds by default) as failed, so a stuck task reaches the on-failure task and pings /fail. Set the timeout to the longest run you consider healthy. If the worker itself stops responding, the success ping never arrives and the grace window catches it.

What this catches

  • A task fails after its retries: the on-failure task POSTs the errors to /fail; you're paged immediately.
  • A task runs past its execution_timeout: Hatchet fails it, and the on-failure task reports it.
  • No worker running: the task is never executed, no ping arrives, and the incident opens.
  • Workflow paused: dropped cron ticks create no run, and queued ones wait; either way the missing check-in is caught.
  • Recovery: the next successful run closes the incident automatically.

Comparing orchestrators? See Monitoring failed workflows across orchestrators. Using Inngest instead? See Inngest functions. Trigger.dev? See Trigger.dev tasks. Prefect? See Prefect flows. For the base pattern, see Monitor Python scripts.

Questions people ask

Front-loaded answers: the most important fact first.

  • Can't Hatchet already tell me when a workflow fails?

    Partly. An on-failure task runs as the last step of a workflow that had a failed task, and you can send any notification from it. The Hatchet engine also has tenant alerting settings for Slack and email. What neither can report is a run that never executes: the on-failure task is itself a task, so it needs a running worker, and a cron tick dropped while the workflow is paused creates no run to fail. A dead-man's-switch expects a check-in on every scheduled run and opens an incident when one is missing.

  • How do I add a dead-man's-switch to a Hatchet workflow?

    Add an on-success task that GETs https://ping.lastping.dev/<id> and an on-failure task that POSTs the upstream errors to the /fail path. In Python these are @workflow.on_success_task() and @workflow.on_failure_task(); in TypeScript, workflow.onSuccess() and workflow.onFailure(). Then create a LastPing monitor with the workflow's cron so a run that never starts also alerts.

  • What happens when no worker is running?

    The cron still enqueues the task, but nothing executes it. Hatchet cancels a task that waits in the queue longer than its schedule_timeout (five minutes by default). No on-success or on-failure task runs, because both need a worker too. LastPing receives no ping, and the incident opens when the grace window ends.

  • Does this work on Hatchet Cloud and on an engine I run myself?

    Yes. Hatchet documents the same SDK and worker model for Hatchet Cloud and for the engine you operate, and on-success and on-failure tasks are part of the SDK, so the setup is identical.

  • Is this free?

    Yes. LastPing is free for individuals. See vs Healthchecks.io and vs Cronitor.

FREE FOR INDIVIDUALS · FULLY HOSTED

No worker, no run, no alert. Unless something is outside.

Two tasks and a matching schedule. First alert in under a minute, and LastPing monitors itself the same way.