Heartbeat Monitoring

Reverse-direction monitoring for cron jobs, scheduled tasks, and processes that should ping us on a schedule.

What heartbeat monitoring is for

Most monitors check that something is up. Signalog pings your URL and verifies the response. Heartbeat monitoring is the opposite: you ping us, and Signalog verifies that the ping arrives on schedule.

Use it for things that should run periodically but don’t expose an HTTP endpoint:

  • Cron jobs: nightly database backups, weekly report generation, log rotation
  • Scheduled tasks: celery beat, Sidekiq cron, Kubernetes CronJobs
  • Long-running batch jobs: eTL pipelines, ML training jobs
  • Background workers: anything that should heartbeat to prove it’s alive

If the ping doesn’t arrive within the expected window, Signalog creates an incident. Same as any monitor going down.

Heartbeat monitoring is available on every plan, including Free.

Step 1: Create a heartbeat monitor

Navigate to Monitors → New monitor: pick Heartbeat as the type.

Configure:

  • Name. nightly-backup or whatever describes the job
  • Expected interval: how often the job should ping (e.g., 24h for nightly, 5m for frequent workers)
  • Grace period: how long after the expected interval before declaring it down (e.g., 30m for a backup that takes ~20m to run)

Save. Signalog generates a unique heartbeat token and shows the ping URL:

https://your-domain/api/v1/heartbeat/hb_abc123def456...

The token is the only auth. Keep it secret. If it leaks, regenerate it from the monitor detail page.

Step 2: Wire your job to ping

Add a single HTTP POST at the end of your job. The job should ping after completing successfully. Pinging at the start defeats the purpose (you’d say “alive” before knowing whether the work succeeded).

Bash / cron

#!/bin/bash
# nightly-backup.sh

pg_dump production > backup.sql
aws s3 cp backup.sql s3://backups/

# Heartbeat — only fires if the above succeeded (script exits non-zero on error)
curl -fsS https://your-domain/api/v1/heartbeat/hb_abc123def456...

Use -fsS (fail silently on errors but show errors). If the heartbeat endpoint itself fails, the script doesn’t loop infinitely.

Python

import requests
import logging

def run_etl():
    extract()
    transform()
    load()

if __name__ == "__main__":
    try:
        run_etl()
        # Heartbeat after success
        requests.post("https://your-domain/api/v1/heartbeat/hb_...", timeout=10)
    except Exception:
        logging.exception("ETL failed — heartbeat NOT sent")
        raise

Node.js

async function runJob() {
  await doTheThing();
  await fetch("https://your-domain/api/v1/heartbeat/hb_...", { method: "POST" });
}

runJob().catch(err => {
  console.error("Job failed:", err);
  process.exit(1);
});

Kubernetes CronJob

Add a sidecar or a final container step:

apiVersion: batch/v1
kind: CronJob
metadata:
  name: nightly-backup
spec:
  schedule: "0 2 * * *"
  jobTemplate:
    spec:
      template:
        spec:
          containers:
            - name: backup
              image: my-org/backup:latest
              command: ["sh", "-c", "/app/backup.sh && curl -fsS https://your-domain/api/v1/heartbeat/hb_..."]
          restartPolicy: OnFailure

Step 3: Verify with a manual ping

Before relying on cron, test the URL:

curl -fsS https://your-domain/api/v1/heartbeat/hb_abc123def456...

Open the monitor detail page. You should see:

  • Last ping: current timestamp
  • Status. “Up” with a green dot
  • Expected next ping: current time + interval

If you see “Last ping: never” the URL is wrong. Double-check the token.

How “down” gets detected

After the first ping, Signalog tracks expected vs actual ping arrival:

  • ✅ On time: ping arrived within interval ± grace period → state stays “up”
  • ⚠️ Late: ping is past interval but within interval + grace period → state stays “up”, marked as “late” in the log
  • ❌ Missed: ping doesn’t arrive within interval + grace period → state transitions to “down” → incident created

When a missed-ping incident fires, it follows the same alert grouping, escalation, and routing rules as any other incident. On-call gets paged, the status page reflects the failure, postmortem can be drafted normally.

Recovery is automatic

When the next ping arrives after a missed-ping incident, the incident auto-resolves. No manual intervention needed. Signalog detects the recovery and posts a resolution update.

Choosing intervals and grace periods

  • Interval: how often the job runs. Use the actual cadence. Every 5min, every hour, daily at 2am. Not your hopes.
  • Grace period: how much variance you tolerate. A common rule of thumb: grace = interval × 0.25, with a minimum of 1 minute.

Examples:

JobIntervalGraceWhy
Health check ping1 min1 minTight tolerance, expect prompt arrival
Frequent worker5 min2 minSome scheduling jitter is normal
Nightly backup24 hours1 hourCron jitter + variable runtime
Weekly report7 days6 hoursLots of variance allowed for weekend cron

If you’re getting false positives (incidents firing for jobs that completed normally), widen the grace period. If you’re getting false negatives (jobs are slow but no incident fires), tighten it.

Multi-step jobs

For jobs with multiple stages, use multiple heartbeat monitors: one per stage:

backup-extract  → pings after pg_dump
backup-upload   → pings after S3 upload
backup-verify   → pings after integrity check

This way you know which stage failed when something goes wrong, not just “the backup failed somewhere.”

Heartbeat tokens vs other auth

Heartbeat tokens are per-monitor: each heartbeat monitor gets its own. They’re meant to be embedded in cron scripts, so we accept them in the URL path (not header) for ease.

Treat them like API keys: if a token leaks, an attacker could send fake heartbeats to make a missed job appear healthy. Rotate them from the monitor detail page if you suspect compromise.

For more sensitive scenarios, run the heartbeat URL behind your VPN or use a webhook with signing instead.

Plan availability

Heartbeat monitoring is available on every plan, counted toward your monitor limit:

PlanTotal monitors
Free5 (any combination)
Pro100
Business500

There’s no separate heartbeat-only limit. Heartbeats count the same as HTTP monitors.

Next steps