What heartbeat monitoring is for
Most monitors check that something is up. Signalog pings your URL and verifies the response. Heartbeat monitoring is the opposite: you ping us, and Signalog verifies that the ping arrives on schedule.
Use it for things that should run periodically but don’t expose an HTTP endpoint:
- Cron jobs: nightly database backups, weekly report generation, log rotation
- Scheduled tasks: celery beat, Sidekiq cron, Kubernetes CronJobs
- Long-running batch jobs: eTL pipelines, ML training jobs
- Background workers: anything that should heartbeat to prove it’s alive
If the ping doesn’t arrive within the expected window, Signalog creates an incident. Same as any monitor going down.
Heartbeat monitoring is available on every plan, including Free.
Step 1: Create a heartbeat monitor
Navigate to Monitors → New monitor: pick Heartbeat as the type.
Configure:
- Name.
nightly-backupor whatever describes the job - Expected interval: how often the job should ping (e.g., 24h for nightly, 5m for frequent workers)
- Grace period: how long after the expected interval before declaring it down (e.g., 30m for a backup that takes ~20m to run)
Save. Signalog generates a unique heartbeat token and shows the ping URL:
https://your-domain/api/v1/heartbeat/hb_abc123def456...
The token is the only auth. Keep it secret. If it leaks, regenerate it from the monitor detail page.
Step 2: Wire your job to ping
Add a single HTTP POST at the end of your job. The job should ping after completing successfully. Pinging at the start defeats the purpose (you’d say “alive” before knowing whether the work succeeded).
Bash / cron
#!/bin/bash
# nightly-backup.sh
pg_dump production > backup.sql
aws s3 cp backup.sql s3://backups/
# Heartbeat — only fires if the above succeeded (script exits non-zero on error)
curl -fsS https://your-domain/api/v1/heartbeat/hb_abc123def456...
Use -fsS (fail silently on errors but show errors). If the heartbeat endpoint itself fails, the script doesn’t loop infinitely.
Python
import requests
import logging
def run_etl():
extract()
transform()
load()
if __name__ == "__main__":
try:
run_etl()
# Heartbeat after success
requests.post("https://your-domain/api/v1/heartbeat/hb_...", timeout=10)
except Exception:
logging.exception("ETL failed — heartbeat NOT sent")
raise
Node.js
async function runJob() {
await doTheThing();
await fetch("https://your-domain/api/v1/heartbeat/hb_...", { method: "POST" });
}
runJob().catch(err => {
console.error("Job failed:", err);
process.exit(1);
});
Kubernetes CronJob
Add a sidecar or a final container step:
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-backup
spec:
schedule: "0 2 * * *"
jobTemplate:
spec:
template:
spec:
containers:
- name: backup
image: my-org/backup:latest
command: ["sh", "-c", "/app/backup.sh && curl -fsS https://your-domain/api/v1/heartbeat/hb_..."]
restartPolicy: OnFailure
Step 3: Verify with a manual ping
Before relying on cron, test the URL:
curl -fsS https://your-domain/api/v1/heartbeat/hb_abc123def456...
Open the monitor detail page. You should see:
- Last ping: current timestamp
- Status. “Up” with a green dot
- Expected next ping: current time + interval
If you see “Last ping: never” the URL is wrong. Double-check the token.
How “down” gets detected
After the first ping, Signalog tracks expected vs actual ping arrival:
- ✅ On time: ping arrived within
interval ± grace period→ state stays “up” - ⚠️ Late: ping is past
intervalbut withininterval + grace period→ state stays “up”, marked as “late” in the log - ❌ Missed: ping doesn’t arrive within
interval + grace period→ state transitions to “down” → incident created
When a missed-ping incident fires, it follows the same alert grouping, escalation, and routing rules as any other incident. On-call gets paged, the status page reflects the failure, postmortem can be drafted normally.
Recovery is automatic
When the next ping arrives after a missed-ping incident, the incident auto-resolves. No manual intervention needed. Signalog detects the recovery and posts a resolution update.
Choosing intervals and grace periods
- Interval: how often the job runs. Use the actual cadence. Every 5min, every hour, daily at 2am. Not your hopes.
- Grace period: how much variance you tolerate. A common rule of thumb: grace = interval × 0.25, with a minimum of 1 minute.
Examples:
| Job | Interval | Grace | Why |
|---|---|---|---|
| Health check ping | 1 min | 1 min | Tight tolerance, expect prompt arrival |
| Frequent worker | 5 min | 2 min | Some scheduling jitter is normal |
| Nightly backup | 24 hours | 1 hour | Cron jitter + variable runtime |
| Weekly report | 7 days | 6 hours | Lots of variance allowed for weekend cron |
If you’re getting false positives (incidents firing for jobs that completed normally), widen the grace period. If you’re getting false negatives (jobs are slow but no incident fires), tighten it.
Multi-step jobs
For jobs with multiple stages, use multiple heartbeat monitors: one per stage:
backup-extract → pings after pg_dump
backup-upload → pings after S3 upload
backup-verify → pings after integrity check
This way you know which stage failed when something goes wrong, not just “the backup failed somewhere.”
Heartbeat tokens vs other auth
Heartbeat tokens are per-monitor: each heartbeat monitor gets its own. They’re meant to be embedded in cron scripts, so we accept them in the URL path (not header) for ease.
Treat them like API keys: if a token leaks, an attacker could send fake heartbeats to make a missed job appear healthy. Rotate them from the monitor detail page if you suspect compromise.
For more sensitive scenarios, run the heartbeat URL behind your VPN or use a webhook with signing instead.
Plan availability
Heartbeat monitoring is available on every plan, counted toward your monitor limit:
| Plan | Total monitors |
|---|---|
| Free | 5 (any combination) |
| Pro | 100 |
| Business | 500 |
There’s no separate heartbeat-only limit. Heartbeats count the same as HTTP monitors.
Next steps
- Add your first monitor. Covers HTTP and heartbeat side-by-side
- Configure alerts. Pipe missed-heartbeat incidents to Slack
- Set up runbooks. Attach a runbook to each heartbeat for “if this fires, here’s how to fix it”