From alert to resolution, tracked

Auto-created incidents with alert grouping, AI-drafted postmortems, and major incident command for the worst days.

Auto-created from monitors

When a monitor goes down, an incident is created automatically with the right severity and linked components.

60-second alert grouping

Flapping monitors don't paging-storm your team. Related alerts collapse into one parent incident automatically.

AI-drafted postmortems

One click drafts the summary, root cause, impact, and action items. You review and publish.

Incident lifecycle

Every incident moves through a clear lifecycle. No ambiguity about where things stand. For your team or your users.

  • 4 stages. Investigating, identified, monitoring, resolved. Match how incidents actually unfold
  • Severity levels from P1 to P4 so teams prioritize correctly from the start
  • Acknowledgment tracking with auto-unacknowledge so silenced pages don't stop escalation
  • Updates split between team-only notes and customer-facing communication
Elevated API Latency Resolved
minor Apr 24, 2026 · 47 min
Resolved 14:47

Latency has returned to normal levels. Root cause was a misconfigured connection pool.

Monitoring 14:32

Fix deployed. Monitoring response times for stability.

Identified 14:15

Root cause identified as a connection pool exhaustion on the primary database.

Investigating 14:00

We are investigating elevated P95 latency on the API Gateway.

Alert grouping

A single bad deploy can fire dozens of alerts in seconds. Signalog detects the pattern and folds related alerts into one parent incident. Your on-call gets one page, not twenty.

  • 60-second window groups related auto-created incidents under a parent
  • Parent incident shows all child alerts with one resolution path
  • Only the parent triggers paging. Child incidents stay quiet
  • Manually break the grouping if a child needs separate handling
Parent

API gateway 5xx spike

3 related incidents grouped under this parent

Auth service · 503
User service · 503
Billing API · 502 intermittent

Major incident command

When an incident graduates to P1 and needs all hands, declare a major. Spin up a war room with a conference bridge, assign roles, and keep stakeholders informed. All from one screen.

  • Declare-major button promotes any incident to a major incident with one click
  • Conference bridge URL (Zoom, Meet, Slack huddle) attached to the incident
  • Assign Incident Commander, Comms Lead, and Scribe roles to specific users
  • Stakeholder updates separate from internal team notes
  • Resolve-major collapses the war room while keeping the incident itself open
Major

Payment processing degradation

Bridge meet.google.com/xyz-abc-def
IC @sarah
Comms @daniel
Scribe @priya

Communication templates

Stop drafting status updates from scratch at 3 AM. Pre-built templates with {{variables}} keep updates consistent, fast, and on-message.

  • Per-team templates for investigating, identified, monitoring, resolved updates
  • Variable substitution. {{service}}, {{started_at}}, {{severity}}, {{component}}
  • Render preview before posting to catch typos
  • Stakeholder vs on-call channels. Update both with one form

Template: Investigating

We're investigating elevated error rates on {{service}}. Our team is actively working on this and will provide an update within 30 minutes.

notify subscribers internal note

AI-drafted postmortems

Stop staring at a blank postmortem template. One click feeds the incident timeline to Gemini or Claude and drafts the summary, root cause, impact, and action items as structured fields.

  • Auto-pulls timeline, severity, duration, and updates as context
  • Drafts 4 structured fields: summary, root cause, impact, action items
  • Powered by Google Gemini and Anthropic Claude. Included on Business, no setup
  • Honest about uncertainty. Flags timeline gaps rather than inventing details
  • Never auto-publishes. You review, edit, and publish
  • Available on the Business plan

Draft with AI

Business

Summary

API gateway returned 5xx for 12 min after deploy 3a4f2b shipped a regression in the auth middleware…

Root cause

Token cache miss in fallback path. Timeline suggests config-related, but logs not attached to confirm.

Stakeholder notifications

Engineers and customers don't need the same updates. Signalog separates noisy team-internal notes from polished customer-facing communication so you can speak to both audiences from one form.

  • Per-update toggle: notify subscribers, or post internally only
  • Subscriber notifications go to email, web push, and the public status page
  • Internal notes show only on the incident detail page for the responding team
  • Major incidents auto-promote select updates to subscribers

Post incident update

Rolled back deploy 3a4f, watching for recovery…
Notify subscribers → email · web push · status page
Page on-call again

Subscribers

847 reached

Internal notes

12 in this incident

Runbooks

Stop wasting minutes searching documentation during an outage. Attach runbooks directly to monitors so responders see exactly what to do when it pages them.

  • Attach step-by-step runbooks to any monitor. Shown the moment an incident fires
  • Responders see instructions immediately on the incident page, no searching required
  • Full markdown support for rich formatting, code blocks, and links

Runbook · API Health

attached to monitor
1 Check API gateway logs for error spikes — `kubectl logs -n prod api-gateway | grep ERROR`
2 Verify database connections and pool usage in Grafana dashboard `db-pool`
3 If memory > 90%, restart pods: `kubectl rollout restart deployment/api-gateway -n prod`
4 Page @platform-lead if not resolved in 15 minutes

Retrospective analytics

Learn from every incident. Track the metrics that matter so you can measurably improve your response process over time.

  • MTTR and MTTA metrics tracked automatically across every incident
  • Incidents by severity and day of week to find patterns in your failure modes
  • Most affected components and longest/fastest resolution times to guide improvements
Analytics
Uptime
99.98%
MTTR
12m
Error Budget
89%
Incidents this week
Mon
Tue
Wed
Thu
Fri
Sat
Sun
SLA Target: 99.9% Meeting target

End-to-end incident flow

Incidents flow through your entire stack. From monitor alert to alert grouping to on-call page to status page update to subscriber notification to AI postmortem. Automatically.

Monitor alertAlert groupingOn-call pageStatus pageSubscribersAI postmortem

Start monitoring in 60 seconds

Free forever for small projects. No credit card required.

Start free