Managing Incidents

Create, update, and resolve incidents with alert grouping, communication templates, AI-drafted postmortems, major incident command, and stakeholder notifications.

Overview

Incidents track service disruptions from detection through resolution. Signalog supports both automatic incident creation from monitor state changes and manual creation for issues that aren’t caught by automated checks. Most incidents resolve quietly through your normal workflow; the few that need more. Major incidents, alert grouping during cascades, AI-drafted postmortems. Get dedicated tooling without changing the core flow.

Automatic vs manual incidents

Auto-created incidents

When a monitor transitions to a Down state and an incident channel is configured, Signalog automatically creates an incident. The incident is linked to the monitor and will auto-resolve when the monitor recovers. The auto_created flag is set to true on these incidents.

Auto-created incidents are also subject to alert grouping. If multiple monitors fail within a 60-second window, related incidents collapse under one parent. Only the parent triggers paging.

Manual incidents

To create an incident manually:

  1. Navigate to Incidents and click New Incident.
  2. Enter a title (e.g., “API latency spike in EU region”).
  3. Select a severity level: P1 through P4.
  4. Choose the affected components on your status page.
  5. Optionally attach affected services from the service catalog.
  6. Write the initial status update message.

Manual incidents are not subject to alert grouping. Declaring one is intentional, and grouping it under another incident wouldn’t make sense.

Posting status updates

Incidents progress through four statuses:

  1. Investigating: you’re aware of the issue and looking into it.
  2. Identified: the root cause has been found.
  3. Monitoring: a fix has been applied and you’re verifying stability.
  4. Resolved: the incident is over.

When posting an update, you choose:

  • New status (or keep current)
  • Message: what to say
  • Notify subscribers toggle. Whether this update reaches customers, or stays internal-only

The notify-subscribers toggle is the most important UX decision per update. See Stakeholder vs on-call notifications for when to flip it on vs off. Default behavior:

  • Status transitions (investigating → identified → monitoring → resolved) default to notify=on
  • Internal notes (“trying X”, “rolled back deploy 3a4f”) default to notify=off

Using templates

Communication templates save time during stressful moments. Templates support variables that auto-populate with context:

We are investigating reports of {{component_name}} being unavailable.
Our monitoring detected an issue at {{started_at}}.
The current status is {{status}} and severity is {{severity}}.
We will provide updates as we learn more.

Available template variables

  • {{component_name}}. Name of the affected component
  • {{monitor_name}}. Name of the triggering monitor
  • {{started_at}}. When the incident began
  • {{status}}. Current incident status
  • {{severity}}. Incident severity level
  • {{service}}. First linked service name

Templates are created under Settings > Incident Templates and can be tagged for stakeholder vs internal use.

Acknowledging incidents (with auto-unack)

When an incident is created (especially via auto-creation), team members can acknowledge it to signal they are actively working on it. Acknowledgment stops the escalation chain from progressing to the next step.

Click Acknowledge on the incident detail page or respond to the alert notification (Slack interaction, SMS reply, voice IVR press-1).

If the escalation policy has auto-unacknowledge enabled, the ack expires after the configured timeout. This catches “silenced” pages where someone acks but then gets pulled into another fire. The escalation re-engages automatically.

Manual escalation: page everyone

Sometimes the on-call rotation is too slow. From the incident detail page, click Page everyone on-call to send a one-time page to every on-call responder across all schedules. Use sparingly. Overuse trains your team to ignore “everyone” pages.

Major incident command

When an incident graduates to P1 and needs coordinated multi-person response:

  1. Click Declare major on the incident
  2. Set conference bridge URL, Incident Commander, Comms Lead, Scribe
  3. Run the response with clear role separation

See Running a Major Incident for the full playbook.

Resolving incidents

To resolve an incident:

  1. Post a final update with the Resolved status (notify subscribers should usually be on for the final update)
  2. Optionally write a resolution summary
  3. The resolved timestamp is recorded and the incident moves to the resolved list

If the incident was linked to a status page, the component status automatically returns to Operational.

For major incidents, you can resolve the major (collapses the war room) separately from resolving the incident itself. Useful when you need a soak period before fully closing.

Generating postmortems

After resolving an incident:

  1. Open the incident and scroll to the Post-Mortem section
  2. Click Generate Post-Mortem: creates the structured shell with timeline pre-filled from incident updates
  3. (Business plan) Click Draft with AI: drafts summary, root cause, impact, and action items via Gemini or Claude. See AI postmortem drafting for details
  4. Review and edit the four fields
  5. Click Publish to make the postmortem public on the status page at /s/{slug}/postmortem/{incident-id}

Published postmortems trigger a notification to status page subscribers. A polished follow-up to the incident’s stakeholder updates.

Plan availability

FeatureFreeProBusiness
Auto-created incidents✅✅✅
Manual incidents✅✅✅
Alert grouping (60s window)✅✅✅
Stakeholder notifications✅✅✅
Communication templates—✅✅
Runbooks attached to monitors—✅✅
Major incident command—✅✅
Postmortems✅✅✅
AI-drafted postmortems——✅

Next steps