Incident Lifecycle

Learn how incidents progress through four states, how they're created automatically or manually, and how they connect to your status page.

The four states

Every incident in Signalog progresses through a defined lifecycle with four states:

1. Investigating

The initial state. The team is aware of the issue and actively diagnosing the root cause. This status is set automatically when an incident is created (whether manually or via auto-creation) and is the signal to users that you’re looking into it.

2. Identified

The root cause has been found. Posting an update with the “Identified” status tells users you know what’s wrong and are working on a fix. This is often where you include technical details about what failed and your plan to resolve it.

3. Monitoring

A fix has been deployed and you’re watching to confirm the issue is fully resolved. This state is important — it signals progress to your users while giving you time to verify stability before declaring resolution.

4. Resolved

The incident is over. Services have returned to normal operation. A resolution update should summarize what happened, what was affected, and what you’re doing to prevent recurrence.

Resolved incidents record the exact resolution timestamp, enabling accurate downtime calculations for SLA reporting.

Automatic creation

When a monitor transitions to the Down state, Signalog can automatically create an incident. Auto-creation is enabled per monitor and requires:

  1. The monitor has the auto-create incidents setting enabled.
  2. At least one component is linked to the monitor on a status page.

Auto-created incidents:

  • Use the monitor name as the incident title.
  • Set the initial status to Investigating.
  • Link the affected status page components.
  • Set the auto_created flag to true.
  • Auto-resolve when the monitor recovers (transitions back to Up).

Manual creation

Not all incidents come from automated monitoring. You can create incidents manually for issues like:

  • Partial degradation that doesn’t trigger a monitor failure.
  • Third-party service issues affecting your users.
  • Planned maintenance (start an incident in advance with a maintenance status).

Severity levels

Incidents have three severity levels:

  • Minor — Limited impact. A small number of users or non-critical feature affected.
  • Major — Significant impact. A core feature is degraded or unavailable for many users.
  • Critical — Service-wide outage. Most or all users are affected.

Severity influences alert routing (e.g., critical incidents might page the on-call engineer, while minor incidents send a Slack message) and is displayed prominently on the status page.

Acknowledgment

When an incident is created, the on-call team member can acknowledge it. Acknowledgment serves two purposes:

  1. Signals to the team that someone is actively handling the issue.
  2. Stops the escalation chain from progressing to the next step.

Unacknowledged incidents continue to escalate based on the configured escalation policy.

Linked changelog entries

After resolving an incident, you can create a linked changelog entry to publicly communicate what happened. This is particularly useful for significant incidents where users expect a detailed explanation. The changelog entry links back to the incident timeline for full transparency.

Next steps