The problem alert grouping solves
When something significant breaks. A database goes read-only, a network partition hits, a bad config rolls out. You don’t get one alert. You get twenty. The auth service throws 503s, the API gateway returns 504s, the worker queue backs up, and three downstream services start flapping. Without grouping, every one of those becomes a separate page. Your on-call gets twenty SMS messages in fifteen seconds, and the actual signal is buried in the noise.
Signalog detects this pattern and folds related alerts into a single parent incident. Your on-call gets paged once.
How grouping works
When a monitor fails and Signalog auto-creates an incident, it checks whether another auto-created incident in the same project started within the last 60 seconds and isn’t yet resolved. If one exists, the new incident becomes a child of that parent.
┌─────────────────────────────────────────────────┐
│ Parent incident: "API gateway 5xx spike" │
│ Started: 14:32:05 │
│ Auto-created: yes │
│ Status: investigating │
├─────────────────────────────────────────────────┤
│ ├─ Child: "Auth service 503" 14:32:08 │
│ ├─ Child: "User service 503" 14:32:11 │
│ └─ Child: "Billing API 502" 14:32:14 │
└─────────────────────────────────────────────────┘
Only the parent triggers paging. Child incidents are still tracked individually. You can see exactly which monitors failed. But their notification side-effects are suppressed. The on-call person gets a single page that says “API gateway 5xx spike” and a single incident page that lists all the affected services.
What’s grouped, what’s not
Signalog groups based on temporal proximity within a project. The 60-second window is intentionally narrow. It’s tuned to catch cascading failures from a single root cause, not to merge unrelated incidents that happened to overlap.
Grouped:
- Multiple auto-created incidents within 60s of each other in the same project
- All become children of the first incident in the burst
Not grouped:
- Manually-created incidents (declaring a major incident, for example). Those are intentional
- Incidents from different projects (cross-project correlation isn’t automatic)
- Anything outside the 60-second window. If a related failure starts 90 seconds later, it’s a separate parent
- Already-resolved parents. A fresh failure 5 minutes after recovery starts a new incident
Why 60 seconds, not longer
A wider window would catch more correlated failures, but also more false positives. Unrelated incidents that happened to fire close together. 60 seconds is long enough to absorb a typical cascade (where downstream services fail within the same minute as the root cause) and short enough that two genuinely separate outages don’t get confused.
If your cascades regularly span longer than 60 seconds, that’s usually a sign the upstream service is flapping. Group the parent incident manually to keep the timeline clean, then file a follow-up to investigate the root cause.
Breaking the grouping
Sometimes a child incident needs separate handling. The same minute, but a different team owns it. From the incident detail page, you can detach a child and promote it to a standalone incident. The original parent stays grouped with whatever children remain.
Interaction with major incidents
If a parent incident is later promoted to a major incident, the children stay grouped under the major. The conference bridge, responder roles, and stakeholder updates all attach to the parent. Children inherit the major status visually but don’t trigger their own war room.
Subscriber notifications during grouping
When you post a customer-facing update on the parent incident with notify subscribers enabled, the update goes out once. Subscribers don’t see the children. To them, the incident is “API gateway 5xx spike,” not twenty separate failures. This keeps the public status page clean and avoids alert fatigue on the customer side.
Next steps
- Understand the incident lifecycle. How parent and child incidents progress through states
- Stakeholder vs on-call notifications. Separate channels for separate audiences
- Run a major incident. When one cascade needs all-hands response