When to declare a major
Most incidents don’t need a war room. A monitor goes down, on-call acks, fixes it, posts a status update, moves on. Major incidents are the exception. Moments when you need a coordinated response with multiple people, a conference bridge, and a clear chain of command.
Declare a major when:
- More than one Tier 1 service is affected
- The incident has been ongoing for 10+ minutes with no clear path to resolution
- Multiple teams need to coordinate (engineering + comms + leadership)
- A customer commitment (SLA, contract) is at risk
If you’re unsure whether to declare, lean toward yes. It’s easier to stand a war room up and tear it down quickly than to coordinate without one.
Step 1: Declare the major
From the incident detail page, click Declare major: you’ll be prompted for:
- Conference bridge URL: zoom, Google Meet, Slack huddle, whatever your team uses. Paste it here so responders know where to go.
- Incident Commander: single person responsible for coordinating the response. Not the person fixing things. The person making sure the right people are working on the right things.
- Communications Lead: owns customer-facing updates. Curates what gets posted to the status page and what stays internal.
- Scribe: keeps the incident timeline detailed and current. The scribe’s notes become the postmortem source material.
You can leave roles unassigned and fill them in later, but try to assign at least the Incident Commander before paging others.
The incident gets a “Major” badge. The conference URL appears prominently on the incident page so anyone joining can dial in immediately.
Step 2: Page additional responders
The Incident Commander’s first job is to make sure the right people are on the call.
- For a database incident, page Data team on-call
- For an auth incident, page Security or Platform on-call
- For revenue impact, page Engineering Manager and Comms Lead
Use the Page everyone on-call button from the incident detail page if you need to escalate widely. This sends a one-time page to every on-call responder across all schedules. Use sparingly.
Step 3: Run the response
The roles are deliberate. Each one has a job that someone needs to be doing without distractions.
Incident Commander
- Decides who’s working on what
- Asks “what’s our hypothesis? what are we doing to confirm or disprove it?” every 5-10 minutes
- Calls for context dumps when new responders join the call
- Decides when to declare resolution
The IC does not debug. The IC’s job is to keep the response coordinated so the people debugging can focus.
Communications Lead
- Posts customer-facing updates to the status page
- Decides what to share publicly vs internally
- Drafts updates for stakeholders (exec team, support team, key customers)
- Coordinates with support so they have answers ready
The Comms Lead writes the polished version of what the engineers are seeing. They translate “deploy 3a4f rolled out a regression in the auth middleware” into “we identified a regression in our recent deployment and are rolling it back.”
Scribe
- Posts an incident update on every meaningful event in the response
- “Started rollback of deploy 3a4f”. Internal note
- “Recovery confirmed for 5 minutes”. Internal note
- “Postmortem will follow within 5 business days”. Stakeholder update at resolution
Detailed scribe notes are the difference between a useful postmortem and a vague one. Quote what people said on the call. Note timestamps. Capture hypotheses that were tried and abandoned.
Step 4: Stakeholder communication
Customer-facing updates flow through the Notify subscribers toggle on each incident update. The Comms Lead writes these; nobody else posts to the public status page during a major.
A typical major incident communication arc:
T+0 "We're investigating elevated error rates. We'll provide an update within 30 minutes."
T+15min "We've identified a regression in our recent deployment. Rolling back now. ETA 15 min."
T+30min "Rollback complete. Monitoring for stability. We'll declare resolved if no further issues in 15 min."
T+45min "Resolved. We'll publish a detailed postmortem within 5 business days. We apologize for the disruption."
Four updates in 45 minutes. Each one is short, factual, and forward-looking. Avoid speculation, avoid technical jargon, never blame (“our cloud provider had an issue”) unless you’re certain.
Step 5: Resolve the major
When the incident is resolved:
- The IC posts the final stakeholder update via the Comms Lead
- The IC clicks Resolve major: this collapses the war room (clears the conference bridge URL and roles) but keeps the incident open
- The incident moves to monitoring state for a soak period (typically 30-60 min)
- Once stable, click Resolve incident to close
Resolving the major is separate from resolving the incident on purpose. Sometimes you need the incident open for follow-up work (a postmortem, a deeper investigation, customer outreach) but the war room is no longer useful.
Step 6: Postmortem
Major incidents always get a postmortem. Within 24 hours of resolution:
- Click Generate Post-Mortem to create the structured shell
- (Business plan) Click Draft with AI to prefill summary, root cause, impact, action items
- The Scribe’s incident notes become the timeline. That’s why detailed scribing during the response is so valuable
- Review with the response team
- Publish within 5 business days
Postmortems for major incidents are usually published. Internally always, externally when there was customer impact.
Roles when you’re a small team
If you have 3-5 engineers and someone goes on call, you don’t have separate IC, Comms Lead, and Scribe. That’s fine.
- For incidents that need a major, the on-call engineer is usually IC + Scribe
- The senior-most available person becomes Comms Lead
- Roles can be reassigned mid-incident as more people come online
The structure exists so you don’t forget the work, not because three separate humans are required.
When NOT to declare a major
- Single-service degradation with a clear path to fix
- Anything where on-call can resolve solo within 30 minutes
- Maintenance-window issues
- “Annoying but not customer-impacting” failures
Declaring majors for non-major incidents trains your team to ignore the war room when it matters. Reserve it for moments when coordinated multi-person response actually adds value.
Plan availability
Major incident command requires Pro or Business. Free-tier teams have full incident management but not the structured war-room features.
Next steps
- Understand stakeholder vs on-call notifications
- Configure communication templates so the Comms Lead has pre-written starts
- AI postmortem drafting. Drafts the postmortem from the Scribe’s timeline