Everything you need to keep services running
Monitoring, status pages, incident management, on-call, service catalog, event ingestion, and AI-drafted postmortems. One platform.
Uptime Monitoring
Monitor HTTP endpoints, SSL certificates, heartbeats, and multi-step API workflows. Detect anomalies with time-of-day baselines before your users notice.
- HTTP, SSL, and heartbeat checks
- Multi-region monitoring (up to 10 regions)
- 30-second check intervals
- Anomaly detection with P95 baselines
- API contract validation (status, headers, JSON schema)
- Multi-step synthetic checks for login + dashboard flows
- Response time charts with deploy markers overlaid
Status Pages
Give your users a beautiful, always-current view of your service health. Fully customizable with your brand.
- 4 tiers of customization (branding to custom CSS injection)
- Dark mode + custom accent colors
- Embeddable widgets, badges, banners, and TV dashboards
- Component hierarchy with groups and per-component subscriptions
- 90-day uptime history bars + RSS / sitemap / status JSON feeds
- Multi-language support (5 languages out of the box)
- Custom domains with SSL · custom error pages on Business
Incident Management
Track incidents from detection to resolution with a complete audit trail. Auto-grouped, auto-paged, auto-postmortemed.
- Auto-created from monitor alerts with the right severity attached
- 60-second alert grouping. Flapping monitors don't paging-storm your team
- 4-stage lifecycle (investigating → identified → monitoring → resolved)
- Communication templates with {{variables}} for fast updates
- Stakeholder notifications. Separate channel from on-call paging
- Runbooks attached to monitors so responders see steps immediately
- Postmortems with timeline, root cause, impact, action items
Latency has returned to normal levels. Root cause was a misconfigured connection pool.
Fix deployed. Monitoring response times for stability.
Root cause identified as a connection pool exhaustion on the primary database.
We are investigating elevated P95 latency on the API Gateway.
AI-Drafted Postmortems
Stop writing postmortems from a blank page. One click drafts the summary, root cause, impact, and action items. You review, edit, and publish.
- Auto-pulls incident timeline, severity, duration, and updates as context
- Drafts summary, root cause, impact, and action items as structured fields
- Powered by Google Gemini and Anthropic Claude. Included on Business, no setup
- Honest about uncertainty. Flags gaps in the timeline rather than inventing details
- Never auto-publishes. Every draft is reviewed and edited by a human first
Draft with AI
BusinessSummary
API gateway returned 5xx for 12 min after deploy 3a4f2b shipped a regression in the auth middleware…
Root cause
Token cache miss in fallback path. Timeline suggests config-related, but logs not attached to confirm.
Action items
• Add cache hit/miss metric · Roll back script · Pre-prod canary on auth path
Major Incident Command
When an incident graduates to P1, declare a major. Spin up a war room with a conference bridge, assign roles, and keep stakeholders informed. All from one screen.
- Declare-major button promotes any incident to a major incident with one click
- Conference bridge URL (Zoom, Meet, Slack huddle) attached to the incident
- Responder roles. Incident Commander, Comms Lead, Scribe. Assigned to specific users
- Stakeholder updates flow to subscribers automatically; team-only updates stay internal
- Resolve-major collapses the war room while keeping the incident itself open for follow-up
Payment processing degradation
On-Call Scheduling
Ensure the right person responds, every time. Auto-rotation, escalation, swaps, override requests, and a metrics layer that tells you whether your rotation is actually working.
- Auto-rotation with pool mode and per-user exclusions
- Multi-step escalation policies with configurable delays
- Auto-unacknowledge after a configurable timeout. Silenced pages don't stop escalation
- Shift swaps (single day or date range) with one-click acceptance
- Override requests. 'cover-me-Friday' workflow with email/SMS handoff
- On-call performance metrics. TTA p50/p90/p95, unacked rate, ack leaderboard
- iCal calendar sync · on-call preview calendar
- SMS paging on Pro · voice paging (real phone call) on Business
Service Catalog
Model your architecture as services and dependencies. When the database goes down, the catalog shows blast radius. Every service that depends on it, with a one-click jump to the right runbook.
- Define services with tier, owner, and free-text description
- Link services to monitors so service status reflects real telemetry
- Map upstream/downstream dependencies between services
- Attach incidents to services to track what was affected
- Status page components can reflect service health automatically
- Available on Pro and Business
Service Catalog · 7 services
Event Ingestion
Pipe alerts in from Sentry, CloudWatch, Datadog, or any custom system. Routing rules turn raw events into structured incidents. Auto-deduped, severity-tagged, and paged.
- Per-source ingest URLs with token-based auth
- Built-in adapters for Sentry, CloudWatch, Datadog, Grafana, custom JSON
- Routing rules with field-level matchers (severity, project, tag)
- Auto-deduplicate within a configurable window. No event-storm pages
- Recent-events log per source so you can debug why a rule didn't fire
- Plan-gated to Pro and Business
Event Routing
level=fatal → P1 incident · auto-page on-call
alarm_state=ALARM → P2 · notify #ops
Alerting & Integrations
Get notified through the channels your team already uses. Every alert is signed, traceable, and retryable.
- Slack, Discord, Microsoft Teams, PagerDuty, Opsgenie, email, webhooks
- Per-project channel routing. Different services, different alerts
- HMAC-SHA256 webhook signing so receivers can verify authenticity
- Webhook delivery logs with one-click retry for failed deliveries
- Web push notifications for browser-based alerts (no SMS cost)
- Test-fire any channel from the UI before relying on it
Integrations
Analytics & Reporting
Understand your reliability posture with SLA tracking, retrospective analytics, on-call performance, and error budgets.
- SLA compliance reports with error budgets
- Incident retrospective analytics (MTTR, MTTA, severity mix)
- On-call performance. Time-to-acknowledge p50/p90/p95, unacked rate
- Deploy markers overlaid on uptime charts to correlate with regressions
- Response code distribution and per-region breakdown
- CSV and JSON data export · downloadable PDF reports
Security & Audit
Enterprise-grade access controls and a complete audit trail of every change. Compliance-friendly out of the box.
- Two-factor authentication (TOTP) with recovery codes. Available on every plan
- GitHub OAuth for one-click signup
- SSO / SAML 2.0 (Okta, Google Workspace, Azure AD, Auth0, OneLogin) on Business
- Audit log with per-action metadata, IP address, and actor
- IP-restricted API keys with CIDR allowlists
- Role-based access control. Owner, admin, member, viewer
- HMAC-signed magic links and ack tokens. No passwords in URLs
Audit Log
@daniel · API Health · 192.0.2.42
@sarah · @priya: member → admin
@daniel · Payment outage · ai_drafted=true
Changelog
Keep your users informed about what's changing. Publish updates, tag releases, and notify subscribers automatically.
- Public changelog page with shareable permalinks
- Tag-based organization (feature, fix, breaking, security)
- Draft → publish workflow with subscriber notification on publish
- Email notifications to status page subscribers
- Integrated into your status page header
Multi-region monitoring
Deploy checks from up to 10 global regions simultaneously. See per-region latency breakdowns and detect geographic outages.
Incident communication templates
Pre-built templates with variables for consistent, fast incident status updates across all channels.