Service Catalog

Model your architecture as services and dependencies so you can see blast radius the moment something breaks.

What the service catalog gives you

Monitors tell you what’s broken. The service catalog tells you what depends on it.

A monitor fails for db-primary.internal. Without a catalog, that’s just a name on the dashboard. Your responder has to remember (or guess) which services use that database. With a catalog, the same alert opens to a service detail page that lists every downstream service: API gateway, auth, billing worker, recommendation engine. The blast radius is in front of you in seconds, not minutes.

The model

A service is a first-class entity with:

  • Name and description: human-readable identification
  • Tier: tier 1 (critical), Tier 2 (important), Tier 3 (best-effort)
  • Owner: single point of accountability (a user or team)
  • Status: operational, degraded, partial outage, major outage, maintenance
  • Linked monitors: monitors whose health drives this service’s status
  • Dependencies: upstream (services this depends on) and downstream (services that depend on this)
  • Linked incidents: incidents that affected this service

Services live at the project level. A typical SaaS team has 10-50 services; large platforms can have hundreds.

Tiers shape response

Service tier is the single most useful field for incident response.

  • Tier 1: customer-facing or load-bearing for revenue. Page on-call immediately, declare a major if degraded for more than 5 minutes.
  • Tier 2: important but degradable. Notify on-call but don’t auto-major.
  • Tier 3: best-effort. File an issue, fix in business hours.

Escalation policies and routing rules can match on tier so the same fatal Sentry error in a Tier 1 service pages on-call but in a Tier 3 service just files a Slack message.

Dependencies = blast radius

The dependency graph is directional and acyclic. When you draw an edge from A → B, you’re saying “A depends on B”. If B goes down, A is at risk.

        ┌──────────┐
        │ Postgres │
        └────┬─────┘
             │
        ┌────▼─────┐         ┌─────────┐
        │ Auth Svc │◄────────┤  Redis  │
        └────┬─────┘         └─────────┘
             │
   ┌─────────┼──────────┐
   │         │          │
 ┌─▼──┐  ┌──▼──┐   ┌────▼────┐
 │API │  │Web  │   │ Mobile  │
 └────┘  └─────┘   └─────────┘

When Postgres goes down, the catalog highlights everything reachable from it: Auth Service is at risk; API, Web, and Mobile are at second-degree risk. You don’t have to recompute this in your head during an incident.

Signalog detects circular dependencies at write time and blocks them. Graphs must be acyclic. If you genuinely have a cycle (rare and usually a sign of architectural debt), model it as two separate services with a clear boundary.

Service status from monitors

Linking a monitor to a service makes the service’s status reactive. If the monitor goes down, the service goes from “operational” to “major outage” automatically. Multiple monitors per service combine. A service is “operational” only if all its critical monitors are passing.

This wires nicely into the public status page: components on the page can mirror service status, so customers see “API Gateway: degraded” the moment your monitors detect it.

Incident attribution

When you declare an incident, attach the affected services. Two things happen:

  1. The service detail page accumulates incident history. Useful for retrospectives (“Auth Service had 4 incidents this quarter, all related to token expiry”)
  2. The incident’s stakeholder updates can name the affected services automatically. No copy-pasting from a runbook

Postmortems pull affected service names into the impact section, so a published postmortem reads “This incident affected Auth Service and API Gateway” instead of “this incident affected several services.”

What it’s not

The catalog isn’t a configuration database (CMDB). It’s not where you store secrets, runbook URLs in long form, or deployment metadata. It’s intentionally minimal. Just enough structure to answer “what’s affected when X breaks” without becoming a maintenance burden of its own.

For runbooks, attach them to monitors. For secrets, use Vault. For deployment metadata, use deploy markers. The service catalog stays focused on dependencies and ownership.

Plan availability

Service catalog is a Pro and Business feature. Free-tier teams can still use components on the status page for customer-facing service rollups, but the structured dependency graph and incident attribution are paid.

Next steps