Service Catalog Setup

Define services, link them to monitors, map dependencies, and use them to route incidents.

Before you start

Service catalog is a Pro/Business feature. You’ll get the most out of it if you have:

  • 5+ monitors already configured
  • A rough mental model of how your services connect
  • An owner per service (a user or team)

You can start small. Define your most critical 3-5 services first and add more as you go. Services don’t need to be exhaustive to be useful.

Step 1: Create your first services

Navigate to Services in the sidebar. Click Add service.

For each service, set:

  • Name: short, recognizable (auth-service, not “Authentication & Authorization Service”)
  • Tier: tier 1 for revenue-load-bearing, Tier 2 for important, Tier 3 for best-effort
  • Owner: the person or team that gets paged for outages
  • Description: one or two sentences. What does this service do? Who uses it?

A useful starting set for a typical SaaS:

Tier 1 services:
- API Gateway       (Platform team)
- Auth Service      (Platform team)
- Postgres Primary  (Data team)
- Stripe Integration (Billing team)

Tier 2 services:
- Recommendation Engine (ML team)
- Search Index         (Data team)
- Email Worker         (Comms team)

Tier 3 services:
- Internal Admin Tool  (Internal Tools)
- Analytics Pipeline   (Data team)

Resist the urge to define every service immediately. Add them as they become relevant to incidents.

Each service can have one or more monitors driving its status. Open a service detail page and click Link monitor.

API Gateway
├─ http-check-prod-api      (HTTP / 200 expected)
├─ ssl-cert-api             (SSL / 30-day warning)
└─ heartbeat-api-deploy-v2  (Heartbeat / 5min interval)

When any of these monitors goes down, the service status reflects it. If you have multiple monitors and want different criticality, use separate services. One for “API Gateway core” (just the HTTP check) and one for “API Gateway certs” (SSL only).

Step 3: Map dependencies

For each service, define what it depends on (upstream) and what depends on it (downstream).

From the service detail page, click Add dependency: pick the upstream service. Signalog will validate that you’re not creating a cycle.

A clean dependency graph for the example above:

Postgres Primary
    └─ Auth Service
         └─ API Gateway
              ├─ Recommendation Engine
              ├─ Search Index
              └─ Stripe Integration
                   └─ Email Worker

Walk through your last 3 incidents. For each one, ask: “What other services were affected?” If the answer was non-obvious, that’s a missing dependency edge. Add it.

Step 4: Use services in incidents

When an incident is created (auto or manual), attach the affected service(s):

  • From the incident detail page, click Edit affected services
  • Select one or more from the dropdown

The service detail page now shows this incident in its history. Postmortems automatically pull affected service names into the impact section.

Step 5: Route by tier in escalation policies

Create separate escalation policies for different tiers:

PolicyPages on-callAuto-unackRepeat
Tier 1Yes, immediately5 minYes, max 3
Tier 2Yes, after 5 min15 minNo
Tier 3No, file Slack message——

Then on each monitor’s settings, point it at the right escalation policy based on the linked service tier.

This way, a fatal Sentry error in your Tier 3 admin tool doesn’t wake anyone at 3 AM, but the same error in a Tier 1 service does.

Step 6: Reflect services on the status page

Components on your public status page can mirror service status. Open a component on the status page → Link to service → pick the service. Now the public status page automatically reflects what your monitors are seeing. No manual updates required.

For private services that shouldn’t be public, just don’t link them to a status page component. The internal catalog stays internal.

Common patterns

Pattern: API surface + downstream workers

API Gateway → Auth → Worker Queue
                   → Email Worker
                   → Search Indexer

The API depends on Auth; workers depend on Auth and the Queue. When Auth fails, all of these are at risk.

Pattern: Read replicas behind a primary

Postgres Primary → Postgres Replica 1
                 → Postgres Replica 2
                 → Postgres Replica 3

If the primary dies, all replicas are at risk. If a replica dies, only services using that specific replica are affected.

Pattern: Third-party dependencies

External: Stripe API → Stripe Integration → Billing Worker → Subscription Renewal

Yes, model third parties. When Stripe has an outage, your catalog tells you exactly what stops working.

What to skip

  • Don’t model every transitive dependency. “Auth depends on Postgres” is enough. You don’t need “Auth depends on Postgres which depends on EBS which depends on AWS us-east-1.”
  • Don’t model deployment topology. That’s what Kubernetes/Terraform/your platform tells you. The catalog is for service-level dependencies, not pod placement.
  • Don’t model build-time dependencies. “Service A imports library B” isn’t a runtime dependency. Skip it.

Maintenance

Catalogs decay if no one tends them. To keep yours useful:

  • After each significant incident, audit the affected services. Were they all in the catalog? Were dependencies right?
  • Quarterly, walk the graph with a fresh set of eyes. Usually one or two edges are stale.
  • Make adding a service part of the launch checklist for new infra. Catalogs that grow with the system stay accurate.

Next steps