the pipeline

One alert in. One investigated case out.

Nightswatch runs five stages in order. Each one has a job, a budget, and an output the next stage can check. Nothing is a black box: every tool call is recorded and every claim in the final report points back at the data that produced it.

01 · Triage

Is this even a new problem?

Incoming alerts are fingerprinted and checked against what's already in flight: repeat firings inside the window attach to the running investigation, flapping monitors are recognised as flapping, and known issues are matched to prior cases. No LLM needed to know it's the same alert again — that's a deterministic check, and it's cheaper and more reliable done in code.

alert.fingerprint = sha256(HighErrorRate|checkout|critical)

→ match: inv_7f31c (running, started 03:12 UTC)

→ action: attach, do not spawn

flap_check: 3 transitions / 15m → below threshold

known_issue: none

02 · Context

Learn the service before theorising

The investigation assembles what a good engineer would open first: the service topology and its dependencies, the metric catalog available for this service, the runbook the alert links to (fetched and indexed automatically), any runbooks you've uploaded, and past incidents that look like this one.

topology: checkout → payments-client → provider-api

metrics: 42 series matched service=checkout

runbook: checkout.md (from alert annotation)

past_incidents: 2 similar (inv_5a02b, inv_31f9e)

03 · Hypothesis

A ranked shortlist, with reasons

Nightswatch proposes candidate causes, ranked, each carrying the initial evidence that suggested it. A hypothesis with nothing behind it doesn't get to look as credible as one the context already supports.

H1 · Payment-provider timeout cascade

runbook §Provider timeouts, similar to inv_5a02b

H2 · Bad deploy / rollout

no supporting signal yet — deploy timeline unchecked

H3 · Database connection saturation

pool metrics elevated in context sweep

04 · Validation

Test each theory. Refuting is a feature.

Each hypothesis gets targeted queries against your telemetry — built from typed, validated intents, never raw query strings written by the model. The result is a verdict with receipts. Ruling something out is worth as much at 3am as ruling something in.

Confirmed ✅

Payment provider p95 latency rose from 180ms → 2.4s at 03:07, saturating the checkout connection pool.

histogram_quantile(0.95, payment_client_duration_seconds) · +1233% vs 1h baseline

Refuted ❌

Bad deploy — no deploys in the window; error onset does not align with any rollout.

deploy_events{service="checkout"} · 0 results in 02:00–03:30 UTC

05 · Report

Delivered where your team already is

The RCA is structured, not a wall of prose: headline diagnosis, stated confidence, the hypotheses with their evidence, and suggested next actions. Feedback buttons close the loop — mark it useful, not useful, or tell it the actual cause, and that verdict becomes part of the memory the next investigation reads.

#incidents03:16 UTC
NightswatchBOT03:16 UTC

🚨 Checkout error rate exceeded 5% — likely payment-provider timeout cascade

Confidence: 🟡 Medium · Investigated in 3m 42s · 14 queries run

Hypothesis 1 — Confirmed ✅

Payment provider p95 latency rose from 180ms → 2.4s at 03:07, saturating the checkout connection pool.

Evidence: rate(payment_client_errors_total[5m]) by pod · runbook: checkout.md §“Provider timeouts”

Hypothesis 2 — Refuted ❌

Bad deploy — no deploys in the window; error onset does not align with any rollout.

Next actions

  1. 1. Fail over to secondary provider (runbook §4)
  2. 2. Raise client timeout from 500ms → 2s as stopgap
👍 Useful👎 Not useful🎯 Mark actual causeView full investigation →

Budgets, not blank cheques

Per-stage timeouts

No stage runs forever. Each one has a wall clock and gets cut off when it expires.

Tool-call caps

The number of queries an investigation may run is bounded before it starts.

Spend caps

Hard limits per investigation and per tenant per day, enforced in the engine.

An investigation that can't finish still reports what it found.

See what it connects to