the pipeline
One alert in. One investigated case out.
Nightswatch runs five stages in order. Each one has a job, a budget, and an output the next stage can check. Nothing is a black box: every tool call is recorded and every claim in the final report points back at the data that produced it.
01 · Triage
Is this even a new problem?
Incoming alerts are fingerprinted and checked against what's already in flight: repeat firings inside the window attach to the running investigation, flapping monitors are recognised as flapping, and known issues are matched to prior cases. No LLM needed to know it's the same alert again — that's a deterministic check, and it's cheaper and more reliable done in code.
alert.fingerprint = sha256(HighErrorRate|checkout|critical)
→ match: inv_7f31c (running, started 03:12 UTC)
→ action: attach, do not spawn
flap_check: 3 transitions / 15m → below threshold
known_issue: none
02 · Context
Learn the service before theorising
The investigation assembles what a good engineer would open first: the service topology and its dependencies, the metric catalog available for this service, the runbook the alert links to (fetched and indexed automatically), any runbooks you've uploaded, and past incidents that look like this one.
topology: checkout → payments-client → provider-api
metrics: 42 series matched service=checkout
runbook: checkout.md (from alert annotation)
past_incidents: 2 similar (inv_5a02b, inv_31f9e)
03 · Hypothesis
A ranked shortlist, with reasons
Nightswatch proposes candidate causes, ranked, each carrying the initial evidence that suggested it. A hypothesis with nothing behind it doesn't get to look as credible as one the context already supports.
H1 · Payment-provider timeout cascade
runbook §Provider timeouts, similar to inv_5a02b
H2 · Bad deploy / rollout
no supporting signal yet — deploy timeline unchecked
H3 · Database connection saturation
pool metrics elevated in context sweep
04 · Validation
Test each theory. Refuting is a feature.
Each hypothesis gets targeted queries against your telemetry — built from typed, validated intents, never raw query strings written by the model. The result is a verdict with receipts. Ruling something out is worth as much at 3am as ruling something in.
Confirmed ✅
Payment provider p95 latency rose from 180ms → 2.4s at 03:07, saturating the checkout connection pool.
histogram_quantile(0.95, payment_client_duration_seconds) · +1233% vs 1h baseline
Refuted ❌
Bad deploy — no deploys in the window; error onset does not align with any rollout.
deploy_events{service="checkout"} · 0 results in 02:00–03:30 UTC
05 · Report
Delivered where your team already is
The RCA is structured, not a wall of prose: headline diagnosis, stated confidence, the hypotheses with their evidence, and suggested next actions. Feedback buttons close the loop — mark it useful, not useful, or tell it the actual cause, and that verdict becomes part of the memory the next investigation reads.
🚨 Checkout error rate exceeded 5% — likely payment-provider timeout cascade
Confidence: 🟡 Medium · Investigated in 3m 42s · 14 queries run
Hypothesis 1 — Confirmed ✅
Payment provider p95 latency rose from 180ms → 2.4s at 03:07, saturating the checkout connection pool.
Evidence: rate(payment_client_errors_total[5m]) by pod · runbook: checkout.md §“Provider timeouts”
Hypothesis 2 — Refuted ❌
Bad deploy — no deploys in the window; error onset does not align with any rollout.
Next actions
- 1. Fail over to secondary provider (runbook §4)
- 2. Raise client timeout from 500ms → 2s as stopgap
Budgets, not blank cheques
Per-stage timeouts
No stage runs forever. Each one has a wall clock and gets cut off when it expires.
Tool-call caps
The number of queries an investigation may run is bounded before it starts.
Spend caps
Hard limits per investigation and per tenant per day, enforced in the engine.
An investigation that can't finish still reports what it found.