on watch while you sleep
Your alerts, investigated. Before you've even opened your laptop.
Nightswatch is an AI SRE that treats every production alert as a case to solve: it queries your telemetry, reads your runbooks, tests its own hypotheses against real data — and delivers an evidence-backed root cause analysis to your team's channel in under five minutes. Whatever you monitor with, wherever your team talks.
Diagnosis only. No auto-remediation — a human still decides what to change.
🔴 FIRING — HighErrorRate
service=checkout · severity=critical
03:12 UTC
Plugs into the monitoring and alerting stack you already run
- Prometheus
- Alertmanager
- New Relic
- Datadog
- Grafana
- CloudWatch
- Slack
- Microsoft Teams
- PagerDuty
- + yours
What actually happens at 3am
3am pages start with 40 minutes of orienteering.
Before anyone forms a theory, they're still finding the dashboard, the right time range, and which pod is angry.
The person on call is never the person who wrote the service.
Rotation math guarantees it. Context lives in someone else's head, and they're asleep.
Your runbooks are great — nobody finds the right one at 3am.
The doc that answers the page exists. It's four wiki hops away from the alert that fired.
Five stages, no hand-waving
01
Triage
Dedup, flap detection and known-issue matching decide whether this alert deserves an investigation at all.
02
Context
Pulls the service topology, the metric catalog, matching runbooks and similar past incidents.
03
Hypothesis
Proposes a ranked shortlist of causes, each with the initial evidence that suggested it.
04
Validation
Every hypothesis is tested against your live metrics — confirmed or refuted with the query as receipts.
05
Report
A structured RCA lands in your incident channel with confidence, evidence and next actions.
This is what lands in your incident channel.
Same investigation, same evidence — rendered wherever your team already talks.
🚨 Checkout error rate exceeded 5% — likely payment-provider timeout cascade
Confidence: 🟡 Medium · Investigated in 3m 42s · 14 queries run
Hypothesis 1 — Confirmed ✅
Payment provider p95 latency rose from 180ms → 2.4s at 03:07, saturating the checkout connection pool.
Evidence: rate(payment_client_errors_total[5m]) by pod · runbook: checkout.md §“Provider timeouts”
Hypothesis 2 — Refuted ❌
Bad deploy — no deploys in the window; error onset does not align with any rollout.
Next actions
- 1. Fail over to secondary provider (runbook §4)
- 2. Raise client timeout from 500ms → 2s as stopgap
Confidence, stated honestly
Medium means medium. It won't dress up a guess as a finding.
Evidence, not vibes
Each hypothesis cites the queries it ran and the runbook section it quoted.
One click to tell it the actual cause
Mark the real cause and the next similar alert is investigated with that memory.
Built by people who've been paged
Evidence-backed RCAs
Every hypothesis cites the exact queries it ran and what came back. There's a full audit trail of each tool call the AI made, inspectable per investigation.
Reads your runbooks automatically
If an alert carries a runbook URL, Nightswatch fetches and indexes it. You can also upload runbooks and postmortems directly.
Remembers past incidents
Every completed investigation becomes a searchable past incident. The next similar alert is diagnosed with the memory of the last one.
Hard cost caps
Per-investigation and per-tenant daily spend limits are part of the engine. When a cap hits you still get a report with whatever was gathered.
Backtest before you trust it
Replay a past alert and see exactly what RCA Nightswatch would have written. Nobody gets notified while you're deciding.
Alert dedup done right
Repeat firings inside a window attach to the running investigation. One flapping monitor is one case, not forty.
Multi-tenant isolation enforced at the query layer · Credentials envelope-encrypted at rest · Signed webhooks · The AI can't write raw queries against your systems — every query is built from typed, validated intents.
How the isolation worksThe next incident is coming. Bring backup.
or replay one of your past alerts in a backtest — first RCA free