on watch while you sleep

Your alerts, investigated. Before you've even opened your laptop.

Nightswatch is an AI SRE that treats every production alert as a case to solve: it queries your telemetry, reads your runbooks, tests its own hypotheses against real data — and delivers an evidence-backed root cause analysis to your team's channel in under five minutes. Whatever you monitor with, wherever your team talks.

Diagnosis only. No auto-remediation — a human still decides what to change.

webhook · alertmanager

🔴 FIRING — HighErrorRate
service=checkout · severity=critical
03:12 UTC

Plugs into the monitoring and alerting stack you already run

  • Prometheus
  • Alertmanager
  • New Relic
  • Datadog
  • Grafana
  • CloudWatch
  • Slack
  • Microsoft Teams
  • PagerDuty
  • + yours

What actually happens at 3am

3am pages start with 40 minutes of orienteering.

Before anyone forms a theory, they're still finding the dashboard, the right time range, and which pod is angry.

The person on call is never the person who wrote the service.

Rotation math guarantees it. Context lives in someone else's head, and they're asleep.

Your runbooks are great — nobody finds the right one at 3am.

The doc that answers the page exists. It's four wiki hops away from the alert that fired.

Five stages, no hand-waving

01

Triage

Dedup, flap detection and known-issue matching decide whether this alert deserves an investigation at all.

02

Context

Pulls the service topology, the metric catalog, matching runbooks and similar past incidents.

03

Hypothesis

Proposes a ranked shortlist of causes, each with the initial evidence that suggested it.

04

Validation

Every hypothesis is tested against your live metrics — confirmed or refuted with the query as receipts.

05

Report

A structured RCA lands in your incident channel with confidence, evidence and next actions.

See the full pipeline

This is what lands in your incident channel.

Same investigation, same evidence — rendered wherever your team already talks.

#incidents03:16 UTC
NightswatchBOT03:16 UTC

🚨 Checkout error rate exceeded 5% — likely payment-provider timeout cascade

Confidence: 🟡 Medium · Investigated in 3m 42s · 14 queries run

Hypothesis 1 — Confirmed ✅

Payment provider p95 latency rose from 180ms → 2.4s at 03:07, saturating the checkout connection pool.

Evidence: rate(payment_client_errors_total[5m]) by pod · runbook: checkout.md §“Provider timeouts”

Hypothesis 2 — Refuted ❌

Bad deploy — no deploys in the window; error onset does not align with any rollout.

Next actions

  1. 1. Fail over to secondary provider (runbook §4)
  2. 2. Raise client timeout from 500ms → 2s as stopgap
👍 Useful👎 Not useful🎯 Mark actual causeView full investigation →

Confidence, stated honestly

Medium means medium. It won't dress up a guess as a finding.

Evidence, not vibes

Each hypothesis cites the queries it ran and the runbook section it quoted.

One click to tell it the actual cause

Mark the real cause and the next similar alert is investigated with that memory.

Built by people who've been paged

Evidence-backed RCAs

Every hypothesis cites the exact queries it ran and what came back. There's a full audit trail of each tool call the AI made, inspectable per investigation.

Reads your runbooks automatically

If an alert carries a runbook URL, Nightswatch fetches and indexes it. You can also upload runbooks and postmortems directly.

Remembers past incidents

Every completed investigation becomes a searchable past incident. The next similar alert is diagnosed with the memory of the last one.

Hard cost caps

Per-investigation and per-tenant daily spend limits are part of the engine. When a cap hits you still get a report with whatever was gathered.

Backtest before you trust it

Replay a past alert and see exactly what RCA Nightswatch would have written. Nobody gets notified while you're deciding.

Alert dedup done right

Repeat firings inside a window attach to the running investigation. One flapping monitor is one case, not forty.

Multi-tenant isolation enforced at the query layer · Credentials envelope-encrypted at rest · Signed webhooks · The AI can't write raw queries against your systems — every query is built from typed, validated intents.

How the isolation works

The next incident is coming. Bring backup.

or replay one of your past alerts in a backtest — first RCA free