Nightly

Incident benchmark
Six staged incidents. Live models. Every number has a trace behind it.

Each incident runs end to end: baseline traffic on the known-good release, a trigger (a bad release written by a real Claude Code session, an injected provider fault, or attack traffic), then Nightly detects, investigates, prices, decides under policy, contains, replays the failed conversations and watches live recovery. Incidents are staged on purpose so they are reproducible.

loading…

#incidentdetecteddiagnosisculpritactionreplayrecoveredat riskLLM cost