For judges
Each criterion, and where to check it.
Nightly is the on-call engineer for AI agents. When a change makes an agent fail quietly, it finds the cause, prices the damage, applies a fix that can be undone, proves it by replaying the failed runs, and pages a human once. Benchmark incidents are staged on purpose so they are reproducible; the Scout incidents ran live through the product.
The agent completes a substantial task
The whole incident lifecycle, on its own: detect → investigate → price → decide → contain → verify → report.
Benchmark of six staged incidents: 6/6 handled correctly; five contained 13–21 s after detection and verified by replay.
Live on a second agent (Scout) connected only through the SDK and neatlogs: detected ~1.5 min after deploy, rolled back, 3/3 replays passed, resolved in under 3 min.
It improved, with before/after numbers
Run 1 → run 3: 4/6 → 6/6 incidents handled correctly, after fixing confidence calibration and impact pricing based on what run 1 showed.
It recovers from failures
If a fix fails verification, Nightly undoes it by itself and escalates. Covered by a test.
It kept working when neatlogs dropped most of Scout’s traces: runs are now also counted from SDK events.
Uses

neatlogs: every agent run and every investigation is traced; Nightly reads traces back over the neatlogs MCP and creates a detection after an attack.
Entire: every bad release was written by a separate Claude Code session with its own checkpoint. Nightly resolves commit → checkpoint → prompt → the coding agent’s own warning.
cfo.ai: the business plan (break-even M08, runway > 24 months) and the customer unit economics Nightly uses to price every incident.
A real product, not a script
GitHub sign-in, a 5-step “Connect an agent” wizard, agent pages, an incident inbox, a Slack app with Approve and Undo, and an SDK on PyPI.
Built during the hackathon, in public
Posted on X with #neatHack as we built. Commit history starts with the plan only; the build log records each phase with timestamps.