The on-call engineer for your AI agents
Nightly watches every run of your agents. When one breaks at 3 AM, it rolls back the bad release, proves the fix, and asks you in Slack only for what can’t be undone.
The same bad release, with and without Nightly
One SDK in your agent. Nightly does the rest.
Works with the tools you already use
One message per incident, updated live. Approve or undo from the channel.
Sign in with GitHub. Approved fixes open as revert PRs on your repo.
Your agent's traces. Nightly reads them back while it investigates.
Checkpoints link a bad deploy to the coding-agent prompt behind it.
Your unit economics put a price on every minute of an incident.
pip install nightly-sdk. Zero dependencies, fails open.
ntfy notifications for the incidents that really need you.
Send pages to PagerDuty, Opsgenie or your own tools.
loading…
It investigates like an engineer.
Reads the failing traces, the deploy log and the diff, and cites every trace it used.
It finds the prompt behind the bug.
Follows the bad deploy to its commit, its Entire checkpoint, and the coding-agent prompt that wrote it.
“Make refunds one step: payments already validates eligibility.”
It only wakes you when it’s worth it.
Prices every minute of the incident. Fixes what’s reversible, then sends one message.
Cause: release r46, written by a Claude Code session. Its Entire checkpoint shows the coding agent warned about it.
Contained: rolled back 18s after detection. Verified: 5/5 failed tickets now pass.
Saved: ~$22,846 that would have been lost before anyone woke up.
INC-1001 · status MONITORINGFixes what’s reversible. Asks about what isn’t.
- ↩︎ Roll back a bad release
- ⇄ Switch to a backup model
- ⏻ Turn off a risky tool
- ⧗ Stop a runaway loop
- ✎ Ship a code fix (opens a PR)
- $ Refund or email customers
- ✕ Delete data
- … anything your policy doesn’t allow
Monitoring tells you. Nightly handles it.
| Agent observability | AI SRE tools | Nightly | |
|---|---|---|---|
| Watches what your AI agent does | ✓ traces, evals, alerts | — watches servers | ✓ traces and outcomes |
| Finds the prompt behind a bad change | — | — | ✓ via the Entire checkpoint |
| Puts a price on the incident | LLM spend only | — | ✓ $/min from your unit economics |
| Fixes it while you sleep | — | infrastructure only | ✓ rollback, model failover, kill switch |
| Proves the fix worked | — | service health | ✓ replays the failed runs |
| Asks before anything irreversible | n/a | varies | ✓ always, in Slack |
Go to sleep. It’s handled.
Connect an agent in five minutes and add Nightly to Slack.