AI SRE

Agents that investigate your production systems during on-call, incidents, and daily production tasks.

On-call

Triages every on-call alert

Resolve AI joins every on-call rotation and starts investigating the moment an alert fires.

Investigations

Agents on every on-call rotation — triaging and investigating alerts in real time.

Investigations108Alerts1658
All teamsLast 3 days
05/1005/1105/1205/13
108 investigations · all triaged by on-rotation agents in the last 3 days
Today · Wed May 13 2026
just now

Scrape Error Rate High: errors > 2% across 2 orgs on same integration

Triaged· 12sChat originOn:monitoring-alerts-systems-prodrotation:systems-oncall
8:46am today

Demo Org not Replying in Slack

Investigating· 4mChat originOn:defaultrotation:support-oncall
6:24am today

High Alert Reception Latency (p90 > 10 mins over 20 min window)

Concluded· 23mPlatformmonitoring-alerts-platform-prodcluster:app0-clusterrotation:platform-oncall
9:49pm yesterday

RDS High Read IOPS: Instance is experiencing high read IOPS

Auto-resolvedPlatformmonitoring-alerts-platform-prodcluster:app0-clusterrotation:db-oncall
Incidents

Investigates incidents alongside engineers

Domain-specialized agents pursue hypotheses in parallel and work with engineers in channels.

Lead
Triager
Investigator
Verifier
Mitigator
+2
more
Root cause: schema drift on pgdb-orders-instance-3 — migrations 0335 and 0337 were silently skipped weeks ago.
07:30
Migration 0335 (event_outcome column) was merged 2026-04-24, but Drizzle's migrator skipped it due to timestamp ordering (PR #27504).
07:30
Confirmed via traces and logs showing column "event_outcome" does not exist and relation "order_doc_state" does not exist errors across 5 services.
07:31
Deployments

Deployment monitoring

Background agents continuously check deployments and existing services to spot anomalies.

Engineering Deploy's Agent

Watches #engineering for deploys; runs post-deploy health checks.

On deploy event

Daily Pulse

Posts daily deploy + incident digest at 7 AM PT.

Daily at 7:00 AM

PostgreSQL Watch

Monitors orders-db-cluster for query regressions and lock contention.

On alert event

Capacity Review

Weekly resource-utilization report across clusters and services.

Weekly · Fridays

Templates · start from a patternBrowse all
Deploy health monitorWatch a channel for deploys; run post-deploy checks.Use template
Daily digestRoll up daily activity from N channels into a morning summary.Use template
Alert triagerWatch an alert channel; classify and route to the right team.Use template
Resource reportPeriodic utilization summary across clusters or services.Use template
Telemetry

Building tests and telemetry

Tells you how your code actually behaves in production directly from Slack, a terminal, or your coding agent.

Wednesday, May 13
VK
You2:14 PM

What's going on with checkout-router p99 today? Is it the deploy at 10:14 PT?

Engineering Deploy's Agent2:14 PM

Yes — drift starts ~2h after the deploy. p99 was steady at ~142ms through 12:00 PT, then climbed to ~387ms as traffic ramped.

Likely candidate: new code path in order-fulfillment.ts hits a slow query at p99 only under load. Error rate and throughput look normal — this is latency-only.

VK
You2:16 PM

Set up an agent to watch checkout-router latency after every deploy going forward. Alert me if drift >25% within 2h.

Engineering Deploy's Agent2:16 PM

Done. I'll watch every deploy that touches checkout-router and alert you if p99 drifts more than 25% within 2 hours.

Background agent created

checkout-router latency watch · Triggers on deploy event · Alerts if p99 drift > 25% within 2h

Send a message…
Reporting

Operational reports

Handoffs, postmortems, and runbooks written while the investigation is still fresh.

WritePreview
Last updated by Maya Lin · 5 days agoHistory
Titledeploy-health
TypeSkillActive
ScopeReads#engineering#engineering-log-prodGitHub ActionsPosts todeploy thread

Deploy Health Skill

Triggered by a deploy notification in #engineering (compare URL, smoke test link, "deploy kicked off"). Post all updates in the deploy thread.

Environments

EnvironmentTypeLog channelCluster
dev0Pre-prod (staging)#engineering-log-devdev0-cluster
app0Production#engineering-log-prodapp0-cluster
rocketProduction (dogfood)#engineering-log-prodapp0-cluster

Environment progression: dev0 → app0 → rocket. Each gates the next.

Conventions

  • Deploy tag: {github_run_id}.{attempt}-{7_char_sha} (e.g. 24893095514.1-dac7506). The github_run_id is the dev0 build run, not the prod deploy run.
  • Tone: Operational. Lead with what broke, what's at risk, what's next.

Pipeline gate

Before running the health check, verify the prod deploy pipeline completed successfully for this deploy tag (Deploy to Cells workflow for app0, Deploy to Rocket from Cells for rocket).

Used and loved by engineers

Removing the toil of investigations, war rooms, and on-call.

“Resolve AI proved it could deliver real results in a constrained environment. It identified dependencies, surfaced accurate root causes 73% faster than our teams, all while integrating cleanly into our existing stack. ”
Angelo Marletta
Angelo MarlettaStaff Software Engineer, Coinbase
“Resolve AI makes our junior on-call engineers as effective as our seniors, flattening the experience curve. We’ve seen a 2x productivity lift while eliminating the runbook gap.”
A.D.Sr. Director of Engineering, Financial Services Company
“Literally everybody said Resolve AI is a ‘game changer’ when it comes to incidents and resolution”
D.R. Production Engineering, NoSQL Database
“Resolve AI feels like a teammate who’s already done half the work. It tells me immediately if something’s critical or can wait, saving countless hours and frustration.”
Andreas Gounaris
Andreas GounarisDirector of Engineering, Blueground
“It correctly identified that all of the errors were from a single transaction being retried rather than a widespread issue. It also provides the exact method from code that the transactions failed in.”
Anon.Software Engineer, Public Consumer Finance Company
“It was able to pinpoint the exact PR that introduced the bug and tell me the event IDs, categories, and how many listings we were over our defined rate limit.”
P.V.Sr. Software Engineer, Mobile Ticketing Marketplace
“We have seen compelling proof that Resolve AI arrived at the same root cause that the humans did, but 4-5 hours before an incident actually happened.”
D.W.Engineering Leader, Cloud Security Provider

Frequently asked questions