Delegate on-call to agents

Your on-call engineers start every shift with answers, not alerts

On-call

Participates in every on-call rotation

Autonomously investigates alerts and builds initial findings before the on-call engineer is paged

Investigations

Agents on every on-call rotation — triaging and investigating alerts in real time.

Investigations108Alerts1658
All teamsLast 3 days
05/1005/1105/1205/13
108 investigations · all triaged by on-rotation agents in the last 3 days
Today · Wed May 13 2026
just now

Scrape Error Rate High: errors > 2% across 2 orgs on same integration

Triaged· 12sChat originOn:monitoring-alerts-systems-prodrotation:systems-oncall
8:46am today

Demo Org not Replying in Slack

Investigating· 4mChat originOn:defaultrotation:support-oncall
6:24am today

High Alert Reception Latency (p90 > 10 mins over 20 min window)

Concluded· 23mPlatformmonitoring-alerts-platform-prodcluster:app0-clusterrotation:platform-oncall
9:49pm yesterday

RDS High Read IOPS: Instance is experiencing high read IOPS

Auto-resolvedPlatformmonitoring-alerts-platform-prodcluster:app0-clusterrotation:db-oncall
Triage

Triages and investigates your alerts

Correlates signals across your observability stack, assesses severity, and identifies blast radius with evidence

Scrape Error Rate HighTriage CompleteStart Deep Investigation
InvestigationThreads
Assessed3 More Steps

Transient metrics provider outage (HTTP 500) causing scrape failures

Root cause5 evidence

Alert firing due to stale NoData KeepLast state

Contributing factor1 evidence

2 orgs with persistent UNAUTHORIZED errors due to missing events_read scope

Open lead1 evidence
Runbook conclusionAlert DetailsImpact

Runbook conclusion

The metrics provider API had a transient outage affecting 27 orgs' alert scrapes in orders-prod-cluster. The error spike was brief and self-resolved; the alert fired late because of stale evaluation. No action required — this is provider-side.

Alert Details

  • Breach: up to 4 orgs exceeded the 2% error rate on the metrics-provider integration in orders-prod-cluster.
  • Timeline: spike at 9:54pm, peaked at 4 orgs at 9:59pm, back to baseline by 10:10pm. Alert fired at 10:26pm.

Impact

  • Blast radius: 27 orgs affected by transient API errors on scrapeType=alerts.
  • Customer-facing: none — spike resolved before page-load impact.
Mitigation actions

Resolves alerts without changing context

Silences noise, executes GitHub Actions, and routes to the right team. Engineers approve or let agents handle known patterns autonomously

Silence alert: Checkout p95 latency above threshold

pending
Why silence Checkout p95 latency above threshold?

Silence Alert

PlatformGrafanaAlertCheckout p95 latency above thresholdDuration60 min
BeforeAlert "Checkout p95 latency above threshold" is firingAfterSilenced for 60 min
Interface

Available in your collaboration tools

Findings, priority lists, and actions surface in Slack, MS Teams, CLI, Resolve AI, or your own agent

#orders-on-call42 members
Confirmed root cause

Missing schema migrations on orders-db causing transaction rollbacks. Rollback ratio crossed 2% threshold at 08:19Z.

View full report
On-Call› #incidents
Resolve AI

Top 3 priorities

  • Apply orders-db migrationshigh
  • Investigate checkout p95 spikemed
  • Review checkout-v2 revert PRlow
resolve
$resolve actions
Revert checkout-v2-routinghigh
Silence Checkout p95 latencymed
Apply orders-db migrationshigh
3 actions pending review
$

Used and loved by engineers

Removing the toil of investigations, war rooms, and on-call.

“Resolve AI allowed us to move from hours to minutes for investigations in many incidents. We pull fewer engineers into war rooms, on-call is materially better, and that translates directly to advertiser trust and revenue protection for a billion-dollar ads business.”
Shahrooz Ansari
Shahrooz AnsariSr. Director of Engineering, DoorDash
“Resolve AI proved it could deliver real results in a constrained environment. It identified dependencies, surfaced accurate root causes 72% faster than our teams, all while integrating cleanly into our existing stack.”
Angelo Marletta
Angelo MarlettaSoftware Engineer, Coinbase
“Resolve AI has changed how our teams work through production incidents. What used to take hours of manual investigation and coordination across teams now gets resolved in a fraction of the time. Our engineers aren't only faster, they're focused on the work that actually drives impact.”
Meir Amiel
Meir AmielChief Trust and Infrastructure Officer, Salesforce
“We’ve seen the value of AI in development, and now we’re applying that same approach to production. We started by partnering with Resolve AI for alert triage, incident investigation, and root cause analysis. We’ve seen positive signs of improvement in mean time to resolve for our critical incidents, and the North Star is self-healing systems.”
Sandeep Contractor
Sandeep ContractorManaging Director, Engineering, MSCI
“What excites me most about Resolve AI's background agents is that I’m no longer starting from zero. The alerts are already investigated. The deployment summaries are already written. The findings are verified and the next steps are waiting for me. A lot of the operational work I used to handle manually is now happening continuously in the background with my oversight. I’m still making the important calls, but I can operate at a scale that just wasn’t possible before.”
Jeff Aronhalt
Jeff AronhaltPrincipal Software Engineer, Gametime
“Resolve AI feels like a teammate who’s already done half the work. It tells me immediately if something’s critical or can wait, saving countless hours and frustration.”
Andreas Gounaris
Andreas GounarisDirector of Engineering, Blueground
“Resolve AI helps my team navigate incidents by correlating signals across logs, metrics, traces, and code automatically. Instead of switching between multiple platforms hunting for clues, we get immediate context. At our deployment velocity, that speed makes all the difference.”
Alex Danilychev Jr
Alex Danilychev JrEngineering Manager, DoorDash
“Incident response at our scale isn't about collecting more signals. It's about understanding why something is failing, quickly enough to limit customer impact.”
Chris Umbel
Chris UmbelAI SRE Lead, Zscaler

Recent updates.

  • May 2026

    Autonomous alert triage

    Every alert investigated automatically, 24/7.

  • May 2026

    Alert resolution

    Agents take action directly, including silencing and GitHub Actions.

  • May 2026

    Deployment monitoring

    Agents watch rollouts and investigate before alerts fire.

  • April 2026

    Adaptive learning

    Triage quality improves with agent teams and engineer corrections.

Frequently asked questions