5 min read

Hero of prod: Duncan Winn

Duncan Winn as a wizard summoning light on a mountaintop
Role
Head of SRE

Sep 29, 2026

“It’s not about who’s on-call; that’s the obvious piece people gravitate towards. It’s about how you inject reliability from the outset, to avoid having the issue in the first place.”

SRE isn’t only about the pager

My time as a field engineer taught me that many companies still have an outdated mindset of SRE exclusively carrying a pager and undertaking production operations. If you only see SRE as an operator role, then it’s fair to say AI will collapse some of that function, because triaging large swaths of complex data is precisely the work AI has become good at.

The more interesting question is what remains for SREs once that type of on-call work is handled effectively. The answer depends entirely on what you believed the SRE job was to begin with.

When the pager goes off

Production operations are challenging. Outages often result in teams traversing oceans of metrics, logs, and traces from large-scale distributed systems that have evolved over many years, sometimes even decades. Knowing exactly where to look and how to debug the system often transcends playbooks, which become stale fast, into siloed, domain-specific knowledge that takes a long time to acquire.

Bad outages also rarely encompass issues you already know about. They usually consist of one or more unforeseen events that are new and novel, and in a large incident the work generally divides into four things. First, verifying the impact: Is this a real outage, how many customers does it affect and in what regions, and what is the nature of that impact? Second, verifying the trigger, which might be a bad binary or config push, or a dependency issue such as a regional network outage. We are not trying to fully root cause anything at that point, only to know enough to stop the bleeding. Third, once the trigger is understood, establishing what the mitigation or fix should be. Fourth, rolling it out.

That last step is the one people underestimate. It can be fast, for example, a binary rollback or a regional traffic drain. But some issues, such as a bad schema push, may be hard to mitigate through any mechanism other than rolling forward a fix. Rolling forward is expensive. Unlike rolling back to a known good state, which you can often do with confidence, it risks breaking something else and demands additional verification, care, and attention.

This is where AI starts to change the work. Verifying the cause and deciding the mitigation, once telemetry is in place, is the sweet spot for AI operations. It can preaggregate the necessary context and lead the expert engineer to the root cause faster. That does two things for engineers: they can get ahead of a much larger surface area of issues, eliminating noisy and distracting alerts, or the classic pager storm, so they can laser focus on what is important; and they get back to engineering what matters, because less time goes into triaging and remediating.

Both strengthen your reliability posture, but they cover the middle of an incident rather than its beginning or its end.

The work that happens before the pager

I don’t think AI closes all of those gaps, at least not for free. Verifying the impact requires rich and accurate telemetry across your service and all of its dependencies. Your ability to roll out a mitigation is largely down to what the fix is, how well you have architected your system to be resilient to failures, and your release engineering practices.

These are not things you can acquire in the middle of an incident. They were either engineered months earlier or they were not. While AI can help you build them, it will not build them for you.

Which gets to what I believe the SRE job actually is. SRE is first and foremost an engineering function, and at least 50% of an SRE’s focus should be spent on engineering projects. Due to the reliability-focused nature of the role, that engineering must be enduring engineering that strengthens service reliability and production health. What AI gives back during an incident is roughly proportional to how much of it you prepared for upfront.

Giving SREs the time back

Assume you get that right and can now fix outages quickly. The next thing that slows the team down is that your release process did not catch these outages in the first place. Shift your detection to the left as far as possible and you will further minimize disruption, and the whole team will move faster. It is the same engineering work, carried out earlier.

On-call is, to some degree, on a par with toil. It also helps us understand the systems we are on-call for, which is as true for developers as it is for SREs, so there is benefit in both teams carrying the pager. But the more we streamline on-call, the more room there is for high-value work.

I see our reliability posture the way I see a security posture: both require constant strengthening, and neither is something you declare finished and walk away from. So the question is not about who’s on-call; that’s the obvious piece people gravitate towards. It’s about how you inject reliability from the outset, to avoid having the issue in the first place.

Automating yourself out of a job

One of SRE’s goals is to try and automate yourself out of a job. We say that with confidence because there is always a new domain or system requiring reliability-oriented engineering.

In the era of AI, SREs still have a job. AI operations can expedite the lower-value parts of the role so that SREs can focus on the high-value work of enduring engineering. If anything, the engineering portion of our work matters more than ever, because it sets the ceiling on how much impact an incident can have, along with how much insight you obtain when something actually breaks.

The best SRE work happens long before the pager goes off.

A knight in full armour typing at a computer, surrounded by fire

Give your heroes of prod a head start

See how Resolve AI investigates incidents and alerts alongside your on-call engineers.

Book a demo