6 min read

Hero of prod: Rita Chiou

Rita Chiou riding a pegasus through the clouds
Company
Blueground

Sep 29, 2026

“Everyone expects things to just work, no questions asked. But when something breaks, that’s when we’re noticed.”

Deferring the Real Cost

I work on the platform team at Blueground. Platform is the layer most people never have to think about, and that’s the whole point. When one of our customers books an apartment on our website, they should never know my team exists. If everything runs smoothly, we stay invisible. It’s a strange kind of success: the better our work, the less anyone knows what we’re working on. Product teams ship the features users see. It’s up to us to make sure users can use those features.

The job has two sides. The first is reactive: when something breaks, the platform team is the first responder. During business hours each team owns its own services, but after hours almost everything routes to whoever is on platform on-call, in rotations running multiple days each month. Anything infrastructure related is ours around the clock, because we own the infra itself.

The second side is proactive, and it is the part outsiders rarely picture. Every incident we have acts as a signal for us. While we quickly fix the immediate problem, we also go a layer deeper to understand why it happened so it doesn’t happen the same way twice. For cases that require more attention, we may even set aside an OKR purely to get ahead of them, hardening security and reliability. If you don't think about how something gets maintained while you're building it, you're just deferring the real cost.

When the pager goes off

Being on-call can feel drastically different depending on how lucky you get. The best case is an alert fires, I already know exactly what’s behind it, so I go straight to the fix and can be back in bed in 10 minutes.

The worst case is the opposite. There’s no obvious cause, no ETA I can give anyone, and it’s the kind of outage my coworkers and our users are feeling, which brings its own kind of pressure. I found that the only way through cases like this was to isolate myself, get into the zone, and methodically check things one by one until something surfaces.

This is hardest late at night, or when I’m away from home. I’ve had plenty of incidents where I was out at dinner with friends and had to step into a corner, block everything else out, and work the problem until it was fixed. People sometimes ask how you get used to that, and I am not sure you fully do. My longtime way of thinking through a hard problem out loud was rubber-ducking with my cat, actually walking her through my theories. But since AI, she (and I) get to sleep a bit more after the sirens go off.

Starting at 3:00 AM with context

Before AI started showing true business value, it was still incredibly helpful when trying to orient myself after an incident struck. We have several layers of security to access our services, so simply logging in to start looking can take five minutes. Picture waking up at 3:00 in the morning, blurry vision, reaching for your phone, and waiting all of that out before you can even see what is wrong.

The earliest real value I got from ResolveAI was that it painted a picture of what had happened before I touched a single system. That matters enormously at 3:00 AM. It also changed how fast I could triage something unfamiliar. When an off-hours alert hits a service I do not have context on, ResolveAI connects the recent releases and changes for me. Before, that meant digging through GitHub and Datadog by hand to piece together what had shipped. Now it is fast enough that the whole first step almost disappears.

Instead of starting from zero and gathering context, I can start with a picture of what has already happened.

A rubber duck that talks back

The bigger shift with ResolveAI has been in how much I trust it for back-and-forth consultation. Early on it would give me a single rough hypothesis, and I’d take it as a starting point rather than an answer. Now it’s more collaborative, like a coworker I can trust. If I have a hunch, I can steer it toward where I think the problem is and we dig into it together.

It’s gone from feeling like a junior engineer’s first guess to a peer I can pressure-test, one that pushes back and sometimes catches something I’d have missed. I’ve always worked step by step, thinking out loud, and now I have something to think out loud with while the investigation and remediation is ongoing. It’s the rubber-ducking again, only this one replies back about telemetry.

The biggest difference for me is how much ResolveAI has sped up my on-call work. It connects tools and context I’d otherwise have to go dig up myself, and performs a deeper pass across integrated systems when there isn’t time to do that manual work by hand. It covers more ground faster than I could alone, and can surface connections I haven’t gotten to yet. Less of my time on-call is spent gathering information and manually working through every possibility, leaving me more time to decide what actually needs to happen next.

What stays human

A while ago I would have said security is the one area that stays firmly human, and AI-generated code is inherently less safe, but now I don’t think that’s quite right anymore. It’s less about the code itself and more about how carefully a team sets things up from the start. Fine-grained permissions, scoped integrations, and security steps built into planning rather than bolted on later still take deliberate effort.

The upside is that AI can also surface security steps that used to get skipped or postponed on early builds, simply because those builds ship faster now and there’s room to do it right from day one. It’s a trade-off either way, not an automatic win or loss. Day-to-day, what I actually notice is smaller habits that pile up and significantly change the way we work: AI can act as a second reviewer catching things a human might miss, CI and syntax issues get resolved faster, and workflows that used to be manual can increasingly be automated end to end.

For the future, if I had to guess, I think there will be a real shift coming toward MLOps, in securely setting up AI workloads is becoming its own discipline, and that need is only growing. Where exactly that leaves platform work day-to-day, I'm curious to see. What I do think holds is that someone still needs to be watching closely enough to catch it when something's quietly going wrong. That judgment stays human. In the end, you can't build if you can't maintain it.

A knight in full armour typing at a computer, surrounded by fire

Give your heroes of prod a head start

See how Resolve AI investigates incidents and alerts alongside your on-call engineers.

Book a demo