← All transcripts

Anthropic Reliability Engineer EXPOSES the Limits of AI SRE Transcript, AI Summary & Key Points

InfoQ · 10 days ago · Science & Technology · 45:18 · EN

Watch on YouTube

Answer

No. Claude does not reliably fix incidents or replace human SREs, although it is highly effective for signal gathering and several supervised incident-management tasks.

AI Summary

Claude and other LLMs can accelerate incident response, especially by gathering signals from metrics, logs, traces, dashboards, and code. They are not reliable autonomous SREs: they can confuse correlation with causation, produce oversimplified root causes, lack historical and tacit system knowledge, and make dangerous recommendations during unfamiliar incidents. AI can handle generic mitigations, rollbacks, handoffs, and documentation while humans focus on prevention, architecture, and system improvements. Wider automation also risks engineering-skill atrophy and increased software complexity through Jevons paradox.

Key Points

  • Anthropic's AI reliability team still hires human reliability engineers because Claude cannot carry a pager reliably.
  • On-call creates cumulative costs in sleep, attention, and personal relationships; it should not be treated as a badge of honor.
  • The highest-value reliability work prevents whole classes of incidents rather than merely fixing individual outages.
  • The OODA Loop—observe, orient, decide, and act—maps well to incident response, but LLM performance is jagged across its stages.
  • LLMs are superhuman at signal gathering: they can query metrics in parallel, write PromQL with few mistakes, and process thousands of log lines without fatigue.
  • December 31 fraud investigation — Claude traced 500 internal errors to an image-processing bug triggered by requests containing exactly 22 images and a PDF, linked roughly 200 active accounts to around 4,000 accounts created in the same window, and identified nine signups per minute as evidence of fraud.
  • Production Rust panic investigation — Claude found the panic location, assertion, segment-ID validation issue, volume pattern, affected hardware platforms, and deployment timing faster than an experienced engineer manually reading logs.
  • During KV-cache incidents, Claude repeatedly interpreted simultaneously rising request volume and errors as a capacity problem, even though the underlying issue was a lost cache; it struggled to distinguish correlation from causation.

AI in practice

Used for

Agents

  • Claude agent — Investigate an on-call alert and prepare an initial incident picture before the engineer reaches a computer. 2 held 30:03

Tools & resources

3 items

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Anthropic Reliability Engineer EXPOSES the Limits of AI SRE — InfoQ (45:18). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by InfoQ. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:10 Hello everyone. Thanks for being here. I'm Alex. Uh I'm on the AI reliability team at Entropic. Uh which means my job is keeping Claude up. Now I want to start by saying saying I know this is another AI talk here. I'm sorry for this. Uh, but I've been using LMS as part of actual incident response and I like to share a candid and humorous discussion about what works and what doesn't.

00:43 Quick context about me. Uh, I'm one of the two people that joined uh the reliability team in London. I did on call for cloud serving stack solo for my first three months. uh which is a very effective and fast way but really stressful to learn about the system. Then I onboarded uh everyone else so I could stop being on call forever. Before that I was at Google uh S on Google Cloud's compute product GC.

01:09 I was on the S sur for S team which is the escalation layer for when an outage is bad enough that even normal SRs want to back up. So I've been carrying pages for a while. Um, and I do have opinions of what makes an incident response process good and what makes one bad. While naturally skeptical in the beginning, uh, since about January this year, I've started doing something that feels slightly transgressive to admit in this room, which is I start reaching out for Claude before I reach out to my monitoring dashboards.

01:52 But then you're asking yourselves the one question uh and I get asked all the time at dinner um someone finds me at this conference hallway at lunchtime my friends at my old company where their VPs are telling them to like do something with AI. Sorry for that. Uh so does Claude really fix your incidents? Is the uncle call rotation just claw right now?

02:16 have you basically automated yourself out of the job just like the original SR book wrote more than 10 years ago and I understand why people ask this like if anyone has this working uh it's the company that makes the model right we have unlimited tokens researchers sit a desk away from me I get involved in training the models surely if it's anyone it's us so let me get the answer out of the Hey, it's a no.

02:51 >> Uh, but I want to sit with the no for a second. Like, it would be genuinely hyperritical for me to stand up here and tell you that cloud fixes everything. My team just shortly will have its one year anniversary. If an LLM could carry a pager, we might not need to hire so much. like the fact that my team exists, uh the fact that we're hiring for many positions uh in like London and Dublin and the US and we have staff positions.

03:18 This should show to you that no, it it doesn't work. However, there's an asterisk over there. Uh and the asterisk is both related to the timelines. Like I think many of us would not be surprised if somewhere in the future we would be able to do such a thing. And today I'll also point out the useful ways in which clot helps me during my on call and but then going back like clot is down more often than any of us would like.

03:48 Like earlier today I was involved in an incident even if I'm at a conference and you may have noticed and you know some of you might have tweeted and I have the same sentiment when when cost goes down full of opinions. People are tagging me, please continue doing these. These are the positive tweets. I also got the the memed ones, but that's social media for you nowadays.

04:11 And uh there's even more like I know at least 10 companies right now whose entire pitch is some version of AIS. Uh since I'm also pretty open about angel investing and somehow people think I would be less skeptical about them and I'm more skeptical about them. There are now benchmarks. There are curated data sets. There's like historical incidents. There's like the game of how much percentage can your uh app solve of these incidents.

04:40 There papers. There's like a whole cottage industry. And the cynical take uh that we could have in in this room full of like senior engineers who've seen hype cycles before is of course this will be like there's VC money slloshing around. Someone was always going to slap AI on ops say that the TAM is billions that will swallow the observability industry and are above average compensation and they go they raise a series A from investors uh those investors are like building a diversified portfolio.

05:12 So there's AI law, there's help, there's AI law, there's AI healthcare and there's AI ops. And yeah, it's a lot. And this is not an exosis like by the time I was like writing the slides, at least two other companies uh appeared. And look, like the cynicism is warranted. Some of the benchmarks are like can the model solve like these prepackaged problem into a clean prompt.

05:36 But this is not like what something uh at 3:00 a.m. when you get page looks at. Actual incidents are not well formatted prompts. Um, but I'm not cynical about the goal. Like I genuinely, unironically, I'm cheering for these people. I want them to crack this. And it's not for the reasons that people typically assume. Uh, the reason is that on call is a tax we levy on humans because our systems are not good enough to look after themselves.

06:05 How many people here have you on call? So, you know the physical reality of it. like your phone buzzes, there's a half a second where it go from, you know, asleep to probably incident commander mode. And it's not that any individual page is bad. Like I like a good incident, but every other day, uh, it's the cumulative weight of this problem. It's never quite relaxing.

06:32 And by on call standards, I'm on the lucky ones. I have follow the sun rotation. London hands off to the west coast at 6 p.m. so I can go to the pub. Uh the worst case for me is I'm up at 6:00 a.m. But like some of you, you're up at 3:00 a.m. or or 1:00 a.m. Um and you're on a 247 on call rotation. There's probably four people. You get paid at 3:00 a.m.

06:57 Uh you go to sleep, get paid at 4:30 a.m. because the other database is broken. And then at 9:00 a.m. you show up at work and you have to be at your standup and look professional and presentable. and and I've been there uh in my career. And also I want to be clear, we treat this should not be a badge of honor. Our industry sometimes treats it like this like we call them war stories uh in panic rooms.

07:21 We use combat at the force but it is a cost in sleep in our attention in our relationships with our with the people and and when someone tells me we're building AI s I my first reaction could be that it'll never work but I'm like god I hope you get it right and there's the second reason which is more cold-blooded some of the best early career advice that I got from someone who was more senior than me uh was that you don't get pro promoted for fixing incidents.

07:55 And I remember thinking that was wrong when the first time I heard it because it feels like it should. Um, you were the hero. You got paged. You figured it out. You fixed it. The service came back up and people said thank you on on the Slack channel. Uh, it does count at first when you're a junior like can I fix a production system on fire? that that's real signal that you're on your way to be a good engineer, but it stops counting quite quickly.

08:25 Um, by the time you're senior, of course, you can fix this. Like that's that's a stable stakes. It's in your job description. What actually moves the needle in in big tech parlance? What gets you promoted? What makes you the tech lead is when people are like, "Huh, they're operating at a different level." It's building the thing that makes a whole class of incidents just not happen.

08:48 not fixing the one fire uh but making it structurally so that those fires will never happen again. And here's where the AI on call dream is it happens for me. If in big if an LLM can handle the generic mitigations, the obvious rollbacks, the have you tried the standard remediation stuff, then us the humans can go and work on the prevention part, the platform part, the stuff that scales and that's a dream and I want it to work.

09:23 Here's the framework uh that I use to explain incident management uh and I'll be using it across for a good chunk of the presentation. It's from a US Air Force colonel uh from uh that used to train fighter pilots. And he was trying to explain why some pilots consistently win against other pilots even though they're flying uh less better aircraft. And the answer was it wasn't the aircraft.

09:48 It was the winner was the person who would cycle the most through this loop of observe what's happening orient what does it mean what's my mental model decide what am I going to do about it act taking the decision and then again taking the feedback because the world changed uh and running this loop again and I feel this maps not perfectly but uh very well honest in response You get paged, you look at your graphs, you look at our logs, you build a mental model of of what might have been broken, uh you decide on a

10:26 mitigation, and then you go to the terminal and and write the command and submit it. The reason I like this frame, uh is that LLMs and incident response, uh they're not uniformly good or uniformly bad. They're just wildly uneven. We use the term jagged in the industry across the stages of this loop. It's genuinely super human at one of them. It's pretty dangerous in another and another part and it's really murky at two of the parts.

10:59 Uh let me walk you through about the observe loop where I think LLMs are superhuman like gathering signal from a system reading the dashboards quering the metrics pulling the logs finding the needle in the haststack. Um this is something that while we train ourselves and we get good at it um it's doesn't come naturally and it's also our limited attention is um not paralyzable not scalable and claude I'm just going to say it and you can replace cloud with any LLM that you use is just better it's not smarter necessarily

11:36 but it's so fast and it can clone itself it can hit every metric endpoint you've got in parallel with PromQL and it can write the syntax with almost no mistakes. Uh, and it doesn't get tired about changing the variables. Uh, it reads the logs at the speed of IO just as they come. And it doesn't get bored uh anymore by 2,000 4,000 lines. And this at scale is something that no one human can match.

12:06 And what we don't like the most than war stories. Uh, it's December 31st, uh, New Year's Eve. Uh, skeleton crew, I don't remember if I was on call. I don't think I was on call, but 500 HTTP internal errors are picking up on cloud opus 4.5. And I open my cloud code and I ask it to have a look with that particular prompt. Now, it's not just that prompt.

12:34 I have written my own skill MD that teaches the model the life of a request. How a request hits the edge servers and then hits the API front ends where we do admission control and quota checking and all the boring business logic and then it goes to the inference backends where we have accelerators like GPUs, TPUs and trains across clouds and the matrixification happens and the tokens go out to the user.

13:04 Once I've taught that there's another file where uh it's common across the company where it teaches cloud where's our data whouse uh what are the most important tables what are the fields and importantly how to introspect and self-s serve uh by itself if it needs more information and the next part of the conversation has been abridged and I have removed proprietary information but uh I have posted internally uh a big chunk chunk of it anantropic.

13:34 So plot goes and pull ups the last hour of server errors. It just writes a SQL query and within seconds it has the answer an unhandled exception in the image processing path. Weird 31st of December. It then goes in the codebase because it's in the monor repo. It finds this error type where it is raised and figures out a possible explanation uh for why it happened.

13:58 It actually posted the entire stack trace, but I won't bore you with Python. Uh, however, it doesn't stop there. This code path has been live for the weeks. It does get log. It looks if someone changes recently, and nobody's hit it until tonight. It pulls three failing requests, reads the raw JSON of the payload, and one has an image array with exactly 22 images, and a PDF is attached.

14:27 uh the second the third it checks even it grabs even more and the bug happens to trigger on 22 and why would that happen? Um so it asks who's sending these? It runs another query where uh it reps the request logs and it sees roughly 200 accounts uh all sending exactly 22 images starting tonight roughly at the same time. Now this as the younger generation says nowadays is pretty sus.

15:09 It runs an air query. It doesn't stop. is relentless to find out how many accounts were created in that timeline. Roughly 4,000 of them, same window, same email template, same provider. Uh to quote the frameless singer songwriter, I knew you were trouble when you walked in. The other 3,800 accounts uh haven't sent a single request. they're dormant, sitting there since December 28th waiting.

15:43 It runs another query to find out how many uh it and then it looks at how uh often did this the these accounts were created and it saw that they were using nine signups a minute. Uh that looks roughly what I would expect someone in account abuse maybe setting a limit. Um so it then says stop looking at the 500s. This is fraud. Not might be suspicious.

16:17 Not worth flagging. I have literally left the claw slop where it gets really excited where it finds something because um I was shocked like I would have only looked at the 500 internal errors. I would have marked this as a bug for my API team and I would not have paged account abuse on the 31st of December to join in with me and start looking on their side of things because I do not have access on PII for what was happening there.

16:48 Um, you would think this would be my keyboard from now on. I went to get a computer science degree and I get three keys. Now, this story in AI, there's this famous move 37 where Alph Go, the deep mind algorithm was playing Lee Settle uh over 10 years ago. The celebration was today's uh ago and all the go players were shocked of like why why was the AI doing that and many people have personal move 37 moments.

17:23 This was my move 37 moment like I knew on that day that this year would be special story number two. This is also a positive one. We had rust panics in production while I was going through the logs manually because this is something uh that that I still do. I don't know why one of my newer teammates not even on call trained uh two to three months in the team just pointed clot code at the logs and it instantly found the root cause like a rasp panic in checkpoint RS with some segment ID validation the file the line

17:59 number the assertion message even before I finish reading my two pages of logs. Um and then he prompts again to get the full volume breakdown like a gazillion panics in 5 minutes then drops then spikes again. Uh claude is very helpful like across two independent hardware platforms so I don't need to go to my CSPs and ask them. Um and it's not reasoning about the cause yet.

18:29 It's just gathering signals way faster than even you know the subject matter experts uh can do it. Um the panics are all in the pre-flow servers. It involves duplicated segment ID. Let me check when the binary was deployed whether it's a rollout change right before the panic started or not. And it just finds it like before me like the new person in the team is way more effective than than the people with experience because they're unencumbered by the fact that they should go manually.

19:06 um and and look at these things. Um again, now we talk about the orient. This is where the UDA loop uh gets interesting for LLMs. Um and this is a failure story finally. Uh but the first thing I need to explain to you is something uh that happens with LLM inference uh so you actually understand the business logic uh and what went wrong. So when you prompt claude of or any transformer the the naive way to generate tokens is to take the entire uh sequence um never going to give uh it's the first one and then put it

19:59 entirely through the transformer and then you would get the first token called um up and then you take that entire sequence again put it again through the transformer and then you would get the next token which is never and you're appending the whole thing. But this is wildly inefficient and no one does the naive one. This is just using academia to explain because by the time you're producing token a thousand, you've reprocessed token one a thousand times for the same uh matrix multiplications.

20:34 The trick uh is called a KV cache uh during what we call the attention step where each token produces a key and a value vector and crucially these don't change so you can save them somewhere and that's the diagram at the bottom. So now inference has two phases. There's prefill where you process the entire prompt in one big parallel pass. you save all the key value pairs uh into this cache uh and this is what we call computebound because you're doing a lot of matrix multiplications and then on the decode or generate

21:11 side you take this KV cache and then you feed one token at a time in the machine and instead of reprocessing the whole prefix you just have one pass. Uh and this is why if if you've ever wondered uh tokens appear like a stream in in claude in chat GPT and Gemini um because they are really generated one by one uh if your inference provider is using a big batch size you can see it even really slow if one machine goes broken uh it's the same but typically we maintain it to so you have a good user experience now this KV

21:53 cache can be gigabytes in size and it's really hard. It's really easy to break it. It's very finicky. It's fragile. Uh so when the KB cache blows up, you suddenly have to repreill a lot of prompts and that's a lot of compute you didn't have it prepared. And this is I would say a class of incidents that happens quite often uh with plot. And when this happens, this is how my graphs look like.

22:28 Um the gray line is filtered requests. The red line is the number of errors. Orange is the incident window. And look at the shape like requests roughly double at exactly the moment errors appear. And then they both drop together. And every single time I would ask Claude what happened here with this graph, Claude will say request volume increased. So this is a capacity problem.

22:58 You just need to add more servers. And I have corrected it on this I don't know six, seven times. you add it to cloud MD uh and it understands about this situation but it will get wrong correlation versus correlation in the other 99 situations. Um and it's not helpful. Like if you have a new joiner in your team, they would immediately be swayed in this direction.

23:26 Uh and they will immediately think, oh, it's a capacity problem. Where actually what you did is you lost your cash. How about you go and fix your cash and understand what was happening and maybe you can salvage it. Um and this is why I think we can't trust LMS for for incident response. if it could take a step back and try discerning between causation and correlation.

23:48 And I know for us humans it is hard as well. Yes, we we fight every day with with this. We we see two lines and we're like, "Oh, it was a roll out obviously or or it was something else. We have bad scars. We have experience. I have seen this outage too many times in the in the past year and I I can know to ignore this uh avenue and I can go and explore the other avenues.

24:15 Um and let's talk about automation from from a higher levels perspective. We in software engineering are not the first ones to have to deal with automated systems. The society for automotive engineers the people who make cars published a framework in 2014 over 10 years ago uh about self-driving cars because that's when we started talking about self-driving cars.

24:37 Still not here in London. Uh but like they tried mapping how would self-driving car cause look like so they could have mental models. So they they wrote about um several stages several parts with several stages. So one thing self-driving cars need to do is like the execution like steering and acceleration. The other thing is monitoring the driving environment.

25:00 Uh and the other part is fallback performance. So something which is got name for like something unexpected happened like an intersection is blocked or you need to overtake a block call. Uh and finally the their uh biggest challenge uh and they call it um the system cap capabil the scope and the system capability uh which actually in my mind reads would this car just self-drive itself when it it's sunny in California or can it actually self-drive itself when it's raining in London uh because you can't just throw one

25:37 into the other as you go up in the level more and more of those columns get filled by uh systems instead of humans. Sometimes it is both at a particular level. Most modern cars you can buy now are actually level two. So we get automated braking because we trust that technology. We have line following. There is cruise control. Level three is where it gets uh interesting.

26:07 Like uh I think here Tesla has full self-driving in the US and in Europe we've got Mercedes and BMW but only in Germany. And sorry that was level three. Level four is is the exciting one. That's Whimo. Uh that's you don't have a driver. It's just a taxi. You you go inside of the car and uh you look at the steering wheel uh moving itself. You watch on the street how cars with no one in are moving from A to B.

26:37 Um ways awe in the streets here in London as well in trial. I have complained to a friend of mine walking there. Why are they taking so long? Like in my mind you take the model weights from San Francisco, flip the signs and the car can drive on the other side of the road. He tells me that's not how research is done. Um but um I wanted to map this to to my job to production engineering and some of you have seen this before maybe in a previous presentation but in my mind the columns for for me are detection who

27:16 identifies that something is wrong either spotting anomalies or correcting signals which is are the fancy ones or actually do I have good alerts when I have internal errors then There's uh the initial response uh who executes the mitigations like restoring the servers doing the cubectl scale commands the failover the rollbacks uh decreasing the user quota for the errand user the stop the bleeding uh that we typically do.

27:46 Then there is prevention uh which someone will do a deeper walk to improve the system like um looking at the contributing factors what are the architecture changes uh the what make sure this doesn't happen again work and then as as the people from sales and cars have the operational domain like does this system uh handle only the carefully instrumented flagship ship service that I have or can I have this AI agent just parachute it in a different team.

28:21 It will learn by itself uh and figure out what's going on. Now I've removed level zero like we were lucky we got automation because computers are pretty good. Level one is just you know some assistance for detection. Uh and in level two is where most system most mature systems get to like observability is fully automated. you have alerts firing reliably even if it's not AI that that's still automation you've correlated some signals um and it's it's where we arrive but the humans still execute the mitigations and the

29:02 humans still write all the fixes so can we get from level three where right now we are human and AI uh in in the response. Can we get to level four, which is what what I've been what I've been talking about like remove the humans from the incident response and level five. Um that's just AGI in my mind. Like if you if you can throw an agent everywhere and you can have it as teammate um that will be really cool.

29:40 Now that was the broad picture. That was the the generic that that was the sitting scene. Um I now have five things uh that I have seen working uh in my team in in my company and and for myself and they've emerged for about one year of trial and error uh across everyone. Pattern one, you have been paged and there's an alert and you're a sensible person.

30:08 So your alert is somewhere committed uh in your infrastructure as a code. Uh and it has a promql or whatever query language is behind it. Uh don't open a blank chat. Don't don't just uh copy and paste things. Just give your AI agent that expression and hand clot some instructions that roughly read as be curious. Pivot this for me. Here's how my system is set up.

30:38 Find things uh that I wouldn't normally look at. And your cloud is parallel or mine is uh it will look at your endpoints. It will look if it's a per region or per cloud thing. It will do uh things by status code. It will look at last week how the traffic compared and it will issue those queries I think faster than you would be able to drill down into your dashboards.

31:06 All if you start this when you get paged and then you arrive 5 to 10 minutes at your computer, you already have a scene set up for you. And we found this extremely helpful. Like as in my previous example with the fraud, Claude is able to catch a lot more things and is very helpful. And this is not a big change just putting the alert in a prompt uh asynchronously for an agent.

31:34 Number two, get an ex simpler, trace it through to your system, but actually uh just ask lots to do it for you. Uh I run desperate systems with different logging patterns and with different log sources uh in in different clouds and the one thing that unites them is thank god a trace ID. So whenever I need to look at what happened at the ingress or at the API or at the at the inference level, previously I I would just open um each of those windows, I would have the bookmarks or I would have a bookmark of bookmarks.

32:12 But if you tell Claude to just do it itself, it actually is very useful because it builds a timeline in its context and you can start talking with the timeline. you can see, oh, that request bounced around three servers, went back to the API where the 10-second timeout expired, and this is why I returned the error. Uh, and and that shows to you that, oh, yeah, it wasn't at the API the timeout, it actually went to the backends.

32:37 And it does this, and uh it's able to with coherence keep um a life of a request. And it's also able to like tell you, oh, it was this server. And you can start talking about, oh, what about that server? did it had a CPU high or what is the condition on the memory and it's it does open a lot of avenues um even if sometimes you do have to reverify the information uh that you got from the model pattern number three is what changed uh in this window uh and and I open a discussion in the conference for this track I joined

33:20 a company where there were not many deploys and there were not many config pushes and while you know rolling out monoliths it it's great uh it's it's also doesn't scale so now I live in a microservices world where like services go go out all the time and config pushes and feature flags and cron jobs and it's really hard for me to correlate what changed with the moment my my incident started.

33:47 So we we've built uh this you know handmade tracker of of all the big changes that happen uh across the company. It is curated in the sense that we're not going to do it with like 100 QPS of events. Uh but if you give claude the time window of when your incident happened uh it will let you know really fast about the deployments and and not only that it will go in and check oh maybe there's a there's a commit there that affects the image prep-processing uh which is very very useful for getting avenues of debugging uh

34:26 and discovering and it's a bit like what I would do by hand previously. Um, now obviously you need to stop it to be really excited by I found the root cause. It's like no, uh, you found three possible issues that could have happened and and like we really need to dive through. But if it says maybe roll this back and big if your roll backs are easy, uh, then yeah, just do the roll back.

34:56 Check if your errors uh are are now down. uh it's much faster and it saves um a lot of effort. Number four, postmort more tense or retrospectives. Cloud is good at the tedious bots. Like I used to hate these. Uh I'm I'm really sorry for for some of the engineers that I worked, but I I sometimes felt it was um picking through fire to get your your more juniors or or medium engineers that join the team.

35:26 I was like, "Yeah, you were part of this incident. you did not fight uh the big fight but as a prize that you're part of it you will now write the two to three page postmortem so you understand the moving parts so you're able to um to learn what could go wrong uh and it was always a chore like no one wants to collate timelines or slack threads uh but now I take the whole Slack channel uh I take my Google meet transcript I put it in text format I give Claude um a prompt with the postmortem template that's like 80% the

36:03 one from the S sur book. Uh and it does produce something. Two issues though. One it gets really really bad at at root causes. It's it's very jumpy. It's like okay this was the thing and we all know it is not one thing. We all it's not one root cause. There are many contributing factors. There are many issues. They're like the Swiss cheese model that you have to go through.

36:30 It was never the rollout. It was never the code change. It was all the processes in the company that allowed you to go uh in in the incident. And CO doesn't know the history of your system. Uh especially if your system is like has been there for 10 years. It doesn't know the reason you didn't test your secondary uh database um fallback and it doesn't know tacet knowledge that you know.

36:54 So, while you get an 80% story that's pretty um it's readable and convincible, I ask you strongly to not share it until you have reviewed it. Um it also if documents like these proliferate across the company and they're not vetted by humans, there's a loss of truth that that happens. And I've seen it. Uh, one trick that we're now doing for postmortems is we tell Claude if you don't feel certain about something, insert a to-do for a human.

37:32 Uh, and it has a capacity of self reflection. Not perfect, but it is there. And and we know that uh a postmortem is not for official consumption until all the to-dos have been solved. Until we've fixed all the issues uh where we're not sure on that text. Because if you feed these postmortems again in a loop, you know, garbage in, garbage out. Uh but you can also have a good source of truth of what were your incidents and when clot summarizes that vetted uh collection of documents, uh it is much better to write your

38:08 road maps. Number five, and I I've left the easy and the fun one. uh shift handoffs like uh just just as opposed uh you can point claude to write your handoff but I have to stress you need to and what the kids say nowadays build in public like all your on call actions are in a channel where the on caller thinks loudly by writing the debug sequence and everything that they've did uh during the the handoff uh and it's good uh it can also pro clot to like uh just stick to the summarizing part.

38:43 Uh, and the person that's coming in doesn't necessarily need to read the 200 or 400 messages that happened during the London daytime. They just can read the one paragraph with links to go to go to the threads. And this is probably like the single most uncomplicated win uh you can get into. Okay, nearly done. But before I wrap up, I want to talk about something that I'm generally not not resolved on.

39:09 Like um it's it's the learning problem. Like if Claude found the stack trace and suggested the roll back and you approved it and it worked, what did you learn? Um senior incident responders aren't smarter. uh they don't know the system beforehand, they have been burnt before, they have seen other people um debug the systems, they have the scar tissue, and if AI starts doing this, will we have our skills atrophied since there's no more feedback loop on the actions that you need to take to know what happens to the system?

40:00 like in in the SR book the reason we had a small team of humans be on call for for a mission critical system and we didn't have a roa of 50 100 senior engineers would do it once a year was because if we would put those engineers on call uh and it was their job when they were not on call they would treat fixing the the issues uh behind the pages as their priority because you don't want to get paged again and this is your whole mission and this is what you do in your entire team but if the AI starts fixing this you

40:41 don't get the paid anymore so you your incentives might not be align and that's on the senior side on the junior and and the mid side you won't get trained in like the easier incidents you won't have to be put on the spot to like I don't know type the roll back command without thinking or like looking at the history in your playbook so when the big thing happens that the model can't fix, you might be miscalibrated on on how to respond to such an incident.

41:09 And I'm I'm genuinely worried about this. um a researcher for OpenAI uh with their name Rune tweeted in 2023 so this is before cloud code uh and before before even people knew about entropic that Jeans's paradox will allow software complexity to drastically increase until it's very hard to do software engineering now this paradox for those of you who don't know It's the favorite paradox in the AI industry.

41:46 Uh it's when technological improvements increase the efficiency of a resources use, but the resulting lower cost actually causes consumption to rise up rather than fall. And in our case specifically, it's easier to write software. Uh so we write much more of it. So the complexity goes up and not down which means things break in more interesting ways which means more incidents which means more on call.

42:18 Um the follow-up reply um is that all improvements in dev tooling that we can do will be cancelled by this evergoing complexity. However, however, and like you know pulls a magic wand, uh we could have agents break this cycle. Uh we could spend arbitrary amounts of compute to simplify and manage the complex systems like we in this room have made recipes.

42:51 We know how to grow teams. We know how to grow organizations. We have microservices. Um, and maybe the AI agents can do what we've collectively learned uh in in our industry, but that's a big if. Um, and and yeah, unsurprisingly, Run is a periphery poster and and this has aged very well. Uh, and it keeps me up at night. Um, I'll let you with the favorite chart in AI.

43:23 I've seen it uh in the presentation uh beforehand. Meter is an organization that evaluates AI capabilities and one of their studies is how long can a model run autonomously. So on the x-axis you can see the release date of models. On the y- axis uh you can see aggregates uh duration task uh and each model is evaluated on a set of tasks where humans were previously asked to do it and measured for their success and timed.

43:52 And if a model manages to complete at least 50% of the tasks, we mark the duration uh on the chart. And you can see the exponential trajectory of like models are getting better and it's not that they're just getting better as we expect linearly. Uh it seems we're accelerating. Um very many uh tweets have been spent debating this. Very many people have been drawing sigmoids that the growth will stop just like growth uh has stopped for for some of these lines.

44:32 People have been claiming since 2022 that uh yes it will. But we have the AI scaling laws. uh we know that models will improve uh once we have more data and and more compute they will become more intelligent. The models are the worst today that they'll ever be and they'll be getting better from from tomorrow onwards uh as as we're working on this. Um, and with that in mind, thank you very much.